SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, relia
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「Agent」の検索結果
267 件While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, relia
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL),
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and a
We study distributed one-dimensional mean estimation under a 1-bit communication constraint. Each agent observ
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous
Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. B
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the cle
Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix: they
Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before t
Large language models are increasingly deployed as agents that plan over long horizons and act through externa
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a mod
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture
Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences an
Modern science and engineering increasingly rely on time-varying data, yet the mathematical tools used to mode
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic proc
Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedu
Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepan
Multi-agent traffic simulation seeks diverse, coordinated, and physically realistic futures from maps and obse
Automated alpha factor discovery searches symbolic trading signals from price-volume panels and order-book dat
Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual ob
Large Language Model (LLM) agents are evolving from single-session tools toward long-term personal assistants
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks us
Recent large language models can emit task-progress signals that agent frameworks use to decide whether a task
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized
KV cache is evolving from a serving optimization into an external memory substrate for long-term LLM agents. I
Personalized memory helps LLM agents deliver stable, tailored assistance by storing and reusing user-specific
Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment
Aerial Vision-and-Language Navigation requires drones to follow natural-language instructions and navigate thr
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-tra
Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmann
The transition from human-centric assistance to Autonomous Software Engineering (ASE) agents has enabled the r
Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherenc
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Long-running language-model agents depend on persistent memory. Many agent-memory systems preserve history thr
Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute,
Modern LLM agents increasingly rely on reusable skills, yet as skill libraries scale to thousands of entries,
This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark that
Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or orchestration in
How can vision-language-action (VLA) models adapt to new environments where world dynamics shift? While recent
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging re
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand
The Indonesian Digital Library of Culture (Perpustakaan Digital Budaya Indonesia, PDBI; budaya-indonesia.org)
Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can cause
Tool-using language agents can delegate and revoke permissions while acting through external services. We show
Accurate citations are the foundation of academic writing, tracing intellectual origins and substantiating cor
In June 2026, thousands of AI agents found that a small public wiki would accept edits from inside their sandb
Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interacti
Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale
Long-term memory is a core capability for personalized LLM agents. To support it, existing memory systems orga
Solving repository-level code tasks requires LLM-based agents to use code search tools to navigate large codeb
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabil
Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the so
Organoids are three-dimensional tissue models whose morphology provides important insights into tumor developm
World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance w
Scientific figures are the interface through which research claims are inspected and reused, but final publish
Maritime target tracking over large distances often requires multi-agent teams without centralized coordinatio
Messages from electronic devices are conventionally received as text, audio, or radio signals. But robots move
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative nav
Large language models (LLMs) and vision-language models (VLMs) have significantly advanced zero-shot task plan
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same env
This note gives an instance demonstrating that the pairwise maximin share (PMMS) property cannot be satisfied
We investigate the query complexity of fairly allocating $m$ indivisible chores among $n$ agents with additive
We study risk-sensitive evolutionary learning dynamics and their long-run equilibrium selection behaviors in c
We study strategyproof mechanisms for building a pathway between two regions of a line segment separated by an
Offline multi-agent payoff models are estimated under a logging distribution but used on distributions induced
We study envy-freeness with subsidies for indivisible items beyond additive valuations. Assuming that every si
The software supply chain has become an increasingly exposed attack surface because of its reliance on intrica
Machine learning has shown its largest gains in the band below 10 ms inside a 3GPP new radio (NR) 5G distribut
Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Per
Most automated brain parcellation tools are developed and validated on T1-weighted (T1w) MRI. Yet, some clinic
In electric delivery fleets, mid-shift charging is non-trivial: each vehicle must decide when, where and how m
Closed-loop AI scientists can generate candidate designs at low marginal computational cost, whereas reliable
AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this
This work introduces an alternative view of efficient exploration and studies its theoretical and empirical im
Tool-augmented language agents are vulnerable to indirect prompt injection (IPI). Unlike direct prompt injecti
We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition
Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforceme
Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve an
Long-running AI agents may read state, reason, wait for tools or human approval, and perform an external actio
Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is
Process mining has long turned event logs into process knowledge: discovered models, conformance evidence, bot
We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and eff
Multi-agent federations need governance that answers three questions under adversarial conditions: who partici
An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Publ
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only parti
The hype around generative AI seems to promise unprecedented productivity (and learning) gains. However, these
Generative and agentic AI are reshaping both the production and evaluation of scientific research. These devel
Mobile GUI agents can execute tasks from natural-language instructions, but their evaluation remains difficult
Transaction-local controls answer whether one financial request may proceed, but market behavior can be distri
Democratizing access to the knowledge held in large corpora of tables such as data lakes is emerging as a cent
Environments are increasingly populated by multiple robots performing independent tasks with limited prior kno
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and a
Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workfl
The application of large language models (LLMs) to personalized medical assistants has garnered growing intere
Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making
Linguistic typology relies on expert analysis of reference grammars across languages, making large-scale cross
Large language models (LLMs) are commonly associated with the distributional hypothesis, according to which (1
Long-running LLM agents rely on external memory to store and reuse information beyond a single context window,
Resource constrained single-board computers including Raspberry Pi, NVIDIA Jetson Nano, Arduino UNO Q, Orange
Converting in-service reinforced-concrete (RC) building blueprints into simulation-ready models---structured f
Large language models (LLMs) increasingly generate Markdown that is consumed by renderers, agents, code extrac
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts
Long-form narrative-to-film generation requires shot-level controllability and cross-clip consistency in both
Agentic systems can interpret user requests, search the live web, and use external tools, but their ability to
Climate change is increasing the severity and unpredictability of natural disasters. In time-critical crises s
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in
Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fide
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first
Reasoning is a promising route to the generalization that autonomous driving requires in the long tail, as it
Localizing a font into new languages is a highly intricate task requiring precise design adaptation of glyphs,
Multimodal continual learning has recently shown great potential for developing agents with human-like intelli
Learning physically plausible dynamics from visual observations is essential for interactive world models and
Autonomous tugboating is central for automating maritime operations such as port logistics and vessel maneuver
Local pedestrian-vehicle forecasting spans heterogeneous physical scales: pedestrians combine root locomotion
Automation in underground mining has the potential to significantly enhance safety, operational efficiency, an
Distributed Dexterous Manipulation (DDM) is a novel paradigm that presents significant control challenges due
Distributed learning control for multirobot systems (MRS) offers significant flexibility in presence of uncert
Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent
Quantum game theory is an extension of classical game theory that uses quantum principles in game theory. The
We study randomized strategyproof mechanisms for strategic obnoxious facility location on a line segment, wher
Generalized Nash Equilibrium Problems (GNEPs) often arise in multi-agent engineering applications that require
The optimal transport (OT) map provides a geometric transformation for aligning probability distributions and
An open, networked web will allow agents to run frozen models from multiple vendors, keep their history privat
LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations o
High-quality structured organic reaction data are essential for developing artificial intelligence for chemist
Sequential memory agents process long documents by reading chunks one after another while maintaining a compac
Enterprise AI is evolving into an Enterprise Operating System where autonomous AI agents can plan, reason, use
Large language models (LLMs) are increasingly used for automated data visualization, yet existing approaches o
Graph-agentic retrieval-augmented generation combines structured evidence with adaptive controllers that can p
Recent advancements in deep learning allow robotic agents to interact with dynamic and unstructured environmen
The generation of safety-critical traffic scenarios is essential for training and evaluating autonomous vehicl
Homes change between a robot's visits. Navigation benchmarks pose their goals in the world the agent currently
Generalist robots promise to transform our society: the same system that prepares a meal or folds laundry migh
Multi-fingered dexterous manipulation remains a frontier for real-world reinforcement learning (RL) due to the
Optimization-based simultaneous localization and mapping (SLAM) makes it possible to reduce accumulated naviga
Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios,
In autonomous driving, Bird's-Eye View (BEV) representations provide a structured, top-down abstraction of the
We study the fair allocation of indivisible goods among agents with strictly positive additive valuations. Pai
We consider a multi-agent Capture the Flag (CtF) scenario in a graph-based environment, where a team of attack
Can interactive vision-and-language agents learn not just what to say but also \textbf{\textit{when}} to say i
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in
This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reason
Diffusion probabilistic models can capture the multi-modal, interaction-rich distribution of joint future traj
Cooperative rehabilitation enhances engagement, task performance, and social-motor interaction, yet it demands
Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred fr
The autonomous localization of fugitive gas emissions using small Unmanned Aircraft Systems (sUAS) constitutes
Building on the pioneering paper of Kearns, Roth, and Ryu (SODA'26), we study information aggregation in a net
We study the online fair division of indivisible items, where items arrive one at a time and must be allocated
Envy-free cake cutting is a central problem in fair division with a striking divide between existence and comp
The strategic facility location problem is defined as follows: $n$ agents report their location in a metric sp
Multi-Task semantic communication (SemCom) prioritizes simultaneous execution of multiple tasks over bit-accur
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart
This technical report is a study of the use of differential game (DG) theory to solve the target-assignment an
This paper presents an energy-based controller for a multiagent robotic system designed to achieve and maintai
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language in
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
We study the allocation of indivisible goods among agents with identical additive valuations, focusing on envy
Gaussian-process Bayesian optimization (GP-BO) excels at black-box optimization of costly functions, e.g., hyp
Multiobjective evolutionary algorithms (MOEAs) naturally expose population-level parallelism, but many mature
This paper is the first in a series on Turn-Based Combat Arena, a configurable framework for turn-based strate
A central challenge in mechanism design is to develop truthful trade mechanisms that maximize the expected gai
We investigate the fair allocation of indivisible items among agents with asymmetric entitlements in mixed man
Language model agents increasingly propose actions, observe external feedback, and explain their own behavior.
Multi-agent combinatorial optimization problems are notoriously challenging due to their NP-hard nature. Recen
Information sharing can improve a pooled estimate while eliminating independent rescue actions. This paper sep
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (fea
We study weighted fair division of indivisible mixed manna under additive valuations. First, we resolve the ge
In this work, we study radically uncoupled learning in discounted general-sum Markov games. Assuming ``$\maths
We study federated online reinforcement learning with linear function approximation. While recent multi-agent
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advant
We study fair allocation of indivisible goods under additive valuations and matroid constraints. A challenging
Bidding games are graph games in which a token is placed on a vertex, each player starts with an initial budge
We study the problem of locating a new homogeneous facility under a prelocated facility. Here, a set of $n$ ag
Envy-freeness up to any good (EFX) and pairwise maximin share (PMMS) are standard local fairness criteria for
Residual maximin share (RMMS) is the largest share threshold that remains guaranteeable throughout dynamic all
We study complete allocations under nonnegative additive valuations through the positive supports of goods, fo
In this paper, we study fair division problems in which resources are structured as graphs and agents must rec
We study fair division of indivisible items when agents have arbitrary two-level preferences: the value of eac
Recent machine learning research has increasingly focused on equilibrium analysis in non-cooperative games rat
Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subga
This paper studies the problem of proportionally fair clustering, where the goal is to select $k$ ``centers''
We study active elicitation of agent preferences for collectively choosing among $m$ alternatives using promin
Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and
We study risk-averse decision making, in which an agent selects actions while being uncertain about the true s
Forecast combination is a reliable way to improve predictive performance when several forecasting models are a
This article introduces peer $k$-oversight, a property of sequential collective decision mechanisms requiring
The VCG family and the AGV mechanism are two classical approaches to efficient implementation in the static so
We consider two-player GR(1) games on graphs, where the system player Eve must satisfy \[ \Box\Diamond A_1\lan
As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them
Coordination is a desirable feature in multi-agent systems, ranging from robotic swarms to socioeconomic netwo
We study contraction properties of non-stationary continuous-time mean-field games (MFGs) under discounting an
In the study of Nash equilibria of finite-player games, one often seeks equilibria that are compatible with pr
As electricity market participants increasingly adopt learning-based agents for their bidding strategies, elec
In the subjective divisibility allocation model, all goods are divisible, every agent $i$ has a non-negative a
People's trust in AI advice diverges as they use it, deepening for some and eroding for others. We study this
Modern multi-agent systems are increasingly deployed at scale over large populations of agents in settings suc
Paradigmatic interaction models explain how collective behaviors can emerge in complex systems from interactio
Recent work in fair division has focused on either simultaneously satisfying closely related fairness notions
Social interaction can improve collective learning but also amplify early mistakes. We study this tension when
We consider fair allocation of indivisible goods in a setting in which agents have subjective valuation functi
Maximizing Nash welfare over indivisible goods is a central problem in resource allocation. For additive valua
We study multilevel fair resource allocation with tree-structured hierarchical relations among agents. At each
In truthful interval covering, each agent has a private interval of unit length, and the goal is to decide whe
This paper studies a class of multi-cluster aggregative games characterized by the coexistence of cooperation
In a deliberative poll, once submissions outnumber what anyone will read, some mechanism chooses which argumen
While contemporary Evolution Strategies handle integer optimization problems effectively, their adaptation mec
Reinforcement learning (RL) algorithms have made strides over the past decade applying them to a wide range of
Standard solution concepts for stochastic games, such as Markov perfect equilibrium and Markov coarse correlat
We design and analyze randomized strategyproof mechanisms for multi-facility location under the utilitarian so
The house allocation problem is a classical one-sided matching problem that concerns the assignment of a set o
We study temporal fair division of indivisible mixed manna. Items arrive over time and must be allocated irrev
Python is widely used in scientific research because it enables rapid development and provides rich ecosystems
We study the classical and parameterized complexity of efficient connected allocation problems on graphs, wher
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fu
Classical seismic data reconstruction relies on manually designed structural priors and iterative operators, w
We introduce and study an online variant of the multi-agent contract model. In our model, agents arrive one-by
Learning algorithms are often used to make decisions in repeated multi-agent environments. When another player
We investigate whether agentic artificial intelligence can automate parts of the process of designing genetic
Peer prediction seeks to incentivize agents to truthfully report an observed signal by rewarding joint sets of
We study the strategyproof placement of \(k\) facilities on the real line for \(n\) agents who privately repor
Shared autonomy requires principled mechanisms for allocating and transferring control between a human and an
Shared autonomy requires principled mechanisms for allocating and transferring control between a human and an
We study the group-fair distortion of metric facility assignment problems, where a set of agents, partitioned
Proportionality (PROP) is one of the simplest fairness criteria for allocating items among agents with additiv
We study fair division of indivisible goods when agents' valuations are accessed only through ordinal comparis
We consider strategic facility location in Euclidean space $\mathbb R^d$, where a mechanism selects a single f
The maximin share (MMS) is a central fairness benchmark for allocating indivisible goods and chores. We study
In Shapley-Scarf housing markets, Ma (1994) shows that top trading cycles (TTC) is the unique mechanism satisf
We study interval scheduling from the perspective of fair allocation. There are $m$ identical machines and a s
Many real-world interactions among self-interested parties can be modeled by game theory, and the rapid advanc
We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each play
We introduce the problem of designing mechanisms that incentivize strategic agents to form self-funded marketp
We study the facility location mechanism design problem where $n$ strategic agents report locations in Euclide
We study equilibrium pricing in oligopolistic data markets with budget-constrained buyers (e.g., machine learn
We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavio
Liquid民主主義の選出問題において、選出プロセスを効率化するための手法を提案。各意思決定者が他者に信頼する方法を考慮し、意思決定者間の協力と意思決定の正当性を考慮する。
Peer-to-peer (P2P) energy trading markets rely on double auction mechanisms to match prosumers and consumers i
本論文では、LLMエージェント間の相互尊重を確立するために、「類似性シグナル」という新しいアプローチを提案します。このアプローチは、エージェント間の類似性を分析することで、相互尊重を促進するという考え方に基づいています。
この研究では、半推測状態の状況でアダプティブ行動を示すために、反射的な組織がどのように機能するかを調査する。これでは、現時点の観察だけを基に、内部状態が情報を保持できる計算プロパティを開発します。
この研究では、LLMへの促進の強化学習を効率化する目的で、検索コストを削減するためのコスト意識のあるクロスタイア転送を提案します。検索コストは、促進の評価に伴うLLMの回答によって大きく異なるためです。このアプローチでは
この研究では、ソフトウェアの開発が複数のエージェントによって長期間にわたって進行する場合の持続可能性を考慮した新しいアプローチであるEvoX Genesisを提案します。
We introduce the study of \emph{multilateral trade}: a mechanism-design problem in which a single potential tr
We study the existence of envy-free up to any item (EFX) allocations of indivisible chores when agents have mo
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instea
We study the agent-wise disjunction of two central fairness notions for indivisible items, where every agent m
We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss en
We study envy elimination by adding goods (EEAG) when the additional pool has bounded supply and no separate b
Inspired by possible future markets of autonomous routing and driving (ARAD), we introduce competitive mediato
Learning dynamics in zero-sum games are typically analyzed under algorithmic symmetry: both agents use the sam
理論オブミンドの評価基準「Avalon-ToM-Bench」を提案。社会的認識を評価するための基準を提供する。
We study fair allocations of indivisible items under general set valuations. We prove that every instance with
This note aims to serve as an entry point to the literature on learning in games, a topic with significant the
We prove that no randomized integral or fractional algorithm for online vertex cover under general vertex arri
We study strategyproof mechanism design without transfers for the two-facility location problem in metric spac
We consider fair allocation of indivisible items among agents with non-negative and additive valuations. The g
The leximin++ proof of Plaut and Roughgarden for agents with identical monotone valuations gives a natural EFX
We prove that every fair-division instance with four agents, additive valuations over the non-negative reals,
Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion?
We study Evolution Strategies (ES) for continual control, where agents must adapt to changing tasks without fo
Automated formulaic alpha discovery aims to generate predictive and interpretable trading signals from large s
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and co
物理的システムの分離方程式を解くためには、ニュラルネットワークの設計、損失関数の定義、および最適化ダイナミクスの手動調整が必要である。研究者は、自動設計のためにLarge Language Models (LLMs)を利
Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or bo