SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, relia
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「Agent」の検索結果
17 件While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, relia
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the cle
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture
Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedu
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherenc
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand
Organoids are three-dimensional tissue models whose morphology provides important insights into tumor developm
The software supply chain has become an increasingly exposed attack surface because of its reliance on intrica
Mobile GUI agents can execute tasks from natural-language instructions, but their evaluation remains difficult
Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workfl
The application of large language models (LLMs) to personalized medical assistants has garnered growing intere
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
Gaussian-process Bayesian optimization (GP-BO) excels at black-box optimization of costly functions, e.g., hyp