NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「Agent」の検索結果
30 件Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate larg
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop inter
Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogene
Predicting a football match before kickoff requires more than knowing past results: a model must use changing
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get co
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task,
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interac
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents ty
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. H
Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex t
As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and i
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in hist
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating compo
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still he
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphas
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedur
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer dire
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs tha
Building assistants that can continually watch the world, remember what they see, and reason over their accumu
Code review helps maintain software quality before code integration, but it also imposes a substantial workloa
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning r
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet
Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurem
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environment