NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「RAG」の検索結果
22 件Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate larg
Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supp
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, whe
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. H
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generat
Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent met
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedur
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diver
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of
Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce waterti
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer dire
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). Howev
When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated dataset
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environment