FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal veri
- 用途
- 技術検証・論文読解補助
- 難易度
- Easy
- コスト
- High
「text」の検索結果
35 件A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal veri
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together wit
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improv
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language mo
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks t
Model compression techniques such as pruning and quantization facilitate the efficient deployment and accelera
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucin
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space,
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-la
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a sev
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recogniza
Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this
Modern Text-to-SQL systems often follow generate-execute-select pipelines, generating multiple candidate queri
A language model's prediction of its next token develops across layers, and lens methods track this process by
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from cur
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behav
Agent performance depends jointly on the model parameters and the executable harness code that manages context
Multimodal models often build on architectures designed for generative vision-language modeling, typically com
Information retrieval (IR) increasingly targets open-ended queries that admit diverse perspectives. Existing I
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Sn
Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scor
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoo
Conversational Recommender Systems (CRS) typically require domain-specific dialogue data, which is costly, sca
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements
On-policy distillation trains a language model on its own generations while a teacher scores them token by tok
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy la
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev