Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) p
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「text」の検索結果
54 件Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) p
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering
We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-se
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generat
We present a continuous, population-scale measurement record of autonomous language-model trading agents opera
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to mod
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustn
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language prob
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation,
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its e
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language m
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressiv
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal veri
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and re
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together wit
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough:
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby pla
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a larg
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic,
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
Large language models achieve superior performance on tasks that require extended reasoning, but long chains o
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a tea
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
Personalized assistants should not only comply with user requests but also assess whether those requests are a
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
We study a governed approach to enterprise analytics: a language model interprets the question, while determin
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improv
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucin
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from cur
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evalua
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoo
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements
On-policy distillation trains a language model on its own generations while a teacher scores them token by tok
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev