Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) p
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「generation」の検索結果
26 件Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) p
Visual generation is evolving from generative models used through a single invocation into agentic control pro
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generat
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its e
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressiv
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrow
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and re
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby pla
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial sim
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models b
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target c
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evalua
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoo
On-policy distillation trains a language model on its own generations while a teacher scores them token by tok
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform