NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「generation」の検索結果
37 件Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate larg
Text-to-video generation has advanced significantly over the past five years through scaling of model size, da
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and syn
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computat
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduc
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, an
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diag
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. Howe
In line with the prevailing direction of vision research, we explore the integration of both generation and ed
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogene
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of nu
Current video generation models achieve impressive results in single-shot generation, yet remain limited in ci
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the cont
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interac
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents ty
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), h
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. H
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of
As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and i
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generat
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diver
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, en
Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce waterti
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer dire
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). Howev
Code review helps maintain software quality before code integration, but it also imposes a substantial workloa
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated dataset
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning r
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing te