Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is evolving from generative models used through a single invocation into agentic control pro
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「Agent」の検索結果
26 件Visual generation is evolving from generative models used through a single invocation into agentic control pro
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavi
Agents can turn shared infrastructure into a channel for coordinated intrusion. The Hugging Face incident and
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
We present a continuous, population-scale measurement record of autonomous language-model trading agents opera
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language prob
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its e
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes,
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language m
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressiv
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together wit
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the fi
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic,
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most lon
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
Personalized assistants should not only comply with user requests but also assess whether those requests are a
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models b
We study a governed approach to enterprise analytics: a language model interprets the question, while determin
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improv
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Vision-language-action (VLA) models map visual observations and language instructions directly to robot action
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long s
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from cur
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns