Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is evolving from generative models used through a single invocation into agentic control pro
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「LLM」の検索結果
34 件Visual generation is evolving from generative models used through a single invocation into agentic control pro
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
We present a continuous, population-scale measurement record of autonomous language-model trading agents opera
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustn
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language prob
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation,
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes,
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language m
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressiv
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrow
We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation pass
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and re
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough:
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the fi
Large language models achieve superior performance on tasks that require extended reasoning, but long chains o
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes con
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models b
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improv
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucin
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long s
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from cur
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evalua
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements