Temporal State Transport in Video Generation: Diagnosing and Correcting Spectral Imbalance
Reliable video generation requires more than high-quality frames to form a coherent story: a model must mainta
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「RAG」の検索結果
22 件Reliable video generation requires more than high-quality frames to form a coherent story: a model must mainta
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
Vision transformers (ViTs) have achieved remarkable generalization across visual domains, yet little is known
Multiple Instance Learning (MIL) is widely used for weakly supervised learning, particularly in digital pathol
The software supply chain has become an increasingly exposed attack surface because of its reliance on intrica
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (L
Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing interm
Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workfl
Taxonomy induction aims to organize concept sets into coherent hierarchical structures. Recent LLM-based metho
Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic.
Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing ext
Out-of-distribution (OOD) detection is critical for safe deployment of medical AI systems. Recently, test-time
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Target-based LiDAR-camera extrinsic calibration is a prerequisite for multi-sensor fusion in robotics. However
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Actio
Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism anal
World-action models guide action generation with predicted future observations, but vision-centric predictions
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial inform
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
Particle swarm optimization (PSO) is a widely used metaheuristic, prized for its simplicity and small paramete