VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines b
- 用途
- 技術検証・論文読解補助
- 難易度
- Easy
- コスト
- High
「rag」の検索結果
26 件Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines b
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavi
Agents can turn shared infrastructure into a channel for coordinated intrusion. The Hugging Face incident and
We present a continuous, population-scale measurement record of autonomous language-model trading agents opera
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal veri
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrow
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough:
Quantization is widely used to reduce the computational and memory demands of neural-network inference. In rec
Large language models achieve superior performance on tasks that require extended reasoning, but long chains o
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a tea
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes con
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a froze
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Vision-language-action (VLA) models map visual observations and language instructions directly to robot action
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long s
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from cur
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoo
Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) b
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev