PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL),
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「video」の検索結果
88 件While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL),
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and a
Reliable video generation requires more than high-quality frames to form a coherent story: a model must mainta
Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposin
Chunk-autoregressive video world models typically condition each generated chunk on one action. An action rece
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, i
The brain uses discrete spikes for dynamic computation, yet, how neural microcircuits (NMCs) solve temporal cr
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model vi
Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge acc
Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and h
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perceptio
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherenc
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Because dense frame-level annotation of colonoscopy videos is costly, we propose WSPolypNet, a weakly supervis
Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate locomot
We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is abl
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent me
Implicit neural representation (INR) has achieved remarkable progress in novel view synthesis and image/video
Partially Relevant Video Retrieval (PRVR) retrieves untrimmed videos when queries describe only short moments.
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimo
Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based
We present FIRE3D, a unified framework that takes a single RGB image or casual RGB video and transforms it int
Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models
World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance w
Video Large Language Models (VideoLLMs) are increasingly deployed in safety-critical applications such as cont
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Multi-object tracking (MOT) is an essential computer vision task that simultaneously tracks multiple objects i
Micro-actions are subtle, low-intensity non-verbal behaviors that provide cues to fine-grained human states, i
Video generation models have recently attracted substantial attention for their ability to generate visually c
Timestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a rando
Driver alerting from dashcam video requires sequential decision-making under partial observability: a system m
Driver motion can provide cues to ongoing behavior, attention, and near-term driving intent. However, most exi
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mo
Recent advances in video generation have made prompt-based control increasingly central to AIGC video generati
Messages from electronic devices are conventionally received as text, audio, or radio signals. But robots move
Large language models (LLMs) and vision-language models (VLMs) have significantly advanced zero-shot task plan
In multi-robot collaboration, task handovers rely on downstream verifiers performing remote attestation, which
Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals soc
Deep learning has achieved state-of-the-art performance in echocardiographic video segmentation, with an incre
We convert black-box clinical prediction models for tabular data into standalone nomograms that can be audited
We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition
How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, w
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed
Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely interve
Short-form video platforms use recommender systems to maximize engagement through highly efficient personalize
Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world
Human pose estimation and keypoint-based action recognition models are increasingly deployed as components of
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Unde
Long-form narrative-to-film generation requires shot-level controllability and cross-clip consistency in both
Agentic systems can interpret user requests, search the live web, and use external tools, but their ability to
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
Privacy-sensitive surveillance systems could benefit from large vision-language models (VLMs), but such models
Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans,
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produc
Morphology of the sub-basal nerve plexus (SNP) reflects peripheral nerve health, and corneal confocal microsco
The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional co
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first
Multimodal sentiment analysis often remains text-dominant due to raw-video noise and insufficient temporal mod
Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a si
Progress in video anomaly understanding (VAU) has long been limited by inherent deficiencies of real-world ano
Synthesizing photorealistic driving videos along specified trajectories is essential for scalable closed-loop
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information tha
Learning physically plausible dynamics from visual observations is essential for interactive world models and
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable contr
Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavi
Research teams and organizations often explore unfamiliar free-text collections, from survey comments and revi
Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failin
Humanoid soccer is a challenging testbed for dynamic whole-body control, requiring robots to coordinate balanc
Can people distinguish between human and AI agency in humanoid teleoperation? To explore this question, we dev
Vision-language-action (VLA) models have shown promising generalization for language-conditioned robot manipul
Human demonstrations contain rich manipulation knowledge, but it remains unclear what information can be trans
Forecasting time series accurately is critical for applications with complex data ranging from energy systems
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-fre
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generaliza
Quadruped robots have demonstrated impressive agility in parkour locomotion across complex terrains. However,
Classical numerical solvers for partial differential equations (PDEs) are computationally expensive to solve r
Industrial recommenders give new content initial views through budgeted exploration, then use early performanc
Network comparison using optimal transport is a growing area of research in network science. Unlike standard g
Evaluating customer creditworthiness is crucial for retail banking operations, as it impacts marketing strateg
We demonstrate a physical mechanism for causal information filtering in a physical reservoir computing (PRC) b
Edge computing systems need to support diverse sensing workloads under tight energy and memory constraints, th
Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, invers
We study temporal fair division of indivisible mixed manna. Items arrive over time and must be allocated irrev
Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or bo