PAC-Bayesian Reconstruction Guarantees for Time Series Variational Autoencoders
Forecasting time series accurately is critical for applications with complex data ranging from energy systems
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「video」の検索結果
89 件Forecasting time series accurately is critical for applications with complex data ranging from energy systems
Multimodal emotion recognition has attracted growing interest due to its importance in human-computer interact
Intelligent extended reality (XR) systems increasingly use eye and head tracking to infer user intent, task, a
Artificial intelligence (AI) is transforming not only what information systems researchers design, but also ho
As the scale of video surveillance data outpaces manual annotation capacities, weakly supervised video anomaly
Artificial intelligence (AI) now supports investment workflows from data and prediction through research, port
Multiplayer Online Battle Arena (MOBA) games rely on matchmaking to maintain competitive balance. Our prior wo
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio
3D perception plays a crucial role in real-world applications such as autonomous driving, robotics, and AR/VR.
Pass receiver selection is a fundamental task in football analytics, aiming to predict the intended receiver u
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions o
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of anno
In the field of computer vision and graphics, high-quality reconstruction of the human body in static scenes h
Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity
Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos that contain moments relevant to a
We present DiT Readout (ReaDiT) Guidance, a lightweight framework for controlling generation with Diffusion Tr
We present ResLearn-XR, a residual learning framework for predicting eXtended Reality (XR) network traffic and
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficie
Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and vid
Real-world dynamics are inherently compositional: multiple entities move simultaneously within a shared scene,
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guide
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos giv
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss label
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice
Video-language benchmarks are usually constructed by the dataset authors without published reliability statist
Eulerian video amplification boosts sub-pixel motion by band-pass filtering per-pixel intensity traces and app
Long-horizon multimodal agents should remember not only what happened but also who participated. This capabili
Object-centric visual representations are important for physical-world perception, but existing visual pretrai
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent meth
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet the
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generaliza
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high me
Noisy annotations pose a significant challenge for supervised deep learning, as neural networks rely on large-
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-co
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and ans
Video generators build long videos by composing shorter parts, either by generating segments one after another
Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its explanat
3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view set
Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion mo
Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities,
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target c
Action-conditioned video models require large-scale visual data paired with control signals that are temporall
In conditional coding-based neural video compression, the quality of temporal context directly affects compres
The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forg
Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suf
Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advan
Camera traps have become an essential tool for wildlife monitoring, motivating the development of computer vis
Quadruped robots have demonstrated impressive agility in parkour locomotion across complex terrains. However,
Mobile robotic platforms offer a flexible alternative to fixed manipulators for non-destructive evaluation (ND
Developing humanoid robots capable of leveraging human behavioral data is essential for general-purpose embodi
Grasping and holding tools while using them presents a considerable challenge not only for robots but also for
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly v
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains exp
World Action Models (WAMs) leverage the capabilities of large-scale pretrained video diffusion models to joint
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach
In camera-controlled video generation, geometry-aware positional encodings condition tokens on camera extrinsi
We present a system of two wearable pneumatic haptic devices that supports continuous, closed-loop, bidirectio
Classical numerical solvers for partial differential equations (PDEs) are computationally expensive to solve r
Extracting scenarios from unlabelled real-world sensor data streams is a critical but challenging task in the
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Cu
Industrial recommenders give new content initial views through budgeted exploration, then use early performanc
Network comparison using optimal transport is a growing area of research in network science. Unlike standard g
Evaluating customer creditworthiness is crucial for retail banking operations, as it impacts marketing strateg
We demonstrate a physical mechanism for causal information filtering in a physical reservoir computing (PRC) b
Edge computing systems need to support diverse sensing workloads under tight energy and memory constraints, th
Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, invers
Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such
Accurate watch-time (WT) prediction is an important requirement for short-video recommendations. Yet WT distri
A rapidly growing range of sequential data tasks, such as identifying trend reversals in financial markets, au
We study temporal fair division of indivisible mixed manna. Items arrive over time and must be allocated irrev
Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or bo
レジリエンシャルコンピューティングでは、非線形ダイナミカル系を使って、時間依存の入力を、高次元の状態空間表現にマッピングする。レジリエンシャルパフォーマンスは、メモリ、非線形性、そしてそれらのトレードオフを反映しているが