onnxruntime — ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
FastVideoは、加速されたビデオ生成用に統一された推論およびポストトレーニングフレームワークです。
- 用途
- クロスプラットフォーム高性能ML推論用エンジンの実現
- 難易度
- Easy
- コスト
- High
「video」の検索結果
121 件FastVideoは、加速されたビデオ生成用に統一された推論およびポストトレーニングフレームワークです。
supervisionは、機械学習技術を活用して、ユーザー独自のコンピュータビジョンツールを作成することができる。
mediapipeは、クロスプラットフォームでカスタマイズ可能なライブおよびストリーミングメディア向けのMLソリューションを提供している。
.diffusion モデルのライブラリ。画像・動画・音声生成に利用可能。
CVATは、機械学習用の業界標準のデータエンジンです。さまざまなスケールのチームが使用し、さまざまなスケールのデータに対応しています。
イメージを注釈するツール。ポリゴン、長方形、円、線、点などを注釈することができる。
SANAは、高解像度画像生成モデルSANAを紹介する本研究であり、低計算コストで優れた高解像度画像を生成できる。
音声認識、声活動検出、テキスト処理などを行う、基盤となる音声認識ツールキットを提供する。
FastVideoは、加速されたビデオ生成用の統合推論とポストトレーニングのフレームワークです。
zenmlは、データパイプラインからエージェントまで、AIプラットフォームです。
統計チャートの生成は、タブラーのデータから生成することが難しい。新しい作成フローでは、データのスクリーン、プロット提案、コード生成、レンダリング、検証による改良が含まれる。
画像やビデオやオーディオディフュージョンモデルのファインチューニングを行うための、汎用的なファインチューニングキット。
このリポジトリはコンピュータサイエンスのビデオコースの一覧を提供しています。
Forecasting time series accurately is critical for applications with complex data ranging from energy systems
Multimodal emotion recognition has attracted growing interest due to its importance in human-computer interact
Intelligent extended reality (XR) systems increasingly use eye and head tracking to infer user intent, task, a
Artificial intelligence (AI) is transforming not only what information systems researchers design, but also ho
As the scale of video surveillance data outpaces manual annotation capacities, weakly supervised video anomaly
Artificial intelligence (AI) now supports investment workflows from data and prediction through research, port
Multiplayer Online Battle Arena (MOBA) games rely on matchmaking to maintain competitive balance. Our prior wo
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio
3D perception plays a crucial role in real-world applications such as autonomous driving, robotics, and AR/VR.
Pass receiver selection is a fundamental task in football analytics, aiming to predict the intended receiver u
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions o
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of anno
In the field of computer vision and graphics, high-quality reconstruction of the human body in static scenes h
Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity
Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos that contain moments relevant to a
We present DiT Readout (ReaDiT) Guidance, a lightweight framework for controlling generation with Diffusion Tr
この論文では、映像 diffuision モデルを用いて変化する動画の差分を予測する方法を説明する。
We present ResLearn-XR, a residual learning framework for predicting eXtended Reality (XR) network traffic and
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficie
Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and vid
Real-world dynamics are inherently compositional: multiple entities move simultaneously within a shared scene,
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guide
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos giv
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss label
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice
Video-language benchmarks are usually constructed by the dataset authors without published reliability statist
Eulerian video amplification boosts sub-pixel motion by band-pass filtering per-pixel intensity traces and app
Long-horizon multimodal agents should remember not only what happened but also who participated. This capabili
Object-centric visual representations are important for physical-world perception, but existing visual pretrai
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent meth
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet the
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generaliza
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high me
Noisy annotations pose a significant challenge for supervised deep learning, as neural networks rely on large-
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-co
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and ans
Video generators build long videos by composing shorter parts, either by generating segments one after another
Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its explanat
3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view set
Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion mo
Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities,
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target c
Action-conditioned video models require large-scale visual data paired with control signals that are temporall
In conditional coding-based neural video compression, the quality of temporal context directly affects compres
The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forg
Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suf
Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advan
Camera traps have become an essential tool for wildlife monitoring, motivating the development of computer vis
Quadruped robots have demonstrated impressive agility in parkour locomotion across complex terrains. However,
Mobile robotic platforms offer a flexible alternative to fixed manipulators for non-destructive evaluation (ND
Developing humanoid robots capable of leveraging human behavioral data is essential for general-purpose embodi
Grasping and holding tools while using them presents a considerable challenge not only for robots but also for
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly v
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains exp
World Action Models (WAMs) leverage the capabilities of large-scale pretrained video diffusion models to joint
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparatio
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach
In camera-controlled video generation, geometry-aware positional encodings condition tokens on camera extrinsi
We present a system of two wearable pneumatic haptic devices that supports continuous, closed-loop, bidirectio
Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this
この研究では、ストリーミングでのビデオ超解像化を実現する一ステップのディフュージョンフレームワーク「FlashVSR」を提案しています。このフレームワークは、局所制限の疎注意と小さい条件的デコーダを組み合わせて、効率的に
Classical numerical solvers for partial differential equations (PDEs) are computationally expensive to solve r
Extracting scenarios from unlabelled real-world sensor data streams is a critical but challenging task in the
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Cu
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from
Awesome-Video-Diffusionは、Recent Diffusion Models for Video Generation, Editing, and Othersのリストを公開しています。
Industrial recommenders give new content initial views through budgeted exploration, then use early performanc
この論文では、Causal-Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive
Network comparison using optimal transport is a growing area of research in network science. Unlike standard g
Evaluating customer creditworthiness is crucial for retail banking operations, as it impacts marketing strateg
We demonstrate a physical mechanism for causal information filtering in a physical reservoir computing (PRC) b
Edge computing systems need to support diverse sensing workloads under tight energy and memory constraints, th
画像認証システムにおける悪用された画像からの画像の認証方法を提示しました。
Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, invers
Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such
Wunjo CE: Face Swap, Lip Sync, Control Remove Objects & Text & Background, Restyling, Audio Separator, Clone V
Accurate watch-time (WT) prediction is an important requirement for short-video recommendations. Yet WT distri
A rapidly growing range of sequential data tasks, such as identifying trend reversals in financial markets, au
長時間のビデオ生成を実現するためのモデルのサポートを紹介している。
OpenWorldLibは、進化する世界モデルを提供する統一されたコードベースです。
Rust言語でCandleライブラリを利用して、PythonやPyTorchを使用せずにDecoder-only LLMを自作した。
We study temporal fair division of indivisible mixed manna. Items arrive over time and must be allocated irrev
医療画像分析で、深層學習モデルが実装されている問題に対する解決策を提示します。治療を導くために、批判的結果に影響を与える変化について特に重点が置かれています。
awesome-artificial-intelligenceは、人工知能に関する教材、アートcles、講義等を集め、提供しているオープンソースプロジェクトです。
画像生成のためのHigh Quality Training Free Inpaintを提供します。このInpaintはStable Diffusionモデルに使用でき、ComfyUIもサポートしています。
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev
Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or bo
レジリエンシャルコンピューティングでは、非線形ダイナミカル系を使って、時間依存の入力を、高次元の状態空間表現にマッピングする。レジリエンシャルパフォーマンスは、メモリ、非線形性、そしてそれらのトレードオフを反映しているが