3D-Aware VLMs with Implicit and Explicit Geometries
3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
- 用途
- 3次元空間理解技術の開発
- 難易度
- Hard
- コスト
- High
「video」の検索結果
92 件3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
都市の電気自動車充電インフラは、可及的速やかに故障を予測・修理することで、耐久性と低炭素化を向上させる必要がある。機械学習を用い、故障を予測するモデルの開発を研究した。
Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as
Archival film restoration is a challenging problem because historical footage contains compound degradations s
GraphVidは、グラフと文本から生成することができ、オブジェクトの複数の移動を正確に制御することができる。グラフではオブジェクトの動きを表す情報を保存し、文から生成の制約を指定することができる。
ElasticTTTは、プログラムがテストのときに動作を調整できるようにした。方法は、テストのときにモデルが前のサンプルの情報と現在の情報を組み合わせて、ビデオを編集する際に正しく動作するようにした。
Video face swapping has no natural paired supervision: no real footage exists of one person's face performing
この研究では、大規模言語モデルを使用して、basketボールの動的理解に基づいて、プレイヤーへの関わりや時間境界を推測するモデルを開発しました。
この研究では、大規模言語モデルを使用して、因果プロセスの理解を進めました。大規模言語モデルを活用することで、因果関係を予測することができました。
この研究では、大規模言語モデルを活用して、因果関係のモデル化を研究しました。大規模言語モデルを活用することで、因果関係を予測することができました。
ビデオLMMの安全性を確認するために、新しい診断フレームワークを提案します。これらのフレームワークは、モデルの挙動、理解、セマンティクスを同時に考慮します。
Semantic-ID-based generative recommendation represents items as sequences of shared semantic tokens, enabling
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognitio
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-en
効率的な多モードの推論は、モデルの性能やFLOPCOuntだけでなく、移動、キャッシュ、変形、量化された表現を保存するコストやメモリ、エネルギーに関する制約にも制限されています。この論文では、最近のビジュアルトークン圧縮
多エージェントのシミュレーションにおいて、共有世界状態がエージェント間で保持され、その世界状態が観測結果に反映されると仮定している。
ビデオ内の物体の空間推論を同時に行うことで、現存するタスク固有の注釈を超えた統一的なビデオ推論システムを構築した。
ビデオ内のキャメラの動きと物体の動きを切り離すことで、モーションの表現学習を改善した。
ビデオ生成モデルの効率性と高品質性を向上させるための新しい方法を提案した。
Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore t
Multi-object tracking in dense crowds requires solving a bipartite assignment problem between detections and t
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with exp
Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at in
Current identity customized video generation methodologies are predominantly limited to single-identity scenar
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and back
Improving video captioning quality typically demands retraining large vision-language models, an expensive and
本論文では、テキストと動画を対応させるDistribution-Alignment Bridge(DAB)を提案します。DABは、テキストと動画のエンティティを確率分布として表現し、両者の間の分布の差異を解決します。この
この論文では、効率的なストリーミングビデオ生成手法であるMs. Forcingを提案します。Ms.フオーシングは、Multi-Scale PatchificationとAttentionを組み合わせた手法です。
この研究では、都市ウォークビデオを分析するために、4つのモダリティの表現(スペース時領域情報、時間平均画像、オーディオ符号化、テキストベースの表現)を使用しました。
この論文では、ViSTR-Benchという手法を提案します。ViSTR-Benchは、MLLMが動的シーンから情報を取得できるかどうかを評価します。
A recent line of work measures causal emergence in reinforcement learning agents through Integrated Informatio
Personal and organizational planning systems maintain two records that drift apart: what was planned (a task's
Effective decision-making in complex and changing environments requires balancing short-term and long-term con
Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving
流動画像生成を扱う研究、HeadCast を用いて流動画像生成を提案する。
この研究では、実世界ロボットのシーケンシャル予測に使用できる、diffusion-based frameworkを提案しました。
連続的に観測された行動を捕捉するためのシーケンシャル推奨モデルを使用すると、期間が長い間隔が発生した場合に、再活性化されたユーザーへのリコールを改善できる提案されている。
The wind energy industry relies on accurate power curve models to make power forecast, evaluate turbine perfor
Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the
オフラインでの短時間の視覚生成が一般的な人間の行動の分析では、人間の行動の長期的な視覚生成は、実践的な長時間の視覚生成では実行不能である。StreamHOI は、人間間の視覚的な行動の生成を生成したいくつの画像を使用して
Open-world video anomaly detection (OWVAD) is expected to detect events that match a user-specified definition
ビデオキャプション生成には、空間と時刻の理解が重要です。PercepCapアルゴリズムは、ビデオ入力を空間時刻認識に分解することで、生成されたキャプションの理解度が向上するとともに、空間時刻の誤差をより正確に検出でき、キ
長時間ビデオエクストラポレーションには、高度な視覚的知能が必要です。Self Gradient Forcingアルゴリズムは、学生モデルを教師モデルから生成される歴史の下で学習させることで、長時間ビデオエクストラポレーシ
Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across divers
Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditi
Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autono
Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic
Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the
Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-
この研究では、ドローンで小さな物体を認識することを目的としたメモリ拡張型大規模言語モデルを開発しました。このモデルは、複雑なドローンの場面で、ユーザーの指示に従って物体を識別できるようになります。
Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most exis
Tracking objects through state transformations is essential for understanding real-world dynamics. However, ex
Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-relat
Pain is a complex and pervasive phenomenon affecting a large percentage of the population, and accurate assess
VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rar
HOST は、ロボットが人間の動作からスキルをすぐに習得できるシステムである。このシステムでは、ロボットは単一の人間の動作ビデオからスキルを習得し、既に習得したスキルを維持する。
Clinical Pathways は、ロボットが実際の環境で安全に動作するためのシステムである。これは、ロボットが病室で安全に作業し、医療スタッフや患者を守る。
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to
A single embedding space that covers text, images, video, and audio lets one index serve every query a user ca
Long-video question answering requires a model to preserve visual evidence over time without repeatedly reproc
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs
While machine learning-based weather models hold significant promise, they struggle to predict the detailed st
Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due
Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise vid
Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasin
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making th
オリジナルのデータとZoom-Inのツールを組み合わせた方法、OmniReasonerを提案する。これにより、オリンモードルLLMsの長いオーディオビデオの論理的推論を改善できる。
記述情報に従って画像や動画データを混ぜ合わせる「対数混合法」を拡張する方法、InstructMixupを提案する。これにより、データを拡張しながらデータの内容とラベルが維持される。
実際の空間知能では、空間に続いて流れるビデオを理解する必要がある。この問題を解決するために、4次元空間を理解することができるモデルを提案する。
この研究では、高空飛行の無信号位置指示のNGPS (Next-Generation Positioning System)というフレームワークを提案しました。NGPSは、GPSの信号を利用せずに位置推定を可能にします。N
World Action Models(WAMs)は、ロボットマニピュレーションをモデル化するパラダイム。WAMsは、視覚ステートトランジションとロボットアクションを同時にモデル化する。しかし、既存のWAMsは、一定の時
行動モーター特徴は、社会認知や人間ロボットインターフェースなどの行動認識の核心です。人間ロボットのNICO用に、2段階のアーキテクチャを提案します。1段階目では、腕の移動を学習するSOMと、手の移動を学習するSOMを使用
Robot-assisted minimally invasive surgery (RMIS) offers major benefits over open and conventional laparoscopic
Reservoir computing exploits nonlinear dynamical systems to encode temporal inputs into high-dimensional state
工場の中の異常が検出されるように設計された機械学習モデルを提案しています。通常の方法では、モデルはビデオ内のすべての内容を考慮し、複雑な問題を解決することは困難です。提案されたモデルのアプローチは、オブジェクトを検出して
ロボット制御を効率化するために、パッチを用いた政策学習を提案し、密集された視覺表現を用いて実装することを目的としている。
Fast planning of novel behaviors in unseen scenarios remains a fundamental challenge in robotics. The high-dim
Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-en
この研究では、2 つのロボット腕を同時に使用することで、緊張組立て操作のパフォーマンスを向上させる end-to-end フレームワークを提供します。ロボット腕は、 CAD モデルの数字、そして望ましい組み立て状態に置か
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although re
Teleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces
Digital twins enable robots to anticipate and adapt to physical interactions, but existing models struggle wit
This paper investigates temporal fair division, a setting where items are allocated over multiple rounds and a
Neural Cellular Automataが複雑な形状を形成するプロセスを研究しました。
多タスク学習はロボティクスの視覚理解系で、セマンティック セグメンテーションと深度推定の統合をサポートします。視覚基底モデル(VFM)は強力な特徴エンコーダとして広く採用されていますが、既存のデコード戦略は重要なボトルネ
Interfacing with Biological Neural Networks (BNNs) requires encoding information into stimulation patterns tha
Self-organized criticality (SOC), a dynamical regime associated with maximal information processing, offers a
Maintaining physical consistency in video generators and world models increasingly relies on vision-language m
Planning contact-rich whole-arm manipulation is challenging because interactions that involve extended robot g
ネットワークダイナミクスシステムを使って時間系列予測を高速化し、ニューロモルフィックコンピューティングを活用した。