3D-Aware VLMs with Implicit and Explicit Geometries
3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
- 用途
- 3次元空間理解技術の開発
- 難易度
- Hard
- コスト
- High
「multimodal」の検索結果
142 件3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
多モーダル理解技術のための新しいアプローチであるMIRROR(Learning from the Other View)を提案しました。MIRRORは、テキスト、図、テキストと図の組み合わせから同等の視点を提供することで
大規模な言語モデルを用いた推論技術のための新しいアプローチであるX$^3$-OPD(Distilling Reasoning into Large Audio-Language Models via On-Policy
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for ena
Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular p
Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as
モデル出力の選択のためのBoN(ベストオブナ)を、部分検証が含まれるビジョン言語タスクに適用する。この方法により、モデル出力を効率化できる。
FL(分散機械学習)におけるパラメータの効率的なフィンテューニングを支援するツールを提案した研究で、TRISHUL(Three-Pronged Spectral Control for Federated Paramet
OpenForgeRLは、ハーネス付きエージェントを訓練するためのフレームワークを提供する。これにより、エージェントが複雑なトラジショナルハーネスを利用して、外部システムと協力し、複数のタスクを同時に解決できるようになっ
GS-Agentは、自然言語から生成することができ、物理的に正しく動作する4次元の世界を生成することができる。方法は、物理的正しさを保つために、生成時に物理的推論を使用した。
A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard t
Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This l
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined
LLMは、人間のアイデンティティのシミュレーションを使用して個人データを削除したり、未均衡なデータを削除したりしますが、これらのアプローチには制限があります。
知識重視の質問応答システム (KI-VQA) を分析するために、新しい評価基準を提案します。これらの基準では、VLMの各タスクを個別に評価することができます。
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognitio
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rar
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-en
効率的な多モードの推論は、モデルの性能やFLOPCOuntだけでなく、移動、キャッシュ、変形、量化された表現を保存するコストやメモリ、エネルギーに関する制約にも制限されています。この論文では、最近のビジュアルトークン圧縮
Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities toget
The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning
Large Language Models (LLMs) は医学教育に大きな可能性を持っていますが、現在のシステムでは、質問に答えるか一時的なフィードバックしか行なわれていません。一方、臨床病例を決定センターへの学習トレ
この論文では、音声字幕の評価手法が提案され、音声字幕の評価において既存の手法の制約を克服することを目指しました。提案されたフレームワークは音声字幕の各側面を評価し、質問回答型の評価手法ではなく字幕の中立性を評価することが
この論文では、記憶システムをサポートするフレームワークMemToolsが構築され、記憶システムの開発を容易にすることを目指しました。これにより開発者は、記憶システムの各コンポーネントを開発およびテストしやすくなり、設計と
複数モーダルデータの統合を支援するための方法とツールを提案し、医療画像認識におけるモーダル間の知見の共有を促進した。
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with exp
Cross-modality image translation offers a route to super-resolution fluorescence microscopy from low-resolutio
Infrared image super-resolution (IISR) mitigates the limitations imposed by low spatial resolution. Existing m
大規模言語モデルはさまざまな画像をテキストに変換する上で優れた性能を示しているが、発生するホログラフィックな診断にはまだ解決策が必要です。この研究では、主流の粗い検出方法の欠点を補うため、細部の診断方法を提案しています。
大規模言語モデルのハロウィーン診断では、対象の 3D 空間関係を推論する際に、視覚化が欠如していることが問題となります。この研究では、これらのハロウィーンを軽減するためのアプローチを提案しています。
大規模言語モデルの圧縮には、モデルのパフォーマンスが低下する可能性があるため、量化の保護が重要です。この研究では、Fisher加重チャネル感受性を用い、MLLMの量化を安定させるためのC-PTQをプロPOSEしています。
パスロジは、現在、パスロジ認識のための画像言語モデルを評価するために広く使用されていますが、この研究では、パスロジ認識において画像言語モデルの視覚知覚が機能していることを疑問に問っています。
感情認識は、現代のアギを促進するために不可欠ですが、大規模
Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query g
Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing thei
Current identity customized video generation methodologies are predominantly limited to single-identity scenar
Improving video captioning quality typically demands retraining large vision-language models, an expensive and
本論文では、テキストと動画を対応させるDistribution-Alignment Bridge(DAB)を提案します。DABは、テキストと動画のエンティティを確率分布として表現し、両者の間の分布の差異を解決します。この
この研究では、マイメイク移植を改善するために、マイメイクの強い地域性を考慮したRegion-Controllable Diffusion Transformer(MagicMakeup)を提案します。
この論文では、DINO-VPTという手法を提案します。DINO-VPTは、Hierarchical Visual Prompt Tuning(HVPT)を使用して、物理的なスポーフィングとデジタルスポーフィングを検出しま
この論文では、ViSTR-Benchという手法を提案します。ViSTR-Benchは、MLLMが動的シーンから情報を取得できるかどうかを評価します。
Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing dat
この論文では、安全な人とロボット間の対話を目的とした、人間の姿勢推定とロボットの動作制御の一連のネットワークが提案されます。
クロアニオトミーの手術を自動化するために、複数のモジュールから形成されるサイバネティックなクローゼッドループのフレームワークを提案します。このフレームワークは、ツールと組織との対話を通じて、ツールと組織の相互作用に対して
人間の知能を模倣するクロアニオトミー手術のフレームワークを提案します。このフレームワークは、前方計画と後方実行を組み合わせて、手術中に手術台の位置を自動的に調整することで、人間と同様の安全で効率的な手順を実現します。
オブジェクト目標のナビゲーションにおける、動的な避け方とマルチフロア環境を考慮した、ゼロショットオブジェクトナビゲーションのフレームワークを提案します。このフレームワークでは、動的な人々とマルチフロア環境を考慮しながら、
Learning-based manipulation policies usually predict robot actions from sensory observations and leave their e
Multimodal learning is a robust approach to improve predictive performance in applications such as medical pro
この研究では、抗原特異性抗体を設計するために、抗原および抗体の間でエピトープレベルでのペアリングが必要であることを考慮した、抗原特異性の抗体多モーダルファンデーションモデル(AAMFM)を提案しました。
分子構造を決定するために、スペクトルデータから自動的な構造解析を実施するための方法を提案している。この方法は、スペクトルデータに基づいてヒントと改良を繰り返すことで、分子構造を決定するもので、分子の可能性の広範な構造スペ
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge fo
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, indep
Small language models and coding agents increasingly generate web front-end code, yet their outputs are typica
異なる順番で画像と質問が提示される場合、視覚言語モデルはモデルのパフォーマンスに大きな影響を受けることが発見された。
Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic fac
Virtual reality (VR) headsets (e.g., Meta Quest, Apple Vision Pro) provide a seamless user experience due to t
Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and dif
多モーダルラージランゲージモデルは、視覚言語タスクに強いですが、高い推論コストで問題となっています。Look Less, Think Fasterアルゴリズムは、単位次元を個別に最適化することで、多モーダルラージランゲー
複数ターンのファッション画像検索は、実世界のファッション検索では重要なタスクです。Diverse-Intent Multi-Turn Fashion Image Retrievalアルゴリズムは、異なる検索用途を扱うこと
画像理解のための多モーダルラージランゲージモデルは、強力ですが、まだ能力と限界については明確な理解が不足しています。この論文では、多モーダルラージランゲージモデルが画像理解においてどの程度の能力と限界を持つか、を分析し、
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existi
Attention Mechanism (AM) selectively focuses on essential information for imaging tasks and captures relations
Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure,
ステレオマッチングは3次元再構成において重要なタスクです。この研究では、ステレオマッチングを確率的生成タスクと組み合わせ、オブジェクト検出の向上を目的として、ステレオマッチングフレームワークと潜在分配を統合する方法を提案
この研究では、スイートペッパーの収穫前期予測を目的として、多モード時系列データを統合するための深層学習フレームワークを提案します。
ETPデザイナはマルチモーダルな電子シアターのデザインを自動化するフレームワークを提案します。
Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by ge
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, seman
Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a sin
この研究では、ドローンで小さな物体を認識することを目的としたメモリ拡張型大規模言語モデルを開発しました。このモデルは、複雑なドローンの場面で、ユーザーの指示に従って物体を識別できるようになります。
この研究では、可視化された質問への対応を評価するために、新しい方法を提案しました。この方法は、質問への回答の正確性だけでなく、質問への回答のパターンや特徴も評価することができます。
自動運転システムには、道路のトポロジー(ドライバブルレーンとその接続性)を理解する機能が必要です。最近の検出モデルは360度の前方視野からボリュームイメージを取得することで、道路上のレーンのトポロジーを推測することができ
地上を表す重力式マップの高解像度版が、多くの用途で役立ちます。たとえば、市区町村の変化を監視したり、エネルギー対策を向上させたり、温室効果ガスの排出量を追跡したりすることができます。4つの主要な全世界建物Rasterデー
Automatic pain localization, which involves identifying the anatomical origin of pain from peripheral physiolo
Language-Guided Grasping は、複雑なシーンで物体の把持を行うために、視覚言語モデル(VLM)を用いる。このアプローチでは、VLM は直接把持を予測するのではなく、3 次元空間における把持の位置を指
ReferTrack は、自然言語で対象の車両に付近する自動車を追従させるシステムである。このシステムでは、対象の車両に付近する自動車を認識する後、自動車の動きを予測する。
SOPD-SocialNav は、学習モデルを小さなロボットに伝える技術であり、ロボットが環境と人間の行動を理解し、ナビゲーションが行えるようにする。
Clinical Pathways は、ロボットが実際の環境で安全に動作するためのシステムである。これは、ロボットが病室で安全に作業し、医療スタッフや患者を守る。
Despite recent advances in general-purpose robotic manipulation, real-world multi-object clutter remains chall
深層学習を用いた形状推定モデルを作成し、オープン平面曲線の形状を推定するための深層学習モデルを提案した。
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether t
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depe
Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully r
A single embedding space that covers text, images, video, and audio lets one index serve every query a user ca
Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combi
The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that
Long-video question answering requires a model to preserve visual evidence over time without repeatedly reproc
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs
Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise vid
Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasin
画像生成において、材料、 객체、領域を制御することが難しい問題がある。 Diffusion Transformers はテキストと画像を組み合わせて処理できるが、どちらをどの程度影響させるか決める仕組みがなかった。 その
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond
オリジナルのデータとZoom-Inのツールを組み合わせた方法、OmniReasonerを提案する。これにより、オリンモードルLLMsの長いオーディオビデオの論理的推論を改善できる。
記述情報に従って画像や動画データを混ぜ合わせる「対数混合法」を拡張する方法、InstructMixupを提案する。これにより、データを拡張しながらデータの内容とラベルが維持される。
無線無人飛行機のルートプランニングでは、視空間と言語モデルを利用して安全なルートを生成する必要がある。この問題を解決するために、テスト時にモデルをスケールアップさせる方法を提案する。
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across
実際の空間知能では、空間に続いて流れるビデオを理解する必要がある。この問題を解決するために、4次元空間を理解することができるモデルを提案する。
この研究では、スパイナル サブアルテラノスパース内の安全な移動、操縦、内視鏡撮影を可能にする医療用ロボットを提案します。
自動運転のための計画とは、状況理解、タイムリーな推論、行動選択というものがあるが、しかし、これらの要素を組み合わせるのは難しい。これを解決するために、シーン理解を分離することによって、計画を安全かつ有効性のあるものにする
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more t
World Action Models(WAMs)は、ロボットマニピュレーションをモデル化するパラダイム。WAMsは、視覚ステートトランジションとロボットアクションを同時にモデル化する。しかし、既存のWAMsは、一定の時
Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, y
Vision-language-action (VLA) models have shown impressive generalization, but often lack interpretability and
Macro placement still requires substantial manual refinement in industrial physical design flows. We present M
ロボット制御を効率化するために、パッチを用いた政策学習を提案し、密集された視覺表現を用いて実装することを目的としている。
existing VLA modelの制約を解決するためのforce-based memory method、FM-VLAを提案する。
existing VLA methodの制約を解決するためのpersistent object token methodを提案し、ロボット制御をより実用的なものにする。
In autonomous driving development, a perception dataset is crucial, as it provides fundamental data for traini
この研究では、混乱のないターゲットに可視化言語アクションモデルを適応させることを目的として、3つのモデルを使用して研究を行った。3つのモデルは、観察から直接行動へのマッピング、テキストチャインオブスロット、潜在的な反復ル
この論文では、ラインダブルロボットのための自発的アクション生成を実現することを目標とし、vision-language 指向性の指令によりロボットが自発的に動作することができることを示します。
Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts
Flow policies can represent multimodal action distributions for robot manipulation, yet a robot must execute o
The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that
Teleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces
Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorpora
Social navigation requires the robot to reason and respond in complex real-world environments. While recent wo
End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor obser
Vision-Language Navigation in dynamic, human-centric environments exposes a fundamental tension: linguistic re
Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet thei
Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and suppor
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazar
Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tigh
この研究では、動的なシナリオを分析するために可視化した地図上にMotion Attributeを付与し、Language QueryによるMotion Attributeフィルタを使用して分析することができます。
この研究では、Vision-とTactile-based ProposalとSimulation-based Inferenceを組み合わせ、物体の位置と姿勢を推定する方法、BayesContactを提案しています。
PIXIEフレームワークは、6次元オブジェクト位置推定を実現し、ロボットハンドの制御と物体の操作を実現します。
Robotic navigation in unstructured environments requires robust situational awareness to safely traverse hazar
スコアベース生成モデルにおけるモード分解能の向上を目的とした研究で、モード分解能がスコア関数に依存しておらず、生成サンプルから混合重みを推測できることを明らかにした。
Longitudinal tumor measurements, dropout information, and genetic covariates provide complementary information
Multimodal optimization aims to locate multiple globally optimal or near-optimal solutions in a single run. Th
非線形オブザーバシオンメカニズムや多次元データには適合しない伝統的なエンサンブルフィルタリングアルゴリズムを導入し、隠蔽データアシミレーションを提案
Circular data, representing angles or directions, are frequently encountered in computer vision, biology, geol
この研究では、マルコフ連鎖モンテカルロ法を改良し、多モーダル分布からサンプリングする能力を高めるための新しいアプローチを提案した。このアプローチでは、微分のノイズパスを使用することで、モデルの収束を高速化し、多モーダル分
We consider the recovery of a pair of sparse vectors from a limited number of nonlinear observations of their
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability
Maintaining physical consistency in video generators and world models increasingly relies on vision-language m
The success of advanced air mobility (AAM) operations is largely contingent on its effective integration with
AIが人間と協力して作り出すアイデアを評価するための新しい手法を提案し、創造性の評価を向上させた。
CMA-ES behaves, per restart, primarily as a local optimizer; multimodal search relies on restart strategies su