3D-Aware VLMs with Implicit and Explicit Geometries
3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
- 用途
- 3次元空間理解技術の開発
- 難易度
- Hard
- コスト
- High
「3d」の検索結果
95 件3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
GS-Agentは、自然言語から生成することができ、物理的に正しく動作する4次元の世界を生成することができる。方法は、物理的正しさを保つために、生成時に物理的推論を使用した。
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However,
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rar
Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather cond
Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore t
We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural r
Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settin
3Dガウシアンスプレイティングによる動的なシーン再構成は、動的なモーションモデリング、構造的安定性とコンパクトな表現のバランスをとることが求められる。実際、既存のprimitive毎に実際に実装されている方法はローカルの
UAVは、高度、ピッチ、ロール、FOVの変動を含む高度なカメラポーズにおいて動作するため、非対称分布の深さが含まれる広範な空中画像におけるモノラル深度推定を実現するには、高度な深度推定手法が必要である。ほとんどの推定手法
一部のGaussianスプレイティングを利用したSL
Digital sewing patterns typically consist of disjoint 2D panels without explicit stitch annotations, making do
Accurate forest inventory and large-scale mapping are essential for ecosystem monitoring and sustainable fores
3D点群のセグメンテーションではクラス不均衡が発生し、有効な解決策が必要です。この研究では、11 つの不均衡対策を 2D のコンピュータビジョンとは異なる 3D の上で評価し、標準的な交差エントロピーと均衡の重み付けが競
大規模言語モデルのハロウィーン診断では、対象の 3D 空間関係を推論する際に、視覚化が欠如していることが問題となります。この研究では、これらのハロウィーンを軽減するためのアプローチを提案しています。
自動化された生理学ラボでは、透明なプラスチック製品を認識、位置付け、操作するために視覚知覚が必要ですが、対象となる高品質のリアルワールドデータセットは現在限られています。この研究では、複雑なマルチオブジェクトのシーンを扱
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and back
この研究では、路上の3Dオブジェクト検出を改善するために、外部性を考慮した地域認識のアラーカンシーを提案します。
この論文では、Focus-Aware Large Avatar Model(FA-LAM)を提案します。FA-LAMは、一時的なGaussian頭の生成に適したモデルです。
この論文では、Lumeraという手法を提案します。Lumeraは、Engine-Native 3D World ReconstructionとLightsを検出するために使用します。
この論文では、ViSTR-Benchという手法を提案します。ViSTR-Benchは、MLLMが動的シーンから情報を取得できるかどうかを評価します。
Pixel-aligned Gaussian splatting enables efficient and generalizable novel-view synthesis. However, high-resol
Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve
この論文では、安全な人とロボット間の対話を目的とした、人間の姿勢推定とロボットの動作制御の一連のネットワークが提案されます。
オブジェクト目標のナビゲーションにおける、動的な避け方とマルチフロア環境を考慮した、ゼロショットオブジェクトナビゲーションのフレームワークを提案します。このフレームワークでは、動的な人々とマルチフロア環境を考慮しながら、
この研究では、注意機構を併用したグラフニューラルネットワーク (Attention Graph Neural Network) を開発し、流体場の予測精度を向上させた。
Physics-based simulations are essential for understanding the electrode-scale discharge behavior of lithium-io
VLSIのグローバルルーティングは、信号ネットワークを 3D グリッド上で割り当てることが目的であり、信号遅れやワ
Latent world models improve sample efficiency in continuous control by optimizing policies over imagined laten
Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the
Slice-to-volume reconstruction (SVR) is the standard method for obtaining high-resolution (HR) 3D fetal brain
MRI画像の強度正規化方法を7つ比較し、3DUネットワークモデルでMeniscusの分割精度を評価。
3D オキュピエンシー予測には、物体の配置と密度を解釈するための視覚的手法が必要です。従来の方法では、計算コストが高くなりすぎていたが、新しく提案されたGaussianSeedアルゴリズムは、層を階層化することで、計算コ
Recent advances in 3D scene editing have leveraged iterative diffusion models to update input views. However,
Impact hammers, also known as rock-breakers, are essential machines in mining operations, where they perform s
Modeling continuous object deformation is important for many computer vision and robotics tasks, such as manip
Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and dif
Novel View Synthesisは、入力画像から新しい視点の画像を生成するタスクです。ATSplatアルゴリズムは、3次元ガウススプラッタリングを Feed-forward に適合させました。これにより、ATSp
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracte
Vision-centric 3D occupancy prediction provides dense scene representations essential for autonomous driving a
ステレオマッチングは3次元再構成において重要なタスクです。この研究では、ステレオマッチングを確率的生成タスクと組み合わせ、オブジェクト検出の向上を目的として、ステレオマッチングフレームワークと潜在分配を統合する方法を提案
Diffusion Magnetic Resonance Imaging (dMRI) is a powerful tool for probing brain microstructure, but clinical
Deep learning frameworks like nnU-Net achieve state-of-theart brain lesion segmentation performance but remain
Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-
3D Gaussian Splatting (3DGS) は 3D セグメント間の接着を実行するために使用され、テキスト ドライブ の 3D シーン エディットには不可欠です。現行の方法では、固定位置撮影から 2D ディ
DRGBTトラッキングの分野では、目標物を変動するセンシングモデリティと観測プラットフォームの条件下で追跡することが求められます。ドリフト、視線、時間の条件変化に関しても検討が必要です。ただし、現在のバenchmarkで
大規模視点合成モデルは、視点間の注意を交差させることで、未知の視点から3Dシーンを推論します。近年、そのようなモデルはRGB情報だけで3Dの空間関係を学習することができたため、近年の研究者たちは、3Dセグメンテーションに
自律ロボットには、障害物や事故の回避能力が必要です。これは、障害物や事故の回避能力が強化されていれば、障害物や事故に対しての対策がより効果的になります。障害物や事故の回避能力が強まることで、ロボットが障害物や事故から安全
Pain is a complex and pervasive phenomenon affecting a large percentage of the population, and accurate assess
Noisy and corrupted points can substantially degrade point cloud recognition performance, especially under cha
Laboratory automation accelerates discovery, yet its adoption is constrained by the high cost, proprietary des
Language-Guided Grasping は、複雑なシーンで物体の把持を行うために、視覚言語モデル(VLM)を用いる。このアプローチでは、VLM は直接把持を予測するのではなく、3 次元空間における把持の位置を指
分子設計を自動化する方法「Boltzmann-Expected Molecular Design with Decoupled Annealing Flows(DECAF)」を提案。分子設計で重要な3次元構造の特性を確率
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs
Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due
難しい環境で運用するためには、自動空飛ブイロード(UAV)が実際に障害物に存在する距離を判断し、安全な軌跡を計画することが求められる。 これを行うために、複数のステージ(マッピングと計画)を連続化した、サイン・ディスタン
実際の空間知能では、空間に続いて流れるビデオを理解する必要がある。この問題を解決するために、4次元空間を理解することができるモデルを提案する。
心臓の囲みの区別は、食道肥厚の測定に重要であるが、しかし、これを正確に区別することは難しい。これを解決するために、周囲の解剖学的構造を利用して囲みの区別を改善する方法を提案する。
Many Blind and Low-Vision (BLV) people rely on guide dogs for moment-to-moment navigation, such as staying on
オープン・バグナビゲーションには、エージェントへの部分観測が含まれます。パフォーマンスの向上のために、内部状態更新が重要です。これを実現するには、ポリシーネットワークの更新が必要です。最近のアプローチでは、トランスフォー
Robot-assisted minimally invasive surgery (RMIS) offers major benefits over open and conventional laparoscopic
The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordin
多元表格予測モデルTabPFNは、条件設定されたサポートセットと、入力クエリーでタスク特指訓練を行うことなく、推論を行います。実行時間における内部挙動を理解する際に、 zig-zag永続ホモロジーを使用することで、Tab
この論文では、アルツハイマー病の予測を行うためにPersistent HomologyとConformal Guaranteeという手法を提案する。この手法は、アルツハイマー病の予測を行うために、時間的な軌道を分析するこ
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other
A line-scanning lidar yields range and azimuth values in a fixed plane. To perceive surrounding objects in 3D,
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but becaus
existing HRI methodの制約を解決するためのgestures imitation methodを提案し、robustなジェスチャー認識を達成する。
existing VLA methodの制約を解決するためのpersistent object token methodを提案し、ロボット制御をより実用的なものにする。
existing robot navigation methodの限界を解決するためのglobal traversability prior extraction methodを提案し、オフロード環境でのロボット移動を実
existing Embodied Foundation Modelの制限を解決するためのcontact-point prediction とnative 3D grounding methodを提案し、更に能力と
Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-en
自律ロボットの位置決めは、ロボットナビゲーションの主要なタスクです。ロボットが予測できない、非静的な障害物、またはロボットが未知の環境に入ることが多い。この研究では、ロボットのオドメトリと距離サンプリングを組み合わせて、
共役ロボットは、人間オペレータと同梱するワークスペースを共有し、機械手のハンドオーバーなどの安全性の高いマイクロイベント頻繁に発生します。但し、従来の静的なハンドオーバーは、非対称の産業工具を取り扱う際、不自然な抓を持つ
In autonomous driving development, a perception dataset is crucial, as it provides fundamental data for traini
この論文では、ラインダブルロボットのための自発的アクション生成を実現することを目標とし、vision-language 指向性の指令によりロボットが自発的に動作することができることを示します。
この論文では、低照明状況のためのSLAM実現を目標とし、LiDAR、深さ、または熱センサなどの補助的なセンサを取り入れることでSLAMを改良します。
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although re
DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-mod
Incorporating prior maps significantly enhances the accuracy and robustness of pose estimation in visual-inert
Global LiDAR-to-BIM initialization must place a robot within an as-designed building model without a prior pos
Precise metric depth estimation is fundamental for autonomous robot navigation, yet monocular systems inherent
This paper presents a method for user-driven robot Learning from Demonstration (LfD) that reduces user effort
LiDAR place recognition supports loop closure, relocalization, and multi-agent map management. As robotic plat
この論文では、インフラリーダーの観点から安全交通のアセスメントを行うためのフレームワーク、PRISAを提案する。このフレームワークは、道路状況の観測と交通安全の評価を提供し、交通渋滞や事故を予測することで安全交通の実現を
Safe model-based reinforcement learning (RL) often bridges control-theoretic analysis and RL for robots to saf
この研究では、動的なシナリオを分析するために可視化した地図上にMotion Attributeを付与し、Language QueryによるMotion Attributeフィルタを使用して分析することができます。
VTLocフレームワークは、視覚情報と触覚情報を統合し、ロボットハンドの位置を推定することで、ロボットハンドの位置推定と動作操作を実現します。
PIXIEフレームワークは、6次元オブジェクト位置推定を実現し、ロボットハンドの制御と物体の操作を実現します。
多タスク学習はロボティクスの視覚理解系で、セマンティック セグメンテーションと深度推定の統合をサポートします。視覚基底モデル(VFM)は強力な特徴エンコーダとして広く採用されていますが、既存のデコード戦略は重要なボトルネ
難しい環境でマルチオブジェクトの検知と追跡が可能なPiVoTを開発、実用的なソリューションを提案した。
We introduce a new conformal prediction method that constructs calibrated prediction sets over collections of
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability
AIが人間と協力して作り出すアイデアを評価するための新しい手法を提案し、創造性の評価を向上させた。
ワーブレートを利用