3D-Aware VLMs with Implicit and Explicit Geometries
3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
- 用途
- 3次元空間理解技術の開発
- 難易度
- Hard
- コスト
- High
「image」の検索結果
252 件3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
印刷品質管理技術のための新しいアプローチであるシンセティック データ生成フレームワークを提案しました。このフレームワークは、ロトグラビューグラビング技術における品質管理のためのシンセティック データを生成することで、印刷
多モーダル理解技術のための新しいアプローチであるMIRROR(Learning from the Other View)を提案しました。MIRRORは、テキスト、図、テキストと図の組み合わせから同等の視点を提供することで
We propose a new approach to two-sample testing for deciding whether two sets of samples are drawn from the sa
Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, be
Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular p
Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as
この研究では、車椅子の位置情報を取得するために、安全な歩道と道路を分類するセグメントを提案し、視覚障害がある人々や盲人の移動を支援する手段になる可能性があります。
Automated detection of vision impairing retina-based ocular conditions from fundus images is important for ear
ダブル量子ドットのCharge Stateを分析するためのMachine Learning方法を提案した研究で、この方法により、量子ドットのCharge Stateが効率的に分析できる。
GraphVidは、グラフと文本から生成することができ、オブジェクトの複数の移動を正確に制御することができる。グラフではオブジェクトの動きを表す情報を保存し、文から生成の制約を指定することができる。
Visual Contrastive Self-Distillationは、セルフディスタンスルールを高速化する方法を提案した。この方法は、入力情報だけで学生と教師の間の情報の不均衡をなくした。
GS-Agentは、自然言語から生成することができ、物理的に正しく動作する4次元の世界を生成することができる。方法は、物理的正しさを保つために、生成時に物理的推論を使用した。
People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language
Video face swapping has no natural paired supervision: no real footage exists of one person's face performing
A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard t
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However,
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined
Replacing an object with one that differs in category or shape requires complete source removal, natural targe
LLMは、人間のアイデンティティのシミュレーションを使用して個人データを削除したり、未均衡なデータを削除したりしますが、これらのアプローチには制限があります。
ProB is a Prolog-based model checker, animator and constraint solver for high-level formal specifications. One
Event-B is a formal method rooted in predicate logic and set theory. We encoded over 600 proof rules in Prolog
知識重視の質問応答システム (KI-VQA) を分析するために、新しい評価基準を提案します。これらの基準では、VLMの各タスクを個別に評価することができます。
ビデオLMMの安全性を確認するために、新しい診断フレームワークを提案します。これらのフレームワークは、モデルの挙動、理解、セマンティクスを同時に考慮します。
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rar
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-en
効率的な多モードの推論は、モデルの性能やFLOPCOuntだけでなく、移動、キャッシュ、変形、量化された表現を保存するコストやメモリ、エネルギーに関する制約にも制限されています。この論文では、最近のビジュアルトークン圧縮
Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities toget
The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning
コーディングエージェントの評価基準を導入し、現実世界のコミットやプルリクエストに基づくタスクを構築した。
多エージェントのシミュレーションにおいて、共有世界状態がエージェント間で保持され、その世界状態が観測結果に反映されると仮定している。
ビデオ内の物体の空間推論を同時に行うことで、現存するタスク固有の注釈を超えた統一的なビデオ推論システムを構築した。
ディフュージョンモデルにおける初期的なNoise Seed の影響が、モデルが生成する高質のイメージに大きく影響していることを提示し、Seed Search 時の時間的負荷を削減するための方法を提案した。
ビデオ内のキャメラの動きと物体の動きを切り離すことで、モーションの表現学習を改善した。
光の伝達の可微分化を用いて、入力が最も影響するシーン要素を特定するための方法を提案した。
アイス認識の精度を向上させるための方法を提案し、視覚認識におけるアイス認識タスクの課題を分析した。
Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather cond
Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore t
We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural r
Underwater image enhancement remains challenging due to wavelength-dependent light absorption, scattering, and
UAVは、高度、ピッチ、ロール、FOVの変動を含む高度なカメラポーズにおいて動作するため、非対称分布の深さが含まれる広範な空中画像におけるモノラル深度推定を実現するには、高度な深度推定手法が必要である。ほとんどの推定手法
Quantitative drug-induced sleep endoscopy (DISE) requires reliable airway boundaries at specific anatomical le
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with exp
Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in h
AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure an
Polarization cues benefit applications such as material detection and de-reflection, yet acquiring them typica
Cross-modality image translation offers a route to super-resolution fluorescence microscopy from low-resolutio
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at in
Infrared image super-resolution (IISR) mitigates the limitations imposed by low spatial resolution. Existing m
Image restoration agents have recently emerged as a flexible paradigm for handling diverse and unpredictable d
LoViF の 2 回目のチャレンジでは、画像修復に新たなアプローチを提案しています。実世界の画像を修復するための包括的な評価基準を提供しており、低光照度、ハッジ、雨、雪などのさまざまな障害に対する解決策を研究者に求めて
大規模言語モデルはさまざまな画像をテキストに変換する上で優れた性能を示しているが、発生するホログラフィックな診断にはまだ解決策が必要です。この研究では、主流の粗い検出方法の欠点を補うため、細部の診断方法を提案しています。
大規模言語モデルのハロウィーン診断では、対象の 3D 空間関係を推論する際に、視覚化が欠如していることが問題となります。この研究では、これらのハロウィーンを軽減するためのアプローチを提案しています。
ドリフス脱失は画像を再構築するために不可欠ですが、再構築画像とドリフス画像のペアリングや標準化されたプロトコルなどの要件を満たすデータセットが不足しているため、評価が難しいです。この研究では、レアルワールドに基づくドリフ
空間理解は、物理世界と静的のセマンティック理解の間でつながるために不可欠です。多くの空間タスクは、場所、領域、パスの自然な表現は、ポインティングやマーキングなど、連続的な視覚的シーンで行われることが多いが、現行の空間推論
自動化された生理学ラボでは、透明なプラスチック製品を認識、位置付け、操作するために視覚知覚が必要ですが、対象となる高品質のリアルワールドデータセットは現在限られています。この研究では、複雑なマルチオブジェクトのシーンを扱
パスロジは、現在、パスロジ認識のための画像言語モデルを評価するために広く使用されていますが、この研究では、パスロジ認識において画像言語モデルの視覚知覚が機能していることを疑問に問っています。
感情認識は、現代のアギを促進するために不可欠ですが、大規模
We present HyperImageNet, a large-scale benchmark for fine-grained hyperspectral land-cover understanding. The
Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing thei
Hyperspectral salient object detection aims to identify visually salient regions from hyperspectral images. Ex
Current identity customized video generation methodologies are predominantly limited to single-identity scenar
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and back
Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the v
Recently, cross-domain few-shot facial expression recognition (CF-FER) has received considerable attention. Ho
この研究では、歯科CBCT画像中のメタルアーティファクトを除去するための循環互換的アドバーサリアルネットワーク(CycleGAN)を提案します。CycleGANを使用すると、メタルアーティファクトを除去した後、CBCT画
この研究では、マイメイク移植を改善するために、マイメイクの強い地域性を考慮したRegion-Controllable Diffusion Transformer(MagicMakeup)を提案します。
この研究では、都市ウォークビデオを分析するために、4つのモダリティの表現(スペース時領域情報、時間平均画像、オーディオ符号化、テキストベースの表現)を使用しました。
この論文では、DINO-VPTという手法を提案します。DINO-VPTは、Hierarchical Visual Prompt Tuning(HVPT)を使用して、物理的なスポーフィングとデジタルスポーフィングを検出しま
この論文では、finger vein画像から年齢と性別を推測するためのMulti-InstanceAge and Gender Estimation(MAGE-Vein)モデルを提案します。
この論文では、Lumeraという手法を提案します。Lumeraは、Engine-Native 3D World ReconstructionとLightsを検出するために使用します。
この研究では、WhereEditという手法を提案します。WhereEditは、Mask-aware Local Latent Editingを使用して、一ステップの画像編集を実行します。
この論文では、Webly Supervised Multi-Label Recognition(WS-MLR)という手法を提案します。WS-MLRは、web画像データセットを使用して、多ラベルを解釈します。
この論文では、ViSTR-Benchという手法を提案します。ViSTR-Benchは、MLLMが動的シーンから情報を取得できるかどうかを評価します。
Pixel-aligned Gaussian splatting enables efficient and generalizable novel-view synthesis. However, high-resol
Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing dat
Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve
人間の知能を模倣するクロアニオトミー手術のフレームワークを提案します。このフレームワークは、前方計画と後方実行を組み合わせて、手術中に手術台の位置を自動的に調整することで、人間と同様の安全で効率的な手順を実現します。
視覚モータリティポリシーを学習する際、人間が視覚アタッチメントを理解し、修正できるようにするため、視覚アタッチメントを明示的にしたフレームワークを提案します。
オートメーションされたマニピュレーションを目的とした、大規模なテーブルトップのデータセットであるTableVerse を提案します。このデータセットには、物理的に可能な実世界のレイアウトを生成する実用的な方法が含まれてお
Predicting how deformable objects evolve under robotic manipulation is a longstanding challenge. Existing appr
Federated learning (FL) enables multiple clinical institutions to collaboratively train a shared disease class
この研究では、ナノポア測定器から得られる複雑な信号を分析するために、多モーダル変換ニューラルネットワーク (Multi-modal Transformer) を提案し、信号分類の精度を向上させた。
Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption th
ネガティブ メトリクスの因数分解を扱う研究、非負行列因数分解 (NMF) を用いて因
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting t
この研究では、人工知能の研究者と神経科学者の間の分野を結びつけるために、脳のシステム構造を研究し、その研究から導かれた新しいアプローチを提案しました。
Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single
Classifier-free guidance (CFG) is the default mechanism for conditional generation in diffusion models, but th
This paper examines the question of whether artificial intelligence (AI) systems can be creative, approached f
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across
Interactive image segmentation is critical for efficient image annotation; however, existing methods often req
We present a privacy-preserving framework for synthetic lung CT slice generation developed for the Image-CLEFm
Traditional query processing engines require continuous development and extensions to support new techniques a
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script langua
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge fo
オフラインでの短時間の視覚生成が一般的な人間の行動の分析では、人間の行動の長期的な視覚生成は、実践的な長時間の視覚生成では実行不能である。StreamHOI は、人間間の視覚的な行動の生成を生成したいくつの画像を使用して
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, indep
MRI画像の強度正規化方法を7つ比較し、3DUネットワークモデルでMeniscusの分割精度を評価。
Multi-object tracking (MOT) plays a fundamental role in visual perception, where accurate trajectory predictio
Small language models and coding agents increasingly generate web front-end code, yet their outputs are typica
異なる順番で画像と質問が提示される場合、視覚言語モデルはモデルのパフォーマンスに大きな影響を受けることが発見された。
AIは、特に当代芸術作品のパスティーシュを作成する能力が高いが、これらの作品はどれだけ実際の作品と似ているかを調べました。
Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic fac
Virtual reality (VR) headsets (e.g., Meta Quest, Apple Vision Pro) provide a seamless user experience due to t
Impact hammers, also known as rock-breakers, are essential machines in mining operations, where they perform s
Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and dif
Novel View Synthesisは、入力画像から新しい視点の画像を生成するタスクです。ATSplatアルゴリズムは、3次元ガウススプラッタリングを Feed-forward に適合させました。これにより、ATSp
多モーダルラージランゲージモデルは、視覚言語タスクに強いですが、高い推論コストで問題となっています。Look Less, Think Fasterアルゴリズムは、単位次元を個別に最適化することで、多モーダルラージランゲー
複数ターンのファッション画像検索は、実世界のファッション検索では重要なタスクです。Diverse-Intent Multi-Turn Fashion Image Retrievalアルゴリズムは、異なる検索用途を扱うこと
画像理解のための多モーダルラージランゲージモデルは、強力ですが、まだ能力と限界については明確な理解が不足しています。この論文では、多モーダルラージランゲージモデルが画像理解においてどの程度の能力と限界を持つか、を分析し、
Housing-level urban physical examination is essential for identifying residential building problems and suppor
Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across divers
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existi
Remote sensing image editing aims to modify remote sensing images according to natural language instructions w
Attention Mechanism (AM) selectively focuses on essential information for imaging tasks and captures relations
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracte
Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure,
Deep gaze estimation works well in controlled capture but degrades in unconstrained settings, where systems mu
セグメンテーションのパフォーマンスの向上と計算的リソースの削減を目的として、Lean-SAM2は対象領域をアサインする対象アンバウンダリーセグメンテーション(SAM2)にターゲットアンチャイニングされたメモリとエンコーダ
OffNadirLocは地学化におけるオフナジアムの視点を考慮するための基準セットを提案します。これにより、ドローンと衛星画像の交差視点地学化プロセスでは重要な構造的シーン理解と内部ドメイン間の関係的制約に重点を置くこと
ステレオマッチングは3次元再構成において重要なタスクです。この研究では、ステレオマッチングを確率的生成タスクと組み合わせ、オブジェクト検出の向上を目的として、ステレオマッチングフレームワークと潜在分配を統合する方法を提案
この研究では、スイートペッパーの収穫前期予測を目的として、多モード時系列データを統合するための深層学習フレームワークを提案します。
OffNadirLocは交差視点地理位置を推定するための基準セットを提案します。これにより、ドローンと衛星画像の交差視点地理位置推定プロセスでは重要な構造的シーン理解と内部ドメイン間の関係制約に焦点を当てることができます
ETPデザイナはマルチモーダルな電子シアターのデザインを自動化するフレームワークを提案します。
Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by ge
Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, seman
Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a sin
Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-
この研究では、糖尿病性黄斑病変の検出を目的としたPRISM-DRシステムを開発しました。このシステムは、医師が見逃す可能性がある小さな低コントラストな病変を見つけるのに役立ちます。
この研究では、ドローンで小さな物体を認識することを目的としたメモリ拡張型大規模言語モデルを開発しました。このモデルは、複雑なドローンの場面で、ユーザーの指示に従って物体を識別できるようになります。
The Segment Anything Model 2 (SAM2) has advanced temporal promptable segmentation, yet its deployment remains
この研究では、可視化された質問への対応を評価するために、新しい方法を提案しました。この方法は、質問への回答の正確性だけでなく、質問への回答のパターンや特徴も評価することができます。
自動運転システムには、道路のトポロジー(ドライバブルレーンとその接続性)を理解する機能が必要です。最近の検出モデルは360度の前方視野からボリュームイメージを取得することで、道路上のレーンのトポロジーを推測することができ
この研究では、地象性AIにおける物理的知識を使用してポーラリメトリック合成開口ラダール画像を分類するための新しいモデルを提案しました。このモデルは、ラダール画像を物理的なプロセスと関連付けることができます。
自律ロボットには、障害物や事故の回避能力が必要です。これは、障害物や事故の回避能力が強化されていれば、障害物や事故に対しての対策がより効果的になります。障害物や事故の回避能力が強まることで、ロボットが障害物や事故から安全
Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-relat
VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rar
The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated sys
We propose a unified variational framework for image segmentation under sparse pixel-level supervision. Our me
The sense of touch is central to manipulation, especially when vision is occluded or ambiguous. Although combi
The training of learned inertial odometry depends on dense, high-precision position ground truth from motion c
ReferTrack は、自然言語で対象の車両に付近する自動車を追従させるシステムである。このシステムでは、対象の車両に付近する自動車を認識する後、自動車の動きを予測する。
このプロジェクトでは、農林業用自動化トラクターのデジタルツインモデリングが行われた。デジタルツインはCAN通信を使用することでトラクターの動きを模倣し、実際のトラクターの動作をシミュレートする。
V2F は、日用消費財をロボットで取扱するためのシステムである。このシステムでは、ロボットが消費財を取扱うときに必要な力を予測し、物体を傷つけたり、物を作れなかったりするのを防ぐ。
深層学習を用いた形状推定モデルを作成し、オープン平面曲線の形状を推定するための深層学習モデルを提案した。
Modeling galaxy-galaxy strong gravitational lenses to infer the brightness of the source galaxy and the mass d
逆問題を解くために、拡散ベースのサンプリングアルゴリズムが提案されていました。これにより、解の特性の正確さが向上することが期待されます。
fMRIデータから視覚情報を解釈するために、スパイクニューラルネットワークを用いた方法を提案し、fMRIデータから視覚情報を解釈する検証を行う。
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether t
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depe
The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms
A single embedding space that covers text, images, video, and audio lets one index serve every query a user ca
Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combi
The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that
Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules
Long-video question answering requires a model to preserve visual evidence over time without repeatedly reproc
Incorrect disposal can contaminate campus recycling streams, and a bin-mounted camera could provide feedback a
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs
While machine learning-based weather models hold significant promise, they struggle to predict the detailed st
Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due
Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasin
画像生成において、材料、 객체、領域を制御することが難しい問題がある。 Diffusion Transformers はテキストと画像を組み合わせて処理できるが、どちらをどの程度影響させるか決める仕組みがなかった。 その
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making th
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond
オリジナルのデータとZoom-Inのツールを組み合わせた方法、OmniReasonerを提案する。これにより、オリンモードルLLMsの長いオーディオビデオの論理的推論を改善できる。
生成モデルの構築のための新しいアプローチが提案されていました。これにより、生成モデルの構築が効率化され、強い表現力が得られるようになります。
記述情報に従って画像や動画データを混ぜ合わせる「対数混合法」を拡張する方法、InstructMixupを提案する。これにより、データを拡張しながらデータの内容とラベルが維持される。
計算機ビジョンと画像認識では、画像の視覚的なリッチネスを評価するために有用な指標が求められるが、これまでの指標は制限があった。この問題を解決するために、チャンネル空間の分散を利用した指標を提案する。
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across
実際の空間知能では、空間に続いて流れるビデオを理解する必要がある。この問題を解決するために、4次元空間を理解することができるモデルを提案する。
心臓の囲みの区別は、食道肥厚の測定に重要であるが、しかし、これを正確に区別することは難しい。これを解決するために、周囲の解剖学的構造を利用して囲みの区別を改善する方法を提案する。
この研究では、スパイナル サブアルテラノスパース内の安全な移動、操縦、内視鏡撮影を可能にする医療用ロボットを提案します。
自動運転のための計画とは、状況理解、タイムリーな推論、行動選択というものがあるが、しかし、これらの要素を組み合わせるのは難しい。これを解決するために、シーン理解を分離することによって、計画を安全かつ有効性のあるものにする
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more t
この研究では、高空飛行の無信号位置指示のNGPS (Next-Generation Positioning System)というフレームワークを提案しました。NGPSは、GPSの信号を利用せずに位置推定を可能にします。N
World Action Models(WAMs)は、ロボットマニピュレーションをモデル化するパラダイム。WAMsは、視覚ステートトランジションとロボットアクションを同時にモデル化する。しかし、既存のWAMsは、一定の時
Robot-assisted minimally invasive surgery (RMIS) offers major benefits over open and conventional laparoscopic
この文書では、閉回路交通シナリオ生成のための変分ベースのアプローチ「E2E-CDiff」を提案しました。これを使用すると、実世界に近い交通ルールを生成したり、交通ルールを操作することができるようになります。
Neural simulation-based inference enables parameter estimation for complex models, but typically requires the
Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the
Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, y
Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete
エッジロボチクスでの画像認識精度を安定させ、その安定性を確保するために、量化後のパフォーマンスを向上させ、分散型データ量化を実現し、分布シフトの影響を緩和する、新しい機械学習アプローチを提案します。
Macro placement still requires substantial manual refinement in industrial physical design flows. We present M
ロボット制御を効率化するために、パッチを用いた政策学習を提案し、密集された視覺表現を用いて実装することを目的としている。
existing VLA modelの制約を解決するためのforce-based memory method、FM-VLAを提案する。
existing robotic grasping methodの限界を解決するためのsim-to-real transfer methodを提案し、成功率を向上させる。
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but becaus
This paper introduces a method for real-time processing and transmission of autonomous underwater vehicle (AUV
existing robot navigation methodの限界を解決するためのglobal traversability prior extraction methodを提案し、オフロード環境でのロボット移動を実
Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-en
ロボット車の視覚システムは、高精度でリアルタイム性能を持つロジスティクス車両の位置検出を実現する必要があります。従来の手法では、複数のモデルが連続してインフェレンズされ、インフェレンスラティシーが増加し、高規模デプロイメ
この論文では、ConceptTreeというフレームワークを提案しています。このフレームワークは、人の見える概念を使用して、マニピュレーションの高位のスキル選択を表現し、透明性を高めます。
In search and rescue operations, there is a period known as the "golden time" during which the probability of
この論文では、ラインダブルロボットのための自発的アクション生成を実現することを目標とし、vision-language 指向性の指令によりロボットが自発的に動作することができることを示します。
Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts
採掘ロボットの性能向上を目指したSeg2Graspを構築し、セグメンテーション、グレイシング、クラスフィルタリングの3つのモジュールで構成されます。セグメンテーションモジュールではTransformerを利用したオブジェ
四足ロボットのナビゲーションのための予測的推論方法が提案されます。ロボットは、現在の観察と短期的な記憶によってアクションを選択しますが、障害物の発展を予測することができないため、このアプローチには課題があります。この課題
The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although re
Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machi
We study the expressivity of shallow polynomial neural networks (PNNs) with monomial activation functions over
DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-mod
Incorporating prior maps significantly enhances the accuracy and robustness of pose estimation in visual-inert
Monocular foundation models provide dense geometry but usually lack a stable metric scale. This paper presents
Precise metric depth estimation is fundamental for autonomous robot navigation, yet monocular systems inherent
Conditional diffusion models have become a powerful and flexible framework for learning complex conditional di
Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous d
End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor obser
With the growing demand for robotics, autonomous drones, and wearable extended reality systems, the deployment
Accurate articulation angle estimation of trucks with trailers is critical for autonomous driving and advanced
This paper proposes an indoor navigation system for the visually impaired, leveraging Ultra-Wideband (UWB) pos
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazar
脳のニューロン同士のつながりを分析する方法を提案する。この方法では、神経伝達の構造を考慮しながら、ニューロン間のつながりを分析できる。
Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tigh
VTLocフレームワークは、視覚情報と触覚情報を統合し、ロボットハンドの位置を推定することで、ロボットハンドの位置推定と動作操作を実現します。
PIXIEフレームワークは、6次元オブジェクト位置推定を実現し、ロボットハンドの制御と物体の操作を実現します。
この研究では、ロボットのナビゲーション時間と注釈時間の制約を考慮したオブジェクト検出フレームワークを提案します。
Robotic navigation in unstructured environments requires robust situational awareness to safely traverse hazar
高次元カテゴリデータを可視化するため、hierarchical optimizing linear assignment (HOMALS)を使用し、可視化に役立つ関連表
Modern generative models are increasingly trained using model-generated signals, creating both opportunities f
難しい環境でマルチオブジェクトの検知と追跡が可能なPiVoTを開発、実用的なソリューションを提案した。
重尾流を見つけるには、Standard diffusionとflowマッチングモデルの欠陥を解決するRandom Clocksを提案した。
Interfacing with Biological Neural Networks (BNNs) requires encoding information into stimulation patterns tha
Spiking Neural Networks (SNNs) trained through unsupervised Spike-Timing-Dependent Plasticity (STDP) have been
ストロチャスティックプロセスに観察値を組み込むことが困難であれば、単に観察値を観察できるものを学習しているという理解を拡張する新しいフレームワークを発表
Circular data, representing angles or directions, are frequently encountered in computer vision, biology, geol
この研究では、新しい特徴融合手法を提案した。この手法は、上からの特徴と下からの特徴の関係性を考慮することで、特徴を効率的に融合し、三次元データを2次元サムライグラフにコンパクトに表現する機能をもたらせる。
The human visual system (HVS) employs foveated sampling and eye movements to achieve efficient perception, con
これは、社会的行動を予測するための新しいフレームワークであるSocial-spatial dependenciesを提案し、個々のエージェントが社会的信号を学習する能力を向上させる。
Pairwise comparisons are fundamental in the analytic hierarchy process. Various consistency indices have been
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability
イベントベースセンシングの活用と生物学的インスピレーションを利用した障害物検出を実現するために、飛行経路を用いた新しいアプローチが提案される。このアプローチは、イベントベースセンシングの活用と生物学的インスピレーションを
A central goal of current Spiking Neural Network (SNN) research is to improve their accuracy toward becoming l
Strategic value can fall when an option becomes visible. A route, signal, bet, or opportunity may be attractiv
この論文では、多目的最適化の解釈を向上させるために用いる Partition-Guided Distance Saliency (PGDS) アルゴリズムを提案しました。これにより、多目的最適化の解釈の向上に役立つものと
This research investigates the optimization of Convolutional and Dense Neural Networks (CNNs and DNNs) for aut
Continual learning (CL), where a model is trained on a sequence of data tasks, is increasingly being adopted a
Current models of representational reliability in neural populations focus on temporal stability: whether popu
Recent breakthroughs in synaptic-resolution network connectomics have revealed that brain circuits feature fin
Hybrid neural networks (HNNs) that integrate artificial neural networks (ANNs) with brain-inspired neural netw
神経網路の設計を目指す本研究では、ANNとSNNを組み合わせたハフマン式設計法
AIが人間と協力して作り出すアイデアを評価するための新しい手法を提案し、創造性の評価を向上させた。
Spiking Vision Transformers (SViTs) have emerged as alternative low-power ViT models, but their large sizes hi
How the wiring and functional organization of cortex shape recurrent computation remains a central question in
Learning in biological multilayer neuronal networks offers insights that extend beyond the classical weighted-
スパイクニューラルネットワークは、エネルギー効率のよいAIモデルです。この研究では、スパイクニューラルネットワークのアクセラレータを実装し、その性能をテストしました。