3D-Aware VLMs with Implicit and Explicit Geometries
3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
- 用途
- 3次元空間理解技術の開発
- 難易度
- Hard
- コスト
- High
「detection」の検索結果
151 件3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
印刷品質管理技術のための新しいアプローチであるシンセティック データ生成フレームワークを提案しました。このフレームワークは、ロトグラビューグラビング技術における品質管理のためのシンセティック データを生成することで、印刷
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for ena
Scaling inference-time computation has emerged as a reliable method to improve the performance of large langua
チェーン・オブ・サウト reasoning モデルの収束不明確さを解決する研究。このモデルの不完全収束は、生成するトークンの数に依存し、モデルには収束しない限り問題を解決する能力がない。これを解決するための予測を終了する
この研究では、自己説明の信頼性を検証するためのRL方法を提案し、自己説明の信頼性を直接最適化するための新しいアプローチを検討します。
この研究では、長期
Automated detection of vision impairing retina-based ocular conditions from fundus images is important for ear
Spectral methods are among the most widely used techniques for community detection, clustering, and graph lear
RFマップ(無線周波数マップ)を推定するためのTransmitter-Aware Diffusion(送信機認識拡張)を提案した研究で、この方法によりRFマップを効率的に推定できる。
LLM(言語モデル)の評価における位置バイアスを分析するための方法を提案した研究で、この方法により、位置バイアスが評価結果にどのような影響を与えるかが明らかにできる。
Advanced Persistent Threats (APTs) remain difficult to detect because only a small fraction of events in large
A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that
A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard t
Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This l
Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, tec
Replacing an object with one that differs in category or shape requires complete source removal, natural targe
この研究では、大規模言語モデルを使用して、basketボールの動的理解に基づいて、プレイヤーへの関わりや時間境界を推測するモデルを開発しました。
この研究では、大規模言語モデルを活用して、信念の共有を組み合わせるモデルを開発しました。大規模言語モデルを活用することで、信念の共有を推測することができました。
LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records,
非真実性の評価において、評価手法が多様な評価基準を捕捉する能力に乏しく、評価者間の偏見が存在する問題を解決するために、CSPF (Constrained Shared-private Fusion) を提案している。
ビデオ内の物体の空間推論を同時に行うことで、現存するタスク固有の注釈を超えた統一的なビデオ推論システムを構築した。
一部のGaussianスプレイティングを利用したSL
Multi-object tracking in dense crowds requires solving a bipartite assignment problem between detections and t
Topological maps are key outputs of autonomous driving perception systems, delivering essential road informati
AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure an
Polarization cues benefit applications such as material detection and de-reflection, yet acquiring them typica
Accurate forest inventory and large-scale mapping are essential for ecosystem monitoring and sustainable fores
大規模言語モデルはさまざまな画像をテキストに変換する上で優れた性能を示しているが、発生するホログラフィックな診断にはまだ解決策が必要です。この研究では、主流の粗い検出方法の欠点を補うため、細部の診断方法を提案しています。
Hyperspectral salient object detection aims to identify visually salient regions from hyperspectral images. Ex
Current identity customized video generation methodologies are predominantly limited to single-identity scenar
Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the v
この論文では、非コントラストCT(NCCT)スキャン中の脳梗塞領域を正確に分割するために、周囲境界を特徴としているFrequency-Spatial Boundary Network(FSB-Net)を開発しました。
この研究では、路上の3Dオブジェクト検出を改善するために、外部性を考慮した地域認識のアラーカンシーを提案します。
この論文では、Lumeraという手法を提案します。Lumeraは、Engine-Native 3D World ReconstructionとLightsを検出するために使用します。
この論文では、安全な人とロボット間の対話を目的とした、人間の姿勢推定とロボットの動作制御の一連のネットワークが提案されます。
人間の知能を模倣するクロアニオトミー手術のフレームワークを提案します。このフレームワークは、前方計画と後方実行を組み合わせて、手術中に手術台の位置を自動的に調整することで、人間と同様の安全で効率的な手順を実現します。
視覚モータリティポリシーを学習する際、人間が視覚アタッチメントを理解し、修正できるようにするため、視覚アタッチメントを明示的にしたフレームワークを提案します。
When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We a
この研究では、クラスター検出アナライザーにおける量子力学の応用を研究し、精度を向上させた。
Neural network misclassifications exhibit characteristic spectral instability in internal activations that is
圧縮流体の解析を目的とした新しいアプローチ、Entropy-Stable Learned Finite Volumes を提案する。
データクリーンシングを扱う研究、CURED を用いてデータクリーンシングを提案する。
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting t
この研究では、Androidマルウェアの検出に使用されるデープラーニングモデルをOptimizeする方法を提案しました。
この研究では、実世界ロボットのシーケンシャル予測に使用できる、diffusion-based frameworkを提案しました。
この研究では、マルチエージェントによるトリージュア
生産的言語モデルの利用による金銭的感情分析に対処するための方法を提案している。複数のエージェントを活用したコミティー方式を使用し、さまざまな粒度のテキストデータに対応できるように、単語レベルのルールベースアプローチ、句節
We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a cont
Harnessing the potential of electroencephalography (EEG) for brain research is fundamentally limited by intrin
Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as l
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge fo
Quantized small autoregressive reasoning models can enter long, repetitive, or unproductive trajectories, yet
回線流量データに対するノイズと漂移を考慮した波列減少アルゴリズムを実装し、静的な波列減少法が漂移のあるシナリオでは効果を低下していると指摘する。
大きな言語モデルは真実の情報を提供できるように見えますが、実際は虚偽情報を提供することが多く、これを検知、検出、および検証するための基準を作成するため、HalluTruthQAが開発されました。
This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as par
Detecting media bias automatically is difficult because biased framing is often subtle, yet in domains such as
Open-world video anomaly detection (OWVAD) is expected to detect events that match a user-specified definition
Housing-level urban physical examination is essential for identifying residential building problems and suppor
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existi
We present a two-stage vision system that detects EEG cap electrodes in a live webcam stream and validates the
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracte
OffNadirLocは地学化におけるオフナジアムの視点を考慮するための基準セットを提案します。これにより、ドローンと衛星画像の交差視点地学化プロセスでは重要な構造的シーン理解と内部ドメイン間の関係的制約に重点を置くこと
OffNadirLocは交差視点地理位置を推定するための基準セットを提案します。これにより、ドローンと衛星画像の交差視点地理位置推定プロセスでは重要な構造的シーン理解と内部ドメイン間の関係制約に焦点を当てることができます
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T
Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autono
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic
Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a sin
Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-
この研究では、糖尿病性黄斑病変の検出を目的としたPRISM-DRシステムを開発しました。このシステムは、医師が見逃す可能性がある小さな低コントラストな病変を見つけるのに役立ちます。
マグネティック リゾナンス イメージング (MRI) のデータ収集には、多くのエネルギーと時間が必要です。アクティブ サンプリングは MRI の速度を増加させる技術ですが、現在のアプローチでは、低周波数部分(解像度)と高
地上を表す重力式マップの高解像度版が、多くの用途で役立ちます。たとえば、市区町村の変化を監視したり、エネルギー対策を向上させたり、温室効果ガスの排出量を追跡したりすることができます。4つの主要な全世界建物Rasterデー
Tracking objects through state transformations is essential for understanding real-world dynamics. However, ex
Automatic pain localization, which involves identifying the anatomical origin of pain from peripheral physiolo
Soft robot exteroception is increasingly being explored for a variety of field applications. In this work, we
ReferTrack は、自然言語で対象の車両に付近する自動車を追従させるシステムである。このシステムでは、対象の車両に付近する自動車を認識する後、自動車の動きを予測する。
Clinical Pathways は、ロボットが実際の環境で安全に動作するためのシステムである。これは、ロボットが病室で安全に作業し、医療スタッフや患者を守る。
ハイパーネットワークを用いた知識付与法を提案し、大規模言語モデルに確実に知識を付与する方法について検討した。
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only fina
We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation u
記号化の種類 (verbatim vs. intended) は、現在の音声認識モデルの評価に影響を与えるが、このような制約はモデルのトレーニングに影響しないことが多い。しかし、ここでは、制約はモデルのトレーニングに影響
The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms
Stance detection aims to identify whether a text expresses a favorable or opposing attitude toward a given tar
As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understan
Incorrect disposal can contaminate campus recycling streams, and a bin-mounted camera could provide feedback a
Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasin
記述情報に従って画像や動画データを混ぜ合わせる「対数混合法」を拡張する方法、InstructMixupを提案する。これにより、データを拡張しながらデータの内容とラベルが維持される。
難しい環境で運用するためには、自動空飛ブイロード(UAV)が実際に障害物に存在する距離を判断し、安全な軌跡を計画することが求められる。 これを行うために、複数のステージ(マッピングと計画)を連続化した、サイン・ディスタン
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across
Many Blind and Low-Vision (BLV) people rely on guide dogs for moment-to-moment navigation, such as staying on
The report envisions a decade in which drones move goods, medical supplies, and information at a scale compara
この研究では、高空飛行の無信号位置指示のNGPS (Next-Generation Positioning System)というフレームワークを提案しました。NGPSは、GPSの信号を利用せずに位置推定を可能にします。N
LLMが広く使用されるようになり、LLMを識別するツールが開発されている。 しかし、識別システムは、使用者の行動に影響を与えている。つまり、識別システムが機能しないと、ユーザが別のシステムを使用することに関連し、最終的な
この論文では、パス値学習を行うためにpath signaturesという
Dynamic graph learning aims to capture evolving structural and semantic patterns in real-world systems, such a
工場の中の異常が検出されるように設計された機械学習モデルを提案しています。通常の方法では、モデルはビデオ内のすべての内容を考慮し、複雑な問題を解決することは困難です。提案されたモデルのアプローチは、オブジェクトを検出して
Eco-Cooperative Adaptive Cruise Control (Eco-CACC) systems rely on accurate localization, signal timing, and i
エッジロボチクスでの画像認識精度を安定させ、その安定性を確保するために、量化後のパフォーマンスを向上させ、分散型データ量化を実現し、分布シフトの影響を緩和する、新しい機械学習アプローチを提案します。
existing robotic grasping methodの限界を解決するためのsim-to-real transfer methodを提案し、成功率を向上させる。
existing fault detection methodの限界を解決するためのadaptive stress testing methodを提案し、商用自動運転システムの故障率を減らす。
existing AUV development methodの制約を解決するためのrobustなオートニモティクス基盤と機械学習アライアンスを開発する。
existing robot control methodの限界を解決するためのmemory-driven orchestration method、RoboHarnessを提案し、長期計画を実現する。
existing Embodied Foundation Modelの制限を解決するためのcontact-point prediction とnative 3D grounding methodを提案し、更に能力と
ロボット車の視覚システムは、高精度でリアルタイム性能を持つロジスティクス車両の位置検出を実現する必要があります。従来の手法では、複数のモデルが連続してインフェレンズされ、インフェレンスラティシーが増加し、高規模デプロイメ
自律ロボットの位置決めは、ロボットナビゲーションの主要なタスクです。ロボットが予測できない、非静的な障害物、またはロボットが未知の環境に入ることが多い。この研究では、ロボットのオドメトリと距離サンプリングを組み合わせて、
In autonomous driving development, a perception dataset is crucial, as it provides fundamental data for traini
In search and rescue operations, there is a period known as the "golden time" during which the probability of
採掘ロボットの性能向上を目指したSeg2Graspを構築し、セグメンテーション、グレイシング、クラスフィルタリングの3つのモジュールで構成されます。セグメンテーションモジュールではTransformerを利用したオブジェ
この論文では、低照明状況のためのSLAM実現を目標とし、LiDAR、深さ、または熱センサなどの補助的なセンサを取り入れることでSLAMを改良します。
Physical artificial intelligence (AI) systems involve distributed sensing agents with embedded AI models that
Linear attention promises constant-time recurrent inference but degrades sharply on associative recall. We for
DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-mod
Indoor robots are increasingly employed for facility management tasks such as cleaning and inspection. These a
LiDAR place recognition supports loop closure, relocalization, and multi-agent map management. As robotic plat
Accurate articulation angle estimation of trucks with trailers is critical for autonomous driving and advanced
This paper proposes an indoor navigation system for the visually impaired, leveraging Ultra-Wideband (UWB) pos
Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and suppor
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazar
オフラインのデータ流れの中で、時間序列のデータに基づいてデータ変化を検知することができる方法が必要。この問題を解決するために、オンラインのデータ流れで検知した変更点をオフラインのデータ流れに適用する方法を提案。
Hallucinations and artificial text in LLM-generated outputs often appear as distributional deviations between
Neural Cellular Automataが複雑な形状を形成するプロセスを研究しました。
VTLocフレームワークは、視覚情報と触覚情報を統合し、ロボットハンドの位置を推定することで、ロボットハンドの位置推定と動作操作を実現します。
この研究では、接触の豊富なマニピュレーションを実現するための、データの収集と学習を改良した方法を提案し、ロボットの制御の精度を
この研究では、ロボットのナビゲーション時間と注釈時間の制約を考慮したオブジェクト検出フレームワークを提案します。
distillationにおける予測のみを扱う学習アプローチを提案し、それをテストした。
この研究では、調整されていない HMC と Langevin Sampler の偏りの解消について議論しました。調整されていないサンプラーは、通常、偏りのあるものであることが知られています。この研究
時系列データを観察し、その中に分割が起こっているかどうかを検知する方法はある。時間系列変化点を検知することができ、変化点の位置がどこにあるかを推定することができる。
この研究では、ウェアラブルデバイスで電気生理学記録(ECG)を分析するために使用される深層学習アルゴリズムを開発することを目的としています。このアルゴリズムは、エネルギー効率が高く、小型化が可能であるため、心臓の病気の検
難しい環境でマルチオブジェクトの検知と追跡が可能なPiVoTを開発、実用的なソリューションを提案した。
多種多様なデータに対して条件的独立性の検定を行う方法です。混合タイプデータに対して統一的な、効率的な、あるいは統計的に有効な解決策は存在しませんでしたが、グラフ上のノード間の距離を比较する方法を提案しています。
金融データの分布の変化を検出し、信用リスクモデルの監視を行う方法です。Jensen-Shannon-DivergenceやKullback-Leibler-Divergenceなどの分散量は異なる種類の変化を検出できます
暗号化された制御システムでは、クラウドがホモモルフィック暗号化された状態を操作し、動物達の動作をプライバシーで管理することができる。安全を確保するために、サイドチャネル攻撃のリスクを考慮しながら、制御機器が信頼できると仮
We show that a single climate realization can be decomposed into forced and internal components by treating ex
分類Forecasterの監視と評価を容易にするために、可変期間のバックテストを導入し、情報依存性を考慮したCalibrationを可能にしました。
A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimiz
この研究では、従来のNAS方法のコストを抑えるための方法を開発します。 この方法では、NASをトランスフォーマーを使用して実行します。
Detecting muscle fatigue via surface electromyography (sEMG) is essential for applications in sports, rehabili
Point-adjustment (PA), for years the default scoring protocol in time-series anomaly detection (TSAD), was sho
Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional struc
Next-generation wireless networksにおける分散型ゲーム理論を用いた6Gのセキュリティを研究します。分散型ゲーム理論は、6Gの通信システムが環境の認識とデータの伝送両方を実現するために必要な
イベントベースセンシングの活用と生物学的インスピレーションを利用した障害物検出を実現するために、飛行経路を用いた新しいアプローチが提案される。このアプローチは、イベントベースセンシングの活用と生物学的インスピレーションを
この研究では、アルツハイマー病の前期診断と生物学的マーカーの検出にAI技術を適用します。AIモデルをトレーニングするために、電気エイセフィログラム(EEG)データを使用し、精度を高めます。また、AIモデルが得た情報を分析
「計算的生命」論文は、ペアが相互作用する複雑なシステムにおいて、自己複製体を容易に発見できることを示しました。ここでは、逆説的には、単純な遺伝子突然変異ウォークを用いた自己複製体の検出に新しいアプローチを提案し、この方法
分散型時間関数記憶体を用いた異常検知システムを開発しました。このシステムは、関連のあるエンティティの予兆行動を共有メモリ空間に保存し、異常検知に役立ちます。このシステムは、異常検知に役立つ新しい方法を提供します。
Large language models (LLMs) increasingly mediate strategic interactions through natural language, making sema
Abnormality detection in complex systems faces two practical barriers: abnormal labels are scarce, and binary
Continual learning that is gradient-free, local, online, and append-only is attractive for edge and streaming
The escalating congestion in orbital space demands advanced monitoring solutions. This work presents a compreh
Efficient processing of continuous audio streams remains a key challenge for real-time and resource-constraine
この研究では、Controlled Dynamics Attractor Transformer (CDAT)を提案しました。このTransformerは、Self-Attention MechanismとAssocia