MLinfo | 機械学習・AI論文まとめ

Discovering Functionally Selective Brain Regions with a Deep Topographic Multimodal Model

この研究では、脳部帯域内のニューロンが同じ反応プロファイルを持つと仮定し、近接な脳部帯域内のニューロンの反応プロファイルを推論し、分野間の結合を特定しました。

自然言語処理RAG画像マルチモーダル

用途: 脳部帯域の研究
難易度: Hard
コスト: High

Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models

Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of ro

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Transition-Based Digital Twin Modelling for Alzheimer's Disease under Sparse Longitudinal Data

Alzheimer's disease (AD) progression is highly heterogeneous and is typically observed through sparse and irre

説明可能深層学習軽量化・量子化分類生成予測

用途: 分類
難易度: Hard
コスト: High

センサ/時系列コンピュータビジョンマルチモーダル時系列

FMplex: Model Virtualization for Serving Extensible Foundation Models

Foundation models (FMs) are increasingly used as backbones for downstream tasks across language, vision, time-

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies

Vision-language-action (VLA) policies provide strong priors for language-conditioned manipulation, but remain

自然言語処理RAGテキストマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

動画大規模言語モデルを使用した質問に対する回答を研究。モデルの能力と限界を調査し、質問に対する答えを生成するための方法を提案した。

深層学習軽量化・量子化テキスト動画マルチモーダル

用途: 動画大規模言語モデルを使用した質問に対する回答
難易度: Hard
コスト: High

LargeMonitor: Monitoring Online Task-Free Continual Learning via Large Pretrained Models

オンライン学習の継続学習では、モデルは非駅性データストリームから知識を継続的に蓄積する必要があります。モデルのパラメータはトレーニング中に効果的に調整される必要がありますが、パラメータ効率的なプロンプトチューニングや

深層学習軽量化・量子化検出テキストマルチモーダル

用途: オンライン学習の継続学習
難易度: Hard
コスト: High

説明可能センサ/時系列深層学習CNN画像テキストマルチモーダル

Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study

この研究では、ゼロショットセマンティック再特定の基準を設定し、画像のセマンティック特定を自動化します。

用途: セマンティック再特定
難易度: Hard
コスト: High

PRISM: Topology-Aware Cross-Modal Imputation for Modality-Deficient Federated Graph Learning

Multimodal federated graph learning (MM-FGL) aims to collaboratively learn from decentralized graphs with text

自然言語処理RAG画像テキストマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

センサ/時系列深層学習Transformer検出生成埋め込み

Multi-View Speech Representation Learning for Parkinson's Disease Detection Using Context-guided Cross-modal Attention

パーキンソン病（PD）の早期検出への取り組みとして、脳の損傷が発症前に生じる話術障害を分析するため、音声分析を用いてパーキンソン病の診断を提唱しています。

用途: パーキンソン病の早期検出
難易度: Hard
コスト: High

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

この論文では、VideoQA が過度に信憑性の

コンピュータビジョンマルチモーダル検出画像動画

用途: ビデオQA に対するカウンターファクタルの推論
難易度: Hard
コスト: High

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation

Multimodal large language models (MLLMs) commonly inherit the deep, symmetric Transformer backbone designed fo

用途: 生成
難易度: Hard
コスト: High

Driving Video Retrieval for Complex Queries with Structured Grounding

Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users

コンピュータビジョンマルチモーダルテキスト動画

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

理論的思考は、最新の基礎モデルシステムが安全かつ効果的に現実世界で動作するには必須のスキルであると考えられています。しかし、理論的思考の進進には、「ショートカット」問題が存在し、タスクは99％の正解率を達成するのに、ただ

自然言語処理RAGテキストマルチモーダル強化学習

用途: 理論的思考の強化問題
難易度: Hard
コスト: High

arxivGitHubあり2026-06-08

Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization

オンポリシーディストリレーションは、近年、重要なポストトレーニングの研究分野となりました。強い教師モデルを使用して学習トレッジを密に細かく指示することで、トピック認識を実現します。しかしなだな的にトークンレベルにおいてデ

深層学習軽量化・量子化マルチモーダル強化学習

用途: オンポリシーディストリレーション問題
難易度: Hard
コスト: High

深層学習軽量化・量子化テキストマルチモーダル強化学習

Stage-1 Controls the Entropy Regime, Not the Outcome

Two-stage post-training -- a Stage-1 warm-start (supervised fine-tuning, SFT, or on-policy distillation, OPD)

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Understanding Quantization-Aware Training: Gradients at Quantized Weights Bias to the Low-Loss Basin

Post-training quantization (PTQ) converts a trained full-precision model into low-bit weights without task-lev

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

コンピュータビジョンセグメンテーション動画マルチモーダル

C$^3$ache: Accelerating World Action Models with Cross Inference Chunk Cache

ワールドアクションモデルを高速化するために、情報のキャッシュと伝達を提案します。

用途: ワールドアクションモデルを高速化するためのキャッシュと伝達
難易度: Hard
コスト: High

自然言語処理大規模言語モデルテキストマルチモーダル

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

この論文では、VLM ゲームエージェントの評価基準が提供され、さまざまなタイプのエージェント間の比較が可能になる。

用途: VLM ゲームエージェントの評価基準
難易度: Hard
コスト: High

自然言語処理大規模言語モデル画像テキストマルチモーダル

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and op

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

品質予測/異常検知コンピュータビジョン動画認識検出画像テキスト

ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

LLMを用いた臨床研究論文の草案作成を支援するために、生成されたテキストを検証するためのアーキテクチャを設計。これにより、虚偽の citaion、数字の不正確な記録、およびガイドライン違反が防がれます。

用途: 医学論文執筆のサポート
難易度: Hard
コスト: High

センサ/時系列深層学習Transformer分類検出テキスト

ATN3D: Density-Aware LiDAR-Radar Early 3D Object Detection Under Extreme Sparsity

自動運転車やインテリジェント輸送システムなどの自動化された車両の感知には3次元オブジェクト検出が必要です。道路での長距離検出は困難ですが、道路ではこの「長距離」に対する感知と決定の時間は約1-2秒です。2つの主な課題が現

用途: 車のデッキの長距離認識に対する3次元オブジェクト検出
難易度: Hard
コスト: High

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multi

深層学習軽量化・量子化画像テキストマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs

Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning t

自然言語処理大規模言語モデルQA画像テキスト

用途: QA
難易度: Hard
コスト: High

CT-VAM: A Cerebello-Thalamic-Inspired Vision-Action Model for Efficient Visuomotor Control

Vision-language-action models have shown strong promise for robot manipulation, yet raw language is primarily

深層学習軽量化・量子化画像マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

MI向き品質予測/異常検知コンピュータビジョンマルチモーダル分類検出画像

Context-Aware Deep Learning for Defect Classification in Atomic-Resolution STEM

マテリアルの非破壊検査を目的としたContext-Aware Deep Learningが提案され、エアロックの欠陥を検出する。

用途: マテリアルの非破壊検査
難易度: Hard
コスト: High

Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer

ボディポーズ認識と行動解釈を目的としたReal-time body pose non-verbal communicationが提案され、人間の動作を認識して行動を解釈する。

コンピュータビジョンマルチモーダル

用途: ボディポーズ認識と行動解釈
難易度: Hard
コスト: High

自然言語処理プロンプトエンジニアリング生成画像テキスト

IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation wi

用途: 生成
難易度: Hard
コスト: High

Vision Language Model Helps Private Information De-Identification in Vision Data

ビジュアル言語モデル（VLM）は、プライバシー保護において有効性の高い能力をもつ。しかし、視覚データを扱う際のプライバシーリスクについては、それまでほとんど注目されていなかった。VLMを使用して、プライバシー保護を確保す

コンピュータビジョン物体検出分類検出画像

用途: ビジョン言語モデルを使用したビジュアルデータのプライバシー保護
難易度: Hard
コスト: High

MI向き深層学習軽量化・量子化テキストマルチモーダル強化学習

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' prefere

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

コンピュータビジョンマルチモーダルQA画像テキスト

Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care

連続的な治療に適した臨床級LLM医系であるBaichuan-M4を導入。臨床的な医療エージェントシステムであるBaichuan-M4は、統合的な医療エージェントシステムをベースとし、医療エージェントと医療エージェントの連

用途: 統合医療医系のためのLLMベースの医療エージェント
難易度: Hard
コスト: High

An Effective Router for Vision-Language Model Selection

Vision-language models (VLMs) with varying performance and resource requirements are widely deployed, making i

自然言語処理大規模言語モデル異常検知画像テキスト

用途: 異常検知
難易度: Hard
コスト: High

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning ov

自然言語処理RAG画像マルチモーダル強化学習

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis

Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but

自然言語処理大規模言語モデル生成画像テキスト

用途: 生成
難易度: Hard
コスト: High

品質予測/異常検知自然言語処理ファインチューニング検出画像テキスト

Failure-Aware Refinement of Vision-Language Model for Lithography Defect Detection

Semiconductor lithography inspection requires reliable detection of small pattern defects such as bridge, burr

用途: 検出
難易度: Hard
コスト: High

A multi-agent system for spine MRI report generation from multi-sequence imaging

Spinal pathology is a leading cause of pain and disability worldwide. Spine MRI is central to clinical evaluat

説明可能自然言語処理埋め込み・検索分類検出生成

用途: 分類
難易度: Hard
コスト: High

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

この研究では、静黙の口承のシンセシスを実現するためのフレームワークを開発します。このフレームワークは、静黙の口承のシンセシスと精度を改善することができます。

センサ/時系列自然言語処理RAG生成音声動画

用途: 静黙の口承のシンセシス
難易度: Hard
コスト: High

Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving

Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer acc

自然言語処理大規模言語モデルQA画像テキスト

用途: QA
難易度: Hard
コスト: High

説明可能品質予測/異常検知コンピュータビジョンセグメンテーションマルチモーダル

Interpretable Crisis Behavior Analysis Using Mobility and Social Media Data

人間は危機時に移動パターンやメディアの投稿のパターンが変化し、分析が難しいようになった。この研究では、運動データやメディアデータの統合を用いて危機時の行動パターンを分析し、危機の状況における行動を予測した。

用途: クライシス時の行動分析
難易度: Hard
コスト: High

自然言語処理大規模言語モデル生成テキストマルチモーダル

H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions

大きな言語モデルには記憶や推論機能があるが、ユーザーとの対話におけるこれらの機能の効果はまだ理解されているわけではない。これを受け、この研究では、人間の相互作用、特に会話における記憶と推論能力を評価するためのマルチモーダ

用途: マルチモーダル記憶の評価
難易度: Hard
コスト: High

品質予測/異常検知自然言語処理大規模言語モデル分類セグメンテーションテキスト

arxivGitHubあり2026-06-08

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

この研究では、低リソース言語や絶滅言語の辞書のデジタル化が重要であるが、マルチモーダル辞書をデジタル化する方法は今まで難しかったが、この研究では、最近のビジョン言語モデルを用いて辞書のデジタル化が容易になり、辞書内の文字

用途: ムルティリンガル辞書のデジタル化
難易度: Hard
コスト: High

品質予測/異常検知コンピュータビジョンマルチモーダル分類画像テキスト

Guide Me Out: A Framework to Benchmark VLM Operators Communication in Crisis Scenarios

危機管理では、コミュニケーションと地理

用途: 危機管理におけるコミュニケーションを評価する
難易度: Hard
コスト: High

センサ/時系列深層学習Transformer分類テキスト音声

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

Large language models (LLMs) provide a powerful reasoning backbone for speech understanding, but integrating c

用途: 分類
難易度: Hard
コスト: High

説明可能自然言語処理RAG画像テキストマルチモーダル

Explicit Representation Alignment for Multimodal Sentiment Analysis

Multimodal affective analysis aims to understand human sentiment and emotion by jointly modeling heterogeneous

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

CRANE: Knowledge Editing for Reasoning MLLMs

The emergence of reasoning multimodal large language models (MLLMs), which generate explicit chain-of-thought

自然言語処理大規模言語モデル異常検知画像テキスト

用途: 異常検知
難易度: Hard
コスト: High

Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level

LLMs are increasingly used to simulate human survey responses, but prior work has mainly evaluated replication

自然言語処理大規模言語モデルマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

表形式向き品質予測/異常検知自然言語処理RAG分類QA画像

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of

用途: 分類
難易度: Hard
コスト: High

arxivGitHubあり2026-06-08

Are Reasoning Vision-Language Models Robust to Semantic Visual Distractions?

Reasoning Vision-Language Models (VLMs) achieve strong performance on complex multimodal tasks, but reliable r

コンピュータビジョンマルチモーダル画像テキスト

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

コンピュータビジョン動画認識テキストマルチモーダル

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models

Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past int

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

HDSL: A Hierarchical Domain-Specific Language for Structured 3D Indoor Scene Generation and Localized Editing with LLM Agents

Text-driven indoor scene generation and editing require an intermediate representation that language models ca

自然言語処理大規模言語モデル生成テキスト3D

用途: 生成
難易度: Hard
コスト: High

品質予測/異常検知コンピュータビジョンセグメンテーション生成画像テキスト

Cranio-Diff: Diffusion-based Cross-domain Craniofacial Reconstruction with 2D X-ray Skull Guidance and Structural Identity Constraints

The state-of-the-art generative models, such as CycleGAN, Pix2Pix, and diffusion models have demonstrated rema

用途: 生成
難易度: Hard
コスト: High

GenEyePose: Patient-Free, Knowledge-Based Saccadic Eye Movement Modeling for Digital Neurophysiologic Biomarker Development

Eye movements, including saccades, are widely regarded as highly sensitive and objective biomarkers of neuroph

深層学習Transformer分類検出生成

用途: 分類
難易度: Hard
コスト: High

GD-MIL: Grade-Disentangled Multiple Instance Learning for Multimodal Biochemical Recurrence Prediction in Prostate Cancer

Biochemical recurrence (BCR) after radical prostatectomy is a critical endpoint in prostate cancer, yet risk s

深層学習CNN画像マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

品質予測/異常検知自然言語処理大規模言語モデル画像テキスト動画

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a crit

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification

Open-domain open-vocabulary detection (ODOVD) requires detectors to generalize to both novel categories and un

深層学習軽量化・量子化分類検出画像

用途: 分類
難易度: Hard
コスト: High

センサ/時系列コンピュータビジョン動画認識画像テキストマルチモーダル

IB-HFN: Information Bottleneck-Driven SAR-Optical Fusion Network for High-Fidelity Cloud Removal

Synthetic aperture radar (SAR)-assisted optical cloud removal aims to recover surface information obscured by

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

深層学習正規化・最適化手法分類生成セグメンテーション

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning

The rapid development of pretrained foundation models has enabled more general image segmentation. Multimodal

用途: 分類
難易度: Hard
コスト: High

Self-supervised Learning Matters: A Simple Ensemble Solution for Micro-Gesture Recognition

In this paper, we present XInsight Lab's solution to the micro-gesture classification track of the 4th MiGA Ch

自然言語処理ファインチューニング分類埋め込み動画

用途: 分類
難易度: Hard
コスト: High

説明可能コンピュータビジョンマルチモーダル生成画像テキスト

MAGIS: Evidence-Based Multi-Agent Reasoning for Interpretable Strabismus Clinical Decision-Making

Strabismus is a common ocular disorder that requires fine-grained subtype diagnosis for individualized treatme

用途: 生成
難易度: Hard
コスト: High

Vision-Language Guided Hyperspectral Object Tracking via Semantics Fusion and Contextual Template Updating

Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (

深層学習軽量化・量子化画像テキスト動画

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

深層学習Transformer検出3Dマルチモーダル

CAMF-Det: Closure-Aware Multimodal Fusion for LiDAR-Camera 3D Object Detection on UAV Platforms

Multimodal 3D object detection based on LiDAR and cameras has demonstrated excellent performance in ground-veh

用途: 検出
難易度: Hard
コスト: High

品質予測/異常検知自然言語処理大規模言語モデル生成画像テキスト

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging

Most existing multi-exposure HDR methods follow a fixed feed-forward reconstruction paradigm, making them pron

用途: 生成
難易度: Hard
コスト: High

コンピュータビジョンセグメンテーション異常検知テキストマルチモーダル

Scaling by Diversified Experience for Vision-Language-Action Models

Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of

用途: 異常検知
難易度: Hard
コスト: High

When Vision Misleads, Let Location Speak: A Worldwide Image Geo-Localization Method via Location Attention Mechanism and Large Multimodal Models

Worldwide image geo-localization aims to determine the capture location of an image on a global scale. Existin

深層学習Transformer検出画像テキスト

用途: 検出
難易度: Hard
コスト: High

深層学習軽量化・量子化セグメンテーションマルチモーダル

DifferSeg: Towards Diverse Multimodal Binary Segmentation via Differential Perception and Frequency Guidance

In many binary segmentation tasks, most multimodal methods rely on fixed feature concatenation for cross-modal

用途: セグメンテーション
難易度: Hard
コスト: High

ProbeAct: Probe-Guided Training-Free Failure Recovery in Vision-Language-Action Models

Vision-Language-Action (VLA) models demonstrate strong perfor-1 mance on language-conditioned robotic manipula

深層学習軽量化・量子化3Dマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation

Vision-Language-Action (VLA) models have become a powerful framework for robotic manipulation, and recent stud

深層学習軽量化・量子化マルチモーダル強化学習

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection

Vision-language-action (VLA) policies can deviate from nominal trajectories during manipulation, even when tas

自然言語処理RAG画像マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on t

自然言語処理RAG画像動画マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

品質予測/異常検知深層学習軽量化・量子化生成マルチモーダル

Intrinsic Selection and Particle Resampling for Inference-Time Scaling Beyond Domain Verifiability

Inference-Time Scaling (ITS) has largely succeeded in verifiable domains like math and coding, where cheap ver

用途: 生成
難易度: Hard
コスト: High

MI向き自然言語処理大規模言語モデル生成テキストマルチモーダル

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it h

用途: 生成
難易度: Hard
コスト: High

FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning

Action-supervised fine-tuning of vision-language-action (VLA) policies fits demonstrations effectively but con

深層学習Transformer画像マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

品質予測/異常検知深層学習軽量化・量子化マルチモーダル強化学習

Reinforcement Learning for Flow-Matching Policies with Density Transport

We present an online reinforcement learning (RL) algorithm for fine-tuning flow-matching policies in continuou

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet exist

自然言語処理ファインチューニングマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

A Resilience-as-a-Service assessment framework for coordinated disruption response in interdependent urban transit systems

Urban public transport disruptions require rapid response strategies, yet existing studies rarely provide a de

深層学習Transformerマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

品質予測/異常検知深層学習Transformer翻訳テキスト音声

HydraQE: OSU's Submission for the IWSLT 2026 Speech Translation Metrics Shared Task

We present HydraQE, our contribution to the IWSLT 2026 Speech Translation Metrics shared task. HydraQE is an e

用途: 翻訳
難易度: Hard
コスト: High

表形式向きコンピュータビジョン動画認識テキスト動画マルチモーダル

Harnessing Streaming Video in the Wild

Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications suc

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

センサ/時系列深層学習Transformer分類検出生成

TRADE: Transducer-Augmented Decoder for Speech LLM

Speech Large Language Models (Speech LLMs) lack a principled mechanism for streaming inference: their label-sy

用途: 分類
難易度: Hard
コスト: High

深層学習Transformer画像テキストマルチモーダル

When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models

Multimodal language models are typically evaluated through external behavior: selecting the correct image--tex

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Vision-Language Work Zone Intelligence for Safety-Critical Speed Regulation of Mixed-Autonomy Vehicles in Dynamic Environments

Temporary work-zone speed limits are communicated through visually inconsistent signage and are often missing

コンピュータビジョン物体検出分類検出画像

用途: 分類
難易度: Hard
コスト: High

センサ/時系列深層学習軽量化・量子化画像マルチモーダル

RGB-S: Image-Aligned Tactile Saliency for Robust Dexterous Manipulation

Effective visuo-tactile integration is critical for robotic dexterous manipulation, especially when visual obs

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping

Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reaso

自然言語処理RAG画像テキストマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback

Achieving fully automated, physically plausible 3D motion synthesis is a core objective in graphics and genera

MI向き深層学習軽量化・量子化生成テキスト3D

用途: 生成
難易度: Hard
コスト: High

BLUE: Toward Better Language Use in Efficient Vision-Language-Action Models for Autonomous Driving

We present BLUE, a minimal method for better language use in vision-language-action (VLA) models for autonomou

深層学習軽量化・量子化生成マルチモーダル

用途: 生成
難易度: Hard
コスト: High

SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creati

自然言語処理ファインチューニング分類検出生成

用途: 分類
難易度: Hard
コスト: High

Facial Expression Recognition in the Deep Learning Era: A Systematic Multi-Criteria Review of Methods, Models, Datasets, Performance, Challenges, and Future Research Directions

Facial Expression Recognition (FER) has advanced rapidly over the last decade, driven by the shift from handcr

深層学習CNN分類マルチモーダル

用途: 分類
難易度: Hard
コスト: High

Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionall

説明可能自然言語処理RAG生成画像動画

用途: 生成
難易度: Hard
コスト: High

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA

Exploratory manipulation often turns an apparent failed attempt into the key evidence for what to do next. For

深層学習軽量化・量子化分類動画マルチモーダル

用途: 分類
難易度: Hard
コスト: High

表形式向きコンピュータビジョン動画認識生成画像テキスト

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving

Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for auto

用途: 生成
難易度: Hard
コスト: High

少数データ向き深層学習Transformer画像テキストマルチモーダル

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs

Multimodal Large Language Models (MLLMs) face a significant inference bottleneck due to the quadratic computat

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

コンピュータビジョンセグメンテーション生成予測画像

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Mo

用途: 生成
難易度: Hard
コスト: High

Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation

Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due

用途: 生成
難易度: Hard
コスト: High

自然言語処理大規模言語モデル画像テキストマルチモーダル

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding

Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models.

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors

Robotic grasping is a fundamental capability in robotic manipulation. Yet grasping remains challenging under p

自然言語処理RAG3Dマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level unders

深層学習CNN検出生成セグメンテーション

用途: 検出
難易度: Hard
コスト: High

コンピュータビジョンセグメンテーション生成マルチモーダル強化学習

Guided Discovery of New Behaviors using Diffusion Policies

Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies exce

用途: 生成
難易度: Hard
コスト: High

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

World action models inherit the predictive capability of world models, enabling action generation to be guided

自然言語処理RAG生成画像マルチモーダル

用途: 生成
難易度: Hard
コスト: High

センサ/時系列コンピュータビジョン3D・点群テキスト3Dマルチモーダル

Language as a Sensor: Calibrated Spatial Belief Estimation in 3D Scenes from Natural Language

Robots deployed in human-centric environments routinely receive natural-language descriptions of spatial infor

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

自然言語処理プロンプトエンジニアリング画像3Dマルチモーダル

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world depl

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

自然言語処理ファインチューニング異常検知画像テキスト

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data

Vision-language models (VLMs) are powerful general-purpose reasoners, yet converting them into robot control p

用途: 異常検知
難易度: Hard
コスト: High

LUNA-AD: Lightweight Uncertainty-Aware Language Model with Lifelong Learning for Autonomous Driving

While large language models (LLMs) offer promising reasoning capabilities, their integration into safety-criti

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs

Compositional priors describe the generic properties of layered functions in deep Bayesian models, where deep

少数データ向きコンピュータビジョンマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding

Multimodal large language models (MLLMs) have made substantial advancements in video understanding, yet the re

自然言語処理大規模言語モデル検出生成テキスト

用途: 検出
難易度: Hard
コスト: High

少数データ向きMI向き条件最適化自然言語処理ファインチューニングテキストマルチモーダル

CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning

Enabling robots to understand and execute tasks from natural language commands while maintaining data efficien

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Aligned but Not Partner-Specific: Distinguishing How Multimodal LLM Agents Succeed in Reference Games Without Human-Like Conventions

Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, pa

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

説明可能品質予測/異常検知自然言語処理大規模言語モデル画像テキストマルチモーダル

arxivGitHubあり2026-06-06

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet the

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via Multimo

深層学習軽量化・量子化生成検索画像

用途: 生成
難易度: Hard
コスト: High

Beyond Raw Signals: Undecoded Generative Latents as Privileged Synthetic Data

While multimodal integration significantly improves computer vision models, deploying them incurs prohibitive

深層学習軽量化・量子化分類生成画像

用途: 分類
難易度: Hard
コスト: High

TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation

Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet thes

用途: 生成
難易度: Hard
コスト: High

コンピュータビジョンセグメンテーション生成マルチモーダル

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning

Test-time Scaling (TTS) has emerged as a pivotal research direction for enhancing model performance by dynamic

用途: 生成
難易度: Hard
コスト: High

MI向きコンピュータビジョンマルチモーダル画像テキスト動画

IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval

Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

センサ/時系列コンピュータビジョンセグメンテーション分類画像テキスト

One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling

Vision-language models (VLMs) enable visual recognition from semantic class descriptions, which makes them att

用途: 分類
難易度: Hard
コスト: High

MI向き品質予測/異常検知自然言語処理大規模言語モデル生成画像テキスト

arxivGitHubあり2026-06-06

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but

用途: 生成
難易度: Hard
コスト: High

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to

自然言語処理RAG生成画像テキスト

用途: 生成
難易度: Hard
コスト: High

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

Humanoid foundation models are advancing faster than we can evaluate them. While real-world testing is expensi

深層学習軽量化・量子化生成マルチモーダル

用途: 生成
難易度: Hard
コスト: High

Learning from Human Driving: A Human-in-the-Loop Online Behavior Cloning Framework for Autonomous Driving

With the evolution of large foundation models (LFMs), data-driven autonomous driving has made significant stri

深層学習軽量化・量子化マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation

自然言語処理大規模言語モデル動画マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

少数データ向き深層学習軽量化・量子化生成マルチモーダル強化学習

Q-VGM: Q-Guided Value-Gradient Matching for Flow-Matching VLA Policies

We propose Q-Guided Value-Gradient Matching (Q-VGM), an off-policy reinforcement learning (RL) method that tac

用途: 生成
難易度: Hard
コスト: High

Combinatorial Landscape Analysis for Dominating Set and Vertex Coloring

We analyze the two combinatorial problems of Dominating Set and Vertex Coloring regarding what kind of local o

コンピュータビジョンマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

コンピュータビジョンセグメンテーション生成回帰テキスト

TBD-VLA: Temporal Block Diffusion Vision Language Action Model

Discrete Vision-Language-Action (VLA) models typically formulate action generation as next-token prediction ov

用途: 生成
難易度: Hard
コスト: High

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex mul

コンピュータビジョンマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Spline Policy: A Structured Representation for Robot Policies

この論文では、ロボット制御の新しい表現方法であるSpline Policy（SP）を提案した。SPは、行動を spline で表現することで、行動をより詳細かつ柔軟に表現することができた。

深層学習Transformerマルチモーダル

用途: ロボット制御の新しい表現方法
難易度: Hard
コスト: High

arxivGitHubあり2026-06-05

RhinoVLA Technical Report

この論文では、VLAモデルをedgeハードウェアにデプロイするための手法を提案しています。この手法は、VLAモデルをedgeハードウェアにデプロイするためのフレームワークです。この手法は、edgeハードウェアを利用してV

深層学習軽量化・量子化画像テキストマルチモーダル

用途: VLAモデルをedgeハードウェアにデプロイするための手法
難易度: Hard
コスト: High

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation

この研究では、自然言語指示を実行するためにもっと実際的なエンベロイメントにおいて、視覚言語航行 (VLN) の問題に対処します。従来の 3 つのステージのアプローチは、目的地に到達するのを困難な場所や、計画と制御間の矛盾

コンピュータビジョンマルチモーダル生成

用途: 自動車のトラクタシー
難易度: Hard
コスト: High

自然言語処理ファインチューニング動画マルチモーダル

Robotic Policy Adaptation via Weight-Space Meta-Learning

Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling ge

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Coarse-to-Control: Action-Token Planning for Vision-Language-Action Models

この論文では、視覚言語行動モデルの改良を実現した。Coarse-to-Controlは、行動に必要な計画の空間を大幅に縮小し、行動の計画を実現するための新しいフレーム

コンピュータビジョンマルチモーダル生成

用途: 視覚言語行動モデルの改良
難易度: Hard
コスト: High

品質予測/異常検知自然言語処理RAG画像動画マルチモーダル

LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Visual-language action (VLA) models enable robots to predict actions directly from observations and language i

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning

World Action Models (WAMs) offer a promising approach to embodied intelligence, yet existing methods rely heav

深層学習軽量化・量子化画像テキスト動画

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

品質予測/異常検知コンピュータビジョンマルチモーダル

LIMMT: Less is More for Motion Tracking

We argue that high-quality motion data can steer tracking policies toward better optimization trajectories ear

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

arxivGitHubあり2026-06-05

ActionMap: Robot Policy Learning via Voxel Action Heatmap

この論文では、ロボットの制御を学習するための、新しいモデルの提案であるactionmapを提示しました。

深層学習軽量化・量子化回帰マルチモーダル

用途: ロボット制御の学習
難易度: Hard
コスト: High

Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation

VLNベンチマークでは、ディシクリットな操作や粗い操作が使われ、UAVのヴィジョンラングジュアクション（VLJ）タスクでは短い操作が中心で、長時間飛行に対応できるfineグラINEDUAVナビゲーション（FLIGHT）ベ

コンピュータビジョンマルチモーダルテキスト動画

用途: ドローンの長時間飛行
難易度: Hard
コスト: High

説明可能コンピュータビジョンセグメンテーション生成埋め込みマルチモーダル

Discrete Causal Representations from Heterogeneous Domains: A Bayesian Approach with Social Survey Applications

この研究では、複数のドメインの複雑なデータを分析するために、Bayesian モデルを使用して因果関係を分析するツールを開発します。主に社会調査に使用できるツールです。

用途: 複数のドメインの因果関係を分析するツールを開発
難易度: Hard
コスト: High

TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

この研究では、ロボット操作のスピードの可変性を扱いました。この研究で提案したTempoVLAは、スピードの変化を可能にする強化学習モデルです。

コンピュータビジョンマルチモーダル強化学習

用途: スピード可変的バージョン言語行動ポリシー
難易度: Hard
コスト: High

品質予測/異常検知自然言語処理RAGセグメンテーションテキスト動画

VOLT: Vision and Language Trajectory Segmentation for Faster-than-Demonstration Policies

この研究では、フェスタースター自動運

用途: フェスタースター自動運転用の高速動作
難易度: Hard
コスト: High

品質予測/異常検知自然言語処理プロンプトエンジニアリングテキストマルチモーダル

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

Vision-Language-Action(バブルラボ、VLアクション)ポリシーが長時間予測と高い不確実性の制御で脆弱であることを認識し、VLアクションポリシーが1パスでのアクションデコードのみを提供し、長時間予測のた

用途: long-horizonおよびhigh-uncertainty ControlでのVLAポリシーが脆弱である問題に対する解決策。
難易度: Hard
コスト: High

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

このリポジトリでは、画像認識モデルにアクション生成能力を付与することを目指したモデルを提案します。このモデルは、画像認識のための事前訓練モデルを用いて、複雑なアクションを生成することができます。

深層学習Transformer検出生成予測

用途: 画像認識とアクションの生成
難易度: Hard
コスト: High

arxivGitHubあり2026-06-04

A Conversational Framework for Human-Robot Collaborative Manipulation with Distributed Generative AI models

この研究では、人間-ロボット協力のためのDistributed Conversational Frameworkを提案します。

自然言語処理大規模言語モデル生成画像テキスト

用途: 人間-ロボット協力
難易度: Hard
コスト: High

深層学習Transformerマルチモーダル強化学習

L-SDPPO: Policy Optimization of Spiking Diffusion Policy for Intra-vehicular Robotic Manipulation

この研究では、L-SDPPO という方法を提案します。これは、連携型ロボット Manipulation に向けたディフュージョンポリシーの最適化を実現するものです。

用途: 連携型ロボットManipulation
難易度: Hard
コスト: High

Robots Need More than VLA and World Models

Generalist robot intelligence is often framed as a policy-scaling problem: collect more robot demonstrations,

コンピュータビジョン3D・点群生成動画3D

用途: 生成
難易度: Hard
コスト: High

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

統合された視覚言語アクションモデルを提案し、これを用いたタスクの性能を向上させることができるようになる。

用途: 統合された視覚言語アクションモデル
難易度: Hard
コスト: High

T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation

Open-vocabulary 3D functionality segmentation enables robots to localize functional object components in 3D sc

自然言語処理RAG分類セグメンテーション画像

用途: 分類
難易度: Hard
コスト: High

コンピュータビジョンマルチモーダル異常検知テキスト動画

Towards a Data Flywheel for Embodied Intelligence in Logistics

Autonomous drivingでは、ロボットが視覚認識した情報に基づいて行動を決定する必要があるが、過去のデータで構築された空間モデルでは、ロボットの行動を予測することが困難であるため、空間モデルを構築することによ

用途: ロボットの行動予測に適した空間を構築
難易度: Hard
コスト: High

arxivPaper only2026-06-03

Worker Utility as Hysteresis: A Preisach Model of Transaction Acceptance in Gig Labour Markets

この研究では、個人の意思決定に対する効率的な解析 (Worker Utility) を提案しており、個人の意思決定を効率的に解析し、それを活用する。

表形式向きCPUで試しやすいコンピュータビジョンマルチモーダル分類

用途: 個人の意思決定に対する効率的な解析
難易度: Hard
コスト: High

arxivPaper only2026-06-02

A Quantitative Approximation Framework for Flow Distillation in Diffusion Models

We develop a quantitative approximation framework for diffusion distillation, viewing few-step sampling as err

深層学習軽量化・量子化マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

arxivPaper only2026-06-02

Multimodal Transformer Based Generic Mixture Density Network for Scattering Timescale Estimation of Fast Radio Bursts

The discovery rate of fast radio bursts (FRBs) continues to increase with the advent of new radio facilities a

センサ/時系列深層学習Transformerマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

arxivPaper only2026-06-02

A Fast Screening Approach for High-dimensional Outcomes and High-dimensional Predictors

Modeling interactions among multimodal, high-dimensional data is intrinsically challenging due to ultra-high d

説明可能品質予測/異常検知深層学習軽量化・量子化マルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

arxivPaper only2026-06-01

Flow-Transformed Implicit Processes for Function-Space Variational Inference

Implicit-process priors define distributions over functions through flexible generative mechanisms, making the

深層学習軽量化・量子化生成マルチモーダル

用途: 生成
難易度: Hard
コスト: High

arxivPaper only2026-05-27

Evolving to the Aesthetics of a Vision-Language Model

Evolutionary systems have demonstrated remarkable results in creative domains, with recent applications in gen

コンピュータビジョンマルチモーダル生成テキスト

用途: 生成
難易度: Hard
コスト: High

arxivPaper only2026-05-22

Planktonzilla: Multimodal dataset and models for understanding plankton ecosystems

Marine plankton underpin aquatic food webs and play a key role in global CO2 sequestration, making reliable sp

少数データ向き深層学習Transformer分類画像テキスト

用途: 分類
難易度: Hard
コスト: High

arxivPaper only2026-05-19

Smooth Partial Lotteries for Stable Randomized Selection

部門間の競争では、評価に基づいて候補者を選択する必要があることが多い。しかし、これまでのランダムな選択メカニズムは、候補の中で微妙な差異のあるデータの不均衡を考慮していなかった。これにより、安定性が低くなる。そのため、今

品質予測/異常検知コンピュータビジョンマルチモーダル

用途: スマートなランダムな選択を促す方法を実現する
難易度: Hard
コスト: High

arxivPaper only2026-05-19

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification

Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorre

自然言語処理大規模言語モデルテキストマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High

arxivPaper only2026-05-18

Mapping the Fitness Landscape: A Structure-Guided Approach to Multi-Modal Optimization

Multimodal optimization requires finding many optima rather than merely keeping a diverse population. Yet most

品質予測/異常検知自然言語処理RAGマルチモーダル

用途: 技術検証・論文読解補助
難易度: Hard
コスト: High