3D-Aware VLMs with Implicit and Explicit Geometries
3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
- 用途
- 3次元空間理解技術の開発
- 難易度
- Hard
- コスト
- High
「text」の検索結果
475 件3次元空間理解技術のための新しいアプローチであるVLM-IE3D(Vision-Language Models with Implicit and Explicit 3D geometry)を提案しました。VLM-IE3
時系列データ分析技術のための新しいアプローチであるTimePNS(Time Series Explanation with Counterfactual Necessity)を提案しました。TimePNSは、時系列データ
多モーダル理解技術のための新しいアプローチであるMIRROR(Learning from the Other View)を提案しました。MIRRORは、テキスト、図、テキストと図の組み合わせから同等の視点を提供することで
大規模な言語モデルを用いた推論技術のための新しいアプローチであるX$^3$-OPD(Distilling Reasoning into Large Audio-Language Models via On-Policy
Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a targe
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for ena
Do independently trained language models come to represent the same thing in the same way? We answer for code,
An initial high-recall stage in an empirical pipeline decides which items pass to later review, labelling, or
Scaling inference-time computation has emerged as a reliable method to improve the performance of large langua
都市の電気自動車充電インフラは、可及的速やかに故障を予測・修理することで、耐久性と低炭素化を向上させる必要がある。機械学習を用い、故障を予測するモデルの開発を研究した。
ディスクリートフロー・マッチングにおけるコンテキストの正しい有用性の利用を検討した。この研究では、ディスクリートフロー・マッチングのモデルの正確さを高めるためにコンテキストの有用性を適切に利用する方法を提案した。
ディスクリート確率模型におけるベイジアン解釈の分析を進めた。この研究では、正確さを高めるために、負のスコア比率を制限したディスクリート確率模型を提案した。
アライメントした言語モデルの偏った表現の理解を進めた。この研究では、アライメントした言語モデルの表現を分析して、偏った表現を理解することができ、これを用いて、偏った表現を正すことができると主張した。
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason throug
Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architec
Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as
The modeling of hydrometeorological time series with limited observations is a key challenge in the monitoring
Inductive Logic Programming (ILP) originated within the Logic Programming community in the Nineties as a frame
この研究では、機械学習モデルを使用して血糖値の変化を予測し、糖尿病管理のためには血糖値データの前処理が重要であることの重要性を強調しています。
この研究では、リスクグラフニューリアルネットワーク(RGNN)を使用して、人口統計学的特性と地域の情報と組み合わせた熱死亡リスクを推定する新しい方法を提案し、DLNMの効果的な方法に代わる可能性があります。
この研究では、自己説明の信頼性を検証するためのRL方法を提案し、自己説明の信頼性を直接最適化するための新しいアプローチを検討します。
この研究では、Bスプライン回帰のために、ナレッジの選択を自動化するための新しい方法であるAutomatic Knot Selectionを提案し、ナレッジの選択とデータ分析を容易にします。
この研究では、長期
ラテン言語モデルを使用すると、言語モデルの内部の計算結果を分析できる。計算結果は、連続ベクトル空間として実行される中間計算であり、これを分析すると、モデルがどのように結果を得ているかを明らかにできる。
Advanced Persistent Threats (APTs) remain difficult to detect because only a small fraction of events in large
Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal poli
GraphVidは、グラフと文本から生成することができ、オブジェクトの複数の移動を正確に制御することができる。グラフではオブジェクトの動きを表す情報を保存し、文から生成の制約を指定することができる。
この研究では、言語モデルが社会的に正しい判断を下すことができる方法について調べた。結果は、模式が対立を保つ能力が高くなり、他の人の視点を受け入れやすくなったことである。
この研究は、リソースフローの動作を表すPetriネットと、APIを操作するためのテストを自動生成する方法を提案した。方法は、APIの機能をテストするためのシナリオを生成し、テストが正しく実行されるようにした。
ElasticTTTは、プログラムがテストのときに動作を調整できるようにした。方法は、テストのときにモデルが前のサンプルの情報と現在の情報を組み合わせて、ビデオを編集する際に正しく動作するようにした。
GS-Agentは、自然言語から生成することができ、物理的に正しく動作する4次元の世界を生成することができる。方法は、物理的正しさを保つために、生成時に物理的推論を使用した。
この研究では、LMOの安全性を調べた。結果は、直面する危険目標に対してモデルが安全なアドバイスを出すことができた。
この研究は、奇数サイクルのシャノン容量の最小限度を検討した。結果は、グラフの独立集合の大きさに基づいて最小限度を計算することができた。
Agentic Context Managementは、エージェントのメモリとコストを管理できるようにした。方法は、エージェントが自己管理できるように、トレーナーが制御できるようにした。
Artificial Epanorthosisは、大規模言語モデルが古典的なルレチックの表現を使用する傾向に注目した。結果は、モデルのトレーニングデータの形状がこの傾向に影響していることができた。
Large language models (LLMs) and agents are now widely used tools in code development, with data typically sen
People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answ
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks rema
Video face swapping has no natural paired supervision: no real footage exists of one person's face performing
In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninf
In automated planning, logical regression is an operation that returns the most general condition necessary fo
Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for
Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based an
A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard t
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However,
Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This l
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack special
Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, tec
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limit
Retrieval-Augmented Generation (RAG) systems increasingly employ multiple LLM agents. Yet, most prior work opt
Replacing an object with one that differs in category or shape requires complete source removal, natural targe
The reliance on unstructured free text for documenting clinical trial protocols creates a significant barrier
LLMは、人間のアイデンティティのシミュレーションを使用して個人データを削除したり、未均衡なデータを削除したりしますが、これらのアプローチには制限があります。
Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just
Event-B is a formal method rooted in predicate logic and set theory. We encoded over 600 proof rules in Prolog
LLM (Large Language Model)とLP (Logic Programming)を組み合わせて、有理数である√2の非有理性を証明します。この証明には、LLMが主観的な論理式を生成し、LPが証明を行うプロ
S2S (Speech-to-Speech) LLMアシスタントを利用して、人間のような話し方をすることができますが、安全対策の実装が困難です。この研究では、S2S LLMアシスタントの安全対策を2つのアプローチで実現し
知識重視の質問応答システム (KI-VQA) を分析するために、新しい評価基準を提案します。これらの基準では、VLMの各タスクを個別に評価することができます。
ビデオLMMの安全性を確認するために、新しい診断フレームワークを提案します。これらのフレームワークは、モデルの挙動、理解、セマンティクスを同時に考慮します。
Semantic-ID-based generative recommendation represents items as sequences of shared semantic tokens, enabling
Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequenti
Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid anal
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognitio
Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summariz
Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of sto
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rar
_guardianAgentBenchBenchmarkは、580のシナリオを6つのドメインで評価し、3つの実稼動フレームワークであるLangChain、LlamaIndex、Vectaraを利用します。このベンチマーク
効率的な多モードの推論は、モデルの性能やFLOPCOuntだけでなく、移動、キャッシュ、変形、量化された表現を保存するコストやメモリ、エネルギーに関する制約にも制限されています。この論文では、最近のビジュアルトークン圧縮
Coding agents ship with one kind of memory: documents. Instruction files, plan artifacts, and auto-written mem
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources.
AI knowledge systems require representations of entity importance for retrieval, recommendation, evidence sele
Large language models (LLMs) have rapidly and significantly entered scientific workflows, but it remains uncle
Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities toget
Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not e
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, ev
LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records,
The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning
Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are
ある単位の人間の認知的難易度は、その単位のスルーパイズアル (Surprisal) と特定の言語モデルによれば線形関数に等しいという理論が存在します。しかし、この理論は実態を反映していないという批判があります。この論文で
Large Language Models (LLMs) は医学教育に大きな可能性を持っていますが、現在のシステムでは、質問に答えるか一時的なフィードバックしか行なわれていません。一方、臨床病例を決定センターへの学習トレ
この論文では、DONDO と呼ばれるアフリカ諸国向けの音声認識ベースモデル (ASR)が構築されました。これらのモデルは、自律学習型スピーチエンコーダーであるw2v-BERT 2.0を使用して構築されています。このエンコ
この論文では、大規模言語モデル (LLMs) が日常的な文化的知識を評価する能力に着目しています。ここで、TriviaRoomQA というクイズスタイルで問題を提示して、LLMs が日常的な文化的知識をどのように評価する
この論文では、音声字幕の評価手法が提案され、音声字幕の評価において既存の手法の制約を克服することを目指しました。提案されたフレームワークは音声字幕の各側面を評価し、質問回答型の評価手法ではなく字幕の中立性を評価することが
この論文では、対称的な位置エンコードにモビアスの対称性を適用しました。これにより、ローテーションの平面での各位置間のホロノミーが -1 となり、シーケンスの両端が決定的に結合されます。この手法により、精度が高額になること
この論文では、会話マンダリンにおける単語の意味と子音の特性の関係を調べました。その結果、単語の
In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft,
Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsist
Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how t
Token cramming compresses sequences into learned embeddings with near-perfect reconstruction, but fixed token
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on ed
Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by levera
Almost every large language model that reaches a broad audience is quantized: trained in full precision, then
Real-world agent learning is often constrained by costly environment interactions, such as running time-consum
Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs o
Distinguishing animate from inanimate concepts in written language requires more than shallow text processing,
Grievance is one of the warning signs analysts look for when assessing threats of violence. It is increasingly
化学物質の構造を予測する言語モデルが信頼性の低い情報を生成する傾向があることを指摘し、原因と解決策について検討している。
ソフトウェア開発の維持フェーズで、ソースコードの自然言語解説を生成するためのモデルの改善を目的とした研究。
コーディングエージェントの評価基準を導入し、現実世界のコミットやプルリクエストに基づくタスクを構築した。
長形推論のための言語モデルが、提供されたコンテキストから乖離した論理を生成する可能性があることを指摘し、コンテキストと推論論理をより適切に融合するため、 REFACT (REstating Facts in Adapti
多エージェントのシミュレーションにおいて、共有世界状態がエージェント間で保持され、その世界状態が観測結果に反映されると仮定している。
ディフュージョンモデルにおける初期的なNoise Seed の影響が、モデルが生成する高質のイメージに大きく影響していることを提示し、Seed Search 時の時間的負荷を削減するための方法を提案した。
分割推論の一般化を促進するためのフレームワークを提案し、instruction factor bias を定式化し、bias を減らすための fine-tuned ポリシーの適切化方法を提示した。
アイス認識の精度を向上させるための方法を提案し、視覚認識におけるアイス認識タスクの課題を分析した。
Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore t
Underwater image enhancement remains challenging due to wavelength-dependent light absorption, scattering, and
ステンレス鋼の表面欠陥検出は、工業的な品質検査において重要であるが、従来の方法は、ひずみ、非対称性や欠陥の境界の不規則性に伴う標準的な共通化の非対称受容領域と剛性のサンプリンググリッドを利用しているため、ひずみ、非対称性
Federated プールトーニングは、視覚言語モデルを軽量なプールで共有することで、視覚的、言語的、音声的なモデルを組み込んだ視覚言語モデル(VLM)の共有の協力的適応を実現する。従来のプールを使用する方法では、個々の
Multi-object tracking in dense crowds requires solving a bipartite assignment problem between detections and t
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with exp
AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure an
Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from
Digital sewing patterns typically consist of disjoint 2D panels without explicit stitch annotations, making do
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at in
Image restoration agents have recently emerged as a flexible paradigm for handling diverse and unpredictable d
大規模言語モデルはさまざまな画像をテキストに変換する上で優れた性能を示しているが、発生するホログラフィックな診断にはまだ解決策が必要です。この研究では、主流の粗い検出方法の欠点を補うため、細部の診断方法を提案しています。
3D点群のセグメンテーションではクラス不均衡が発生し、有効な解決策が必要です。この研究では、11 つの不均衡対策を 2D のコンピュータビジョンとは異なる 3D の上で評価し、標準的な交差エントロピーと均衡の重み付けが競
大規模言語モデルのハロウィーン診断では、対象の 3D 空間関係を推論する際に、視覚化が欠如していることが問題となります。この研究では、これらのハロウィーンを軽減するためのアプローチを提案しています。
大規模言語モデルの圧縮には、モデルのパフォーマンスが低下する可能性があるため、量化の保護が重要です。この研究では、Fisher加重チャネル感受性を用い、MLLMの量化を安定させるためのC-PTQをプロPOSEしています。
空間理解は、物理世界と静的のセマンティック理解の間でつながるために不可欠です。多くの空間タスクは、場所、領域、パスの自然な表現は、ポインティングやマーキングなど、連続的な視覚的シーンで行われることが多いが、現行の空間推論
パスロジは、現在、パスロジ認識のための画像言語モデルを評価するために広く使用されていますが、この研究では、パスロジ認識において画像言語モデルの視覚知覚が機能していることを疑問に問っています。
感情認識は、現代のアギを促進するために不可欠ですが、大規模
Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query g
Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing thei
Current identity customized video generation methodologies are predominantly limited to single-identity scenar
Improving video captioning quality typically demands retraining large vision-language models, an expensive and
本論文では、テキストと動画を対応させるDistribution-Alignment Bridge(DAB)を提案します。DABは、テキストと動画のエンティティを確率分布として表現し、両者の間の分布の差異を解決します。この
この研究では、マイメイク移植を改善するために、マイメイクの強い地域性を考慮したRegion-Controllable Diffusion Transformer(MagicMakeup)を提案します。
この研究では、都市ウォークビデオを分析するために、4つのモダリティの表現(スペース時領域情報、時間平均画像、オーディオ符号化、テキストベースの表現)を使用しました。
この論文では、DINO-VPTという手法を提案します。DINO-VPTは、Hierarchical Visual Prompt Tuning(HVPT)を使用して、物理的なスポーフィングとデジタルスポーフィングを検出しま
この論文では、Lumeraという手法を提案します。Lumeraは、Engine-Native 3D World ReconstructionとLightsを検出するために使用します。
この研究では、WhereEditという手法を提案します。WhereEditは、Mask-aware Local Latent Editingを使用して、一ステップの画像編集を実行します。
この論文では、ViSTR-Benchという手法を提案します。ViSTR-Benchは、MLLMが動的シーンから情報を取得できるかどうかを評価します。
Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls
Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve
強制制約に基づく強化学習を利用し、低コストで高精度の組み立てが可能になると同時に、組み立てに失敗してもロボットが安全に回避できるように、ロボットの制御のための強化学習を提案します。
オブジェクト目標のナビゲーションにおける、動的な避け方とマルチフロア環境を考慮した、ゼロショットオブジェクトナビゲーションのフレームワークを提案します。このフレームワークでは、動的な人々とマルチフロア環境を考慮しながら、
オートメーションされたマニピュレーションを目的とした、大規模なテーブルトップのデータセットであるTableVerse を提案します。このデータセットには、物理的に可能な実世界のレイアウトを生成する実用的な方法が含まれてお
Learning-based manipulation policies usually predict robot actions from sensory observations and leave their e
Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying qu
This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment inform
Personal and organizational planning systems maintain two records that drift apart: what was planned (a task's
Language models are thought to exhibit the phenomenon of superposition, representing many more features than d
We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with profess
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature
この研究では、分布に変化がない場合の時間軸への適応性を
Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption th
Transformers の効率化を目的とした新しいアプローチ、ELSA (Efficient Low-Rank and Sparse Attention Approximation) を提案する。
パラメータ効率の確保を目的とした Low-Rank Adaptation (LoRA) のランクの確立を扱う研究を紹介する。
ニューラル ネットワークの解釈可能性を向上させるために Quadrilateral Loss を提案する。
OLED 材料の開発を目指す新しいアプローチ、causal language models を用いて optoelectronic プロパティを予測するフレームワークを提案する。
生分子注文を扱う研究、Plausibility-Driven Prioritization を用いて生分子注文を提案する。
Modern power systems increasingly require probabilistic forecasts amid interacting uncertainties from renewabl
データクリーンシングを扱う研究、CURED を用いてデータクリーンシングを提案する。
Sample-based Quantum Diagonalization (SQD), an extension of Quantum Selected Configuration Interaction (QSCI),
ネガティブ メトリクスの因数分解を扱う研究、非負行列因数分解 (NMF) を用いて因
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving larg
この研究では、AIベースの採用システムを使用して、過去の履歴データに基づいて機械学習モデルをトレーニングすることで、社会的偏見を永続させたり強化したりするリスクを評価および緩和することを目的とした。
この研究では、抗原特異性抗体を設計するために、抗原および抗体の間でエピトープレベルでのペアリングが必要であることを考慮した、抗原特異性の抗体多モーダルファンデーションモデル(AAMFM)を提案しました。
この研究では、Consumer Wearablesに基づく Short-term Heart Rate Variability(HRV)予測を目的とした、Time Series Foundation Modelsの評価を
この研究では、学習前の時系列ベースの学習模型を、トレーニング後の適応を使用して、目的のタスクに適応させる方法を提案しました。
この研究では、人工知能の研究者と神経科学者の間の分野を結びつけるために、脳のシステム構造を研究し、その研究から導かれた新しいアプローチを提案しました。
この研究では、マルウェア検出に使用されるデープラーニングモデルにおける、位置依存性KVキャッシュ(Key-value Cache)を改善する方法を提案しました。
この研究では、実世界ロボットのシーケンシャル予測に使用できる、diffusion-based frameworkを提案しました。
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has bec
We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processe
Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-a
分子構造を決定するために、スペクトルデータから自動的な構造解析を実施するための方法を提案している。この方法は、スペクトルデータに基づいてヒントと改良を繰り返すことで、分子構造を決定するもので、分子の可能性の広範な構造スペ
大規模言語モデルを制御するために活性化制御を使用するときに生じる可能性のある外部性、つまり安全性を低下させる可能性と、禁止された要素を誘発する可能性を軽減するために、制御ベクトルの洗浄を実施する提案されている。
生産的言語モデルの利用による金銭的感情分析に対処するための方法を提案している。複数のエージェントを活用したコミティー方式を使用し、さまざまな粒度のテキストデータに対応できるように、単語レベルのルールベースアプローチ、句節
language モデルを前訓練するために、Muon などの矩式最適化を使用するが、これらのモデルの内部幾何学を保持する方法についてはよくわかっていない。仮定から、モデルの内部幾何学を安定化するために、SGDに内在するb
VLSIのグローバルルーティングは、信号ネットワークを 3D グリッド上で割り当てることが目的であり、信号遅れやワ
In HEP data analyses, finding an adequate function to model binned data has largely relied on a manual process
この研究では、多マスク分散言語モデルを提案します。このモデルには、複数のマスクがあり、それぞれが異なる生成タスクを実行することになります。このモデルは、生成タスクの多様性を高めることができ、生成された文がより多様性の高い
この研究では、核量子効果をシミュレートするために、画像時刻パス積分を利用した分散機械学習アルゴリズムを提案します。このアルゴリズムは、分散機械学習を利用して核量子効果をシミュレートすることに成功し、核量子効果に関連する問
Do quantum kernels improve cross-sectional stock return prediction? We run a controlled horse race on the Chin
High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances i
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across
Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys t
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain,
Large language models increasingly use search tools to retrieve up-to-date information, introducing a new atta
A record system declares when two records refer to the same entity, occurrence, scope, or rule. Its disclosed
Harnessing the potential of electroencephalography (EEG) for brain research is fundamentally limited by intrin
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow grap
TextGrad improves language-model systems by revising text from feedback. Its core thesis is that natural-langu
Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their
Traditional query processing engines require continuous development and extensions to support new techniques a
Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script langua
Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as l
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge fo
最新言語モデル(LLM)が危険な生成を防ぐための確信的な安全な限界を計算するための新しいフレームワークを提案した。Clopper-Pearsonの信頼区間の新しい応用として、PAC(可能性が最も近い)の境界を得るためのア
モデルの脆弱性を解決するために、四つのエージェントに分割される多様なフレームワークPoTREを導入した。モデルの推論能力を強化し、単一のストリーミングアプローチよりも複雑な理論的制約とアブストラクションに抵抗できるように
ある知識関係がマスクスタイルのパラミトリックパラメータ化方法で適切かどうかを計算したメトリックとしてMaskability Index (MI)を導入した。DepthRankの違いを用いて、与えられたパラメータ化方法で知
侵攻テストツールが異なっている点、決定主義的な性質、狭く特定されたスコープ、専門技術の操作を用いたものと異なり、LLM駆動の自治的セキュリティツールは3つの次元で不確実性を示した。政策決定への説明が困難、影響の開放性、行
3つのタスクをサポートする音曲生成フレームワークを提示した。これらのタスクには、歌詞、テキストの説明、音楽的特性を利用して、歌詞の生成、バンドの音楽の生成、カバー曲の生成などが含まれる。
文化的意味の表現が表現された翻訳には、翻訳システムが表現する意味を理解するために、表現の文化的背景を考慮する必要があることを指摘した。文化的背景が表現されている表現された翻訳には、いくつかの課題があり、LLMベースの翻訳
分析報告の迅速な解釈が求められるときに行われるマルウェア分析を実現するために、閉じた重みの大きい言語モデルを使用しないことが多い。オープン重みの言語モデルは、マルウェア分析のために適切な言語能力と、閉じた重みの大きい言語
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggl
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challe
Quantized small autoregressive reasoning models can enter long, repetitive, or unproductive trajectories, yet
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, indep
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with mislea
The function of many genes is still unknown, and conventional driver-discovery methods, which rely on how freq
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge fo
Large language models can answer scientific questions, yet a correct output does not reveal whether the model
Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspe
ITオペレーションにおける安全なリメディエーションを確保するためのリスクを制約した介入決定問題として定式化し、安全なリメディエーションを確実に行うためのConstrained Markov Decision Proces
We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive
Large language models trained and aligned within different linguistic and regional ecosystems may frame the sa
Small language models and coding agents increasingly generate web front-end code, yet their outputs are typica
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. T
Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply the
異なる順番で画像と質問が提示される場合、視覚言語モデルはモデルのパフォーマンスに大きな影響を受けることが発見された。
危険な問題に対する正しい答えを提供する大きな言語モデルと費用の効率が良い、小さな言語モデルを協力させる技術が開発されました。
LLMは、状況から価値観を判断できるかどうか、という研究が調査されました。LLMは、状況に応じて真の価値観を推測することができました。
同じカテゴリを組み合わせることしかできないという考えに対抗して、異なるカテゴリを組み合わせることができるかどうかについて、言語モデルが調査されました。
Persona simulation involves utilizing large language models (LLMs) to anticipate human choices or interactions
大きな言語モデルは真実の情報を提供できるように見えますが、実際は虚偽情報を提供することが多く、これを検知、検出、および検証するための基準を作成するため、HalluTruthQAが開発されました。
大きな言語モデルは、さまざまな価値観に対して異なる反応を示すことがあり、これらの反応がどのように影響するかを調べました。
大きな言語モデルは、ユーザーの信念と事実的な正しさを合わせる傾向があるが、これらの傾向は多様であることを明らかにしました。
AIは、特に当代芸術作品のパスティーシュを作成する能力が高いが、これらの作品はどれだけ実際の作品と似ているかを調べました。
大きな言語モデルのエージェントは、第三者のスキルによる実際的な危険を認識し回避する能力を評価します。
大きな言語モデルの答えは、質問や入力の形態に応じて異なる傾向があることを認識しました。
職業コード付けは、職業タイトルから職業分類を識別することであり、二つのステップで実行される二つのアプローチのうちのどちらかが最も効果的であることを示しました。
大きな言語モデルは自発報告を提供するが、これらの報告は人間のインタビュー問診や不明確な提示に基づいており、このモデルの自発報告能力とその心理的意味を理解することが求められました。
文語の簡素化は、言語学習者の理解を促進するための有効な手法ですが、現在実際に有効であるかどうかが確立されていません。ルーマニア語で文語の簡素化に関する基準とリソースが作成されました。
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, s
This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as par
小さな言語モデルに対するドロップインコーパス、tiny_schillerを導入し、単一ファイルで利用できるようにし、言語モデルを簡単にprototyping、fine-tuning、教育、研究に利用できるようにする。
Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lag
FinMMEval 2026 タスク 2 は、英語で提出された短答式の金融問題を解決することを目的としています。英語以外の言語による証拠も使用されます。
FinMMEval 2026 タスク 1 は、英語、中国語、アラビア語、ヒンディー語で行われる多言語的な金融質問に答えるものを評価します。
There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, c
この研究ではSentence Splitterシステムを提案し、自然言語処理の精度を高めることができました。このシステムは、自然言語を句点で分割することができます。
With the wide application of large language models (LLMs) in real-world scenarios, the value implication of th
Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic fac
3D オキュピエンシー予測には、物体の配置と密度を解釈するための視覚的手法が必要です。従来の方法では、計算コストが高くなりすぎていたが、新しく提案されたGaussianSeedアルゴリズムは、層を階層化することで、計算コ
人名と場所の関係を抽出するタスクは、歴史的ニュース記事の解釈において重要です。従来の方法では、言語モデルの前処理が必要でしたが、Lightweightアルゴリズムは、依存グラフと近接特性を使って、歴史的ニュース記事から人
この研究では、AI生成物の論理的評価に必要なものとして、生成物がどうやって結果を得るのかを明らかにすることの重要性を強調しています。この研究では生成物を分解し、その論理的な構造を理解するために自然言語推論を利用し、生成物
Recent advances in 3D scene editing have leveraged iterative diffusion models to update input views. However,
Open-world video anomaly detection (OWVAD) is expected to detect events that match a user-specified definition
長時間ビデオエクストラポレーションには、高度な視覚的知能が必要です。Self Gradient Forcingアルゴリズムは、学生モデルを教師モデルから生成される歴史の下で学習させることで、長時間ビデオエクストラポレーシ
多モーダルラージランゲージモデルは、視覚言語タスクに強いですが、高い推論コストで問題となっています。Look Less, Think Fasterアルゴリズムは、単位次元を個別に最適化することで、多モーダルラージランゲー
複数ターンのファッション画像検索は、実世界のファッション検索では重要なタスクです。Diverse-Intent Multi-Turn Fashion Image Retrievalアルゴリズムは、異なる検索用途を扱うこと
画像理解のための多モーダルラージランゲージモデルは、強力ですが、まだ能力と限界については明確な理解が不足しています。この論文では、多モーダルラージランゲージモデルが画像理解においてどの程度の能力と限界を持つか、を分析し、
Housing-level urban physical examination is essential for identifying residential building problems and suppor
Remote sensing image editing aims to modify remote sensing images according to natural language instructions w
Attention Mechanism (AM) selectively focuses on essential information for imaging tasks and captures relations
Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure,
Vision-centric 3D occupancy prediction provides dense scene representations essential for autonomous driving a
ステレオマッチングは3次元再構成において重要なタスクです。この研究では、ステレオマッチングを確率的生成タスクと組み合わせ、オブジェクト検出の向上を目的として、ステレオマッチングフレームワークと潜在分配を統合する方法を提案
OffNadirLocは交差視点地理位置を推定するための基準セットを提案します。これにより、ドローンと衛星画像の交差視点地理位置推定プロセスでは重要な構造的シーン理解と内部ドメイン間の関係制約に焦点を当てることができます
ETPデザイナはマルチモーダルな電子シアターのデザインを自動化するフレームワークを提案します。
Synthesizing native 2K multi-garment virtual try-on is a formidable frontier in digital fashion, critically bo
Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by ge
Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, seman
この研究では、ドローンで小さな物体を認識することを目的としたメモリ拡張型大規模言語モデルを開発しました。このモデルは、複雑なドローンの場面で、ユーザーの指示に従って物体を識別できるようになります。
The Segment Anything Model 2 (SAM2) has advanced temporal promptable segmentation, yet its deployment remains
自動運転システムには、道路のトポロジー(ドライバブルレーンとその接続性)を理解する機能が必要です。最近の検出モデルは360度の前方視野からボリュームイメージを取得することで、道路上のレーンのトポロジーを推測することができ
この研究では、地象性AIにおける物理的知識を使用してポーラリメトリック合成開口ラダール画像を分類するための新しいモデルを提案しました。このモデルは、ラダール画像を物理的なプロセスと関連付けることができます。
3D Gaussian Splatting (3DGS) は 3D セグメント間の接着を実行するために使用され、テキスト ドライブ の 3D シーン エディットには不可欠です。現行の方法では、固定位置撮影から 2D ディ
DRGBTトラッキングの分野では、目標物を変動するセンシングモデリティと観測プラットフォームの条件下で追跡することが求められます。ドリフト、視線、時間の条件変化に関しても検討が必要です。ただし、現在のバenchmarkで
VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rar
The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated sys
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to
Laboratory automation accelerates discovery, yet its adoption is constrained by the high cost, proprietary des
Language-Guided Grasping は、複雑なシーンで物体の把持を行うために、視覚言語モデル(VLM)を用いる。このアプローチでは、VLM は直接把持を予測するのではなく、3 次元空間における把持の位置を指
ReferTrack は、自然言語で対象の車両に付近する自動車を追従させるシステムである。このシステムでは、対象の車両に付近する自動車を認識する後、自動車の動きを予測する。
SOPD-SocialNav は、学習モデルを小さなロボットに伝える技術であり、ロボットが環境と人間の行動を理解し、ナビゲーションが行えるようにする。
分子設計を自動化する方法「Boltzmann-Expected Molecular Design with Decoupled Annealing Flows(DECAF)」を提案。分子設計で重要な3次元構造の特性を確率
この研究では、Oceanモデルを使用して、オーシャンで不完全な観測を使用する可能性と、生成的ステートスペースモデルと最適化フレームワークを使用して直接不完全な観測から学習する能力を評価します。
Algebraic statistics characterizes statistical models through polynomial constraints, but it has mainly been u
In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a poster
Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vu
Instruction tuning is meant to make language models follow user requests, yet it is unclear whether small mode
ハイパーネットワークを用いた知識付与法を提案し、大規模言語モデルに確実に知識を付与する方法について検討した。
Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untr
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex
分散型言語モデル(LLM)やコンテキストを活用するエージェントは、製品開発やファイナンス分野で活用されている。エージェントを実用化するには、堅牢性、安全性、信頼性を確保することが大切となる。 このチュートリアルでは、エー
Low-rank adaptation introduces a static learned update applied identically to every input. The update provides
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether t
機械学習モデルの解釈のためのツールが提案されていました。これにより、モデルがどのように機能しているかが理解できるようになります。
Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format
Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge re
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to
人間が与えるラベル相違を研究した研究では、主に分異議のあるデータを選んで分析したところ、非上昇漸近性オペレータを持つ仮説がChaosNLIで低いラベル相互に関連性があると結論づけた。しかしながら、データが選択されていない
この研究では、学生チームのテーブル演習(TTX)における評価方法を提案し、複雑でオープンエンドな状況にあるチームの行動とコミュニケーションを記録できるTTX学習プラットフォームを使用します。
Offline再調整学習(RL)で、アクション偏好キューを使用し、エキスパートのフィードバックを利用してポリシーを向上させます。
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback
この研究では、平衡方程式を満たすPINNs(物理基準付きニューラルネットワーク)を使用して、平均脱出時間の計算を目的とした椭球型境界条件付きPINNsを提案し、PINNsを使用した計算と実験室データを比較します。
Reliable Text Difficulty Assessment is a prerequisite for valid text simplification workflows and personalized
この研究では、Chain-of-Thought (CoT) スーパーヒバテーションでは、最終的な答えに到達するまでの理由を公開することで、中間的には提供される理由の質を強力にし、しかし、多くの場合には、前のステージに到達
ある単語を複数のスピーカーや環境の異なる条件下で言語モデルが使用できるようにしたい場合は、単語の抽出を実現する必要がある。しかし、現在の言語モデルでは、スピーカーの特性や環境の特性が単語に含まれていることが多い。ここでは
We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation u
この研究では、固定化された言語モデルを強化するために、自律進化する対話スキルを開発しています。このシナリオでは、ユーザーの反応がモデルの進化に影響を受けないため、対話の対称性を維持する必要があります。
この研究では、強化学習の報酬探求を量化するために、新しい測定方法を提案しています。この方法は、モデルが報酬を取得する際にどのように操作しようとしているかを示すことができます。
Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model
アラビア語の発音記号化は重要な問題だが、データが不足していることが難点の一つである。この問題を解決するために、ここでは「Connectionist Temporal Classification (CTC)」を使った制約
人間は、大きな言語モデルを使って長い論理的推論を行うが、このような推論の結果は正しくない可能性がある。ここでは、これらの発言を検証する手法を提唱する。
Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotat
大規模言語モデルは、決定タスクを遂行する過程で、実行された事実を含むパラメトリックな知識を漏らす傾向にある。大規模言語モデルが実際にどのような意思決定タスクを遂行したかを検証するのは困難であるものの、これが確かに事実であ
Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully r
This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (A
Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large
This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Mode
Large language models (LLMs) have driven rapid progress in electronic design automation (EDA), yet their appli
The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms
Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural
Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premis
Translating brain signals into text could restore communication for people with severe paralysis, yet practica
Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choo
Stance detection aims to identify whether a text expresses a favorable or opposing attitude toward a given tar
大判言語モデル(LLM)における感情の解釈を研究し、感情表現は内在する主観的変数によってどのように説明されるかを問う。
A single embedding space that covers text, images, video, and audio lets one index serve every query a user ca
ここでは、危険な知識を持つモデルにコントロールトークンを追加し、コントロールトークンに基づいてモデルが危険な知識を操作することを目標としていました。
Multi-vector dense retrieval models, such as ColBERT, achieve strong retrieval effectiveness by modelling fine
Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT):
Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combi
リカバリーのためのプログラムを学習するフレームワークを提案し、そのプログラムを用いて、文書にラベルを付与する検索システムを構築する。
As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understan
Long-video question answering requires a model to preserve visual evidence over time without repeatedly reproc
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs
Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise vid
Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasin
画像生成において、材料、 객체、領域を制御することが難しい問題がある。 Diffusion Transformers はテキストと画像を組み合わせて処理できるが、どちらをどの程度影響させるか決める仕組みがなかった。 その
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond
無線無人飛行機のルートプランニングでは、視空間と言語モデルを利用して安全なルートを生成する必要がある。この問題を解決するために、テスト時にモデルをスケールアップさせる方法を提案する。
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across
実際の空間知能では、空間に続いて流れるビデオを理解する必要がある。この問題を解決するために、4次元空間を理解することができるモデルを提案する。
心臓の囲みの区別は、食道肥厚の測定に重要であるが、しかし、これを正確に区別することは難しい。これを解決するために、周囲の解剖学的構造を利用して囲みの区別を改善する方法を提案する。
自動運転のための計画とは、状況理解、タイムリーな推論、行動選択というものがあるが、しかし、これらの要素を組み合わせるのは難しい。これを解決するために、シーン理解を分離することによって、計画を安全かつ有効性のあるものにする
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more t
World Action Models(WAMs)は、ロボットマニピュレーションをモデル化するパラダイム。WAMsは、視覚ステートトランジションとロボットアクションを同時にモデル化する。しかし、既存のWAMsは、一定の時
オープン・バグナビゲーションには、エージェントへの部分観測が含まれます。パフォーマンスの向上のために、内部状態更新が重要です。これを実現するには、ポリシーネットワークの更新が必要です。最近のアプローチでは、トランスフォー
行動モーター特徴は、社会認知や人間ロボットインターフェースなどの行動認識の核心です。人間ロボットのNICO用に、2段階のアーキテクチャを提案します。1段階目では、腕の移動を学習するSOMと、手の移動を学習するSOMを使用
Pneumatic artificial muscles have wide applications in robotics and industrial fields. Conventional pneumatic
The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordin
Gold-standard phenotype labels are often unavailable at scale in electronic health record (EHR) studies becaus
多元表格予測モデルTabPFNは、条件設定されたサポートセットと、入力クエリーでタスク特指訓練を行うことなく、推論を行います。実行時間における内部挙動を理解する際に、 zig-zag永続ホモロジーを使用することで、Tab
この論文では、パス値学習を行うためにpath signaturesという
Neural simulation-based inference enables parameter estimation for complex models, but typically requires the
Language models produce probabilities over words, but professional decisions require uncertainty over meaningf
Dynamic graph learning aims to capture evolving structural and semantic patterns in real-world systems, such a
Discourse relations provide document structure, critical to language understanding and enabling language model
Persona prompting is widely used to steer LLM agent behavior, yet the narrative framing of a task can matter m
Reasoning-specialized language models show large performance gains over base models, yet the internal changes
Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relatio
When a language model must choose one answer from a large space of equally valid options, a format clause -- "
Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, y
Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have
Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete
Automated analyses of privacy policies enable large-scale assessments of transparency in digital ecosystems, y
Large language models (LLMs) largely rely on Transformers, where self-attention provides global token interact
Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should acc
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly label
Pruning long context for coding agents has been a vital technology for efficient context management. While exi
訓練データの品質を高めるために、付与ラベルに基づいてデータを選択する方法を提案し、訓練データから選択されないデータは排除することを目指している。
Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other
工場の中の異常が検出されるように設計された機械学習モデルを提案しています。通常の方法では、モデルはビデオ内のすべての内容を考慮し、複雑な問題を解決することは困難です。提案されたモデルのアプローチは、オブジェクトを検出して
この研究では、ルビック評価を含む非確認タスクの最適化を目的とします。従来のRLには、モデル評価の情報が使われるだけですが、モデル自身は反省や自己改善はすることがありません。ここでは、LJMをコーチとみなして、モデルが反省
金融質問回答を実行するには、長い標準化されて高度に冗長な説明書に分散する証券取引委員会(SEC)の証拠を取得する必要がある。既存の取得を拡張するおよび多要素システムの多くの選択肢は、モデルの先行事項と目的のファイルリング
大きな言語モデルは実用システムで増えているため、費用対効果のあるモデルを選択することが重要になる。モデルを割り当てるためにLKM路線が提案された。しかし、既存の路線方法は入力問に基づいてモデルを選択し、モデルに適合しない
Vision-language-action (VLA) models have shown impressive generalization, but often lack interpretability and
ロボット制御を効率化するために、パッチを用いた政策学習を提案し、密集された視覺表現を用いて実装することを目的としている。
existing VLA modelの制約を解決するためのforce-based memory method、FM-VLAを提案する。
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but becaus
existing robotic control methodの限界を解決するためのbackward dynamics extractionとunpaired domain translation methodを提案し、
Fast planning of novel behaviors in unseen scenarios remains a fundamental challenge in robotics. The high-dim
Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by
In autonomous driving development, a perception dataset is crucial, as it provides fundamental data for traini
この研究では、混乱のないターゲットに可視化言語アクションモデルを適応させることを目的として、3つのモデルを使用して研究を行った。3つのモデルは、観察から直接行動へのマッピング、テキストチャインオブスロット、潜在的な反復ル
Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts
We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using
Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machi
Analytical placers rely on differentiable objective functions to guide placement, typically combining intermed
Teleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces
Humanoid robots have become increasingly popular in applications such as social interaction, education, and se
This paper investigates temporal fair division, a setting where items are allocated over multiple rounds and a
This paper proposes a causal independence principle for value -- the value Causal Markov Condition (v-CMC) --
How deep does a graph neural network need to be on a sparse graph? We study its purest statistical form: node
Backpropagation makes training deep networks memory intensive because it must store intermediate activations.
Social navigation requires the robot to reason and respond in complex real-world environments. While recent wo
Safe and socially compliant navigation in open human-robot environments requires robots to reason about hetero
Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and suppor
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazar
A large body of Semi-supervised Learning~(SSL) algorithms encounter the threshold $τ$ to select pseudo-labels.
Predictive models deployed at scale influence future data, a phenomenon called performativity. And there is al
Hallucinations and artificial text in LLM-generated outputs often appear as distributional deviations between
Physics-informed neural networks (PINNs) are unusually sensitive to interacting choices of architecture, activ
Fish-like swimming has inspired the design of several dozens if not hundreds of bioinspired robots in the last
For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days
この研究では、Robotのヘッドレスアンドメインアームの両方を1台のロボットに組み込み、両機能を切り替えれるようにする技術、Handroidを開発しています。
VTLocフレームワークは、視覚情報と触覚情報を統合し、ロボットハンドの位置を推定することで、ロボットハンドの位置推定と動作操作を実現します。
PIXIEフレームワークは、6次元オブジェクト位置推定を実現し、ロボットハンドの制御と物体の操作を実現します。
この研究では、ブロック交換ポリシーを使用して $N$ 個の独立した同型マシンを維持するために、オペレーターが $K^*$ の最適な期間を決定するためのデータ駆動型法を開発しました。この法は、オペレーターが選択した期間に基
この研究では、時系列データの予測を目的に、新しいフレームワークを提案しました。このフレームワークは、時系列データの予測を容易にするために、グループの注意とニューハウクスプロセスを組み込んでいます。
この研究では、サンプル協方差行列の精度を向上させる方法を検討しました。特に、サンプルサイズが小さく、収集されたデータの特性から、正解率の期待値は小さい場合に、問題が最も発生しやすくなる可能性があります。この研究では、正解
Layer-wise post-training quantization of large language models minimizes each layer's reconstruction error in
Large Language Models(LLM)の個別化はモデルを適応させることができるが、計算リソースが限られている状況では、コストがかかるSupervised Fine-Tuning法か、軽量なIn-Contex
Helmholtz方程式は、時間共伴振波の伝播を記述する重要な方程式であり、媒質が損失した場合複素係数を持ちます。ここでは、空間での波場から波方程式を推測するために、物理知識に基づくGaussian Process(GP
Lipschitz continuity is a fundamental property of neural networks that characterizes their sensitivity to inpu
述語学習における歴史的類推を推測し、歴史的類推を評価するためのアナロジー ディープ リサーチという新しいタスクを提案し、述語学習における歴史的類推が重要な役
最小限マージン定理を仮想環境における公平性と利益のトレードオフの分析に導入します。この定理により、公平性と利益のトレードオフを最小限の公平性確保することで解決することができます。
Large language models (LLMs) have made automated heuristic design (AHD) increasingly practical by generating e
Advertisers delegate bidding to autobidders; users delegate tasks to language-model agents. A person describes
ルート選択に関する研究を進め、人間の選び方を再現することで、スマートなナビゲーションシステムを開発します。
分類Forecasterの監視と評価を容易にするために、可変期間のバックテストを導入し、情報依存性を考慮したCalibrationを可能にしました。
分布型強化学習のリスク評価を容易にするために、分布型強化学習におけるリスク評価を分析しました。
健康状態の推測は、人間の健康状態を推定するために、生物学的および行動的なデータを使用します。この研究では、健康状態を予測するための基金モデルとしてのDAG-FMを提案しました。
Decision support systems (DSS) increasingly run retention what-if analysis on synthetic customer populations,
Crafa's algorithmic analysis of the Italian electoral law (the "Rosatellum") showed that the statutory text de
We study robust repeated contextual pricing, where valuations depends linearly on the features. At each round
Standard learning rate schedules such as cosine annealing are tied to a fixed training horizon, limiting their
We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on $\mathbb R^d$ whose discrete a
In contrast to most studies on neural network approximation theory that characterize results through a single
Pseudo-labeling based on Optimal Transport (OT) has become an effective mechanism for enhancing short text clu
While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typica
Within-class variance in language-model representations is commonly read as incomplete neural collapse. We arg
Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative reca
The escalating demand for Machine Learning (ML) training resources in recent years has resulted in a substanti
これは、Exploratory Landscape Analysisにおけるランダムサブスペースのサンプリングを使用するためのフレームワークであるSampling on Random Subspacesを提案している。
これは、社会的行動を予測するための新しいフレームワークであるSocial-spatial dependenciesを提案し、個々のエージェントが社会的信号を学習する能力を向上させる。
複数のエージェントの行動を分析するための方法を提案した。複数のエージェントの行動を
We revisit the complexity of deciding whether a graphical game admits a pure Nash equilibrium (PNE) parameteri
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability
これは、プレイヤーが勝つゲームの勝利条件の強制とパロディーを目的としています。カードプレーヤーのゲームで特に興味を持っています。
複数の買い手を持つ市場における交渉システムを構築します。マーケットの規模を知り切れていない場合、セラーの損失が生じます。セラーは市場の規模を測る必要がありますが、これは複数の買い手を持つ場合に困難です。
買い手が情報を共有すると交渉が困難になる可能性があります。この問題に対処するために、貢献者が情報を共有するリスクを減らすための新しいアプローチ、「contextual auction with bandit learni
This article is about the development of a fuzzy cognitive map using a local large language model. In the ligh
Designing effective multi-objective Bayesian optimization (MOBO) algorithms requires balancing many interdepen
The integration of Large Language Models (LLMs) with evolutionary computation has emerged as a powerful paradi
本研究では、推論プロセスの検証を目的とした Heaviside 不連続性の考慮を提案する。これにより、推論プロセスにおける潜在的なミスを検出した上で、正しい出力を生成することができる。
オンライン購入の最適化を目的とするストラテジックビーイングアージェントフレームワークを発表する。
Language models increasingly mediate paid advice: agents submit open-ended forecasts, recommendations, plans,
Large language models remain limited as continual learning systems, motivating renewed interest in Sparse Dist
We investigate whether structured reasoning interventions improve the strategic economic reasoning of large la
In-context learning (ICL) は、現代の AI アーキテクチャのフワードパスの内側に埋め込まれた、潜在的なグレーディエント降下です。ICL を生物学的に可能性がある Spiking Neural N
この研究では、メタボリック マルチエージェント最適化 (MMAO) が動的最適化に適用できるようにする必要がありました。MMAO-Dyn は、環境の変化によって元の有効な局所的構造を無効にした非stationary な設
AIを援助するための意思決定者によるオーバーサイトの研究。AIが提案した行動の評価と決定を行うために、意思決定者とAIが情報を交流するオーバーサイトの実現を研究する。
LLMの不正行為に対する防御。この研究では、LLMの不正行為を防ぐための防御の枠組みを開発し、LLMの不正行為の危険性を分析する。
予測情報を組み合わせて正確な予測を得るための枠組みの開発。この研究では、個人的予測情報を組み合わせて正確な予測を得るための枠組
Mechanism design increasingly faces heterogeneous environments containing both traditional utility maximizers
This paper introduces an original approach to an underexplored issue: the integration of a new member into an
Continual learning (CL), where a model is trained on a sequence of data tasks, is increasingly being adopted a
Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs
Large language models (LLMs) demonstrate broad reasoning abilities but struggle with accuracy and reliability
Dreams splice together people, places, and times that never met. Neuroscience suggests this recombination is n
Large language models (LLMs) increasingly mediate strategic interactions through natural language, making sema
Traditional meta-heuristics often rely on fixed population sizes, manually chosen search scales, and externall
これは、LARGE LANGUAGE MODELS (LLM) の理論心の評価を拡張し、三重なるWerewolfゲームを追加しました。
再発生モデルに効率的な記憶の管理を提案しており、記憶の除去はデータの新規記憶と共存する必要があります。
この研究では、多様性のあるトキシックテストを検索します。
大規模言語モデルの戦略1対1ベンチマークであるAge of LLMを紹介。マインスイーパーゲームを想定し、フォーゲットオブラーサー、マインスイーパー対戦、JSONスキーマへの従属性という三つのストレスアウトを設定。
Backpropagation-trained dense neural networks are powerful function approximators, but they couple learning ac
Electroencephalography (EEG) foundation models increasingly rely on multi-dataset training and evaluation, yet
Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM
Maintaining physical consistency in video generators and world models increasingly relies on vision-language m
この研究では、モデルの行動を分析し、モデルの行動を他の環境に適応させる能力を評価する方法であるBehavioral Portability Testを開発しました。
Building on resistive communication, this paper presents a physics-based design of an on-chip neural network w
What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Insp
Empirical economists often start their projects with a toolbox. Shared packages, replication archives, and cir
Abnormality detection in complex systems faces two practical barriers: abnormal labels are scarce, and binary
Attentionメカニズムをフラストレーションされた同期の観点から研究した。この方法では、トークンの状態を相関する相のフェーズとして設定することで、Attentionメカニズムがどのようにして計算されるかを理解すること
この論文では、スパイクニューロンの生物学的合理性を評価するためのオプティマイズフレームワークを提案します。このフレームワークは、Izhikevichの生物学的合理性の定義に基づいており、スパイクニューロンをモデル化するた
多オブジェクトの最適化における偏差修正の効果を分析し、偏差を削減するための方法を提案した。
この研究では、PFSPのIterated Greedy (IG) アルゴリズムのパフォーマンスを改善するために、Large Language Model-Driven Cooperative Operator Ensem
この研究では、自動補助関数設計(AHD)についての研究を行った。AHDは、マシン学習が可能になる以前から研究されていたトピックであり、マシン学習によって、AHDがさらに活用可能になった。この研究では、AHDにおけるメタ認