MIRROR: Learning from the Other View for Multi-Modal Reasoning
多モーダル理解技術のための新しいアプローチであるMIRROR(Learning from the Other View)を提案しました。MIRRORは、テキスト、図、テキストと図の組み合わせから同等の視点を提供することで
- 用途
- 多モーダル理解技術の開発
- 難易度
- Hard
- コスト
- High
「LLM」の検索結果
250 件多モーダル理解技術のための新しいアプローチであるMIRROR(Learning from the Other View)を提案しました。MIRRORは、テキスト、図、テキストと図の組み合わせから同等の視点を提供することで
Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a targe
Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for ena
Scaling inference-time computation has emerged as a reliable method to improve the performance of large langua
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason throug
Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architec
Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent
この研究では、機械学習モデルを使用して血糖値の変化を予測し、糖尿病管理のためには血糖値データの前処理が重要であることの重要性を強調しています。
この研究では、反対称関数を用いて、機械学習モデルが状態のどの点からどの点への値の差を予測できるような相対的な値学習(RV)を提案し、制御や推定を向上させる可能性があります。
この研究では、自己説明の信頼性を検証するためのRL方法を提案し、自己説明の信頼性を直接最適化するための新しいアプローチを検討します。
CUDAカーネルの生成を支援するCudaPerfを提案した研究で、この方法により、高性能のCUDAカーネルを効率的に生成できる。
LLM(言語モデル)の評価における位置バイアスを分析するための方法を提案した研究で、この方法により、位置バイアスが評価結果にどのような影響を与えるかが明らかにできる。
オフラインRL(非実時学習)におけるタスクの分割を支援するOffline RL with Hierarchical Action Chunkingを提案した研究で、この方法により、タスクの分割が効
Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal poli
Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality
この研究では、言語モデルが社会的に正しい判断を下すことができる方法について調べた。結果は、模式が対立を保つ能力が高くなり、他の人の視点を受け入れやすくなったことである。
この研究は、リソースフローの動作を表すPetriネットと、APIを操作するためのテストを自動生成する方法を提案した。方法は、APIの機能をテストするためのシナリオを生成し、テストが正しく実行されるようにした。
この研究では、LMOの安全性を調べた。結果は、直面する危険目標に対してモデルが安全なアドバイスを出すことができた。
この研究は、奇数サイクルのシャノン容量の最小限度を検討した。結果は、グラフの独立集合の大きさに基づいて最小限度を計算することができた。
Large language models (LLMs) and agents are now widely used tools in code development, with data typically sen
People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks rema
In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninf
Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However,
For years, supply chain planning at e-commerce firms has operated as a collection of isolated projects. Each p
Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, tec
Retrieval-Augmented Generation (RAG) systems increasingly employ multiple LLM agents. Yet, most prior work opt
The reliance on unstructured free text for documenting clinical trial protocols creates a significant barrier
LLMは、人間のアイデンティティのシミュレーションを使用して個人データを削除したり、未均衡なデータを削除したりしますが、これらのアプローチには制限があります。
このプロセスでは、大規模言語モデルを使用して、ダイナミックプロセスモデルに基づいて自動化された制御戦略を生成します。
この研究では、大規模言語モデルを活用して、Greekに基づく書籍検索システムの評価を行いました。大規模言語モデルを活用することで、検索精度が高まりました。
この研究では、大規模言語モデルを活用して、経済学の研究活動をサポートするシステムを開発しました。このシステムは、学者が理論モデル開発を自動化することができます。
この研究では、大規模言語モデルを使用して、GREEKのスラングを研究しました。このスラングは大規模言語モデルを活用することで推測することができました。
Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just
LLM (Large Language Model)とLP (Logic Programming)を組み合わせて、有理数である√2の非有理性を証明します。この証明には、LLMが主観的な論理式を生成し、LPが証明を行うプロ
S2S (Speech-to-Speech) LLMアシスタントを利用して、人間のような話し方をすることができますが、安全対策の実装が困難です。この研究では、S2S LLMアシスタントの安全対策を2つのアプローチで実現し
老年者は移動が困難になることが多いため、この研究では老年者の安全な移動支援システムを開発します。このシステムでは、LLMと予測モデルを組み合わせて、老年者の安全な移動を支援します。
知識重視の質問応答システム (KI-VQA) を分析するために、新しい評価基準を提案します。これらの基準では、VLMの各タスクを個別に評価することができます。
ビデオLMMの安全性を確認するために、新しい診断フレームワークを提案します。これらのフレームワークは、モデルの挙動、理解、セマンティクスを同時に考慮します。
再発防止を目指す会話助言の評価基準である RegretBench を提案します。这一基準评估了會話助言の多輪交互式決定における後悔を最小化すること。
代理記憶の学習は、LGMが効果的に情報を保持・更新・処理できることを意味します。この研究では、アトリビューテッド グラフィックフィードバックを使用して、代理記憶を最適化する方法を提案します。
Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequenti
Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid anal
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognitio
Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summariz
_guardianAgentBenchBenchmarkは、580のシナリオを6つのドメインで評価し、3つの実稼動フレームワークであるLangChain、LlamaIndex、Vectaraを利用します。このベンチマーク
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources.
Large language models (LLMs) have rapidly and significantly entered scientific workflows, but it remains uncle
Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not e
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, ev
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correc
Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone d
LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records,
The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning
Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are
シミュレーション駆動設計では、高精度なシミュレーションを少なくすることで設計を実現しています。既存の手法では、その問題に取り組むために最適化アルゴリズムが改善されてきましたが、問題の定義自体は検討されていません。この論文
Large Language Models (LLMs) は医学教育に大きな可能性を持っていますが、現在のシステムでは、質問に答えるか一時的なフィードバックしか行なわれていません。一方、臨床病例を決定センターへの学習トレ
この論文では、大規模言語モデル (LLMs) が日常的な文化的知識を評価する能力に着目しています。ここで、TriviaRoomQA というクイズスタイルで問題を提示して、LLMs が日常的な文化的知識をどのように評価する
この論文では、音声字幕の評価手法が提案され、音声字幕の評価において既存の手法の制約を克服することを目指しました。提案されたフレームワークは音声字幕の各側面を評価し、質問回答型の評価手法ではなく字幕の中立性を評価することが
In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft,
Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsist
Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how t
Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by levera
Almost every large language model that reaches a broad audience is quantized: trained in full precision, then
Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs o
Distinguishing animate from inanimate concepts in written language requires more than shallow text processing,
ソフトウェア開発の維持フェーズで、ソースコードの自然言語解説を生成するためのモデルの改善を目的とした研究。
コーディングエージェントの評価基準を導入し、現実世界のコミットやプルリクエストに基づくタスクを構築した。
Chinese language の長形法律研究報告における出典の信頼性を評価し、信頼性が低い出典を検出および評価する目的で LegalCiteTrust を提案している。
長形推論のための言語モデルが、提供されたコンテキストから乖離した論理を生成する可能性があることを指摘し、コンテキストと推論論理をより適切に融合するため、 REFACT (REstating Facts in Adapti
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at in
大規模言語モデルはさまざまな画像をテキストに変換する上で優れた性能を示しているが、発生するホログラフィックな診断にはまだ解決策が必要です。この研究では、主流の粗い検出方法の欠点を補うため、細部の診断方法を提案しています。
大規模言語モデルのハロウィーン診断では、対象の 3D 空間関係を推論する際に、視覚化が欠如していることが問題となります。この研究では、これらのハロウィーンを軽減するためのアプローチを提案しています。
大規模言語モデルの圧縮には、モデルのパフォーマンスが低下する可能性があるため、量化の保護が重要です。この研究では、Fisher加重チャネル感受性を用い、MLLMの量化を安定させるためのC-PTQをプロPOSEしています。
空間理解は、物理世界と静的のセマンティック理解の間でつながるために不可欠です。多くの空間タスクは、場所、領域、パスの自然な表現は、ポインティングやマーキングなど、連続的な視覚的シーンで行われることが多いが、現行の空間推論
パスロジは、現在、パスロジ認識のための画像言語モデルを評価するために広く使用されていますが、この研究では、パスロジ認識において画像言語モデルの視覚知覚が機能していることを疑問に問っています。
感情認識は、現代のアギを促進するために不可欠ですが、大規模
この論文では、Lumeraという手法を提案します。Lumeraは、Engine-Native 3D World ReconstructionとLightsを検出するために使用します。
この論文では、ViSTR-Benchという手法を提案します。ViSTR-Benchは、MLLMが動的シーンから情報を取得できるかどうかを評価します。
Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls
強制制約に基づく強化学習を利用し、低コストで高精度の組み立てが可能になると同時に、組み立てに失敗してもロボットが安全に回避できるように、ロボットの制御のための強化学習を提案します。
Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choo
Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying qu
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated
We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with profess
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature
この研究では、分布に変化がない場合の時間軸への適応性を
OLED 材料の開発を目指す新しいアプローチ、causal language models を用いて optoelectronic プロパティを予測するフレームワークを提案する。
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving larg
この研究では、マルウェア検出に使用されるデープラーニングモデルにおける、位置依存性KVキャッシュ(Key-value Cache)を改善する方法を提案しました。
Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-a
生産的言語モデルの利用による金銭的感情分析に対処するための方法を提案している。複数のエージェントを活用したコミティー方式を使用し、さまざまな粒度のテキストデータに対応できるように、単語レベルのルールベースアプローチ、句節
VLSIのグローバルルーティングは、信号ネットワークを 3D グリッド上で割り当てることが目的であり、信号遅れやワ
Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of moder
High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances i
Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys t
AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they
Large language models increasingly use search tools to retrieve up-to-date information, introducing a new atta
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow grap
Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their
Traditional query processing engines require continuous development and extensions to support new techniques a
This study empirically analyzed generative AI as an emerging discovery pathway to academic library resources.
最新言語モデル(LLM)が危険な生成を防ぐための確信的な安全な限界を計算するための新しいフレームワークを提案した。Clopper-Pearsonの信頼区間の新しい応用として、PAC(可能性が最も近い)の境界を得るためのア
モデルの脆弱性を解決するために、四つのエージェントに分割される多様なフレームワークPoTREを導入した。モデルの推論能力を強化し、単一のストリーミングアプローチよりも複雑な理論的制約とアブストラクションに抵抗できるように
侵攻テストツールが異なっている点、決定主義的な性質、狭く特定されたスコープ、専門技術の操作を用いたものと異なり、LLM駆動の自治的セキュリティツールは3つの次元で不確実性を示した。政策決定への説明が困難、影響の開放性、行
文化的意味の表現が表現された翻訳には、翻訳システムが表現する意味を理解するために、表現の文化的背景を考慮する必要があることを指摘した。文化的背景が表現されている表現された翻訳には、いくつかの課題があり、LLMベースの翻訳
組み合わせ方程式の最適化を解くための新しいフレームワークを提示した。分布される量子アルゴリズムの局所的な制限に直面する際、最適化の解を導けるために、分布される量子近似最適化アプローチと深層学習アルゴリズムを組み合わせた。
分析報告の迅速な解釈が求められるときに行われるマルウェア分析を実現するために、閉じた重みの大きい言語モデルを使用しないことが多い。オープン重みの言語モデルは、マルウェア分析のために適切な言語能力と、閉じた重みの大きい言語
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challe
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with mislea
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge fo
Large language models can answer scientific questions, yet a correct output does not reveal whether the model
Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspe
Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolvin
Design Rule Check closureを促進するための自動修正フレームワーク、EvoDRCを開発し、複雑な幾何学的相互作用を考慮した修正を実行する。
We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive
Large language models trained and aligned within different linguistic and regional ecosystems may frame the sa
Small language models and coding agents increasingly generate web front-end code, yet their outputs are typica
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. T
Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply the
危険な問題に対する正しい答えを提供する大きな言語モデルと費用の効率が良い、小さな言語モデルを協力させる技術が開発されました。
LLMは、状況から価値観を判断できるかどうか、という研究が調査されました。LLMは、状況に応じて真の価値観を推測することができました。
同じカテゴリを組み合わせることしかできないという考えに対抗して、異なるカテゴリを組み合わせることができるかどうかについて、言語モデルが調査されました。
Persona simulation involves utilizing large language models (LLMs) to anticipate human choices or interactions
大きな言語モデルは真実の情報を提供できるように見えますが、実際は虚偽情報を提供することが多く、これを検知、検出、および検証するための基準を作成するため、HalluTruthQAが開発されました。
大きな言語モデルは、さまざまな価値観に対して異なる反応を示すことがあり、これらの反応がどのように影響するかを調べました。
大きな言語モデルは、ユーザーの信念と事実的な正しさを合わせる傾向があるが、これらの傾向は多様であることを明らかにしました。
AIは、特に当代芸術作品のパスティーシュを作成する能力が高いが、これらの作品はどれだけ実際の作品と似ているかを調べました。
大きな言語モデルのエージェントは、第三者のスキルによる実際的な危険を認識し回避する能力を評価します。
大きな言語モデルの答えは、質問や入力の形態に応じて異なる傾向があることを認識しました。
小さな言語モデルに対するドロップインコーパス、tiny_schillerを導入し、単一ファイルで利用できるようにし、言語モデルを簡単にprototyping、fine-tuning、教育、研究に利用できるようにする。
FinMMEval 2026 タスク 1 は、英語、中国語、アラビア語、ヒンディー語で行われる多言語的な金融質問に答えるものを評価します。
With the wide application of large language models (LLMs) in real-world scenarios, the value implication of th
Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic fac
この研究では、偏見が蓄積されることが多くのLLMで問題となります。一方、この研究によって、LLMの偏見を解決する新しいアプローチが提案されました。
3D オキュピエンシー予測には、物体の配置と密度を解釈するための視覚的手法が必要です。従来の方法では、計算コストが高くなりすぎていたが、新しく提案されたGaussianSeedアルゴリズムは、層を階層化することで、計算コ
この研究では、AI生成物の論理的評価に必要なものとして、生成物がどうやって結果を得るのかを明らかにすることの重要性を強調しています。この研究では生成物を分解し、その論理的な構造を理解するために自然言語推論を利用し、生成物
ビデオキャプション生成には、空間と時刻の理解が重要です。PercepCapアルゴリズムは、ビデオ入力を空間時刻認識に分解することで、生成されたキャプションの理解度が向上するとともに、空間時刻の誤差をより正確に検出でき、キ
多モーダルラージランゲージモデルは、視覚言語タスクに強いですが、高い推論コストで問題となっています。Look Less, Think Fasterアルゴリズムは、単位次元を個別に最適化することで、多モーダルラージランゲー
複数ターンのファッション画像検索は、実世界のファッション検索では重要なタスクです。Diverse-Intent Multi-Turn Fashion Image Retrievalアルゴリズムは、異なる検索用途を扱うこと
画像理解のための多モーダルラージランゲージモデルは、強力ですが、まだ能力と限界については明確な理解が不足しています。この論文では、多モーダルラージランゲージモデルが画像理解においてどの程度の能力と限界を持つか、を分析し、
Remote sensing image editing aims to modify remote sensing images according to natural language instructions w
Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure,
ETPデザイナはマルチモーダルな電子シアターのデザインを自動化するフレームワークを提案します。
Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by ge
この研究では、ドローンで小さな物体を認識することを目的としたメモリ拡張型大規模言語モデルを開発しました。このモデルは、複雑なドローンの場面で、ユーザーの指示に従って物体を識別できるようになります。
Despite recent advances in general-purpose robotic manipulation, real-world multi-object clutter remains chall
自動変換モデルで使用されるLLMの同定の精度の評価に役立つ「Total Variation Distance Estimation」を行った研究。この研究では3種類のアクセス モデルと異なる推定方法を提案し、実験で推定方
決定関数とコストの関数間の共変性により、損失関数を最適化することで、適切な行動決定を可能にすることができます。また、これに基づいて、共変性の傾向を最適化する方向性を考察し、正確に予測された結果を持つモデルを導出するのに役
Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vu
ハイパーネットワークを用いた知識付与法を提案し、大規模言語モデルに確実に知識を付与する方法について検討した。
Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untr
知識を重視した自己向上の研究を実施し、自己向上を知識を重視することにより効果的に行う方法を提案した。
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex
分散型言語モデル(LLM)やコンテキストを活用するエージェントは、製品開発やファイナンス分野で活用されている。エージェントを実用化するには、堅牢性、安全性、信頼性を確保することが大切となる。 このチュートリアルでは、エー
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether t
Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format
Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge re
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to
この研究では、学生チームのテーブル演習(TTX)における評価方法を提案し、複雑でオープンエンドな状況にあるチームの行動とコミュニケーションを記録できるTTX学習プラットフォームを使用します。
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depe
We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation u
この研究では、強化学習の報酬探求を量化するために、新しい測定方法を提案しています。この方法は、モデルが報酬を取得する際にどのように操作しようとしているかを示すことができます。
人間は、大きな言語モデルを使って長い論理的推論を行うが、このような推論の結果は正しくない可能性がある。ここでは、これらの発言を検証する手法を提唱する。
大規模言語モデルは、決定タスクを遂行する過程で、実行された事実を含むパラメトリックな知識を漏らす傾向にある。大規模言語モデルが実際にどのような意思決定タスクを遂行したかを検証するのは困難であるものの、これが確かに事実であ
This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (A
Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large
This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Mode
Large language models (LLMs) have driven rapid progress in electronic design automation (EDA), yet their appli
Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that c
Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choo
大判言語モデル(LLM)における感情の解釈を研究し、感情表現は内在する主観的変数によってどのように説明されるかを問う。
A single embedding space that covers text, images, video, and audio lets one index serve every query a user ca
ここでは、危険な知識を持つモデルにコントロールトークンを追加し、コントロールトークンに基づいてモデルが危険な知識を操作することを目標としていました。
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs
オリジナルのデータとZoom-Inのツールを組み合わせた方法、OmniReasonerを提案する。これにより、オリンモードルLLMsの長いオーディオビデオの論理的推論を改善できる。
The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordin
LLMが広く使用されるようになり、LLMを識別するツールが開発されている。 しかし、識別システムは、使用者の行動に影響を与えている。つまり、識別システムが機能しないと、ユーザが別のシステムを使用することに関連し、最終的な
Neural simulation-based inference enables parameter estimation for complex models, but typically requires the
Persona prompting is widely used to steer LLM agent behavior, yet the narrative framing of a task can matter m
Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relatio
When a language model must choose one answer from a large space of equally valid options, a format clause -- "
Purpose: Understanding how much of routine policing involves vulnerable people could inform resourcing, traini
Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete
This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European M
Automated analyses of privacy policies enable large-scale assessments of transparency in digital ecosystems, y
Large language models (LLMs) largely rely on Transformers, where self-attention provides global token interact
Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses.
Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should acc
Pruning long context for coding agents has been a vital technology for efficient context management. While exi
Battery-free Internet of Things (IoT) requires iterative design of vibration energy harvesters (VEHs) under co
Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other
研究では、LLMの不正回答を起こす根本原因を探りました。モデルを5つの家族と7つのBCTバイアスのタイプで検討すると、モデル内の特定のパターンが見つかりました。このパターンが不正回答の根本原因となります。
この研究では、ルビック評価を含む非確認タスクの最適化を目的とします。従来のRLには、モデル評価の情報が使われるだけですが、モデル自身は反省や自己改善はすることがありません。ここでは、LJMをコーチとみなして、モデルが反省
大きな言語モデルは実用システムで増えているため、費用対効果のあるモデルを選択することが重要になる。モデルを割り当てるためにLKM路線が提案された。しかし、既存の路線方法は入力問に基づいてモデルを選択し、モデルに適合しない
Fast planning of novel behaviors in unseen scenarios remains a fundamental challenge in robotics. The high-dim
共役ロボットは、人間オペレータと同梱するワークスペースを共有し、機械手のハンドオーバーなどの安全性の高いマイクロイベント頻繁に発生します。但し、従来の静的なハンドオーバーは、非対称の産業工具を取り扱う際、不自然な抓を持つ
Linear attention promises constant-time recurrent inference but degrades sharply on associative recall. We for
We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using
Analytical placers rely on differentiable objective functions to guide placement, typically combining intermed
This paper proposes a causal independence principle for value -- the value Causal Markov Condition (v-CMC) --
In The Algebraic Mind, Marcus identified three cognitive components: operations over variables, recursively st
Two critiques of connectionist cognition converge on one missing capacity. In The Algebraic Mind, Marcus isola
空間分布関数は、データ分析における重要な手法であるが、正確な空間分布関数を評価する方法が必要。この問題を解決するために、空間分布関数を評価する方法を提案。
Hallucinations and artificial text in LLM-generated outputs often appear as distributional deviations between
Physics-informed neural networks (PINNs) are unusually sensitive to interacting choices of architecture, activ
この研究では、動的なシナリオを分析するために可視化した地図上にMotion Attributeを付与し、Language QueryによるMotion Attributeフィルタを使用して分析することができます。
弾性的深層ネットワークの安定性に関連する問題を解決するための新しい原理が提案されました。この原理は、入力量のエクスポネンシャルに基づく安定性しきい値が得られます。この安定しきい値は、各残差ブロックの速度場の入力量のエクス
Large Language Models(LLM)の個別化はモデルを適応させることができるが、計算リソースが限られている状況では、コストがかかるSupervised Fine-Tuning法か、軽量なIn-Contex
述語学習における歴史的類推を推測し、歴史的類推を評価するためのアナロジー ディープ リサーチという新しいタスクを提案し、述語学習における歴史的類推が重要な役
Large language models (LLMs) have made automated heuristic design (AHD) increasingly practical by generating e
Multi-task model merging combines separately trained expert models into a single model that handles all tasks
大型言語モデル(LLM)は最近急速に普及していますが、その推論に際してはAI加速器が必要になります。トークンフェーズはLSTMなどのニューラルネットワークで処理される分野ですが、現在AI加速器におけるこの分野の効率を向上
This paper proposes the certainty-equivalent first-order learning (CEFOL) algorithm, a deep learning algorithm
複数のエージェントの行動を分析するための方法を提案した。複数のエージェントの行動を
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability
複数の買い手を持つ市場における交渉システムを構築します。マーケットの規模を知り切れていない場合、セラーの損失が生じます。セラーは市場の規模を測る必要がありますが、これは複数の買い手を持つ場合に困難です。
This article is about the development of a fuzzy cognitive map using a local large language model. In the ligh
Designing effective multi-objective Bayesian optimization (MOBO) algorithms requires balancing many interdepen
The integration of Large Language Models (LLMs) with evolutionary computation has emerged as a powerful paradi
本研究では、推論プロセスの検証を目的とした Heaviside 不連続性の考慮を提案する。これにより、推論プロセスにおける潜在的なミスを検出した上で、正しい出力を生成することができる。
It is increasingly common to aggregate predictions from multiple LLMs, each with domain expertise or access to
Large language models remain limited as continual learning systems, motivating renewed interest in Sparse Dist
We investigate whether structured reasoning interventions improve the strategic economic reasoning of large la
LLMの不正行為に対する防御。この研究では、LLMの不正行為を防ぐための防御の枠組みを開発し、LLMの不正行為の危険性を分析する。
データ市場におけるデータの取引の促進。この研究では、データの取引を促進するための経済的インセンティブを開発する。
LLM-assisted evolutionary search (LES) has emerged as a promising paradigm for automated algorithm design. How
Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs
Large language models (LLMs) demonstrate broad reasoning abilities but struggle with accuracy and reliability
Large language models (LLMs) increasingly mediate strategic interactions through natural language, making sema
これは、LARGE LANGUAGE MODELS (LLM) の理論心の評価を拡張し、三重なるWerewolfゲームを追加しました。
Evaluating LLM agents requires dynamic environments that go beyond static reasoning and zero-sum games. Real-w
この研究では、多様性のあるトキシックテストを検索します。
大規模言語モデルの戦略1対1ベンチマークであるAge of LLMを紹介。マインスイーパーゲームを想定し、フォーゲットオブラーサー、マインスイーパー対戦、JSONスキーマへの従属性という三つのストレスアウトを設定。
Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM
Maintaining physical consistency in video generators and world models increasingly relies on vision-language m
この研究では、モデルの行動を分析し、モデルの行動を他の環境に適応させる能力を評価する方法であるBehavioral Portability Testを開発しました。
Building on resistive communication, this paper presents a physics-based design of an on-chip neural network w
What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Insp
Deploying multi-agent reinforcement learning (MARL) in the real world is often limited by model mismatches bet
Empirical economists often start their projects with a toolbox. Shared packages, replication archives, and cir
In this work we present a LLM powered, evolutionary code synthesis system for structured data translation in a
この研究では、PFSPのIterated Greedy (IG) アルゴリズムのパフォーマンスを改善するために、Large Language Model-Driven Cooperative Operator Ensem
この研究では、自動補助関数設計(AHD)についての研究を行った。AHDは、マシン学習が可能になる以前から研究されていたトピックであり、マシン学習によって、AHDがさらに活用可能になった。この研究では、AHDにおけるメタ認