paperless-ngx — A community-supported supercharged document management system: scan, index and archive all your documents
paperless-ngxは、コミュニティによってサポートされたスーパーチャージドのドキュメント管理システムで、ドキュメントのスキャン・インデックス・アーカイブが可能である。
- 用途
- ドキュメント管理
- 難易度
- Easy
- コスト
- High
「LLM」の検索結果
391 件paperless-ngxは、コミュニティによってサポートされたスーパーチャージドのドキュメント管理システムで、ドキュメントのスキャン・インデックス・アーカイブが可能である。
rayは、core分布ランタイムとAIライブラリで構成されたAI計算エンジンで、スケーラブルなAI計算をサポートする。
さまざまなLLMのゲートウェイとして使えるライブラリ。
ユーザーの行動を認識し、オートエージェントを構築するためのツール。
Unsloth Studioは、オープンモデルのトレーニングと実行を支援するWebUIです。このライブラリは、Gemma4、Qwen3.5などのオープンモデルのテストとトレーニングを支援するために使われます。
LLMの推論 Transparency を高めるために、DiffusionGemmaの計算を分離しVariable Transparency とAlgorithmic Transparencyを評価します。
このリポジトリでは、AIアプリケーションをローカルに実行できるツールキットであるRunAnywhere-sdksを提供しています。
マルチモーダルAIに適したオープンレイクハウスフォーマットです。このフォーマットでは、パレットからデータを2行のコードで変換することができ、100倍速くなります。また、ベクトルインデックスやデータバージョニングが可能です
このリポジトリでは、私的なAIプラットフォームであるDocGPTを提供しています。
Rustを使ってモジュラーLLMアプリケーションを構築することができるライブラリです。
LLMを評価するプラットフォームであり、さまざまなモデルとデータセットをサポートする。
Apple Silicon上でLLM推論サービスをシステムエンジニアが作成するチュートリアル。
セキュリティゲートウェイを提供するクラウドネイティブなプラットフォームです。
大規模言語モデルのテスト時間調整に関する調査のリポジトリ。
この論文では、現在のVision-Language-Benchmark(VLB)を超える、MLLMがアクティブな観察を実演できるようにするためのバenchmark、ActiveVisionを提案する。このActiveVi
このリポジトリでは、Lecture Learning Modelsに対してReinforcement Learningを実行するライブラリを提供しています。
ARTは、多段強化学習トレーナーです。このトレーナーは、GRPOを使用して、現実世界のタスクに対して、多段強化学習を行うことができます。
この論文では、LLM を提供するために使用される Mooncake サービス プラットフォームについて説明しています。Mooncakeは、Kimi というリーディングのLLMサービスを提供するサービスです。Kimiは、M
このリストは、金融市場で使用できる強化言語モデル(LLM)と深層学習の戦略やツールに関するawesomeリストです。
このリポジトリでは、高性能で大規模なベクトルデータベースとベクトル検索エンジンを提供しています。
xtunerは、超大規模MoEモデルを高速にトレーニングするためのトレーニングエンジンです。
このリポジトリでは、AIワークロードを管理するためのシステムであるSkypilotを提供しています。
TensorZeroは、LLMゲートウェイ、オブザーバビリティ、評価、最適化、実験を統一したオープンソースのLLMOpsプラットフォームです。
metaflowは、AI/MLシステムを構築・管理・ディプロイするために使用できるプラットフォームです。
flyteは、高度に動的で堅牢なAIオーケストレーションプラットフォームであり、データ、モデル、コンピューティングを統合してAIワークフローを作成することができます。
giskard-ossは、LLMエージェントの評価とテストライブラリを提供します。
aimは、利用しやすく強力なオープンソースのエクスペリメントトラッカーです。
このリポジトリでは、AIモデルの互換性を確保するためのオープンスタンダードであるONNXを提供しています。
音声認識、声活動検出、テキスト処理などを行う、基盤となる音声認識ツールキットを提供する。
Pytorchの教科書としては非常に詳細で、実践的な方法を提示しています。
CVV または CWE への分類を実現し、バグ修正のために重要な手順となるCVEへの CWE 分類を自動化する。
このリポジトリでは、中文LLaMA & Alpaca LLMsを提供しています。
LLMやVLMのFine-Tuningを簡素化したライブラリ。
ドキュメントを構造化するために使えるオープンソースのETLソリューション。
オープンソースのGPT/LLMエージェント作成ツールです。
LLMを利用するために、セマンティック検索やLLMのオーケストレーションなどを行えるフレームワーク。
マシン学習、統計学習などに関する統計的エンジンです。
AIを使ったwebスクレイピングツールです。
ゼネレーティブAIに関連するリソースの一覧。
(この項は、MiniOneRec — Minimal reproduction of OneRecのリポジトリで説明しているため、このリポジトリは削除しました)
マルチモーダル理解技術のための新しいアプローチであるMIRRORを提案しました。MIRRORは、テキスト、図、テキストと図の組み合わせから等価な視点を提供することで、視覚的な推論や複雑な推論力を向上し、さまざまなモデルの
Ludwigは、LLM (Large Language Model) のカスタム化と構築のための低コストフレームワークです。このフレームワークは、ユーザーがカスタム LLM を構築し、トレーニングするのを容易にします。
このリポジトリでは、LLMベースのエージェントアプリケーションのための強化学習の橋渡しを提供しています。
OpenAIに互換性があり、Cloud APIとして利用できるLLM。
モデルをサービングするためのライブラリを紹介している。
LLMを使用して、自然言語処理における情報抽出を行うためのPythonライブラリです。
prompts.chatは、コミュニティが共有したChatGPT用のプロンプットを発見・収集できる場所で、無料でオープンソースで提供されている。
Implementation for MatMul-free LM.
home-llmは、ローカルLIMを使ってスマートホームの制御を可能にするHome Assistantの統合モデルです。
このリポジトリでは、AIエンジニアリングのためのリソースを提供しています。
PyTorchベースのリージョニングLMMを作成するためのチュートリアルです。
AIエージェントの開発と実装を行うためのエンドツーマンド、コードファーストのチュートリアル。
このリポジトリでは、高スループットと低メモリ消費のLLMインフェレンザエンジンであるVLLMを提供しています。
Trade-up recommendation identifies higher-quality alternatives that preserve a customer's purchase intent whil
LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are
Cryptocurrency markets exhibit extreme volatility and non-stationary dynamics that challenge conventional fore
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, we
Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot
When a customer adds a professional camera to their cart, should the system suggest a matching lens, a generic
Large language models become consequential agents when surrounding systems let outputs change external state.
Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity a
Distillation is common in LLM post-training, where on-policy knowledge distillation (OPKD) uses student-genera
Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across re
Designing viable drug candidates requires searching a combinatorially large and rugged chemical space for mole
Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the
A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is follo
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM ag
LLM decision components that can operate within agent workflows often produce action-relevant recommendations
Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot
Decompilation recovers high-level source from compiled machine code and serves as a foundation for security ta
Public vulnerability databases collect rich information about known software flaws, including their weakness t
Quantum circuits are central to implementing quantum algorithms on quantum devices, where quantum gates must b
Building automation systems generate rich sensor data yet remain insight-poor because heterogeneous point nami
Automated reference-based evaluation methods play a critical role in assessing natural language generation sys
Recent years have witnessed great advances in the reasoning ability of Large Language Models (LLMs). However,
Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is in
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustn
LLM agent systems increasingly combine provenance tracking, authorization, policy enforcement, protocol adapte
Large language model agents increasingly rely on execution traces to master complex interactive tasks. However
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language prob
Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of it
A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We
Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, comput
Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs),
Recent reports during the AAAI-27 review cycle highlight the risk of reviewers coordinating bids for reciproca
Formalizing mathematics in a proof assistant, where a machine checks every definition, statement and proof, ha
The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational a
This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reason
Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompt
Biological agents navigate familiar environments not by re-solving routes for each new goal, but by reusing a
We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively evolve o
AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is cont
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write res
Large language models (LLMs) show potential for medical tasks, but their single-turn question-answer format do
As LLMs take on roles requiring moral advice, understanding how they attribute moral agency becomes critical.
AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there
Hallucination-where a language model generates outputs that are factually incorrect or unsupported by the sour
Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are observ
LLM-based chatbots are increasingly used as everyday confidants. Because they are designed to maximize user sa
Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being
Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-conte
Automotive infotainment validation still relies on manual testing, slow, costly, and incompatible with agile r
Large language models (LLMs) have significantly advanced automated program repair (APR), yet existing evaluati
Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memor
Improving an industrial recommender is an iterative research-and-engineering process rather than a direct path
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio
Recent advances in large language models (LLMs) create opportunities to enrich simulation-based energy policy
Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large languag
Cross-cultural mediation by large language models (LLMs) requires deciding both when to intervene and how to r
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time
Media bias in news articles operates through subtle linguistic cues---loaded language, selective framing, and
Although automatic text simplification (ATS) is critical for accessibility, its progress has not matched the r
Ransomware detection and family attribution require analysis of different modalities because it can use packin
Financial large language models are increasingly deployed for summarization of reports and disclosures, where
Activation steering has shown promise for controlling LLM generation along well-defined attributes, but it rem
Autoregressive (AR) large language models formulate reasoning as token-level probabilistic sampling, which ind
Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial int
Large language models (LLMs) increasingly rely on information retrieval (IR) systems, such as Retrieval-Augmen
Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite th
In this paper, we study the problem of personalized survey response prediction using fine-tuned large language
Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individu
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language m
Background. Large language models (LLMs) are being adopted in biomedical research at a rapid and accelerating
Training a competitive Text-to-SQL agent usually depends on human-annotated natural-language/SQL pairs, which
Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations
This paper studies adaptive recommendation under intent drift, where feedback from each recommendation outcome
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes,
We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge a
Leveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising pa
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet
Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometri
Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID
Whether repeated identical buying questions exhaust a language model's brand recommendations depends on retrie
Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages
This paper describes the participation of the BIT.UA team from the University of Aveiro in the 14th edition of
Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions,
Constructive feedback is crucial for creative writers to refine their storytelling abilities. Since receiving
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation,
Early identification of Alzheimer's disease (AD) remains challenging because established assessment methods ca
In persona-based dialogue generation (PDG), LLMs often overuse persona attributes by incorporating them regard
Multilingual large language models often struggle to reason in low- to mid-resource languages. Prior work has
Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large languag
Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potenti
Multi-vehicle cooperative autonomous driving enhances the safety and reliability of autonomous driving systems
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generatio
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
このライブラリは、空間情報を扱うためのコンピュータビジョンライブラリです。
デバイス上のLLM推論をXビット量化を使用したもの。
微舆は人人可用的多Agent舆情分析助手であり、情報茧房を打破して舆情の原貌を還元し、未来の走向を予測し、決策を助けることができます。
医学画像に対する疾患検出モデルを開発し、臨床現場で早期検出と迅速な介入を容易にすることを目的としたフレームワークを提案します。
Merging a LoRA adapter into its base model is standard deployment practice: it removes the runtime adapter's p
We present Hakken, a domain-agnostic prediction and explanation system performing knowledge prediction, i.e.,
A conformal certificate can be valid when an LLM answers alone and invalid when the same LLM sees peers that u
Large language models (LLMs) are increasingly used to annotate cultural texts at scales that are impractical f
Learning rich medical concept representations is essential for EHR prediction. Text-attributed knowledge graph
We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation pass
Large-scale logistics networks require synthetic data generation capabilities to support scenario-based planni
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a m
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dom
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than a
Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressiv
Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model
Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) throu
Continual knowledge-updating methods are often declared superior from one final checkpoint and one conventiona
We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models tha
Despite increasing reliance on LLMs that reason with external evidence supplied by tools, retrieval-augmented
Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data.
Personally Identifiable Information (PII) detection is a foundational component of data protection infrastruct
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is u
LLM-based code generation fails when correctness depends on execution-dependent coupling: the meaning of one r
Self-driving laboratories (SDLs) combine automated experimentation with adaptive decision-making to accelerate
Accurate city-level IP Geolocation is an important enabler for the modern digital ecosystem, underpinning serv
A key challenge in reliable LLM deployment is recognizing when uncertainty reflects irreducible variability in
Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted ext
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Pause-token methods improve LLM reasoning by inserting special tokens into sequences. Prior work explains thes
We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from Chin
Performance modeling is central to hardware design and software optimization, yet constructing these models re
In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input f
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the fi
Large language models deployed in high-stakes settings frequently generate plausible but ungrounded claims. St
Chinese online comments often convey social meaning through indirect and playful language that is hard to inte
Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrat
Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial m
Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an i
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos giv
A critical challenge in deploying Large Language Models (LLMs) is developing reliable mechanisms to estimate t
Latent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities w
Intent-based networking realization starts by translating high-level intents into low-level network configurat
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrow
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes con
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough:
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require comple
Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large lan
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, thei
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly
Several studies have evaluated the ability of Large Language Models (LLMs) for meal planning, yielding positiv
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect the
In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stan
Large language models (LLMs) have become ubiquitous tools for code generation and editing. However, developmen
How do the methods used to train language models to refuse harmful requests shape how that refusal actually wo
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However
Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluat
Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream
Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD be
BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may
Computer-aided engineering (CAE) simulation is among the largest and most demanding areas of engineering, wher
Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechani
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agent
Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system wi
The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assi
Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to se
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with c
Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating gro
An agent that inherits six one-line memories may pull at most one archived source record before acting; a dire
Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legib
Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to bias
Large language models achieve superior performance on tasks that require extended reasoning, but long chains o
Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique
Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy oft
Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two
In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background kno
Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armeni
Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existi
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes
The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializin
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and re
Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly grow
Long-horizon multimodal agents should remember not only what happened but also who participated. This capabili
Generative image models can now produce high-quality images, follow complex instructions, and support precise
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in cu
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and ans
Motivated by numerous parallelizable stochastic search problems, most notable and timely among them being LLM
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models b
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content
What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a sing
Multiobjective evolutionary algorithms (MOEAs) naturally expose population-level parallelism, but many mature
Heuristic design for combinatorial optimization remains heavily reliant on expert knowledge, while existing la
Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically disti
Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to perso
Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language.
Large language models (LLMs) exhibit in-context learning capabilities, where they can learn new tasks from pro
Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficient
Long-term LLM agents must preserve information across interactions while distinguishing repeated evidence, his
Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both
Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided c
This study examines the ability of large language models (LLMs) to predict the risk of weather-related forced
Multi-agent LLM pipelines increasingly assign roles, including execution and verification, to models of differ
Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmark
Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large La
Competitive programming has become a key test of large language model reasoning, with international competitio
Sign language processing systems have traditionally operated at the sentence level, ignoring critical discours
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an arbitrary target autho
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its interven
Large language models now answer medical questions with expert-level performance. However, the context these s
Selecting a retrieval model for a production RAG system requires reliable comparative evaluation, but obtainin
Conventional task-level evaluation asks whether a robot policy completes a specified action, but can miss fail
Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improv
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language mo
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks t
Model compression techniques such as pruning and quantization facilitate the efficient deployment and accelera
NestJSベースのAIチャットボット開発ツールです。
最適なAIモデルを効率的に学習するためのオーサリングツール。Agent Lightningを使用して、トレーナーをセットアップし、データをトレーニングしてモデルを学習することができる。
DEFault++は、Transformerアーキテクチャでの内部コンポーネントの不正常な動作を認識するために、3つのレベルでハイエラルキーの学習ベースの診断手法を実装しました。
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ g
Sampling from distributions conditioned on desired semantic properties is an emerging challenge in modern gene
Price extraction from websites is a key task for market monitoring, price comparison, and business analytics i
Language model agents increasingly propose actions, observe external feedback, and explain their own behavior.
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach
Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucin
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-exte
Retrieval is the first stage of modern search and advertising systems, selecting a candidate set from a large
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space,
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a sev
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recogniza
Modern Text-to-SQL systems often follow generate-execute-select pipelines, generating multiple candidate queri
この研究では、COVID-19臨床パスウェイズの予測監視を支援するために、パイプラインを構築しました。このパイプラインには、データリフティング、時間的再構成、イベントログの構築、プリフィックスベースの表現、予測モデルの整
Vane is an AI-powered answering engine.
分析システムの性能を向上するための学習モデル開発を行う。
An easy-to-use Python framework to generate adversarial jailbreak prompts.
Large Language Models (LLMs)を高速化するためには、Transformerの構造を改善する必要があります。この研究では、早期・中期のTransformer層を繰り返し使用することで、Langua
学習中のアイデアや知識を整理するための日記。
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can tr
Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache an
Quality-diversity (QD) algorithms have been gaining traction in robot learning, where diverse motion primitive
Auto-bidding is a long-horizon sequential decision problem for maximizing conversion value under budget and ke
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long s
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from cur
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improv
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behav
Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarel
AIパワードシステムやデータパイプラインの監視と評価を行うための、オープンソースのMLとLLMの観測性フレームワーク。
A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handle
Large language models are increasingly required to generate responses that satisfy multiple competing objectiv
We study fair division of indivisible items when agents have arbitrary two-level preferences: the value of eac
Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scor
This paper studies the problem of proportionally fair clustering, where the goal is to select $k$ ``centers''
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic sam
Automated bidding (autobidding) is a core component of modern online advertising systems. Within this componen
As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities
エージェントRRLに関連するアワーショットリスト。
trafilaturaはPythonとコマンドラインツールで、Webからテキストとメタデータを取得して、CSV、JSON、HTML、MD、TXT、XML形式で出力を提供します。
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in t
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements
Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, a
The functional annotation of genes in non-model organisms remains a significant challenge in computational bio
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
Large language models access knowledge inconsistently across languages, but to what extent do they differ in t
このリポジトリには、LLM、RAG、およびオーソリティの認識を含む、AIエンジニアリングのための深いドキュメントがあります。
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with i
Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a respon
Lipschitz constants are a standard way to quantify the sensitivity of neural networks to small input perturbat
販売データを分析するために、機械学習モデルが使用されるリソースが提供されていました。
Data mixing is a central design problem in large language model pretraining: given a fixed token budget, pract
Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central
Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While lar
Analog circuit topology synthesis remains challenging because useful designs occupy a tiny fraction of a combi
AIエージェントを組み立てるためのライブラリ。
本研究では、生成推奨システムにおけるアイテムIDの構築、調整、生成の手法について、アイテムIDの構築方法を分析しています。
Q-learning with linear function approximation can be unstable because an arbitrary approximation architecture
Rust言語でCandleライブラリを利用して、PythonやPyTorchを使用せずにDecoder-only LLMを自作した。
GUI操作自動化に伴う停止判定、復讐、再検索に関する問題を解決し、 GUI操作自動化を実現するためのフレームワークを開発します。
Python is widely used in scientific research because it enables rapid development and provides rich ecosystems
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fu
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries
Classical seismic data reconstruction relies on manually designed structural priors and iterative operators, w
Large language models can generate executable programs, which makes it possible to search directly over proced
We give a two-player zero-sum repeated game between a learner and nature whose value identity generates Bayesi
We investigate whether agentic artificial intelligence can automate parts of the process of designing genetic
Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, le
Most LLM-based automated algorithm design methods optimize a designated component within a human-specified sca
Large language models process large amounts of information but usually lack an explicit mechanism for maintain
AIドライブのマルチエージェント研究アシスタント。仮説の生成、データ分析、およびレポートの生成を自動化する。
Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity
Many real-world interactions among self-interested parties can be modeled by game theory, and the rapid advanc
Embodied AIやロボットとLarge Language Modelを組み合わせた研究のリポジトリ。
解決策候補の多様化を向上させる進化戦略を開発するために、LLMを用いて解決策候補を生成するアプローチを提案している。
この作品は、Large Language Model (LLM) サービスにおけるデフォルト設計とトークンの価格設定を行う為の機械学習手法を提案しています。この手法は、LLM サービスにおけるデフォルト設計とトークンの価
大規模言語モデルによる質問のルーティングをオープン化するために、リバーサルオークションを用いて質問のルーティングを実現する。
OpenRLHFは、Ray上に構築された強化学習フレームワークです。このフレームワークは、PPO、DAPO、REINFORCE++など、様々な強化学習アルゴリズムをサポートしています。
本論文では、LLMエージェント間の相互尊重を確立するために、「類似性シグナル」という新しいアプローチを提案します。このアプローチは、エージェント間の類似性を分析することで、相互尊重を促進するという考え方に基づいています。
成功した突然変異戦略の中には、単一の実行で利用可能な知識が存在し、その知識は複数のタスク間で移行することができる。しかし、既存のLLM-ベースの進化的フレームワークでは、再利用可能な知識は捨てられ、同じアイディアの再発見
この研究では、LLMへの促進の強化学習を効率化する目的で、検索コストを削減するためのコスト意識のあるクロスタイア転送を提案します。検索コストは、促進の評価に伴うLLMの回答によって大きく異なるためです。このアプローチでは
VERDICT は、マルチモーダル論理の確認と検証をサポートするフレームワークである。このフレームワークでは、多くの場合、確認と検証は、エージェントが提供するさまざまなスコアを合計すると考えられてきたが、このフレームワー
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instea
静的評価から、動的評価に進むことができるようにする方法を提案した研究。この研究では、大規模言語モデルを用いてデータ汚染を防ぐための評価方法について検討している。
We study envy elimination by adding goods (EEAG) when the additional pool has bounded supply and no separate b
理論オブミンドの評価基準「Avalon-ToM-Bench」を提案。社会的認識を評価するための基準を提供する。
ポーカーの対局戦略を最適化する研究です。この研究では、独立チップモデル(ICM)を超える戦略的継続性最適化(SCO)アルゴリズムを提案しました。
We study fair allocations of indivisible items under general set valuations. We prove that every instance with
Restaking-based protocols enable verifiable LLM inference without the high proving cost of zkML or the hardwar
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously u
Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges t
Condorcet's paradox is a foundational result in social choice theory, showing that no matter which candidate w
混合方程式戦略の推論を可能にする方法であるSolver-Guided Reasoningが提案されました。この方法では、ゲーム理論における推論をできるだけ効果的に行うことができます。
Relayは、計算コストを削減しながらLLMエボリューションのための並列プロセスの一時的なハンドオフを実現します。
エージェント評価のプロセスを迅速化するために、AV-AIVATは、イレギュラー情報ゲームにおいて、停止条件を確実に確立し、結果が明らかとなって停止できる手法を提案します。
Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion?
We give a negative solution to MAIS-O60. We first construct an example in which an initially active ReLU neuro
LLMが外部アクションを取り続けている場合、エージェントのメモリーが古くなったり誤った情報を持ったりする可能性があります。この問題を解決するために、この研究ではSafeCommitという技術を提案しています。SafeCo
この研究では、LLMを用いてマルチハイスティック エンサンブルを進化させる技術を提案するものです。エンサンブルの各ハイスティックは単独で効果的なものではなく、他のハイスティックと協力して最高の結果を達成するものと想定され
この研究では、3D MRIと臨床記録を活用した大規模言語モデルの開発を提唱。提案されたNeuroMosaicは、医療画像を解剖学的情報に基づく地域情報に変換し、臨床記録と分子情報に調整を行い、MRI領域との接続を確実にし
この研究では、記憶を維持し更新する能力を強化するため、繰り返し神経ネットワークを用いて記憶の特性を分析しました。記憶を維持するための神経計算の一種として、ダイショビュールノーマリゼーション(Divisive Normal
Automated formulaic alpha discovery aims to generate predictive and interpretable trading signals from large s
The rapid development of Large Language Models (LLMs) has opened new avenues for Automated Heuristic Design (A
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and co
Automated heuristic design (AHD) with large language models (LLMs) has produced strong heuristics for combinat
In dynamic multi-mode project scheduling, activities have alternative execution modes and uncertain durations,
物理的システムの分離方程式を解くためには、ニュラルネットワークの設計、損失関数の定義、および最適化ダイナミクスの手動調整が必要である。研究者は、自動設計のためにLarge Language Models (LLMs)を利
Parent selection significantly affects exploration, exploitation, and complexity control in genetic programmin
シミュレーション駆動設計では、高精度なシミュレーションを少なくすることで設計を実現しています。既存の手法では、その問題に取り組むために最適化アルゴリズムが改善されてきましたが、問題の定義自体は検討されていません。この論文
Analytical placers rely on differentiable objective functions to guide placement, typically combining intermed