ExecCritic: Learn to Test, Test to Improve for Coding Agents
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「LLM」の検索結果
25 件Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Large language models (LLMs) have recently emerged as a promising tool for automating robot design from high-l
The software supply chain has become an increasingly exposed attack surface because of its reliance on intrica
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (L
Citation-based measures of scientific influence typically treat citations as uniform signals, ignoring the dif
The application of large language models (LLMs) to personalized medical assistants has garnered growing intere
Taxonomy induction aims to organize concept sets into coherent hierarchical structures. Recent LLM-based metho
Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic.
This paper presents our system for CCL2026-Eval Task 5: Minor-Grain Breeding Information Extraction (MGBIE), w
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their r
Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing ext
We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations
Large language models achieve strong performance across diverse tasks, but deployment remains costly because o
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Steering vectors have rapidly emerged as a popular and effective method for guiding the output of LLMs in very
The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need fo
Research teams and organizations often explore unfamiliar free-text collections, from survey comments and revi
Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languag
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
Most LLM-based automated algorithm design methods optimize a designated component within a human-specified sca
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously u