Do Reasoning Representations Help Humans Evaluate LLM Outputs?
Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are
- 用途
- 検出
- 難易度
- Hard
- コスト
- High
「Calibration」の検索結果
62 件Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are
Approximate machine unlearning aims to remove the influence of specific training data from a trained model wit
When an external reference set (an anchor) is used to decompose an LLM-judge panel's error into a quality sign
Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observe
Several geometry-aware approaches to low-rank adaptation have emerged for parameter-efficient fine-tuning of l
To mitigate the scalability bottleneck in the radio access network (RAN) in federated edge learning (FEEL), ov
Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix: they
How far a pushed object slides depends on its mass and friction, which no single image reveals. Pretrained vis
Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. T
Generative AI inverts the typical periodization of literary history: the periodizing tag Victorian can now com
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts req
Partially Relevant Video Retrieval (PRVR) retrieves untrimmed videos when queries describe only short moments.
Arm pose estimation enables applications in fitness, extended reality input, rehabilitation, and life logging.
We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2
Offline multi-agent payoff models are estimated under a logging distribution but used on distributions induced
Shortcut learning denotes the widespread situation in which a classifier exploits spurious correlations rather
Automated prediction of Enzyme Commission (EC) numbers plays a central role in functional annotation and compu
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how loca
Deployment of Large Language Models (LLMs) on memory-constrained edge devices relies heavily on aggressive pos
We convert black-box clinical prediction models for tabular data into standalone nomograms that can be audited
Inference energy per token drives the cost and carbon footprint of deployed transformers. It is dominated by d
LLM-based digital twins promise to reduce repeated human data collection by generating person- specific respon
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs,
Multi-agent federations need governance that answers three questions under adversarial conditions: who partici
Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: d
Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity
Bone-selective digitally reconstructed radiograph (DRR) synthesis depends on high-resolution encoder detail, y
Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, in
We present our solution to the LUMPI track of the UCF UrbanTwin Sim2Real LiDAR Challenge at the 6th DriveX Wor
3D reconstruction typically strives for geometric fidelity or visual plausibility. Radio frequency digital twi
Out-of-distribution (OOD) detection is critical for safe deployment of medical AI systems. Recently, test-time
Target-based LiDAR-camera extrinsic calibration is a prerequisite for multi-sensor fusion in robotics. However
Visual-Inertial (VI) fusion is fundamental to accurate and robust state estimation, where camera and IMU measu
We derive an exact gradient-step representation of the RoPE-softmax forward pass. For every deterministic RoPE
In many scientific disciplines, weak signals of interest are obscured by dominant nuisance signals that are se
Linear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigend
Training strategy, namely whether to retrain from scratch or fine-tune from the previous checkpoint, is an ove
Although professional workflows leverage large language models widely, the interpretation for auditing unconst
Rollcast is a probabilistic forecasting method for univariate time series that combines a compact set of rolli
Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigatio
Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach fo
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite
Generative modeling directly on geometric manifolds can avoid errors introduced by flattening non-Euclidean da
Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their pe
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improv
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated
Many statistical models involve parameter-dependent normalizing constants that are computationally intractable
At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determi
Language model agents increasingly propose actions, observe external feedback, and explain their own behavior.
We develop a marginal coordinate test for regression with Euclidean predictors and a random-object response in
Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is i
Missing data, measurement error, and population heterogeneity are pervasive challenges in analyzing data arisi
We study risk-averse decision making, in which an agent selects actions while being uncertain about the true s
Bayesian inference in compound loss models must often be repeated across policies, market scenarios, and prior
Reliable prediction of time-varying channel state information (CSI) is essential for efficient wireless commun
Recommendation impressions are a finite resource, hence delivering a recommendation to a user who would discov
Calibration requires probabilistic reports to be conditionally unbiased and reliably interpretable as probabil
As an extension of existing Bayesian persuasion framework with inadequate message mechanism, we study direct r
We show with experiments and system-level simulations that it is possible to successfully mitigate the impact
Echo-state networks enable efficient temporal learning by fixing the recurrent dynamics and training only a li