When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay
Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback l
- 用途
- 回帰
- 難易度
- Hard
- コスト
- High
「text」の検索結果
51 件Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback l
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, relia
Test-time adaptation (TTA) has emerged as a prominent strategy for adapting vision-language models to distribu
Table detection is a core task in document analysis, supporting downstream applications such as information re
Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-graine
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, i
Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedu
Long context language models now advertise windows of one million tokens, but two habits limit how much of tha
Automating filament tracing in Cryo-Electron Microscopy (Cryo-EM) is essential for 3D helical reconstruction b
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert util
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
We describe the Snugi-AI-v2 submission to eRisk 2026 Task 2, the second edition of contextualized early depres
Organoids are three-dimensional tissue models whose morphology provides important insights into tumor developm
Pancreatic tumor segmentation in 3D CT volumes is challenged by extreme scale variability across both the panc
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Large language models (LLMs) have recently emerged as a promising tool for automating robot design from high-l
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (L
Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing interm
LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to l
Citation-based measures of scientific influence typically treat citations as uniform signals, ignoring the dif
The application of large language models (LLMs) to personalized medical assistants has garnered growing intere
Taxonomy induction aims to organize concept sets into coherent hierarchical structures. Recent LLM-based metho
Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic.
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their r
Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing ext
We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations
Large language models achieve strong performance across diverse tasks, but deployment remains costly because o
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Spatio-temporal context has become increasingly crucial for visual tracking. However, most existing approaches
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Ensuring effective transfer learning for vision-language models without compromising their generalization perf
Existing SLAM systems lack modeling of the functional relations required for fine-grained robotic interaction.
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Actio
Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism anal
The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need fo
Research teams and organizations often explore unfamiliar free-text collections, from survey comments and revi
Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languag
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness,
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained b
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial inform
Driver behavior is heterogeneous, context-dependent, and changes over time, and these properties shape the tra
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
Population optimizers such as CMA-ES, DE, and multi-objective evolutionary algorithms drive search mainly thro
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously u
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background