Task-Directed Residual AddUNet:Perfect-Reconstruction Routing for Full-Rate Representations
This paper establishes a perfect-reconstruction (PR) interpretation of AddUNet and its full-rate realization,
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「embedding」の検索結果
63 件This paper establishes a perfect-reconstruction (PR) interpretation of AddUNet and its full-rate realization,
A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named
Recent approaches to 30-day hospital readmission prediction rely on pre-trained language models applied to dis
Molecular property prediction requires representations that generalize from limited labeled data to structural
We study the adaptation of pretrained language models to univariate time-series forecasting through a paramete
Urban region representation learning commonly combines heterogeneous data sources, such as mobility flows, poi
The Gaussian kernel is a widely used similarity measure underlying kernel methods such as kernel PCA and spect
Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient simi
3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answ
Mixed-curvature representation learning seeks to capture rich geometric structures that cannot be adequately m
We introduce document embedding geometry as a quantitative observable of conceptual reorganization and develop
Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive a
Amid the rapid advancement of physical-world intelligence, cloud-edge collaborative large language models (LLM
CounterFactual Examples (CFEs) are a cornerstone of eXplainable Artificial Intelligence (XAI), offering local,
Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro
Universal multimodal embedding (UME) learns unified representations across modalities, enabling a single model
Speaker recognition neural networks learn latent representations (i.e. speaker embeddings) from input utteranc
Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts
Procedural audio has emerged as a viable source for transferable audio representation learning, but its design
Long-text image--text congruence scoring is increasingly important for vision-language systems that must evalu
Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported la
Large language models (LLMs) with human-like performance on linguistic tasks have transformed the study of lan
As large language models become capable translators of classical texts, a key challenge is deciding which outp
Open-vocabulary camouflaged object segmentation (OVCOS) aims to segment unseen camouflaged objects under text
Volume-based multimodal retrieval jointly scores a text query with a candidate's video, audio, and subtitle em
Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disease in which early assessment remai
Unpaired cross-modal distillation transfers grade structure from histopathology into a micro-ultrasound (micro
Light detection and ranging (LiDAR) remains less explored than RGB-D sensing for perceptive legged locomotion,
Although aerial-legged robots offer combined agility and efficiency, controlling high-speed hopping under comp
Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a si
Semantic image communication seeks to preserve task-relevant scene structure under limited channel resources,
Predictive Process Monitoring (PPM) aims at predicting at runtime and as early as possible the future states o
Multi-Hop Knowledge Graph Question Answering (KGQA) tasks require models to assemble relational evidence along
Modern semi-supervised learning (SSL) couples pseudo-label generation and classifier training, using the class
Recent work has shown that a teacher model can transfer a bias to a student through a dataset from which every
Temporal knowledge graph embedding (TKGE) models infer missing facts in knowledge graphs that evolve over time
Robots need touch to manipulate objects safely and reliably, as many properties, such as softness, texture, an
Automated radiology report generation has advanced rapidly in diagnostic accuracy, yet generated reports frequ
Deepfake detection systems often exhibit significant performance degradation when deployed on unseen manipulat
Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult
This work addresses decentralized online Riemannian optimization on Hadamard manifolds. Prior work under geode
Knowledge graph question answering usually assumes that one system can reach the whole graph. In practice, fac
Human motion may be viewed as a combination of action content, style, and body morphology. Existing motion sty
Earth-observation (EO) foundation models have become exceptionally effective at learning se mantic, high-dimen
Personal robots should adapt their co-speech gesture style to a new user without requiring model retraining. W
World models trained with joint-embedding predictive architectures learn compact, structured latent representa
Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallu
We investigate the effect of pretrained T5 model scale and explicit motion features on pose-to-text Indian Sig
Fine-grained robotic manipulation depends on understanding parts, not only whole objects. Existing 3D foundati
Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning
Medical image classification is frequently complicated by transitional categories whose feature distributions
Learning robust representations for time-series signals under noise and distribution shifts remains challengin
The digitalisation of electrical distribution networks has increased the exposure of power-grid infrastructure
Leveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising pa
Calibrating an agent-based model (ABM) is difficult because its objective landscape is stochastic and rugged,
Self-supervised pretraining has transformed language and vision, but its value for molecular graph neural netw
Dynamic networks are being applied in many domains, from social media to logistics systems, each with their ow
This chapter reconstructs the Hopfield network as a physical theory of memory rather than merely an early neur
Most LLM-based automated algorithm design methods optimize a designated component within a human-specified sca
Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity
Finding a representative description of graph entities that captures their structural roles and homophily is a
データドライブの日常生活時間の最適化を促進するために、Quality Diversity を使用して時間と健康の関係を考慮したオプティミゼーションテクニックを提案した。
フリーザントランソーメールの特性化を改善し、非線形のプローブを使用した。