UniMate: One Unified Model to Animate Diverse Skeletons
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「Fine-tuning」の検索結果
74 件Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
Cryptocurrency markets exhibit extreme volatility and non-stationary dynamics that challenge conventional fore
Tabular foundation models (TabFMs) achieve strong performance on structured data, particularly for standard cl
Bidirectional discrete diffusion model appears naturally suited to genomic modeling because it can reconstruct
Passive cooling eliminates the energy overhead and mechanical failure modes of fans, making it attractive for
Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets w
Large language models are now trained and evaluated under a diverse set of paradigms: supervised fine-tuning (
Hallucination-where a language model generates outputs that are factually incorrect or unsupported by the sour
Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex
Financial large language models are increasingly deployed for summarization of reports and disclosures, where
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individu
We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge a
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external
Acquiring high quality annotated medical image data is critical for training deep learning models; however, an
Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals
Text-guided diffusion editing raises disinformation concerns, making reliable image provenance essential. Whil
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generatio
We introduce Mitra-v2, a tabular foundation model that delivers state-of-the-art performance on real-world cla
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its
As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-o
RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation,
Pause-token methods improve LLM reasoning by inserting special tokens into sequences. Prior work explains thes
We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from Chin
Latent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities w
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic,
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough:
Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternativ
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full
How do the methods used to train language models to refuse harmful requests shape how that refusal actually wo
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss label
The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializin
Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained
Anatomic tracer studies reveal how axon bundles project from an injection site, branch into smaller groups of
Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigatio
While data-driven 3D shape correspondence estimation has recently seen substantial progress, robust matching u
Real-world image super-resolution (SR) increasingly relies on Diffusion Transformer (DiT) backbones, whose int
Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in e
Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a
Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction c
Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinf
Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benc
Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditi
Learned graph simulators provide an efficient alternative to high-fidelity solvers for granular dynamics. Howe
Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language.
Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficient
Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In d
We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines th
Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates dom
Competitive programming has become a key test of large language model reasoning, with international competitio
World Action Models (WAMs) leverage the capabilities of large-scale pretrained video diffusion models to joint
The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictio
Emerging edge AI workloads increasingly require arithmetic units that can trade computational accuracy for eff
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach
Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to
Quality-diversity (QD) algorithms have been gaining traction in robot learning, where diverse motion primitive
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions acro
Off-road navigation can fail when physical structures induce irrecoverable states such as high-centering or en
In animals such as elephants and octopuses, acquiring non-visual information about an object and physically en
Industrial recommenders give new content initial views through budgeted exploration, then use early performanc
Automated bidding (autobidding) is a core component of modern online advertising systems. Within this componen
Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teach
Discrete diffusion models have become a strong, widely adopted class of generators for sequence data, and stee
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through t
Population optimizers such as CMA-ES, DE, and multi-objective evolutionary algorithms drive search mainly thro
Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for
Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smo
複合商品の販売では、商品を組み合わせることができないことがある。Large-Market Disciplineは、この問題を解決するためのフレームワークを提供し、価格と商品の組み合わせを最適化する。
LLMが外部アクションを取り続けている場合、エージェントのメモリーが古くなったり誤った情報を持ったりする可能性があります。この問題を解決するために、この研究ではSafeCommitという技術を提案しています。SafeCo
この研究では、記憶を維持し更新する能力を強化するため、繰り返し神経ネットワークを用いて記憶の特性を分析しました。記憶を維持するための神経計算の一種として、ダイショビュールノーマリゼーション(Divisive Normal
ディナミカルシステムを化学的なパラフレーズで説明する手法を提案し、システムの
Constrained Optimization Problems are crucial in fields such as engineering, economics, and robotics, where hi