UniMate: One Unified Model to Animate Diverse Skeletons
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「3d」の検索結果
102 件Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
AI-driven de novo molecular design offers a promising route to accelerate early-stage drug discovery by genera
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to esti
Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR uns
3D perception plays a crucial role in real-world applications such as autonomous driving, robotics, and AR/VR.
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds
When there is not enough labeled data to properly train deep learning models, transfer learning can help. We s
Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-v
Explicit primitive-based radiance fields such as 3D Gaussian Splatting typically model view-dependent appearan
Recent hybrid Structure-from-Motion (SfM) systems combine the robustness of feed-forward 3D reconstruction wit
Semantic 3D maps are increasingly constructed automatically for aerial robotics by integrating learned semanti
In the field of computer vision and graphics, high-quality reconstruction of the human body in static scenes h
Cross-view localization (CVL) estimates the pose of a ground image by matching it to a geo-referenced satellit
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. U
Weakly supervised 3D occupancy prediction reduces the reliance on costly 3D annotations by learning from 2D ps
Monocular depth estimation foundation models, such as the Depth Anything series, have achieved remarkable perf
Multi-modal medical images and clinical reports provide complementary anatomical, functional, and semantic inf
3D surface cutting and UV unwrapping are fundamental problems in computer graphics. Traditional geometric opti
To enable accurate and rapid photon control point and proton beamlet dose calculation in the DoseRAD2026 chall
Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D
Structure-from-Motion (SfM) is a fundamental tool for sparse 3D reconstruction with broad impact in robotics a
Recent generative image-editing Diffusion Transformers (DiTs) demonstrate impressive semantic editing capabili
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Achieving robust SLAM in large-scale underground coal mines with complex structures and severe degeneracies re
Reliable autonomous mapping, environmental sampling, last-mile logistics, and infrastructure deployment depend
Robotic dressing assistance is a promising solution for supporting older adults with physical impairments in d
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounde
Fixed 3D Gaussian Splatting (3DGS) reconstructions provide realistic novel views but lack the traversability c
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a larg
The geometric morphology of deposited filaments can significantly influence the structural performance and sta
Medical image inpainting has the potential to improve automated brain MRI analysis by reconstructing healthy t
Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which r
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with hu
Digital materials fabricated by multi-material 3D printing are designed as controlled mixtures of stiff and co
Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss label
Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD be
Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained
Automated aortic segmentation in 4D flow MRI is essential for reproducible hemodynamic assessment but is limit
Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigatio
Object-centric visual representations are important for physical-world perception, but existing visual pretrai
While data-driven 3D shape correspondence estimation has recently seen substantial progress, robust matching u
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial sim
3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent meth
Bundle Adjustment (BA) is a cornerstone of 3D computer vision and has benefited from decades of advances in sp
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-co
Dust in agriculture presents a significant challenge for autonomous agricultural machinery. Dust can impair th
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in cu
3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene,
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Implicit neural representations (INRs) can model continuous 3D shapes with a shared coordinate decoder and per
The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in
3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view set
Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion mo
3D Gaussian Splatting has become a de facto scene representation for novel view synthesis, yet robustly learni
Large-scale 3D surface reconstruction from aerial imagery is fundamental to geospatial mapping and urban model
Automatic generation of hand gestures is essential for the transmission of Indian classical dance and critical
Advances in neural rendering have enabled high-fidelity multi-view reconstruction of 3D scenes. However, free-
We present PointGT, a point-based 3D representation that enables simultaneous editing of object geometry and a
A key bottleneck in 3D Gaussian Splatting training is the continual growth of Gaussian primitives, which incre
Accurate identification of weld seam geometries is essential for automated robotic post processing operations
Mobile robotic platforms offer a flexible alternative to fixed manipulators for non-destructive evaluation (ND
This paper presents an end-to-end computational pipeline that converts a selected object mesh, a measured obje
Autonomous navigation in space requires reliable terrain assessment for safe operations, especially in undergr
Autonomous unmanned aerial vehicles (UAVs) increasingly operate in cluttered environments where global planner
QLAUN Bot (Quad-Legged Adaptive Unmanned Navigator Robot) is a torque-controlled quadruped robot that is resea
Existing radar-LiDAR fusion methods rely on fixed residual weights, even though the informativeness of radar D
Air-ground collaborative Vision-and-Language Navigation (VLN) pairs an unmanned aerial vehicle (UAV) with a gl
Learned graph simulators provide an efficient alternative to high-fidelity solvers for granular dynamics. Howe
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained b
Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains exp
LiDAR-based semantic segmentation is a core perception module for autonomous vehicles and mobile robots. Despi
Large-scale training and refined optimization techniques have greatly improved sparse multi-view 3D reconstruc
World Action Models (WAMs) leverage the capabilities of large-scale pretrained video diffusion models to joint
Multi-axis 3D printing enables support-free fabrication and improved part quality, but robustly processing rea
Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen
Accurate joint encoder offsets are essential for kinematic consistency in humanoid lower limbs, yet existing c
Humanoid learning increasingly relies on transforming vast and diverse human motion data into high-quality rob
Robotic processing of irregular steel scrap requires dense 3-D measurement to replace manual visual assessment
Executing long-term tasks in dynamic environments requires embodied agents to maintain robust and adaptive 3D
Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to
High-precision insertion remains a fundamental challenge in robotic manipulation due to the strict alignment r
Conventional soft robot actuators excel in compliance, but their uncontrolled deformations compromise accuracy
Cooperative perception allows a drone fleet to combine observations from multiple viewpoints. However, existin
In indoor environments, object positions frequently change due to human activities or embodied-agent interacti
Language-driven dexterous grasp models, such as DextER, perform well when instructions specify where to grasp,
Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibilit
Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations
Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed
Tactile sensing is essential for physical interaction in robotics and human--machine systems. However, combini
Most shape reconstruction methods assume measurements defined over planar sensing domains, such as RGB images
Residual maximin share (RMMS) is the largest share threshold that remains guaranteeable throughout dynamic all
Point-cloud data routinely captured by modern imaging and sensor technologies provide detailed geometric descr
The simple exclusion process (SEP) is a paradigmatic model for nonequilibrium transport, yet the rich dynamics
Gromov-Wasserstein (GW) compares distributions through relations within each space. This pointwise comparison
Training controllers that are safe and robust in simulation, and systematically assessing their readiness for
Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distin
この研究では、3D MRIと臨床記録を活用した大規模言語モデルの開発を提唱。提案されたNeuroMosaicは、医療画像を解剖学的情報に基づく地域情報に変換し、臨床記録と分子情報に調整を行い、MRI領域との接続を確実にし