Learning to build covering structures with continuous adjustments
Robotic construction offers the potential to use materials more efficiently and create complex geometries, but
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「3d」の検索結果
108 件Robotic construction offers the potential to use materials more efficiently and create complex geometries, but
In proton therapy, plans are typically optimized on a single planning CT, making robustness evaluation essenti
Class disentanglement (the separation of a representation's class-conditional point clouds along depth and ove
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-trainin
In molecular discovery, molecule size is coupled to composition, structure, and other target properties. Yet m
Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposin
Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applicat
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional me
Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D. This embeds points in a
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic proc
Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and h
Object detection and tracking are fundamental components of perception systems for autonomous driving. Achievi
Automating filament tracing in Cryo-Electron Microscopy (Cryo-EM) is essential for 3D helical reconstruction b
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Because dense frame-level annotation of colonoscopy videos is costly, we propose WSPolypNet, a weakly supervis
We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is abl
Autonomous 3D active mapping requires a space robot to choose where to sense while building the geometry neede
Spherical observations provide global visual context for 3D scene understanding. However, visual information i
Three-dimensional femoral reconstruction from radiographs supports surgical planning, implant sizing, and post
We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation mode
Object 6D pose estimation formulations have progressively reduced reliance on object-specific priors, evolving
We present FIRE3D, a unified framework that takes a single RGB image or casual RGB video and transforms it int
High-fidelity vehicle assets are essential for controllable traffic scene generation, particularly for synthes
Person identification from millimeter-wave (mmWave) point clouds has mainly relied on gait. Indoor walking, ho
While 3D Gaussian Splatting (3DGS) has emerged as a powerful representation for real-time novel view synthesis
A simultaneous localization and mapping (SLAM) method using a monocular camera and a low-cost inertial measure
We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2
Observing objects grasped by a robot hand is challenging due to severe visual occlusions. Although in-hand man
Gaussian splats provide a fast, high-fidelity representation for 3D objects but are often constructed from inc
Floorplans are compact, appearance-invariant maps ideal for indoor localization, yet existing methods rely on
Text-to-motion (T2M) generation maps natural language to human joint movements, aiding gaming, VR, and robotic
We present EdMCGS (Event-driven Markov chain Gaussian Splatting), an end-to-end method for reconstructing dyna
Pancreatic tumor segmentation in 3D CT volumes is challenged by extreme scale variability across both the panc
Video generation models have recently attracted substantial attention for their ability to generate visually c
Reconstructing a three-dimensional left-ventricular (LV) endocardial surface from cardiac magnetic resonance (
Tensegrity robots offer lightweight, compliant mobility over challenging terrain but remain difficult to model
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative nav
We present DCLP++, a local navigation frameworkthat uses footprint clearance as the geometric basis for studyi
Long-horizon target navigation requires a robot to sustain task execution across evolving observations, decisi
Many pedestrian trajectory prediction algorithms have been proposed to improve the safety of navigation for mo
Bringing multiscale geometric analysis directly to irregular point clouds remains difficult: quantities such a
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial c
3D Gaussian Splatting has recently revolutionised novel view synthesis as well as many other 3D vision methods
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed
Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely interve
LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to l
Converting in-service reinforced-concrete (RC) building blueprints into simulation-ready models---structured f
Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial r
We address the problem of discovering repeated elements from a single image. In contrast to existing approache
Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to
Gaussian Splatting has been effective in inferring scene representations that excel in novel view synthesis. M
Accurate 3D plant organ segmentation is fundamental to automated phenotyping. Existing approaches rely on anno
Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, in
Agentic systems can interpret user requests, search the live web, and use external tools, but their ability to
We present our solution to the LUMPI track of the UCF UrbanTwin Sim2Real LiDAR Challenge at the 6th DriveX Wor
We present a self-supervised approach, Self-MVGTE, for estimating 3D gaze targets from multiple camera views.
Privacy-sensitive surveillance systems could benefit from large vision-language models (VLMs), but such models
Radiance field representations such as 3D Gaussian Splatting (3DGS) enable high-quality novel view synthesis b
3D reconstruction typically strives for geometric fidelity or visual plausibility. Radio frequency digital twi
SLAM systems based on 3D Gaussian Splatting (3DGS) have recently demonstrated promising reconstruction accurac
Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust mul
Deep learning models for CT scan analysis are often limited by the scarcity of precise pixel-level annotations
Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fide
Single-view 3D reconstruction, also known as image-to-3D, is a persistently challenging task due to the extrem
For learning generalizable motion representations from large-scale unlabeled data, Self-supervised learning (S
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
3D point cloud semantic segmentation is essential for real-world spatial understanding, yet the prohibitive co
Predicting hand--object interaction fields requires locating the nearest object-surface point for each hand jo
Synthesizing photorealistic driving videos along specified trajectories is essential for scalable closed-loop
Learning physically plausible dynamics from visual observations is essential for interactive world models and
Target-based LiDAR-camera extrinsic calibration is a prerequisite for multi-sensor fusion in robotics. However
Existing SLAM systems lack modeling of the functional relations required for fine-grained robotic interaction.
Estimating the 6D pose of textureless objects without prior CAD models remains a critical challenge due to the
Planning six-degree-of-freedom (6-DoF) grasps for unseen objects in cluttered tabletop scenes from a single-vi
Autonomous mobile robots performing person-following tasks often suffer from temporary occlusions and sensor t
Active mapping requires a robot to select camera viewpoints that efficiently reconstruct an unknown 3D scene.
Can exploratory UAV waypoint sequences be generated from multimodal onboard observations and a fixed-dimension
Behavior-cloned visuomotor policies can remain accurate near their training distribution yet fail when object
Developing unified physics-based humanoid controllers that can navigate complex 3D scenes and manipulate objec
Real-time 3D mapping is fundamental for autonomous robotic navigation, with Euclidean Signed Distance Fields (
Learning-based manipulation requires supervision that is both semantically meaningful and physically executabl
This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments us
Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios,
Human demonstrations contain rich manipulation knowledge, but it remains unclear what information can be trans
Reconstructing continuous terrain manifolds from massive, unstructured airborne LiDAR point clouds remains cha
Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-v
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Achieving robust SLAM in large-scale underground coal mines with complex structures and severe degeneracies re
Reliable autonomous mapping, environmental sampling, last-mile logistics, and infrastructure deployment depend
Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to esti
Robotic dressing assistance is a promising solution for supporting older adults with physical impairments in d
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounde
Fixed 3D Gaussian Splatting (3DGS) reconstructions provide realistic novel views but lack the traversability c
The autonomous localization of fugitive gas emissions using small Unmanned Aircraft Systems (sUAS) constitutes
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with hu
Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigatio
Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service
Accurate identification of weld seam geometries is essential for automated robotic post processing operations
Learned graph simulators provide an efficient alternative to high-fidelity solvers for granular dynamics. Howe
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained b
Residual maximin share (RMMS) is the largest share threshold that remains guaranteeable throughout dynamic all
Point-cloud data routinely captured by modern imaging and sensor technologies provide detailed geometric descr
The simple exclusion process (SEP) is a paradigmatic model for nonequilibrium transport, yet the rich dynamics
Training controllers that are safe and robust in simulation, and systematically assessing their readiness for
Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distin
この研究では、3D MRIと臨床記録を活用した大規模言語モデルの開発を提唱。提案されたNeuroMosaicは、医療画像を解剖学的情報に基づく地域情報に変換し、臨床記録と分子情報に調整を行い、MRI領域との接続を確実にし