Sound-based Multi-Person 3D Pose Estimation
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to esti
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「audio」の検索結果
52 件Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to esti
Multimodal emotion recognition has attracted growing interest due to its importance in human-computer interact
Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the
As generative audio models grow in complexity, the computational and ecological costs of synthesizing everyday
A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is follo
Diffusion TV is an interactive AI art installation that offers a tangible and embodied experience of diffusion
LLM agent systems increasingly combine provenance tracking, authorization, policy enforcement, protocol adapte
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio
While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing
Language models are trained on tokenized text that obscures the sound structure of words, yet they reliably pr
We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. U
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Cooperative rehabilitation enhances engagement, task performance, and social-motor interaction, yet it demands
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stoc
Temporal integration gives continuous-time recurrent networks memory, but in deep stacks it also delays bottom
Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) throu
Whisper is a widely used foundation model for automatic speech recognition (ASR), but its generative decoder c
Recent work on controllable music generation has focused on autoregressive models, leaving diffusion-based sys
Cardiac, neural, behavioral, and speech measurements from wearable and mobile devices provide partial, noise-s
Decompensation represents a critical transition in the course of cirrhosis, yet clinicians have limited non-in
Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an i
Modern misinformation is often heard before it is read, yet fact-checking systems are still evaluated mainly o
ASR systems sometimes produce fluent text that is unrelated to the speech they receive. We view these hallucin
Music audio-language models are evaluated almost entirely by accuracy on multiple-choice questions. This proto
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full
In numerical signal processing for electroacoustic composition, the progressive loss of specific development a
LETHE (Latent-parameter Evolution with Temporal Hierarchical quasi-Equilibrium) is a self-referential sonic-ob
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with c
Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existi
The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializin
Video-language benchmarks are usually constructed by the dataset authors without published reliability statist
Metabolic dysfunction-associated steatotic liver disease (MASLD) affects approximately 30% of the general popu
Robots operating in physical environments make control decisions based on uncertain sensor measurements, which
Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In d
Monolithic world models predict the entire next state at every step, spending capacity re-predicting the stati
Fully-powered knee prostheses, unlike traditional passive knees, can perform controlled positive work, reducin
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial inform
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ g
Road traffic noise remains a major environmental challenge, yet most speed management strategies are static an
For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autono
Obtaining data from neuromorphic sensors and processing it with Spiking Neural Networks is a promising solutio
Cochlear implants (CI) restore hearing for individuals with severe to profound hearing loss. However, CI users
Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start reg
Prior analyses by Derezinski and Warmuth established all-size sampling identities, selected-OLS unbiasedness,
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through t
Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, invers
Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such
A rapidly growing range of sequential data tasks, such as identifying trend reversals in financial markets, au
Calibration requires probabilistic reports to be conditionally unbiased and reliably interpretable as probabil
Persistent acoustic monitoring can detect machine faults without physical contact, but always-on inference is
この研究では、人工知能の研究者と神経科学者の間の分野を結びつけるために、脳のシステム構造を研究し、その研究から導かれた新しいアプローチを提案しました。