screenpipe — YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)
ユーザーの行動を認識し、オートエージェントを構築するためのツール。
- 用途
- オートエージェント構築
- 難易度
- Easy
- コスト
- High
「audio」の検索結果
77 件ユーザーの行動を認識し、オートエージェントを構築するためのツール。
Unsloth Studioは、オープンモデルのトレーニングと実行を支援するWebUIです。このライブラリは、Gemma4、Qwen3.5などのオープンモデルのテストとトレーニングを支援するために使われます。
🤗 Transformersは、テキスト・ビジョン・音声など複雑なモデル定義をサポートするフレームワークで、インフェレンスターやトレーニングに使用できる。
mediapipeは、クロスプラットフォームでカスタマイズ可能なライブおよびストリーミングメディア向けのMLソリューションを提供している。
.diffusion モデルのライブラリ。画像・動画・音声生成に利用可能。
AI用のデータセットを提供するプラットフォームです。
オープンソースのAI推論最適化と展開用ツールキットです。
統計チャートの生成は、タブラーのデータから生成することが難しい。新しい作成フローでは、データのスクリーン、プロット提案、コード生成、レンダリング、検証による改良が含まれる。
電気生理信号から表現を学習し、脳コンピューターインターフェースの開発を支援する。
この研究では、自然言語処理の負担を減らすモジュラリティを目指しています。モジュラリティとは、システムを小さくて独立した部分に分割して、それぞれを簡素化することです。この研究では、文脈に応じてモジュラリティを変更できるメカ
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and a
Dense self-attention treats all token pairs as equally plausible before learning, an interaction-isotropic pri
Text-to-speech systems often face a trade-off between natural prosody and efficient inference: higher perceptu
Machine learning surrogates based on neural operators have shown broad applicability in solving forward PDE pr
Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perceptio
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherenc
Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model'
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a commo
Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. T
Large Language Models are increasingly deployed as information intermediaries, yet measuring their political b
Political texts are rarely authored by the nominal speaker alone. Tweets, speeches, reports, and official stat
Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the so
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backcha
Messages from electronic devices are conventionally received as text, audio, or radio signals. But robots move
This paper presents a soft robotic drummer for accurate and efficient drum rolls. High-frequency drum rolls re
デベロッパー向けのモデロプティミゼーションフレームワークです。モデルの高速化と効率化を実現することができます。
テキスト分析、センチメント分析や単語分割などを行えるライブラリ。
ModelScopeは、モデルをサービス化するためのプラットフォームです。モデルを作成し、ホスティングし、管理し、配信することができます。
Physics-Informed Neural Networks (PINNs) have recently emerged as a promising approach for solving Partial Dif
Audio provenance attribution - which system produced a synthetic utterance - is reported at near-ceiling accur
Deep learning is an efficient technique to monitor the real time welding process, reducing post-welding repair
Fibre optic sensing, such as distributed acoustic sensing (DAS), has become a widespread technology for geophy
Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cro
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (
This paper presents a small-scale quantitative experiment that links syntactic structure to stylistic function
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three
Language models compute over tokens: language is their input, their output, and increasingly their internal re
Large audio-language models (LALMs) can hallucinate audio objects, answering "yes" to absent sound events, thu
Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversat
Voice phishing detection faces three critical challenges: real criminal recordings are unavailable due to priv
Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a si
This paper proposes an open-set ego-noise separation framework for legged-robot audition via annotation-free a
画像やビデオやオーディオディフュージョンモデルのファインチューニングを行うための、汎用的なファインチューニングキット。
Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failin
Can people distinguish between human and AI agency in humanoid teleoperation? To explore this question, we dev
Principal component analysis (PCA) can rotate away from its population target when a covariance matrix is esti
A sound runtime admission gate executes only actions it can certify, and certifies only what its observations
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to esti
Cooperative rehabilitation enhances engagement, task performance, and social-motor interaction, yet it demands
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions
Temporal integration gives continuous-time recurrent networks memory, but in deep stacks it also delays bottom
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stoc
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path toward
マルチラギングスピーチ生成やクリエイティブボイスデザイン、ルートライフクライミングなど、テクスチャファリーTTSの最新技術を実現するためのフレームワークです。
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial inform
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ g
Obtaining data from neuromorphic sensors and processing it with Spiking Neural Networks is a promising solutio
Matcha-TTSは、高速で条件付き流のマッチングを実現するTTSアーキテクチャであり、話者の特徴を考慮する。
Cochlear implants (CI) restore hearing for individuals with severe to profound hearing loss. However, CI users
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities
Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start reg
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through t
Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, invers
Wunjo CE: Face Swap, Lip Sync, Control Remove Objects & Text & Background, Restyling, Audio Separator, Clone V
Drench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learni
Calibration requires probabilistic reports to be conditionally unbiased and reliably interpretable as probabil
Persistent acoustic monitoring can detect machine faults without physical contact, but always-on inference is
画像生成のためのHigh Quality Training Free Inpaintを提供します。このInpaintはStable Diffusionモデルに使用でき、ComfyUIもサポートしています。
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev
この研究では、人工知能の研究者と神経科学者の間の分野を結びつけるために、脳のシステム構造を研究し、その研究から導かれた新しいアプローチを提案しました。