screenpipe — YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)
ユーザーの行動を認識し、オートエージェントを構築するためのツール。
- 用途
- オートエージェント構築
- 難易度
- Easy
- コスト
- High
「multimodal」の検索結果
165 件ユーザーの行動を認識し、オートエージェントを構築するためのツール。
SGLangは、大規模言語モデルのサービングフレームワークです。このライブラリは、高性能なサービスフレームワークで、大規模言語モデルのサービングをサポートしています。
マルチモーダルAIに適したオープンレイクハウスフォーマットです。このフォーマットでは、パレットからデータを2行のコードで変換することができ、100倍速くなります。また、ベクトルインデックスやデータバージョニングが可能です
🤗 Transformersは、テキスト・ビジョン・音声など複雑なモデル定義をサポートするフレームワークで、インフェレンスターやトレーニングに使用できる。
データをロギング・ストーリング・クエリして視覚化できるSDKです。
この論文では、現在のVision-Language-Benchmark(VLB)を超える、MLLMがアクティブな観察を実演できるようにするためのバenchmark、ActiveVisionを提案する。このActiveVi
xtunerは、超大規模MoEモデルを高速にトレーニングするためのトレーニングエンジンです。
統計チャートの生成は、タブラーのデータから生成することが難しい。新しい作成フローでは、データのスクリーン、プロット提案、コード生成、レンダリング、検証による改良が含まれる。
CVV または CWE への分類を実現し、バグ修正のために重要な手順となるCVEへの CWE 分類を自動化する。
The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetim
Universal multimodal embedding (UME) increasingly demands encoder's capacity for handling a broad range of tas
Test-time adaptation (TTA) has emerged as a prominent strategy for adapting vision-language models to distribu
A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it stil
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional me
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual,
Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing sig
Combining search with function approximation has driven major advances in game-playing programs, making self-p
Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D. This embeds points in a
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual ob
In this study, we identify depth-dependent prefix redundancy in final-readout LLM embedding models, notably ac
Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmann
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perceptio
Moving-object perception must decide which image regions correspond to real motion and keep every instance ide
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherenc
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Learning direct current circuit concepts requires learners to connect invisible physical quantities, such as c
Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to exter
Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described targ
How can vision-language-action (VLA) models adapt to new environments where world dynamics shift? While recent
Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate locomot
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a commo
The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of re
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation mode
Object 6D pose estimation formulations have progressively reduced reliance on object-specific priors, evolving
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimo
Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based
Preoperative evaluation of trigeminal neuralgia (TN) often requires joint interpretation of structural MRI, wh
Despite recent advances in surgical vision-language models (VLMs), temporal reasoning remains limited because
Out-of-distribution (OOD) detection is crucial for safe deployment of medical AI systems, where domain shifts
World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance w
Objective assessment of Freezing of Gait (FoG) in Parkinson's disease (PD) relies predominantly on wearable In
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Unified multimodal models integrate visual understanding and generation within a single network, yet the two c
Driver alerting from dashcam video requires sequential decision-making under partial observability: a system m
An embodied assistant working beside a person must track task state, recognize help seeking, choose how to int
Autonomous navigation in unstructured environments requires robust scene understanding, yet Vision-Language Mo
Large language models (LLMs) and vision-language models (VLMs) have significantly advanced zero-shot task plan
このリポジトリでは、AIモデルの互換性を確保するためのオープンスタンダードであるONNXを提供しています。
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how loca
Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data r
Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its ver
Vision-language models are increasingly used in settings where some input modalities may be unavailable, yet w
How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, w
Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cro
Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely interve
Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforceme
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability
Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text
Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world
Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason
Language models compute over tokens: language is their input, their output, and increasingly their internal re
Long-running LLM agents rely on external memory to store and reuse information beyond a single context window,
Converting in-service reinforced-concrete (RC) building blueprints into simulation-ready models---structured f
Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial r
Automated drone surveillance has become increasingly important for public safety, critical infrastructure prot
Long-form narrative-to-film generation requires shot-level controllability and cross-clip consistency in both
Contrastive learning approaches achieve strong performance by training models to bring similar samples closer
Training-free collaborative pipelines that integrate Vision Foundation Models such as CLIP, SAM, and DINO achi
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
Privacy-sensitive surveillance systems could benefit from large vision-language models (VLMs), but such models
Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans,
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produc
3D reconstruction typically strives for geometric fidelity or visual plausibility. Radio frequency digital twi
Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust mul
Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to obje
The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional co
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelin
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Natural adversarial examples (NAEs) reveal that vision models can fail under realistic semantic changes beyond
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Cranial nerves (CNs) play essential roles in sensory, motor, and autonomic functions. Accurate CN parcellation
Behavior policies are often formulated as continuous generative models, whose iterative denoising processes ar
Vision-language models (VLMs) have recently shown excellent progress in open-ended image-to-text generation. H
Multimodal sentiment analysis often remains text-dominant due to raw-video noise and insufficient temporal mod
Multimodal continual learning has recently shown great potential for developing agents with human-like intelli
Compared with relying solely on initial observations and language instructions, predicting goal images with ge
Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a si
Progress in video anomaly understanding (VAU) has long been limited by inherent deficiencies of real-world ano
Ensuring effective transfer learning for vision-language models without compromising their generalization perf
Deep learning models have been increasingly applied to Time Series Forecasting (TSF) in recent years. Transfor
Executing contact-rich tasks efficiently requires the seamless integration of whole-body coordination and phys
Connected robotics is an emerging 6G application where mobile robots follow natural-language instructions to m
Vision-Language-Action (VLA) policies are commonly adapted to new manipulation settings through additional gra
Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable inte
Under ubiquitous teleoperation environments with optically challenging conditions, an interface for tele-opera
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do
Sparse outcome feedback limits what robots can learn from unsuccessful attempts at complex manipulation. Faile
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Actio
Generalizable and robust dexterous in-hand manipulation requires a policy to infer object pose, geometry, cont
ゼネレーティブAIに関連するリソースの一覧。
モデルをサービングするためのライブラリを紹介している。
AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior w
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than expl
Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by
Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failin
Although highly effective in vision and language domains, applying in-context learning to robotics remains cha
Can exploratory UAV waypoint sequences be generated from multimodal onboard observations and a fixed-dimension
Long-horizon robot manipulation with Vision-Language-Action (VLA) policies remains vulnerable to execution-tim
Vision-language-action (VLA) models have shown promising generalization for language-conditioned robot manipul
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge f
Vision-language-action (VLA) models have substantially advanced language-guided robot manipulation, yet reliab
Vision-language-action (VLA) models adapted through supervised fine-tuning (SFT) inherit a structural asymmetr
Robots that follow open-ended language instructions need to connect semantic intent to visual scene understand
Integrating visuomotor policies or Vision-Language-Action (VLA) models with force/torque (F/T) perception has
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
Can interactive vision-and-language agents learn not just what to say but also \textbf{\textit{when}} to say i
Vision-language models are increasingly used as reward functions for robotic learning, but this role requires
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail wh
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Reliable robot-to-human handover requires the robot to infer when the person is ready to receive the object, a
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in
Vision-language-action (VLA) models are trained by imitation and capture what action to take but not why; addi
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounde
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with hu
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demand
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynam
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generaliza
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrai
Quadruped robots have demonstrated impressive agility in parkour locomotion across complex terrains. However,
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language in
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial sim
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ g
Vision-language-action (VLA) models map visual observations and language instructions directly to robot action
分析システムの性能を向上するための学習モデル開発を行う。
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform
Rust言語でCandleライブラリを利用して、PythonやPyTorchを使用せずにDecoder-only LLMを自作した。
GUI操作自動化に伴う停止判定、復讐、再検索に関する問題を解決し、 GUI操作自動化を実現するためのフレームワークを開発します。
VERDICT は、マルチモーダル論理の確認と検証をサポートするフレームワークである。このフレームワークでは、多くの場合、確認と検証は、エージェントが提供するさまざまなスコアを合計すると考えられてきたが、このフレームワー
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev
この研究では、3D MRIと臨床記録を活用した大規模言語モデルの開発を提唱。提案されたNeuroMosaicは、医療画像を解剖学的情報に基づく地域情報に変換し、臨床記録と分子情報に調整を行い、MRI領域との接続を確実にし
Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimoda
In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language M