screenpipe — YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)
ユーザーの行動を認識し、オートエージェントを構築するためのツール。
- 用途
- オートエージェント構築
- 難易度
- Easy
- コスト
- High
「multimodal」の検索結果
47 件ユーザーの行動を認識し、オートエージェントを構築するためのツール。
SGLangは、大規模言語モデルのサービングフレームワークです。このライブラリは、高性能なサービスフレームワークで、大規模言語モデルのサービングをサポートしています。
マルチモーダルAIに適したオープンレイクハウスフォーマットです。このフォーマットでは、パレットからデータを2行のコードで変換することができ、100倍速くなります。また、ベクトルインデックスやデータバージョニングが可能です
🤗 Transformersは、テキスト・ビジョン・音声など複雑なモデル定義をサポートするフレームワークで、インフェレンスターやトレーニングに使用できる。
データをロギング・ストーリング・クエリして視覚化できるSDKです。
この論文では、現在のVision-Language-Benchmark(VLB)を超える、MLLMがアクティブな観察を実演できるようにするためのバenchmark、ActiveVisionを提案する。このActiveVi
xtunerは、超大規模MoEモデルを高速にトレーニングするためのトレーニングエンジンです。
統計チャートの生成は、タブラーのデータから生成することが難しい。新しい作成フローでは、データのスクリーン、プロット提案、コード生成、レンダリング、検証による改良が含まれる。
CVV または CWE への分類を実現し、バグ修正のために重要な手順となるCVEへの CWE 分類を自動化する。
Test-time adaptation (TTA) has emerged as a prominent strategy for adapting vision-language models to distribu
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherenc
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
Preoperative evaluation of trigeminal neuralgia (TN) often requires joint interpretation of structural MRI, wh
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
このリポジトリでは、AIモデルの互換性を確保するためのオープンスタンダードであるONNXを提供しています。
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Ensuring effective transfer learning for vision-language models without compromising their generalization perf
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Actio
ゼネレーティブAIに関連するリソースの一覧。
モデルをサービングするためのライブラリを紹介している。
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial sim
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Vision-language-action (VLA) models map visual observations and language instructions directly to robot action
分析システムの性能を向上するための学習モデル開発を行う。
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform
Rust言語でCandleライブラリを利用して、PythonやPyTorchを使用せずにDecoder-only LLMを自作した。
GUI操作自動化に伴う停止判定、復讐、再検索に関する問題を解決し、 GUI操作自動化を実現するためのフレームワークを開発します。
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev