label-studio — Label Studio is a multi-type data labeling and annotation tool with standardized output format
データラベル化と注釈化を行うためのツールです。
- 用途
- データラベル化ツール
- 難易度
- Easy
- コスト
- Low
「text」の検索結果
144 件データラベル化と注釈化を行うためのツールです。
マシンラーニングシステムの理論と実装に関する本。
ユーザーの行動を認識し、オートエージェントを構築するためのツール。
Unsloth Studioは、オープンモデルのトレーニングと実行を支援するWebUIです。このライブラリは、Gemma4、Qwen3.5などのオープンモデルのテストとトレーニングを支援するために使われます。
SGLangは、大規模言語モデルのサービングフレームワークです。このライブラリは、高性能なサービスフレームワークで、大規模言語モデルのサービングをサポートしています。
ドキュメントを構造化するために使えるオープンソースのETLソリューション。
🤗 Transformersは、テキスト・ビジョン・音声など複雑なモデル定義をサポートするフレームワークで、インフェレンスターやトレーニングに使用できる。
paperless-ngxは、コミュニティによってサポートされたスーパーチャージドのドキュメント管理システムで、ドキュメントのスキャン・インデックス・アーカイブが可能である。
.diffusion モデルのライブラリ。画像・動画・音声生成に利用可能。
大規模言語モデルのテスト時間調整に関する調査のリポジトリ。
ノードベースのビジュアルプログラミングツールです。
この論文では、現在のVision-Language-Benchmark(VLB)を超える、MLLMがアクティブな観察を実演できるようにするためのバenchmark、ActiveVisionを提案する。このActiveVi
SANAは、高解像度画像生成モデルSANAを紹介する本研究であり、低計算コストで優れた高解像度画像を生成できる。
統計チャートの生成は、タブラーのデータから生成することが難しい。新しい作成フローでは、データのスクリーン、プロット提案、コード生成、レンダリング、検証による改良が含まれる。
電気生理信号から表現を学習し、脳コンピューターインターフェースの開発を支援する。
このリポジトリでは、トークナイザーの最適化を提供しています。
presidioは、テキスト、画像、構造化データを含む敏感データを検出、削除、マスク、アノニマイズするオープンソースフレームワークです。自然言語処理、パターンマッチング、カスタマイズ可能なパイプラインをサポートします。
Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback l
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, relia
Test-time adaptation (TTA) has emerged as a prominent strategy for adapting vision-language models to distribu
Table detection is a core task in document analysis, supporting downstream applications such as information re
Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-graine
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, i
Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedu
Long context language models now advertise windows of one million tokens, but two habits limit how much of tha
Automating filament tracing in Cryo-Electron Microscopy (Cryo-EM) is essential for 3D helical reconstruction b
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert util
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
We describe the Snugi-AI-v2 submission to eRisk 2026 Task 2, the second edition of contextualized early depres
Organoids are three-dimensional tissue models whose morphology provides important insights into tumor developm
Pancreatic tumor segmentation in 3D CT volumes is challenged by extreme scale variability across both the panc
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Large language models (LLMs) have recently emerged as a promising tool for automating robot design from high-l
このリポジトリでは、Lecture Learning Modelsに対してReinforcement Learningを実行するライブラリを提供しています。
LLMを利用するために、セマンティック検索やLLMのオーケストレーションなどを行えるフレームワーク。
テキスト分析、センチメント分析や単語分割などを行えるライブラリ。
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (L
Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing interm
LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to l
Citation-based measures of scientific influence typically treat citations as uniform signals, ignoring the dif
The application of large language models (LLMs) to personalized medical assistants has garnered growing intere
Taxonomy induction aims to organize concept sets into coherent hierarchical structures. Recent LLM-based metho
Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic.
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their r
Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing ext
We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations
Large language models achieve strong performance across diverse tasks, but deployment remains costly because o
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Spatio-temporal context has become increasingly crucial for visual tracking. However, most existing approaches
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Ensuring effective transfer learning for vision-language models without compromising their generalization perf
Existing SLAM systems lack modeling of the functional relations required for fine-grained robotic interaction.
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Actio
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) p
ゼネレーティブAIに関連するリソースの一覧。
LLMを使用して、自然言語処理における情報抽出を行うためのPythonライブラリです。
このリポジトリは自然言語処理(NLP)に関するリソースをまとめたものです。
Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism anal
The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need fo
Research teams and organizations often explore unfamiliar free-text collections, from survey comments and revi
Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languag
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering
We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-se
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generat
We present a continuous, population-scale measurement record of autonomous language-model trading agents opera
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to mod
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustn
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language prob
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation,
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its e
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language m
この論文では、映像 diffuision モデルを用いて変化する動画の差分を予測する方法を説明する。
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness,
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressiv
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal veri
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and re
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together wit
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough:
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby pla
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a larg
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic,
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
Large language models achieve superior performance on tasks that require extended reasoning, but long chains o
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a tea
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
Personalized assistants should not only comply with user requests but also assess whether those requests are a
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained b
We study a governed approach to enterprise analytics: a language model interprets the question, while determin
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improv
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
マルチラギングスピーチ生成やクリエイティブボイスデザイン、ルートライフクライミングなど、テクスチャファリーTTSの最新技術を実現するためのフレームワークです。
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial inform
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucin
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from cur
Matcha-TTSは、高速で条件付き流のマッチングを実現するTTSアーキテクチャであり、話者の特徴を考慮する。
Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation an
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evalua
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoo
この論文では、Causal-Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive
trafilaturaはPythonとコマンドラインツールで、Webからテキストとメタデータを取得して、CSV、JSON、HTML、MD、TXT、XML形式で出力を提供します。
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements
Driver behavior is heterogeneous, context-dependent, and changes over time, and these properties shape the tra
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
On-policy distillation trains a language model on its own generations while a teacher scores them token by tok
Vowpal Wabbitは、機械学習を進歩させるためのオンライン学習、ハッシュ、reduceなどの強力なアルゴリズムを含むシステムです。その結果、さまざまな問題に応じて、高品質な解決策を提供できます。
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform
Wunjo CE: Face Swap, Lip Sync, Control Remove Objects & Text & Background, Restyling, Audio Separator, Clone V
Population optimizers such as CMA-ES, DE, and multi-objective evolutionary algorithms drive search mainly thro
spaCyはPythonで動くIndustrial-strength Natural Language Processing(Language理解の高度なライブラリです。文脈理解、構文解析、名詞の抽出など、複雑なNLPタ
This repository provides an implementation of "A Simple yet Effective Training-free Prompt-free Approach to Ch
長時間のビデオ生成を実現するためのモデルのサポートを紹介している。
Rust言語でCandleライブラリを利用して、PythonやPyTorchを使用せずにDecoder-only LLMを自作した。
医療画像分析で、深層學習モデルが実装されている問題に対する解決策を提示します。治療を導くために、批判的結果に影響を与える変化について特に重点が置かれています。
画像エディティング用推論モデルの改良方法についての公式実装であるFlowEdit。
Embodied AIやロボットとLarge Language Modelを組み合わせた研究のリポジトリ。
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously u
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background