label-studio — Label Studio is a multi-type data labeling and annotation tool with standardized output format
データラベル化と注釈化を行うためのツールです。
- 用途
- データラベル化ツール
- 難易度
- Easy
- コスト
- Low
「classification」の検索結果
159 件データラベル化と注釈化を行うためのツールです。
🤗 Transformersは、テキスト・ビジョン・音声など複雑なモデル定義をサポートするフレームワークで、インフェレンスターやトレーニングに使用できる。
ultralyticsはYOLO(You Only Look Once)の技術を使用したオブジェクト検出ライブラリで、高い精度を提供している。
supervisionは、機械学習技術を活用して、ユーザー独自のコンピュータビジョンツールを作成することができる。
paperless-ngxは、コミュニティによってサポートされたスーパーチャージドのドキュメント管理システムで、ドキュメントのスキャン・インデックス・アーカイブが可能である。
CVATは、機械学習用の業界標準のデータエンジンです。さまざまなスケールのチームが使用し、さまざまなスケールのデータに対応しています。
イメージを注釈するツール。ポリゴン、長方形、円、線、点などを注釈することができる。
FiftyOneは、データセットの精査とAIモデル可視化を支援するライブラリです。このライブラリは、データセットの品質を高め、AIモデルを可視化するのを支援するために使用できます。
オープンソースのAI推論最適化と展開用ツールキットです。
PyTorchで使用できる画像エンコーダとバックボーンの最大のコレクションです。トレーニング、評価、推論など様々なスクリプトや事前の重み付きデータが含まれます。
このプロジェクトは2Dおよび3D顔の分析を実現するための基盤プロジェクトであり、最先端の技術を導入して顔の分析を実現します。
電気生理信号から表現を学習し、脳コンピューターインターフェースの開発を支援する。
この研究では、自然言語処理の負担を減らすモジュラリティを目指しています。モジュラリティとは、システムを小さくて独立した部分に分割して、それぞれを簡素化することです。この研究では、文脈に応じてモジュラリティを変更できるメカ
presidioは、テキスト、画像、構造化データを含む敏感データを検出、削除、マスク、アノニマイズするオープンソースフレームワークです。自然言語処理、パターンマッチング、カスタマイズ可能なパイプラインをサポートします。
CoreNLPはJavaで開発されたNLPツールのセットであり、分割、文分割、名詞認識、パーシング、コorefence、感情分析などを行える。
マシン学習、統計学習などに関する統計的エンジンです。
The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetim
Federated global navigation satellite system (GNSS) monitoring distributes a proprietary classifier to partly
Designing foundation models for graphs is challenging due to the irregular structure of graphs and the differe
Standard open-set node classification methods rely on the homophily assumption, where connected nodes share la
Predictive process monitoring aims at forecasting various aspects of running processes. Among the different ta
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they he
Understanding why individuals respond differently to psilocybin requires modeling the drug's transcriptional p
Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a
Advances in tactile sensing have made contact-rich perception possible, accelerating progress in robotic manip
Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide
Onboard object detection in Earth observation is constrained by limited computational resources and the absenc
Moving-object perception must decide which image regions correspond to real motion and keep every instance ide
Temporal graph learning models the evolution of dynamic systems, where both structural interactions and semant
Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate locomot
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a commo
The Ising model is extended to the Potts model for multinomial data. We introduce a Rater Ising-Potts model th
We describe the Snugi-AI-v2 submission to eRisk 2026 Task 2, the second edition of contextualized early depres
We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation mode
Cancer segmentation models can fail silently, generating plausible but incorrect masks that risk missed findin
Person identification from millimeter-wave (mmWave) point clouds has mainly relied on gait. Indoor walking, ho
Table Structure Recognition (TSR) aims to extract the bounding boxes of cells and table structure (e.g., HTML)
Microscopic image analysis has long been recognized as a promising approach for monitoring activated sludge. I
Multiple Instance Learning (MIL) is widely used for weakly supervised learning, particularly in digital pathol
Objective assessment of Freezing of Gait (FoG) in Parkinson's disease (PD) relies predominantly on wearable In
Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment
Micro-actions are subtle, low-intensity non-verbal behaviors that provide cues to fine-grained human states, i
The learning-rate schedule is a consequential choice in training deep networks, yet the policies in common use
CLIPという画像認識モデルをオープンソースとして実装したライブラリ。
デベロッパー向けのモデロプティミゼーションフレームワークです。モデルの高速化と効率化を実現することができます。
stanzaは、さまざまな言語を処理するための言語処理用ライブラリです。
Acidophilic proteins that remain stable and functional under highly acidic conditions, are important for indus
The software supply chain has become an increasingly exposed attack surface because of its reliance on intrica
Purpose: Accurate CT protocol selection is critical for diagnostic quality and patient safety, yet the current
Bringing multiscale geometric analysis directly to irregular point clouds remains difficult: quantities such a
Shortcut learning denotes the widespread situation in which a classifier exploits spurious correlations rather
We study the training dynamics of multiclass logistic regression on high-dimensional Gaussian mixture models w
Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its ver
Automatic modulation classification (AMC) models are frequently trained and validated on synthetic or channel-
Time series data is very common in many real-world applications and in numerous domains, with increasing inter
In this paper, we propose a class-wise dimension (channel) selection framework for Multivariate Time Series Cl
Deformability cytometry (DC) is a type of imaging flow cytometry, which uses a camera-equipped device to measu
Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cro
This study investigates the classification of individuals as healthy or at risk of Parkinson's disease using m
Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely interve
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability
Persistent AI assistants are intended to extend human attention, memory, and coordination across changing digi
Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs to di
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked b
Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making
Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason
This paper presents a small-scale quantitative experiment that links syntactic structure to stylistic function
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three
Voice phishing (vishing) unfolds in real time; by the time a call has ended and post-hoc classification is pos
The ability to generate counterarguments is important for critical thinking and balanced discourse, yet existi
Natural language processing systems underperform on code-mixed text, particularly for low-resource language pa
Human pose estimation and keypoint-based action recognition models are increasingly deployed as components of
Tree species recognition supports forest inventory and biodiversity monitoring but still depends on scarce tax
Contrastive learning approaches achieve strong performance by training models to bring similar samples closer
Incomplete multi-view multi-label learning requires not only robust semantic aggregation from partially observ
Standard Convolutional Neural Networks (CNNs) exhibit severe performance degradation due to a strong inductive
Post-hoc explanation methods are widely used to inspect image classifiers, but their reliability depends on de
Automated fingermark identification is the foundation of forensic investigation, yet progress in the field is
Privacy-sensitive surveillance systems could benefit from large vision-language models (VLMs), but such models
Modern Visual Place Recognition (VPR) methods excel on standard benchmarks yet remain brittle in feature-poor
Deep learning models for CT scan analysis are often limited by the scarcity of precise pixel-level annotations
Natural adversarial examples (NAEs) reveal that vision models can fail under realistic semantic changes beyond
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Cranial nerves (CNs) play essential roles in sensory, motor, and autonomic functions. Accurate CN parcellation
Under ubiquitous teleoperation environments with optically challenging conditions, an interface for tele-opera
Trust-Hub reuses host circuits: several files differ mainly in the inserted Trojan. When gates from sibling va
We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk
We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with
This paper examines the theory of Invariant Structural Learning (ISL), which proposes a non-optimization appro
When a user asks an assistant to forget a record, the test is whether the memory now matches the state it woul
AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior w
An open, networked web will allow agents to run frozen models from multiple vendors, keep their history privat
Training strategy, namely whether to retrain from scratch or fine-tune from the previous checkpoint, is an ove
LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations o
Choosing when to translate multilingual documents is a central routing problem in text classification: transla
Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failin
Learning with class-conditional label noise often relies on a transition model from latent clean classes to ob
感覚変換、すなわちバーチャルエキスパートが可能なPytorch実装。
Reconstructing continuous terrain manifolds from massive, unstructured airborne LiDAR point clouds remains cha
Late Gadolinium Enhancement (LGE) on cardiac magnetic resonance is a key marker of myocardial scar, but its li
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Robotic systems increasingly operate in dynamic, uncertain, and open-ended environments, where design-time ass
Building on the pioneering paper of Kearns, Roth, and Ryu (SODA'26), we study information aggregation in a net
Multi-Task semantic communication (SemCom) prioritizes simultaneous execution of multiple tasks over bit-accur
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart
Robots interacting with people must recognize not only explicit commands, but also social cues such as invitat
Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service
Verifying that manufactured batches of milling tools or carbide rotary burrs conform to production order sheet
Accurate identification of weld seam geometries is essential for automated robotic post processing operations
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby pla
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-
Graph neural networks (GNNs) are a class of neural networks suitable for learning on graph-structured data. Th
A fundamental quantity in machine learning is the optimal performance achievable by any model on a given task.
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated
Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstr
The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictio
We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifi
We consider semi-supervised classification from a partially classified sample arising from a two-component Wei
YOLOv5という物体検出アルゴリズムをPyTorchから他の言語に変換できるライブラリ。
This paper develops a general convolutional neural network (CNN) framework for detecting heterogeneous event-d
Scientific discovery often requires reasoning over competing hypotheses that are consistent with experimental
Informative label missingness can change the usual efficiency ordering between completely and partially labell
We propose a new design of fair classifiers for multi-class classification problems in the presence of vector-
Neuroevolution of Augmenting Topologies (NEAT) and its advanced version, Evolvable-Substrate HyperNEAT (ES-Hyp
Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, lim
Obtaining data from neuromorphic sensors and processing it with Spiking Neural Networks is a promising solutio
We study complete allocations under nonnegative additive valuations through the positive supports of goods, fo
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns
Token prediction is a central pre-training objective for modern language models. Despite its empirical success
This paper introduces a family of multiclass linear Perceptron classifiers with a multiplicative margin mechan
While modern generative models excel at modeling complex data, precise inference-time control in conditional g
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evalua
Thermodynamic computers are stochastic physical devices designed to perform calculations at the thermal energy
Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN impos
Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel,
Object classification in event-based computer vision is a task that is attracting considerable research attent
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through t
ES-HyperNEAT evolves substrate topology through adaptive quadtree subdivision; to our knowledge, no implementa
Analog circuit topology synthesis remains challenging because useful designs occupy a tiny fraction of a combi
This paper studies multiple fixed points in a discrete-time hysteresis neural network. The network consists of
spaCyはPythonで動くIndustrial-strength Natural Language Processing(Language理解の高度なライブラリです。文脈理解、構文解析、名詞の抽出など、複雑なNLPタ
Drench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learni
Deep Learning models can include billions of parameters or more, making it difficult to explain their internal
Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, le
An invariant behavioral profile is the defining vulnerability of traditional honeypot installations: a skilled
Evolutionary feature construction has shown strong promise in symbolic regression by automatically discovering
このライブラリは、コンピューター ビジョンのための高度なAI解釈と可視化ソリューションです。このライブラリは、CNN、ビジョン トランスフォーム、分類、物体検出、分割、画像類似度など、さまざまなコンピューター ビジョンの
Finding a representative description of graph entities that captures their structural roles and homophily is a
この研究では、半推測状態の状況でアダプティブ行動を示すために、反射的な組織がどのように機能するかを調査する。これでは、現時点の観察だけを基に、内部状態が情報を保持できる計算プロパティを開発します。
Spiking Transformers model token interactions primarily through spiking self-attention (SSA). However, binary
Spiking Transformers provide a promising paradigm for efficient visual processing with spike-driven computatio
Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distin
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background
This paper explores the challenges and the methodologies associated with learning quality representations in s
Many learning problems require representations that reconcile direct input, nearby structure, and broader cont
EEG foundation-model gains may depend on cohort, montage, or probe design. We evaluated five models on five ta
この研究では、人工知能の研究者と神経科学者の間の分野を結びつけるために、脳のシステム構造を研究し、その研究から導かれた新しいアプローチを提案しました。