ultralytics — Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
ultralyticsはYOLO(You Only Look Once)の技術を使用したオブジェクト検出ライブラリで、高い精度を提供している。
Category
画像分類、検出、セグメンテーション、動画認識など、視覚AIの実装と評価に関係する技術群です。
ultralyticsはYOLO(You Only Look Once)の技術を使用したオブジェクト検出ライブラリで、高い精度を提供している。
YOLOv5という物体検出アルゴリズムをPyTorchから他の言語に変換できるライブラリ。
supervisionは、機械学習技術を活用して、ユーザー独自のコンピュータビジョンツールを作成することができる。
データラベル化と注釈化を行うためのツールです。
ultralyticsはYOLO(You Only Look Once)の技術を使用したオブジェクト検出ライブラリで、高い精度を提供している。
supervisionは、機械学習技術を活用して、ユーザー独自のコンピュータビジョンツールを作成することができる。
データラベル化と注釈化を行うためのツールです。
ultralyticsはYOLO(You Only Look Once)の技術を使用したオブジェクト検出ライブラリで、高い精度を提供している。
supervisionは、機械学習技術を活用して、ユーザー独自のコンピュータビジョンツールを作成することができる。
コンピュータビジョンのデータセット、変換、モデルのライブラリ。
CVATは、機械学習用の業界標準のデータエンジンです。さまざまなスケールのチームが使用し、さまざまなスケールのデータに対応しています。
イメージを注釈するツール。ポリゴン、長方形、円、線、点などを注釈することができる。
ノードベースのビジュアルプログラミングツールです。
このライブラリは、3次元幾何学とモーションの解析のためのオープンソースライブラリです。このライブラリは、複数の視点からの画像を扱い、構造計算とマルチビューステレオの解析をサポートしています。
データをロギング・ストーリング・クエリして視覚化できるSDKです。
このリポジトリでは、金融分野に適したLarge Language Modelsを提供しています。
3D点群処理のためのライブラリであるPoint Cloud Library(PCL)。
このプロジェクトは2Dおよび3D顔の分析を実現するための基盤プロジェクトであり、最先端の技術を導入して顔の分析を実現します。
stanzaは、さまざまな言語を処理するための言語処理用ライブラリです。
OpenCVを用いて画像処理の学習方法を紹介している。
CARLAは、オープンソースのシミュレータで、主に自動運転研究のために使われます。このシミュレータを使うことで、車両などのロボットをシミュレートし、様々なシナリオを実行できます。
Deepfake detection models often rely on high-quality inputs, fixed inference paths, and computationally expens
Building on the pioneering paper of Kearns, Roth, and Ryu (SODA'26), we study information aggregation in a net
Forecasting time series accurately is critical for applications with complex data ranging from energy systems
Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in sa
Time series data is very common in many real-world applications and in numerous domains, with increasing inter
Self-supervised learning relies on so-called data augmentations $φ(x)$ of unlabeled datapoints $x$ --- for exa
Gaussian kernel sums are the computational core of maximum mean discrepancies (MMDs), kernel gradient flows, S
The representation chosen for a mathematical operation can affect both its algebraic form and its empirical le
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Artificial intelligence (AI) now supports investment workflows from data and prediction through research, port
Multiplayer Online Battle Arena (MOBA) games rely on matchmaking to maintain competitive balance. Our prior wo
Computer-use agents can execute increasingly complex tasks in graphical interfaces, but their interaction expe
3D perception plays a crucial role in real-world applications such as autonomous driving, robotics, and AR/VR.
Comparing intelligent systems under deployment constraints requires more than predictiveaccuracy.This paper de
Sparse autoencoder (SAE) features are increasingly used to explain and steer language-model behavior, but it r
Pass receiver selection is a fundamental task in football analytics, aiming to predict the intended receiver u
Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We
A merchant's payment processor, ledger, ERP and bank feed are updated by messages that get delayed, duplicated
Catheterisation image processing requires segmentation models that are fast, accurate and explainable. While m
Grounded language-model pipelines can be divided into three stages: selecting an object, retrieving passages f
Memory-based evolutionary algorithms for dynamic optimization often carry a redundant second copy of the genot
This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphol
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds
We present a system that uses a Vision-Language Model (VLM) as a diagnostic agent for adapting a detect-to-tra
Semantic 3D maps are increasingly constructed automatically for aerial robotics by integrating learned semanti
In the field of computer vision and graphics, high-quality reconstruction of the human body in static scenes h
We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. U
Single domain generalization (SDG) aims to learn a model from one labeled source domain that generalizes to un
3D surface cutting and UV unwrapping are fundamental problems in computer graphics. Traditional geometric opti
Visual counting is commonly formulated at the instance level, aiming to estimate how many objects of a queried
Reliable robot-to-human handover requires the robot to infer when the person is ready to receive the object, a
Robotic dressing assistance is a promising solution for supporting older adults with physical impairments in d
Fixed 3D Gaussian Splatting (3DGS) reconstructions provide realistic novel views but lack the traversability c
The FitzHugh-Nagumo (FHN) system serves as a simplified model of neuronal voltage dynamics, capturing the acti
We present ResLearn-XR, a residual learning framework for predicting eXtended Reality (XR) network traffic and
In this paper, the problem of data-driven discovery of nonlinear ordinary differential equations (ODEs) is rec
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a larg
Parameterised graph theory studies how the complexity of graph-theoretic problems depends on structural parame
This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with
Ensuring the security of the power system is essential for stability and reliability, especially in the event
Medical image inpainting has the potential to improve automated brain MRI analysis by reconstructing healthy t
Reconstruction-based anomaly detectors are accurate but opaque: a deep autoencoder flags a sample without tell
We study the allocation of indivisible goods among agents with identical additive valuations, focusing on envy
Unsupervised anomaly detection scores each point of an unlabelled, contaminated sample in a single pass, and i
Institutions practising outcome-based education compute learning outcome attainment routinely, while reviews o
The computation of the Bures-Wasserstein (BW) barycenter of an ensemble of positive definite matrices arises t
Human mobility predictability concerns the best prediction performance attainable from a given target and inpu
Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded
In this paper we study linear non-Gaussian acyclic models (LiNGAM) when used in federated environments. These
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness,
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficie
Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and vid
Semantic ID (SID) generative recommendation predicts the next item by generating a short tuple of discrete tok
Restricted eigenvalue (RE) bounds govern stable recovery by norm-regularized estimators. For isotropic sub-Gau
This paper examines the relationship between the parameters of autoencoder models and the statistical properti
Let a finite population of n labelled examples carry a class-weighted loss, with pi*n in a rare positive class
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart
Digital materials fabricated by multi-material 3D printing are designed as controlled mixtures of stiff and co
Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness reci
This paper investigates structural priming in language model (LM) production, examining how preceding structur
Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guide
In this work, a novel evaluation scheme built on a generalized variant of the Rand Index measure, namely, the
Procedural instruction following is a basic requirement for controllable language-model systems, especially wh
Long-term human-AI interaction is difficult because the information that guides inference is updated implicitl
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice
Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multim
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby pla
Object-centric visual representations are important for physical-world perception, but existing visual pretrai
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet the
Bundle Adjustment (BA) is a cornerstone of 3D computer vision and has benefited from decades of advances in sp
Dust in agriculture presents a significant challenge for autonomous agricultural machinery. Dust can impair th
Mirrors are common in real-world images, yet producing geometrically consistent reflections with generative mo
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in
Labelling vision datasets, especially for segmentation tasks, is a laborious and costly process that stymies n
Video generators build long videos by composing shorter parts, either by generating segments one after another
Action-conditioned video models require large-scale visual data paired with control signals that are temporall
3D Gaussian Splatting has become a de facto scene representation for novel view synthesis, yet robustly learni
In conditional coding-based neural video compression, the quality of temporal context directly affects compres
Although document OCR systems perform increasingly well on routine documents, complex formulas, structured tex
Answering what-if queries about a scene with a VLM usually means injecting the assumption as text or repaintin
We present PointGT, a point-based 3D representation that enables simultaneous editing of object geometry and a
Camera traps have become an essential tool for wildlife monitoring, motivating the development of computer vis
Sampling-based motion planning algorithms are a popular class of trajectory planning algorithm due to their sp
The planning algorithms inside an Autonomous Vehicle (AV) rely on information from on-board sensors whose line
Quadruped robots have demonstrated impressive agility in parkour locomotion across complex terrains. However,
Accurate identification of weld seam geometries is essential for automated robotic post processing operations
Semantic mapping plays a crucial role in the ability of a robot to interact with objects, operate and navigate
Mobile robotic platforms offer a flexible alternative to fixed manipulators for non-destructive evaluation (ND
QLAUN Bot (Quad-Legged Adaptive Unmanned Navigator Robot) is a torque-controlled quadruped robot that is resea
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality
Air-ground collaborative Vision-and-Language Navigation (VLN) pairs an unmanned aerial vehicle (UAV) with a gl
This paper presents an experimental design for constructing a multimodal dataset to analyze user engagement in
Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite
Data-consistent inversion (DCI) constructs probability measures whose push-forward distributions agree with ob
Gaussian-process Bayesian optimization (GP-BO) excels at black-box optimization of costly functions, e.g., hyp
Gaussian graphical models (GGMs) are essential tools for interpretable structure learning. However, in high-di
A fundamental quantity in machine learning is the optimal performance achievable by any model on a given task.
Learning graph structures from data is a fundamental problem that spans a wide range of signal processing and
Generative modeling directly on geometric manifolds can avoid errors introduced by flattening non-Euclidean da
We study a variant of the Thompson Sampling (TS) algorithm, called $α$-TS, for solving stochastic generalized
LiDAR-based semantic segmentation is a core perception module for autonomous vehicles and mobile robots. Despi
Humans can perform complex manipulations given a simple intent through an overall instruction, while continuou
Ultra-low-altitude unmanned aerial vehicles (UAVs) require surround vision near buildings, vegetation, and oth
Generating safety-critical scenarios is essential for evaluating autonomous driving systems. However, existing
Physical reservoir computing (PRC) refers to the use of a physical dynamical system as a computational resourc
Humanoid learning increasingly relies on transforming vast and diverse human motion data into high-quality rob
Monolithic world models predict the entire next state at every step, spending capacity re-predicting the stati
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Unityを使用してマシンラーニングエージェントを訓練して訓練できるツールです。
Feature-based newsvendor models use observable covariates to tailor inventory decisions, aiming to balance hol
Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstr
Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable.
We consider semi-supervised classification from a partially classified sample arising from a two-component Wei
High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionall
Multi-Unmanned Aerial Vehicle (UAV) disaster-response systems require coordinated task assignment and local tr
A long-standing open problem in robot manipulator control is whether global regulation can be achieved by clas
Robotic processing of irregular steel scrap requires dense 3-D measurement to replace manual visual assessment
High-precision insertion remains a fundamental challenge in robotic manipulation due to the strict alignment r
We present a system of two wearable pneumatic haptic devices that supports continuous, closed-loop, bidirectio
YOLOv5という物体検出アルゴリズムをPyTorchから他の言語に変換できるライブラリ。
CoreNLPはJavaで開発されたNLPツールのセットであり、分割、文分割、名詞認識、パーシング、コorefence、感情分析などを行える。
This work shows that diffusion models learned with standard denoising loss can provide effective global MCMC p
We develop a marginal coordinate test for regression with Euclidean predictors and a random-object response in
Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is i
Classical numerical solvers for partial differential equations (PDEs) are computationally expensive to solve r
We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the traini
End-to-end autonomous driving models plan future trajectories from raw sensor input. While earlier driving ben
Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibilit
The downwash wake of a hovering quadrotor governs both the vehicle's own performance and the safe spacing of m
Visual SLAM is commonly evaluated on clean trajectories, although deployment failures are often caused by adve
Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations
Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success
Tactile sensing is essential for physical interaction in robotics and human--machine systems. However, combini
Most shape reconstruction methods assume measurements defined over planar sensing domains, such as RGB images
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Cu
Test-time reasoning has significantly improved performance in domains ranging from games to language models. H
We study the problem of locating a new homogeneous facility under a prelocated facility. Here, a set of $n$ ag
Envy-freeness up to any good (EFX) and pairwise maximin share (PMMS) are standard local fairness criteria for
Residual maximin share (RMMS) is the largest share threshold that remains guaranteeable throughout dynamic all
この作品は机械学習の論文を100行のコードで実装する方法について説明しています。
Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation an
Deployed decisions are often optimized once and retained because updates impose operational, regulatory, or sw
Kernel density estimation (KDE) is one of the most fundamental statistical estimators of density functions. It
Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structure
We study a class of product-reference diffusion algorithms for sampling from a discrete distribution. We show
We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power
We extend the FLOP (fast learning of order and parents) algorithm recently proposed by Wienöbst et al. (2026)
Coordination is a desirable feature in multi-agent systems, ranging from robotic swarms to socioeconomic netwo
Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN impos
Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start reg
Joint analyses across multiple institutions are increasingly important in biomedical and epidemiological resea
Model selection becomes particularly challenging under strong predictor dependence and model-class uncertainty
In this paper, we study the problem of augmenting a tiny target sample with a massive auxiliary sample. Utiliz
Network comparison using optimal transport is a growing area of research in network science. Unlike standard g
Two of the most fundamental questions in statistical learning theory are the following: which prediction probl
Point-cloud data routinely captured by modern imaging and sensor technologies provide detailed geometric descr
Evaluating customer creditworthiness is crucial for retail banking operations, as it impacts marketing strateg
We demonstrate a physical mechanism for causal information filtering in a physical reservoir computing (PRC) b
Edge computing systems need to support diverse sensing workloads under tight energy and memory constraints, th
We study pursuit-evasion games on graphs with a single pursuer and an invisible evader. The pursuer may assign
Driver behavior is heterogeneous, context-dependent, and changes over time, and these properties shape the tra
Exact Bayes prediction enjoys fast predictive regret guarantees, but exact posterior updating or representatio
Reconstructing continuous physical fields from sparse measurements is central to scientific monitoring, invers
Functional data analysis is an important statistical field that treats data as random functions. In practice,
The OBABO and BAOAB schemes and the other standard Strang splittings of kinetic (underdamped) Langevin dynamic
We study the sample complexity of learning near-optimal bilateral trade mechanisms. Unlike previous work on le
In this paper, we study alternating regret in online convex optimization (OCO), motivated by the success of al
Gromov-Wasserstein (GW) compares distributions through relations within each space. This pointwise comparison
Machine learning is usually formalized through samples, while the persistent individual to which multiple obse
Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID),
ES-HyperNEAT evolves substrate topology through adaptive quadtree subdivision; to our knowledge, no implementa
Discrete diffusion models offer a promising alternative to autoregressive generation by enabling parallel upda
Accurate watch-time (WT) prediction is an important requirement for short-video recommendations. Yet WT distri
One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local m
Standard solution concepts for stochastic games, such as Markov perfect equilibrium and Markov coarse correlat
OpenWorldLibは、進化する世界モデルを提供する統一されたコードベースです。
Metric distortion has primarily been studied for social choice functions, which select a single winner from or
As an extension of existing Bayesian persuasion framework with inadequate message mechanism, we study direct r
Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing ex
We study temporal fair division of indivisible mixed manna. Items arrive over time and must be allocated irrev
Spatial aliasing occurs when two or more distinct locations produce highly similar place-cell representations,
Persistent acoustic monitoring can detect machine faults without physical contact, but always-on inference is
Reinforcement learning (RL) theory fundamentally depends on probability theory through the Markov chain. There
Peer prediction seeks to incentivize agents to truthfully report an observed signal by rewarding joint sets of
Motivated by modern marketplaces, where the platform or the seller routinely gathers detailed user profiles, w
Budget-constrained advertisers commonly rely on two control mechanisms: pacing scales bids, whereas throttling
We study fair division of indivisible goods when agents' valuations are accessed only through ordinal comparis
In Shapley-Scarf housing markets, Ma (1994) shows that top trading cycles (TTC) is the unique mechanism satisf
We study interval scheduling from the perspective of fair allocation. There are $m$ identical machines and a s
An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the
Battery swapping is a rapid way to recharge electric vehicles (EVs). As more and more entities are involved in
We study the existence of envy-free up to any item (EFX) allocations of indivisible chores when agents have mo
機械学習とデータ分析のためのC++のツールキット。
In this work, we introduce analytical replay experiments to the evolutionary computing community. Replay exper
ネ
Characterising optimisation problem instances is a fundamental part of understanding the behaviour and perform
The single-selection prophet inequality is a canonical Bayesian online selection problem in which independent
A bidder can quietly buy a stake in a company before making an offer for it. That stake, a toehold, is suppose
この論文では、Dynamic Multi-Objective Optimizationの問題を解くために、Special Point Skeleton Reconstruction (SPSR)アルゴリズムを提案し、Pa
Generative models can reproduce an observational distribution while encoding an incorrect causal structure. We
Streaming data-driven dynamic multi-objective optimization requires algorithms to track time-varying Pareto fr
Many learning problems require representations that reconcile direct input, nearby structure, and broader cont
In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language M
Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or bo
この研究では、リカレントニューラルネットワークの構造とメモリの使用量の関係を調べた。結果は、メモリの使用量が減少し、モデルがより効率的に学習することができるというものであり、これは、リカレントニューラルネットワークのパフ
レジリエンシャルコンピューティングでは、非線形ダイナミカル系を使って、時間依存の入力を、高次元の状態空間表現にマッピングする。レジリエンシャルパフォーマンスは、メモリ、非線形性、そしてそれらのトレードオフを反映しているが