diffusers — 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
.diffusion モデルのライブラリ。画像・動画・音声生成に利用可能。
- 用途
- 画像・動画・音声生成
- 難易度
- Easy
- コスト
- High
「generation」の検索結果
109 件.diffusion モデルのライブラリ。画像・動画・音声生成に利用可能。
Rustを使ってモジュラーLLMアプリケーションを構築することができるライブラリです。
医学画像分析は、医療の診断や治療を支援するために画像に記載されたデータから情報を抽出する研究分野です。この研究では、foundation modelsを用い、医療画像分析のための新しいアプローチを提案しました。found
SANAは、高解像度画像生成モデルSANAを紹介する本研究であり、低計算コストで優れた高解像度画像を生成できる。
このリポジトリでは、データとAIアルゴリズムを製品化するためのプラットフォームであるTaipyを提供しています。
ベクトル検索と構造化されたフィルタリングを組み合わせたベクターデータベースです。
TensorZeroは、LLMゲートウェイ、オブザーバビリティ、評価、最適化、実験を統一したオープンソースのLLMOpsプラットフォームです。
flyteは、高度に動的で堅牢なAIオーケストレーションプラットフォームであり、データ、モデル、コンピューティングを統合してAIワークフローを作成することができます。
オープンソースのAI推論最適化と展開用ツールキットです。
Awesome-Video-Diffusionは、Recent Diffusion Models for Video Generation, Editing, and Othersのリストを公開しています。
FastVideoは、加速されたビデオ生成用の統合推論とポストトレーニングのフレームワークです。
zenmlは、データパイプラインからエージェントまで、AIプラットフォームです。
この論文では、ディフュージョンモデルの高速化を目的としたNVIDIA FastGenについて説明しています。FastGenは、ディフュージョンモデルから高速に生成することが可能です。
オープンソースのAIオーケストレーションフレームワークです。LLMアプリケーションの構築に必要なパイプラインやエージェントワークフローの設計ができるようになっています。
医学画像に対する疾患検出モデルを開発し、臨床現場で早期検出と迅速な介入を容易にすることを目的としたフレームワークを提案します。
この研究では、大規模言語モデルを活用して、経済学の研究活動をサポートするシステムを開発しました。このシステムは、学者が理論モデル開発を自動化することができます。
長形推論のための言語モデルが、提供されたコンテキストから乖離した論理を生成する可能性があることを指摘し、コンテキストと推論論理をより適切に融合するため、 REFACT (REstating Facts in Adapti
ディフュージョンモデルにおける初期的なNoise Seed の影響が、モデルが生成する高質のイメージに大きく影響していることを提示し、Seed Search 時の時間的負荷を削減するための方法を提案した。
Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settin
Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in h
Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from
Infrared image super-resolution (IISR) mitigates the limitations imposed by low spatial resolution. Existing m
Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the v
NestJSベースのAIチャットボット開発ツールです。
xtunerは、超大規模MoEモデルを高速にトレーニングするためのトレーニングエンジンです。
音声認識、声活動検出、テキスト処理などを行う、基盤となる音声認識ツールキットを提供する。
この論文では、Causal-Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive
LLMを利用するために、セマンティック検索やLLMのオーケストレーションなどを行えるフレームワーク。
流動画像生成を扱う研究、HeadCast を用いて流動画像生成を提案する。
この研究では、抗原特異性抗体を設計するために、抗原および抗体の間でエピトープレベルでのペアリングが必要であることを考慮した、抗原特異性の抗体多モーダルファンデーションモデル(AAMFM)を提案しました。
We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive
分散式推論には、高解像度ビデオ生成のためにコストが高いという問題があります。Evolving Cache Schedulesアルゴリズムは、コストと効率性のトレードオフを最適化することで、キャッシュで推論コストを削減しま
Attention Mechanism (AM) selectively focuses on essential information for imaging tasks and captures relations
Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling
Noisy and corrupted points can substantially degrade point cloud recognition performance, especially under cha
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
デバイス上のLLM推論をXビット量化を使用したもの。
販売データを分析するために、機械学習モデルが使用されるリソースが提供されていました。
OpenWorldLibは、進化する世界モデルを提供する統一されたコードベースです。
CVPRに基づくAIを取り入れるための資料集を提供します。CVPR 2026、2025、2024、およびECCV 2024に基づくAIGCに関する研究論文とソフトウェアコードを含みます。
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback
Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large
Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural
この研究では、高空飛行の無信号位置指示のNGPS (Next-Generation Positioning System)というフレームワークを提案しました。NGPSは、GPSの信号を利用せずに位置推定を可能にします。N
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate larg
Text-to-video generation has advanced significantly over the past five years through scaling of model size, da
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and syn
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computat
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduc
AIエージェントをGoogle Cloudに展開することが可能で、CI/CD、評価、観察など、プロダクションリードテンプレートが事前に用意されています。
人工DNAシーケンスを生成するモデルを提案し、DNAシーケンスを扱える機械学習的手法を開発することを目的としている。
The global competition for developing robotic foundation models is intensifying. Among the data collection sys
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, an
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diag
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. Howe
In line with the prevailing direction of vision research, we explore the integration of both generation and ed
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogene
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of nu
Current video generation models achieve impressive results in single-shot generation, yet remain limited in ci
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the cont
モデルをサービングするためのライブラリを紹介している。
Open-dLLMはOpen diffusion language modelを公開しており、コード生成の前トレーニング、評価、推論、チェックポイントを公開しています。
Analytical placers rely on differentiable objective functions to guide placement, typically combining intermed
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interac
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the
Over the past few years, diffusion-based Schrödinger bridge models have been proposed to approximate optimal t
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents ty
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), h
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. H
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of
As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and i
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generat
ゼネレーティブAIに関連するリソースの一覧。
LLMのマージに関してのマニュアルです。理論、方法、応用などについての概要が記載されています。
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diver
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, en
画像認証システムにおける悪用された画像からの画像の認証方法を提示しました。
Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce waterti
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer dire
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). Howev
Code review helps maintain software quality before code integration, but it also imposes a substantial workloa
AIエージェントの開発と実装を行うためのエンドツーマンド、コードファーストのチュートリアル。
LakonLabは、AsymFlow、pi-Flow、GMFlowなどの生成型流体力学を実装するためのオープンソースプロジェクトです。
MemVidは、サーバーレスで単一ファイルの記憶層を提案し、AIエージェントが即時検索と長期的な記憶を持つようにする記憶層です。
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification
Emotion-driven Style Controlを使用してテキストから声の変換が実行され、感情のあるテキストをエモタイザブルな声に変換することが可能になります。
UniPicは、オープンソースの最先端の画像編集モデルの実装です。
この研究では、COVID-19臨床パスウェイズの予測監視を支援するために、パイプラインを構築しました。このパイプラインには、データリフティング、時間的再構成、イベントログの構築、プリフィックスベースの表現、予測モデルの整
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated dataset
本研究では、生成推奨システムにおけるアイテムIDの構築、調整、生成の手法について、アイテムIDの構築方法を分析しています。
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning r
マルチラギングスピーチ生成やクリエイティブボイスデザイン、ルートライフクライミングなど、テクスチャファリーTTSの最新技術を実現するためのフレームワークです。
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing te
AIドライブのマルチエージェント研究アシスタント。仮説の生成、データ分析、およびレポートの生成を自動化する。
Magic123は、画像を1枚入力し、画像と3Dデータ双方の情報を利用して高質の3Dオブジェクトを生成することができる。
この論文では、RAG、AIパイプライン、企業検索を含むクラウド テンプレートを提供するアプリケーション「llm-app」を紹介します。 llm-app は Docker で動作し、Sharepoint、Google Dr
学習中のアイデアや知識を整理するための日記。
Operad理論を用いて、モデルが組み合わせ式に対する複合的な回答の合致性を検証する手法が提案された。
医療画像分析で、深層學習モデルが実装されている問題に対する解決策を提示します。治療を導くために、批判的結果に影響を与える変化について特に重点が置かれています。
画面の生成モデルであるHunyuanVideoを開発した。HunyuanVideoは、複雑なシーケンスを生成する能力を持つ。
分析システムの性能を向上するための学習モデル開発を行う。
画像生成のためのHigh Quality Training Free Inpaintを提供します。このInpaintはStable Diffusionモデルに使用でき、ComfyUIもサポートしています。
このリポジトリでは、AIエンジニアリングのためのオープンソースプラットフォームであるMLflowを提供しています。
Train high-quality text-to-image diffusion models in a data & compute efficient manner