K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
- 用途
- QA
- 難易度
- Easy
- コスト
- High
「text」の検索結果
51 件Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate larg
Text-to-video generation has advanced significantly over the past five years through scaling of model size, da
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop inter
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and syn
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computat
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduc
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of langu
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on
Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supp
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, an
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Un
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diag
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, whe
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. Howe
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogene
Predicting a football match before kickoff requires more than knowing past results: a model must use changing
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task,
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Exis
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer a
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents ty
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), h
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. H
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. A
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in hist
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating compo
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generat
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language mode
Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent met
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedur
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diver
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has beco
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing r
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer dire
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs tha
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). Howev
Building assistants that can continually watch the world, remember what they see, and reason over their accumu
Code review helps maintain software quality before code integration, but it also imposes a substantial workloa
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical
Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize of
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware
Training-free in-context segmentation enables new object categories to be introduced at inference time from a
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing te
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environment