K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
- 用途
- QA
- 難易度
- Easy
- コスト
- High
「LLM」の検索結果
36 件Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that t
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diag
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, whe
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. Howe
In line with the prevailing direction of vision research, we explore the integration of both generation and ed
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogene
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task,
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interac
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Exis
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents ty
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), h
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. H
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. A
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in hist
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generat
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still he
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language mode
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the co
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs tha
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned on
Code review helps maintain software quality before code integration, but it also imposes a substantial workloa
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical
Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize of
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing te
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environment