K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
- 用途
- QA
- 難易度
- Easy
- コスト
- High
「multimodal」の検索結果
21 件Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduc
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, whe
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. Howe
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogene
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the cont
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. A
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robo
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generat
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language mode
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedur
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diver
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has beco
Building assistants that can continually watch the world, remember what they see, and reason over their accumu
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions.
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing te