K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
- 用途
- QA
- 難易度
- Easy
- コスト
- High
「image」の検索結果
29 件Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam questi
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computat
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduc
Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supp
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, an
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diag
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. Howe
In line with the prevailing direction of vision research, we explore the integration of both generation and ed
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of nu
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents ty
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. A
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generat
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language mode
Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent met
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedur
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has beco
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing r
Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce waterti
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs tha
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). Howev
When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often
Building assistants that can continually watch the world, remember what they see, and reason over their accumu
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions.
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated dataset
Training-free in-context segmentation enables new object categories to be introduced at inference time from a
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing te