Skip to content
Papers.

Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens research paper by UC Berkeley, 2025

UC Berkeley · Nov 24, 2025 · Reasoning · 62 citations · 29 upvotes · unverified

Read on arXiv

What it shows

Chain-of-Visual-Thought (COVT) enables Vision-Language Models to reason through visual tokens, improving their performance on perceptual tasks by capturing dense visual information.

UnverifiedHugging Face's summary; not yet checked by hand.

More from UC Berkeley

All 12
PaperCitations
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive ExecutionFreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on personal machines.Inference and efficiency · Aug 2026 · Unverified2
Playful Agentic Robot LearningEmbodied robots learn reusable skills through self-directed play and exploration, then apply these skills to improve performance on downstream tasks without additional training.Agents and evaluation · Jun 2026 · Unverified7
Agents' Last ExamOver 1,000 real, checkable professional tasks written with 250+ industry experts; the hardest tier is far from solved.Agents and evaluation · Jun 202615
Flash-KMeans: Fast and Memory-Efficient Exact K-MeansFlash-kmeans enables efficient online k-means clustering on GPUs through novel kernel-level optimizations that eliminate I/O bottlenecks and reduce atomic write contention.Inference and efficiency · Mar 2026 · Unverified6
dLLM: Simple Diffusion Language ModelingA unified open-source framework is presented that standardizes core components of diffusion language modeling for reproduction, customization, and accessible development of both large and small models.Foundation models · Feb 2026 · Unverified12
SLA2: Sparse-Linear Attention with Learnable Routing and QATSLA2 improves sparse-linear attention in diffusion models by introducing a learnable router, direct attention formulation, and quantization-aware fine-tuning for enhanced efficiency and quality.Inference and efficiency · Feb 2026 · Unverified20
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache QuantizationQuant VideoGen addresses KV cache memory limitations in autoregressive video diffusion models through semantic-aware smoothing and progressive residual quantization, achieving significant memory reduction with minimal latency impact.Inference and efficiency · Feb 2026 · Unverified16
Residual Context Diffusion Language ModelsResidual Context Diffusion (RCD) enhances diffusion large language models by recycling discarded token information through contextual residuals, improving accuracy with minimal computational overhead.Training and scaling · Jan 2026 · Unverified8
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 202222.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20255,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified172
About this paper
Authors
Yiming Qin, Bomin Wei, Jiaxin Ge and 4 more
arXiv
2511.19418 · PDF
Venue
arXiv.org
Citations
62, 7 influential · Semantic Scholar
Upvotes
29 · Hugging Face
Code
github.com/Wakals/CoVT
Lab
UC Berkeley

Changes

What changed
Influential citationsfirst count: 7Sep 25, 2026
Citationsfirst count: 62Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.