Skip to content
Papers.

Flash-KMeans: Fast and Memory-Efficient Exact K-Means research paper by UC Berkeley, 2026

UC Berkeley · Mar 10, 2026 · Inference and efficiency · 6 citations · 64 upvotes · unverified

Read on arXiv

What it shows

Flash-kmeans enables efficient online k-means clustering on GPUs through novel kernel-level optimizations that eliminate I/O bottlenecks and reduce atomic write contention.

UnverifiedHugging Face's summary; not yet checked by hand.

More from UC Berkeley

All 12
PaperCitations
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive ExecutionFreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on personal machines.Inference and efficiency · Aug 2026 · Unverified2
Playful Agentic Robot LearningEmbodied robots learn reusable skills through self-directed play and exploration, then apply these skills to improve performance on downstream tasks without additional training.Agents and evaluation · Jun 2026 · Unverified7
Agents' Last ExamOver 1,000 real, checkable professional tasks written with 250+ industry experts; the hardest tier is far from solved.Agents and evaluation · Jun 202615
dLLM: Simple Diffusion Language ModelingA unified open-source framework is presented that standardizes core components of diffusion language modeling for reproduction, customization, and accessible development of both large and small models.Foundation models · Feb 2026 · Unverified12
SLA2: Sparse-Linear Attention with Learnable Routing and QATSLA2 improves sparse-linear attention in diffusion models by introducing a learnable router, direct attention formulation, and quantization-aware fine-tuning for enhanced efficiency and quality.Inference and efficiency · Feb 2026 · Unverified20
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache QuantizationQuant VideoGen addresses KV cache memory limitations in autoregressive video diffusion models through semantic-aware smoothing and progressive residual quantization, achieving significant memory reduction with minimal latency impact.Inference and efficiency · Feb 2026 · Unverified16
Residual Context Diffusion Language ModelsResidual Context Diffusion (RCD) enhances diffusion large language models by recycling discarded token information through contextual residuals, improving accuracy with minimal computational overhead.Training and scaling · Jan 2026 · Unverified8
Latent Implicit Visual ReasoningWhile Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning modality.Reasoning · Dec 2025 · Unverified10
Topic
PaperCitations
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20225,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified688
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified507
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified294
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified249
DeepSeek-OCR: Contexts Optical CompressionDeepSeek-OCR uses optical 2D mapping to compress long contexts, achieving high OCR precision with reduced vision tokens and demonstrating practical value in document processing.DeepSeek · Oct 2025 · Unverified189
About this paper
Authors
Shuo Yang, Haocheng Xi, Yilong Zhao and 10 more
arXiv
2603.09229 · PDF
Venue
arXiv.org
Citations
6, 0 influential · Semantic Scholar
Upvotes
64 · Hugging Face
Code
github.com/svg-project/flash-kmeans
Lab
UC Berkeley

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 6Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.