Skip to content
Papers.

Vibe Checker: Aligning Code Evaluation with Human Preference research paper by Google DeepMind, 2025

Google DeepMind · Oct 8, 2025 · Agents and evaluation · 2 citations · 34 upvotes · unverified

Read on arXiv

What it shows

Vibe Checker evaluates LLMs by combining functional correctness and instruction following to better align with human coding preferences.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Google DeepMind

All 12
PaperCitations
DiffusionGemma Technical ReportDiffusionGemma is a fine-tuned mixture-of-experts language model that uses discrete diffusion to generate text blocks in parallel, achieving high speed while preserving capabilities like multimodal inputs and reasoning.Foundation models · Jul 2026 · Unverified1
Gemma 4 Technical ReportOpen multimodal models from 2.3B to 31B parameters, with a thinking mode and image and audio input.Foundation models · Jul 2026110
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMsLiteFrame, a lightweight video encoder with Compressed Token Distillation training method, reduces latency and increases frame processing capacity for long-form video understanding in Video LLMs while maintaining accuracy.Inference and efficiency · May 2026 · Unverified1
Understanding the Challenges in Iterative Generative Optimization with LLMsGenerative optimization using large language models faces challenges due to implicit design decisions about artifact modification and learning evidence that significantly impact success across different applications.Foundation models · Mar 2026 · Unverified8
LoGeR: Long-Context Geometric Reconstruction with Hybrid MemoryLoGeR enables long-term 3D video reconstruction by combining bidirectional priors with a hybrid memory system that includes parametric Test-Time Training and non-parametric sliding window attention mechanisms.Reasoning · Mar 2026 · Unverified35
SIMA 2: A Generalist Embodied Agent for Virtual WorldsSIMA 2, built on a Gemini foundation model, interacts in 3D virtual worlds, reasons about goals, handles complex instructions, and autonomously learns new skills through open-ended self-improvement.Agents and evaluation · Dec 2025 · Unverified18
Robot Learning from a Physical World ModelPhysWorld integrates video generation and physical world modeling to enable accurate robotic manipulation from visual demonstrations without real robot data.Multimodal and robotics · Nov 2025 · Unverified21
Video models are zero-shot learners and reasonersVeo 3, a generative video model, exhibits zero-shot capabilities across various visual tasks, suggesting a trajectory towards becoming a unified, generalist vision foundation model.Multimodal and robotics · Sep 2025 · Unverified215
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified4,134
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified761
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified77
About this paper
Authors
Ming Zhong, Xiang Zhou, Ting-Yun Chang and 9 more
arXiv
2510.07315 · PDF
Citations
2, 0 influential · Semantic Scholar
Upvotes
34 · Hugging Face
Lab
Google DeepMind · on Companies · on Acquisitions

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 2Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.