Skip to content

Verification Limits Code LLM Training research paper by Cohere, 2025

Cohere · Sep 25, 2025 · Training and scaling 1 year ago

Read on arXiv

What it shows

Verification design and strategies impact code generation performance, showing that richer and more diverse test suites improve capabilities while calibrated verification thresholds enhance data usability.

Hugging Face's summary; not yet checked by hand.

More from Cohere

All 30
PaperCitations
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language ReasoningOptimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language.Reasoning · Sep 20262 weeks ago-
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic ReasoningAn open competition on unseen Linguistics Olympiad puzzles, graded by the official jury; small 14B systems beat models twice their size.Reasoning · Aug 20265 weeks ago-
CALIBER: Calibrating Confidence Before and After Reasoning in Language ModelsTrains reasoning models to state their confidence twice, before thinking and after answering, halving calibration error against the best single-estimate method.Reasoning · Jun 20263 months ago-
AI Exposure Scores: what they measure, what they miss, and what comes nextReviews how widely cited scores of which jobs AI can assist are used in policy, what they miss, and newer measures that address it.Foundation models · Jun 20263 months ago-
The Culture Funnel: You Can't Align What isn't in the DataModern LLM pipelines experience a cultural data funnel where explicit cultural signals diminish during post-training, necessitating shifts in training data approaches for better cultural alignment.Alignment and safety · Jun 20263 months ago-
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft RewardsTurns each prompt into a checklist scored item by item, giving reinforcement learning partial-credit rewards for tasks that cannot be checked automatically.Foundation models · May 20264 months ago-
Agents Explore but Agents Ignore: LLMs Lack Environmental CuriosityLLM-based agents fail to exploit discovered unexpected information despite recognizing it, indicating a lack of environmental curiosity that depends on tools, compute, and training data distribution.Agents and evaluation · Apr 20265 months ago-
Tiny Aya: Bridging Scale and Multilingual DepthTiny Aya demonstrates high-quality multilingual capabilities with 3.35 billion parameters through region-aware posttraining and balanced language performance.Training and scaling · Mar 20266 months ago-
Topic
PaperCitations
LoRA: Low-Rank Adaptation of Large Language ModelsLoRA fine-tunes a large model by training small low-rank matrices, cutting trainable parameters by 10,000 times.Microsoft · Jun 20215 years ago23.3k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.OpenAI · Jan 20206 years ago9,065
Training Compute-Optimal Large Language ModelsChinchilla: for a fixed compute budget, train a smaller model on more data; parameters and tokens should grow together.Google DeepMind · Mar 20224 years ago3,756
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismMegatron-LM splits each Transformer layer across GPUs to train models with billions of parameters.NVIDIA · Sep 20197 years ago3,179
DeepSeek LLM: Scaling Open-Source Language Models with LongtermismDeepSeek LLM, an open-source language model project, develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.DeepSeek · Jan 2024 · Unverified2 years ago855
LoRA Learns Less and Forgets LessLoRA learns less than full fine-tuning on code and math, but forgets less of what the model already knew.Databricks · May 20242 years ago407
About this paper
Authors
Srishti Gureja, Elena Tommasone, Jingyi He and 3 more
arXiv
2509.20837 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
0 · Hugging Face
Lab
Cohere · on Companies · on Acquisitions · on Releases

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.