Skip to content

Cohere research papers

Company · 30 papers · 0 citations · latest Sep 2026 2 weeks ago

Research page

30 of 30 papers, newest first

PaperCitations
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language ReasoningOptimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language.Reasoning · Sep 20262 weeks ago-
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic ReasoningAn open competition on unseen Linguistics Olympiad puzzles, graded by the official jury; small 14B systems beat models twice their size.Reasoning · Aug 20265 weeks ago-
CALIBER: Calibrating Confidence Before and After Reasoning in Language ModelsTrains reasoning models to state their confidence twice, before thinking and after answering, halving calibration error against the best single-estimate method.Reasoning · Jun 20263 months ago-
AI Exposure Scores: what they measure, what they miss, and what comes nextReviews how widely cited scores of which jobs AI can assist are used in policy, what they miss, and newer measures that address it.Foundation models · Jun 20263 months ago-
The Culture Funnel: You Can't Align What isn't in the DataModern LLM pipelines experience a cultural data funnel where explicit cultural signals diminish during post-training, necessitating shifts in training data approaches for better cultural alignment.Alignment and safety · Jun 20263 months ago-
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft RewardsTurns each prompt into a checklist scored item by item, giving reinforcement learning partial-credit rewards for tasks that cannot be checked automatically.Foundation models · May 20264 months ago-
Agents Explore but Agents Ignore: LLMs Lack Environmental CuriosityLLM-based agents fail to exploit discovered unexpected information despite recognizing it, indicating a lack of environmental curiosity that depends on tools, compute, and training data distribution.Agents and evaluation · Apr 20265 months ago-
Tiny Aya: Bridging Scale and Multilingual DepthTiny Aya demonstrates high-quality multilingual capabilities with 3.35 billion parameters through region-aware posttraining and balanced language performance.Training and scaling · Mar 20266 months ago-
CIRCLE: A Framework for Evaluating AI from a Real-World LensA six-stage framework for measuring how deployed AI behaves and affects organisations, linking field tests and red teaming to usable metrics.Agents and evaluation · Feb 20267 months ago-
Unlocking Reasoning Capability on Machine Translation in Large Language ModelsFinds that step-by-step reasoning makes machine translation worse, then trains a model on structured draft-and-revise traces that improve it.Reasoning · Feb 20267 months ago-
SimMerge: Learning to Select Merge Operators from Similarity SignalsPredicts which way of merging language models will work best from cheap similarity signals, avoiding costly merge-and-test searches.Foundation models · Jan 20268 months ago-
The Art of Asking: Multilingual Prompt Optimization for Synthetic DataPrompt-space optimization enhances multilingual large language models by systematically transforming prompts for naturalness, cultural adaptation, and difficulty, leading to improved performance across various metrics.Retrieval and data · Oct 202511 months ago-
EAGER: Entropy-Aware GEneRation for Adaptive Inference-Time ScalingEAGer, a training-free method, uses token-wise entropy to optimize computational resources and improve performance on complex reasoning tasks.Inference and efficiency · Oct 202511 months ago-
Making, not Taking, the Best of NFusion-of-N (FusioN) method improves LLM generation quality by synthesizing elements from multiple samples, outperforming Best-of-N in various settings and tasks.Foundation models · Oct 202512 months ago-
Verification Limits Code LLM TrainingVerification design and strategies impact code generation performance, showing that richer and more diverse test suites improve capabilities while calibrated verification thresholds enhance data usability.Training and scaling · Sep 20251 year ago-
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMsThe study examines and proposes new sampling and selection strategies to enhance inference-time compute for multilingual and multi-task large language models, demonstrating significant improvements in win-rates across various languages and tasks.Inference and efficiency · Jun 20251 year ago-
Aya Vision: Advancing the Frontier of Multilingual MultimodalityMulti-modal language models are enhanced through synthetic data creation and cross-modal merging techniques, achieving superior performance in multilingual settings.Multimodal and robotics · May 20251 year ago-
Command A: An Enterprise-Ready Large Language ModelCommand A, a multilingual large language model, uses decentralized training with self-refinement and model merging to achieve efficient and top-performing Retrieval Augmented Generation for enterprise use.Agents and evaluation · Apr 20251 year ago-
Aya Expanse: Combining Research Breakthroughs for a New Multilingual FrontierA new family of 8B and 32B parameter multilingual language models, Aya Expanse, achieves state-of-the-art performance across multiple languages and outperforms larger models in evaluations.Alignment and safety · Dec 20241 year ago-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationGlobal-MMLU addresses cultural biases and translation artifacts in multilingual datasets by involving professional annotators and offering culturally distinct subsets for more accurate model evaluation.Agents and evaluation · Dec 20241 year ago-
M-RewardBench: Evaluating Reward Models in Multilingual SettingsA systematic evaluation of reward models in multilingual settings reveals significant performance gaps and dependencies on translation quality and resource availability.Agents and evaluation · Oct 20241 year ago-
To Code, or Not To Code? Exploring Impact of Code in Pre-trainingThe inclusion of code data during pre-training significantly improves LLMs' performance across various natural language and world knowledge tasks, beyond code generation.Training and scaling · Aug 20242 years ago-
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsBAM enhances the Mixture of Experts framework by fully recycling dense model parameters, including attention layers, leading to better performance and efficiency than previous methods.Inference and efficiency · Aug 20242 years ago-
Self-Improving Robust Preference OptimizationSRPO, a self-improving offline RLHF framework, achieves robustness to out-of-distribution tasks by optimizing a min-max objective that jointly enhances self-improvement and generative policies, leading to superior performance compared to DPO.Alignment and safety · Jun 20242 years ago-
Aya 23: Open Weight Releases to Further Multilingual ProgressAya 23 is a powerful multilingual language model for 23 languages, outperforming previous models on a wide range of tasks.Inference and efficiency · May 20242 years ago-
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse ModelsEvaluating Large Language Models using a Panel of smaller evaluators outperforms single large evaluators with reduced bias and cost.Foundation models · Apr 20242 years ago-
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsReinforcement Learning from Human Feedback (RLHF) benefits from simpler REINFORCE-style optimization methods over computationally expensive PPO or "RL-free" methods, enhancing performance and reducing cost in large language models.Alignment and safety · Feb 20242 years ago-
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Agents and evaluation · Feb 20242 years ago-
Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningThe initiative builds a human-curated instruction-following dataset spanning 65 languages and creates the largest multilingual collection of instruction-following instances through templating and translating existing datasets across 114 languages, contributing datasets and platforms for participatory research.Retrieval and data · Feb 20242 years ago-
When Less is More: Investigating Data Pruning for Pretraining LLMs at ScalePerplexity outperforms more complex methods for data quality estimation in pruning large language model training datasets, resulting in performance improvement with reduced data.Training and scaling · Sep 20233 years ago-
About and links

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.