Skip to content

Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning research paper by Cohere, 2026

Cohere · Sep 9, 2026 · Reasoning · 33 upvotes 2 weeks ago

Read on arXiv

What it shows

Optimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language.

Hugging Face's summary; not yet checked by hand.

More from Cohere

All 30
PaperCitations
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic ReasoningAn open competition on unseen Linguistics Olympiad puzzles, graded by the official jury; small 14B systems beat models twice their size.Reasoning · Aug 20265 weeks ago-
CALIBER: Calibrating Confidence Before and After Reasoning in Language ModelsTrains reasoning models to state their confidence twice, before thinking and after answering, halving calibration error against the best single-estimate method.Reasoning · Jun 20263 months ago-
AI Exposure Scores: what they measure, what they miss, and what comes nextReviews how widely cited scores of which jobs AI can assist are used in policy, what they miss, and newer measures that address it.Foundation models · Jun 20263 months ago-
The Culture Funnel: You Can't Align What isn't in the DataModern LLM pipelines experience a cultural data funnel where explicit cultural signals diminish during post-training, necessitating shifts in training data approaches for better cultural alignment.Alignment and safety · Jun 20263 months ago-
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft RewardsTurns each prompt into a checklist scored item by item, giving reinforcement learning partial-credit rewards for tasks that cannot be checked automatically.Foundation models · May 20264 months ago-
Agents Explore but Agents Ignore: LLMs Lack Environmental CuriosityLLM-based agents fail to exploit discovered unexpected information despite recognizing it, indicating a lack of environmental curiosity that depends on tools, compute, and training data distribution.Agents and evaluation · Apr 20265 months ago-
Tiny Aya: Bridging Scale and Multilingual DepthTiny Aya demonstrates high-quality multilingual capabilities with 3.35 billion parameters through region-aware posttraining and balanced language performance.Training and scaling · Mar 20266 months ago-
CIRCLE: A Framework for Evaluating AI from a Real-World LensA six-stage framework for measuring how deployed AI behaves and affects organisations, linking field tests and red teaming to usable metrics.Agents and evaluation · Feb 20267 months ago-
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 20224 years ago22.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified2 years ago8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20251 year ago5,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified2 years ago502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified1 year ago185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified1 year ago172
About this paper
Authors
Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca and 5 more
arXiv
2609.10445 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
33 · Hugging Face
Lab
Cohere · on Companies · on Acquisitions · on Releases

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.