Skip to content

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation research paper by Cohere, 2024

Cohere · Dec 4, 2024 · Agents and evaluation · 21 upvotes 1 year ago

Read on arXiv

What it shows

Global-MMLU addresses cultural biases and translation artifacts in multilingual datasets by involving professional annotators and offering culturally distinct subsets for more accurate model evaluation.

Hugging Face's summary; not yet checked by hand.

More from Cohere

All 30
PaperCitations
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language ReasoningOptimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language.Reasoning · Sep 20262 weeks ago-
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic ReasoningAn open competition on unseen Linguistics Olympiad puzzles, graded by the official jury; small 14B systems beat models twice their size.Reasoning · Aug 20265 weeks ago-
CALIBER: Calibrating Confidence Before and After Reasoning in Language ModelsTrains reasoning models to state their confidence twice, before thinking and after answering, halving calibration error against the best single-estimate method.Reasoning · Jun 20263 months ago-
AI Exposure Scores: what they measure, what they miss, and what comes nextReviews how widely cited scores of which jobs AI can assist are used in policy, what they miss, and newer measures that address it.Foundation models · Jun 20263 months ago-
The Culture Funnel: You Can't Align What isn't in the DataModern LLM pipelines experience a cultural data funnel where explicit cultural signals diminish during post-training, necessitating shifts in training data approaches for better cultural alignment.Alignment and safety · Jun 20263 months ago-
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft RewardsTurns each prompt into a checklist scored item by item, giving reinforcement learning partial-credit rewards for tasks that cannot be checked automatically.Foundation models · May 20264 months ago-
Agents Explore but Agents Ignore: LLMs Lack Environmental CuriosityLLM-based agents fail to exploit discovered unexpected information despite recognizing it, indicating a lack of environmental curiosity that depends on tools, compute, and training data distribution.Agents and evaluation · Apr 20265 months ago-
Tiny Aya: Bridging Scale and Multilingual DepthTiny Aya demonstrates high-quality multilingual capabilities with 3.35 billion parameters through region-aware posttraining and balanced language performance.Training and scaling · Mar 20266 months ago-
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,134
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified9 months ago761
WebWatcher: Breaking New Frontier of Vision-Language Deep Research AgentWebWatcher, a multimodal agent with enhanced visual-language reasoning, outperforms existing agents in complex visual and textual information retrieval tasks using synthetic trajectories and reinforcement learning.Alibaba (Qwen) · Aug 2025 · Unverified1 year ago117
SkillOpt: Executive Strategy for Self-Evolving Agent SkillsSkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.Microsoft · May 2026 · Unverified4 months ago85
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic ReasoningWe present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model.NVIDIA · Dec 2025 · Unverified9 months ago81
AgentFold: Long-Horizon Web Agents with Proactive Context ManagementAgentFold, a novel proactive context management paradigm, enhances long-horizon task performance through dynamic context folding, achieving superior results on benchmarks compared to larger models and proprietary agents.Alibaba (Qwen) · Oct 2025 · Unverified11 months ago77
About this paper
Authors
Shivalika Singh, Angelika Romanou, Clémentine Fourrier and 20 more
arXiv
2412.03304 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
21 · Hugging Face
Lab
Cohere · on Companies · on Acquisitions · on Releases

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.