Skip to content

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models research paper by AMD, 2025

AMD · Oct 16, 2025 · Agents and evaluation · 4 citations · 2 upvotes · unverified 11 months ago

Read on arXiv

What it shows

XModBench evaluates cross-modal consistency in OLLMs, revealing challenges in spatial and temporal reasoning, modality disparities, and directional imbalance.

UnverifiedHugging Face's summary; not yet checked by hand.

More from AMD

All 6
PaperCitations
Instella-MoE Technical ReportInstella-MoE is an open Mixture-of-Experts language model trained on AMD GPUs using sparse activation, Gated Multi-head Latent Attention, and multi-stage post-training to achieve strong benchmark performance with full reproducibility.Foundation models · Sep 2026 · Unverified3 weeks ago0
Stabilizing Efficient Reasoning with Step-Level Advantage SelectionShort-context post-training induces reasoning compression but causes instability; Step-level Advantage Selection addresses this by selectively adjusting reasoning steps based on confidence and verification outcomes, improving accuracy-efficiency trade-off in reasoning tasks.Reasoning · Apr 2026 · Unverified5 months ago1
Dynamic Chunking Diffusion TransformerDynamic Chunking Diffusion Transformer adapts token sequence length based on image content and diffusion timestep, improving efficiency and performance over fixed-token approaches.Architectures · Mar 2026 · Unverified6 months ago2
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceDUET-VLM presents a dual compression framework that reduces visual tokens while maintaining high accuracy in vision-language models through coordinated vision and language backbone processing.Inference and efficiency · Feb 2026 · Unverified7 months ago0
Instella: Fully Open Language Models with Stellar PerformanceInstella, a family of fully open large language models, achieves state-of-the-art performance using open data and is competitive with leading open-weight models, with specialized variants for long context and mathematical reasoning.Alignment and safety · Nov 2025 · Unverified10 months ago6
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,157
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified10 months ago766
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Cohere · Feb 20242 years ago410
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.ServiceNow · Mar 20242 years ago359
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Scale AI · Jul 20251 year ago311
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Scale AI · Sep 20251 year ago246
About this paper
Authors
Xingrui Wang, Jiang Liu, Chao Huang and 7 more
arXiv
2510.15148 · PDF
Venue
arXiv.org
Citations
4, 0 influential · Semantic Scholar
Upvotes
2 · Hugging Face
Code
github.com/XingruiWang/XModBench
Lab
AMD · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 28, 2026today
Citationsfirst count: 4Sep 28, 2026today
New paperFound by the weekly scan, unverifiedSep 28, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.