Skip to content

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation research paper by Amazon, 2026

Amazon · Sep 29, 2026 · Agents and evaluation · 35 upvotes · unverified 6 days ago

Read on arXiv

What it shows

Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style.

By Haibo Jin, Xinjie Li, Najmeh Sadoughi and 4 more · arXiv 2609.38660 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from Amazon

All 8
PaperCitations
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual EvidenceLatent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words.Reasoning · Sep 2026 · Unverified7 days ago-
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RLAgent trajectories record what an agent does and what happens next.Agents and evaluation · Sep 2026 · Unverified2 weeks ago0
VIDEOP2R: Video Understanding from Perception to ReasoningVideoP2R, a process-aware reinforcement fine-tuning framework, improves video reasoning and understanding by modeling perception and reasoning separately, achieving state-of-the-art results on multiple benchmarks.Reasoning · Nov 2025 · Unverified10 months ago10
Adaptive Multi-Agent Response Refinement in Conversational SystemsA multi-agent framework enhances conversational quality by refining responses through agents responsible for factuality, personalization, and coherence, outperforming existing methods on challenging datasets.Agents and evaluation · Nov 2025 · Unverified10 months ago3
Chronos-2: From Univariate to Universal ForecastingChronos-2, a pretrained model with a group attention mechanism, achieves state-of-the-art performance in zero-shot univariate, multivariate, and covariate-informed forecasting tasks.Inference and efficiency · Oct 2025 · Unverified11 months ago200
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMsThought templates enhance long-context language models by structuring evidence combination and guiding multi-hop inference, leading to consistent performance improvements across various benchmarks.Reasoning · Oct 2025 · Unverified12 months ago4
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular ReasoningTaTToo, a novel table-grounded Process Reward Model, enhances tabular reasoning by explicitly addressing table-specific operations and integrating tool-based verification, leading to significant performance improvements over existing PRMs.Reasoning · Oct 2025 · Unverified12 months ago13
Topic
PaperCitations
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGemini 2.X model family, including Gemini 2.5 Pro and Flash, offers superior coding, reasoning, and multimodal understanding capabilities across a range of computational efficiencies.Google · Jul 2025 · Unverified1 year ago4,244
DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek-V3.2 introduces DeepSeek Sparse Attention and a scalable reinforcement learning framework, achieving superior reasoning and performance compared to GPT-5 and Gemini-3.0-Pro in complex reasoning tasks.DeepSeek · Dec 2025 · Unverified10 months ago784
Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAya, a multilingual generative language model supporting over 50% lower-resourced languages, excels in both generative and discriminative tasks across 99 languages, and provides extensive evaluations and open-source resources.Cohere · Feb 20242 years ago410
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?WorkArena and BrowserGym evaluate large language model-based agents' ability to perform enterprise software tasks, revealing gaps in current agent capabilities and differences between open and closed-source LLMs.ServiceNow · Mar 20242 years ago387
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsRubrics as Rewards (RaR) framework uses structured rubrics as interpretable reward signals for on-policy training, improving performance in real-world reinforcement learning tasks with subjective criteria.Scale AI · Jul 20251 year ago343
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?SWE-Bench Pro is a challenging benchmark for coding models, featuring complex, enterprise-level problems that require substantial code modifications, with performance evaluations showing significant limitations in current models.Scale AI · Sep 20251 year ago270
About this paper
Authors
Haibo Jin, Xinjie Li, Najmeh Sadoughi and 4 more
arXiv
2609.38660 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
35 · Hugging Face
Lab
Amazon · on Companies · on Quarterly · on Paydays

Changes

What changed
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.