Skip to content

Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains research paper by Salesforce, 2025

Salesforce · Oct 20, 2025 · Reasoning · 4 upvotes 11 months ago

Read on arXiv

What it shows

FARE, a family of large-scale parameter evaluators, surpasses specialized RL-trained evaluators in both static benchmarks and real-world tasks through data-driven development and iterative rejection-sampling supervised finetuning.

Hugging Face's summary; not yet checked by hand.

More from Salesforce

All 33
PaperCitations
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM AgentsKeeps an agent's raw past runs and writes a task-specific memory only when a new task arrives, rather than deciding up front what to keep.Agents and evaluation · Sep 20263 days ago-
Flattening Every Memory Peak in Long-Context Mixture-of-Experts TrainingBounds the four memory peaks that crash long-context mixture-of-experts training, from expert dispatch to optimizer state, with fixed-size GPU schedules.Architectures · Sep 202613 days ago-
RISE: Recursive Improvement via Self-Extrapolating Policy DistillationRISE improves language model post-training by recursively generating dense token-level supervision from the model's own reinforcement learning trajectory via self-extrapolation, avoiding external teachers.Training and scaling · Sep 20263 weeks ago-
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?The study introduces a benchmark to evaluate LLM agents under evolving tool, skill, and agent harnesses, revealing persistent gaps in retention, adaptation, and harness-induced forgetting.Agents and evaluation · Sep 20263 weeks ago-
Random Attention: Rethinking KV Cache Eviction for Efficient ReasoningRandom eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved.Reasoning · Sep 20263 weeks ago-
DarwinX: Evolving Agent Harnesses Through Natural SelectionDarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.Agents and evaluation · Jul 20268 weeks ago-
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use AgentsStateAct improves computer-use agents by grounding actions, verification, and memory in direct program state rather than screenshots, boosting success rates while reducing cost.Agents and evaluation · Jul 20262 months ago-
Evidence-Backed Video Question AnsweringEvidence-Backed Video Question Answering requires models to provide answers with precise spatio-temporal segmentation evidence, revealing gaps between reasoning and visual grounding that are improved by large-scale instruction tuning.Multimodal and robotics · Jul 20262 months ago-
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 20224 years ago22.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified2 years ago8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20251 year ago5,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified2 years ago502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified1 year ago185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified1 year ago172
About this paper
Authors
Austin Xu, Xuan-Phi Nguyen, Yilun Zhou and 3 more
arXiv
2510.17793 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
4 · Hugging Face
Lab
Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.