Skip to content

Multi-task retriever fine-tuning for domain-specific and efficient RAG research paper by ServiceNow, 2025

ServiceNow · Jan 8, 2025 · Inference and efficiency · 10 upvotes · unverified 1 year ago

Read on arXiv

What it shows

Instruction-tuned retrieval encoder addresses domain-specific challenges for efficient, scalable, and fast Retrieval-Augmented Generation (RAG) applications.

UnverifiedHugging Face's summary; not yet checked by hand.

More from ServiceNow

All 20
PaperCitations
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-CallingAgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects.Agents and evaluation · Aug 2026 · Unverified4 weeks ago-
StarHarness: Evolving Harnesses with Stratified Search for Enterprise EnvironmentsStarHarness evolves fixed-weight agent harnesses via stratified task pools and hidden selection to improve enterprise tool-use performance and cross-model transfer.Retrieval and data · Aug 2026 · Unverified4 weeks ago-
SynthDocBench: Controlled Benchmark for Long-Context Visual Document UnderstandingSynthDocBench is a synthetic long-context visual document benchmark that isolates failure modes in vision-language models, revealing sharp length degradation, positional sensitivity, and chart comprehension breakdowns.Agents and evaluation · Jul 2026 · Unverified2 months ago-
EVA-Bench: A New End-to-end Framework for Evaluating Voice AgentsEVA-Bench presents a comprehensive evaluation framework for voice agents that simulates realistic conversations and measures performance across multiple voice-specific failure modes using novel accuracy and experience metrics.Agents and evaluation · May 2026 · Unverified4 months ago-
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer DynamicsEnterprise discovery agents that read system configuration at runtime outperform traditional world models in configurable environments where dynamics change over time.Agents and evaluation · May 2026 · Unverified4 months ago-
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient ReasoningApriel-Reasoner is a 15B-parameter language model trained with reproducible multi-domain reinforcement learning to improve reasoning efficiency and accuracy across diverse tasks while reducing inference costs.Reasoning · Apr 20265 months ago-
Therefore I am. I ThinkReasoning models appear to encode action choices before beginning textual deliberation, as evidenced by early decision detection and activation steering effects.Reasoning · Apr 2026 · Unverified5 months ago-
Terminal Agents Suffice for Enterprise AutomationSimple terminal-based coding agents using programmatic interfaces and foundation models can effectively perform enterprise tasks comparable to or better than complex tool-augmented agents.Agents and evaluation · Mar 2026 · Unverified5 months ago-
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20233 years ago8,481
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20224 years ago5,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified1 year ago688
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified1 year ago507
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified1 year ago294
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified1 year ago249
About this paper
Authors
Patrice Béchard, Orlando Marquez Ayala
arXiv
2501.04652 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
10 · Hugging Face
Lab
ServiceNow · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperFound by the weekly scan, unverifiedSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.