Skip to content

Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT research paper by eBay, 2026

eBay · Jan 28, 2026 · Inference and efficiency · 0 citations · 2 upvotes · unverified 8 months ago

Read on arXiv

What it shows

OptiKIT is a distributed framework that automates LLM optimization processes, enabling efficient GPU resource utilization and consistent performance improvements across heterogeneous infrastructure without requiring specialized optimization expertise.

By Nicholas Santavas, Kareem Eissa, Patrycja Cieplicka and 6 more · arXiv 2601.20408 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from eBay

All 2
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20233 years ago8,815
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20224 years ago5,460
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified1 year ago722
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified1 year ago531
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified1 year ago305
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified1 year ago251
About this paper
Authors
Nicholas Santavas, Kareem Eissa, Patrycja Cieplicka and 6 more
arXiv
2601.20408 · PDF
Venue
arXiv.org
Citations
0, 0 influential · Semantic Scholar
Upvotes
2 · Hugging Face
Lab
eBay · on Companies · on Quarterly

Changes

What changed
Influential citationsfirst count: 0Oct 5, 2026today
Citationsfirst count: 0Oct 5, 2026today
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.