Skip to content

RDMA Point-to-Point Communication for LLM Systems research paper by Perplexity, 2025

Perplexity · Oct 31, 2025 · Inference and efficiency · 8 upvotes 11 months ago

Read on arXiv

What it shows

TransferEngine provides a uniform interface for flexible point-to-point communication in large language models, supporting disaggregated inference, reinforcement learning, and Mixture-of-Experts routing across different hardware.

Hugging Face's summary; not yet checked by hand.

Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20233 years ago8,481
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20224 years ago5,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified1 year ago688
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified1 year ago507
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified1 year ago294
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified1 year ago249
About this paper
Authors
Nandor Licker, Kevin Hu, Vladimir Zaytsev and 1 more
arXiv
2510.27656 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
8 · Hugging Face
Code
github.com/perplexityai/pplx-garden
Lab
Perplexity · on Companies · on Acquisitions · on TechConf

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.