Skip to content

CoDA: Coding LM via Diffusion Adaptation research paper by Salesforce, 2025

Salesforce · Sep 27, 2025 · Inference and efficiency · 43 upvotes 12 months ago

Read on arXiv

What it shows

CoDA, a 1.7B-parameter diffusion coder, achieves competitive performance with smaller models through confidence-guided sampling and is released with open-source tools.

Hugging Face's summary; not yet checked by hand.

More from Salesforce

All 33
PaperCitations
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM AgentsKeeps an agent's raw past runs and writes a task-specific memory only when a new task arrives, rather than deciding up front what to keep.Agents and evaluation · Sep 20263 days ago-
Flattening Every Memory Peak in Long-Context Mixture-of-Experts TrainingBounds the four memory peaks that crash long-context mixture-of-experts training, from expert dispatch to optimizer state, with fixed-size GPU schedules.Architectures · Sep 202613 days ago-
RISE: Recursive Improvement via Self-Extrapolating Policy DistillationRISE improves language model post-training by recursively generating dense token-level supervision from the model's own reinforcement learning trajectory via self-extrapolation, avoiding external teachers.Training and scaling · Sep 20263 weeks ago-
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?The study introduces a benchmark to evaluate LLM agents under evolving tool, skill, and agent harnesses, revealing persistent gaps in retention, adaptation, and harness-induced forgetting.Agents and evaluation · Sep 20263 weeks ago-
Random Attention: Rethinking KV Cache Eviction for Efficient ReasoningRandom eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved.Reasoning · Sep 20263 weeks ago-
DarwinX: Evolving Agent Harnesses Through Natural SelectionDarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.Agents and evaluation · Jul 20268 weeks ago-
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use AgentsStateAct improves computer-use agents by grounding actions, verification, and memory in direct program state rather than screenshots, boosting success rates while reducing cost.Agents and evaluation · Jul 20262 months ago-
Evidence-Backed Video Question AnsweringEvidence-Backed Video Question Answering requires models to provide answers with precise spatio-temporal segmentation evidence, revealing gaps between reasoning and visual grounding that are improved by large-scale instruction tuning.Multimodal and robotics · Jul 20262 months ago-
Topic
PaperCitations
Efficient Memory Management for Large Language Model Serving with PagedAttentionvLLM manages the KV cache like pages of virtual memory, serving models with 2 to 4 times the throughput.UC Berkeley · Sep 20233 years ago8,481
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessFlashAttention computes exact attention with far fewer GPU memory reads and writes, making long sequences faster.Stanford University · May 20224 years ago5,334
Group Sequence Policy OptimizationGroup Sequence Policy Optimization (GSPO) is a reinforcement learning algorithm that improves training efficiency and performance of large language models by using sequence-level importance ratios and operations.Alibaba (Qwen) · Jul 2025 · Unverified1 year ago688
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionNSA, a trainable sparse attention mechanism, enhances long-context modeling efficiency without sacrificing performance, achieving improvements in speed and accuracy over full attention models.DeepSeek · Feb 2025 · Unverified1 year ago507
SmolVLM: Redefining small and efficient multimodal modelsSmolVLM, a series of compact multimodal models, achieves high performance with minimal GPU memory usage, making efficient deployment on mobile and edge devices possible.Hugging Face · Apr 2025 · Unverified1 year ago294
Inference-Time Scaling for Generalist Reward ModelingSelf-Principled Critique Tuning enhances pointwise generative reward modeling for large language models, improving scalability and quality compared to existing methods.DeepSeek · Apr 2025 · Unverified1 year ago249
About this paper
Authors
Haolin Chen, Shiyu Wang, Can Qin and 12 more
arXiv
2510.03270 · PDF
Citations
Not counted yet · Semantic Scholar
Upvotes
43 · Hugging Face
Code
github.com/SalesforceAIResearch/CoDA
Lab
Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
New paperAdded to the listSep 26, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.