Skip to content
Papers.

X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests research paper by Microsoft, 2026

Microsoft · Jan 11, 2026 · Reasoning · 1 citations · 48 upvotes · unverified

Read on arXiv

What it shows

Code LLMs trained on fully synthetic data using a feature-based synthesis pipeline achieve superior performance on competitive programming benchmarks while reducing dependence on real-world coding datasets.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Microsoft

All 45
PaperCitations
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon TasksLLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run.Agents and evaluation · Sep 2026 · Unverified0
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy DistillationWe study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget.Foundation models · Sep 2026 · Unverified0
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning ModelsLarge Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones.Reasoning · Sep 2026 · Unverified0
BI-Agent and BI-Bench: Towards Automating End-to-End Business IntelligenceBusiness intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau.Agents and evaluation · Sep 2026 · Unverified0
StudentSim: Training LLM-based Student SimulatorsStudentSim trains per-student simulators that answer like a given learner and change their answers under a tutor's guidance.Applied AI · Sep 20260
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution TracesAutoSaddler automatically improves LLM agent harnesses via offline failure-driven optimization, boosting performance on long-horizon benchmarks.Agents and evaluation · Aug 2026 · Unverified2
Agent Lightning v1.0: Towards Harnessed Agentic RLAgent Lightning v1.0 enables reproducible reinforcement learning for arbitrary agent harnesses, substantially improving coding-agent performance with minimal data and compute.Agents and evaluation · Aug 2026 · Unverified2
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse PrefetchingOasisKV improves LLM inference throughput by storing full KV caches in lower memory tiers and prefetching only relevant entries into HBM using speculative-decoding lookahead predictions.Inference and efficiency · Aug 2026 · Unverified0
Topic
PaperCitations
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsAsking a model to write out its intermediate steps (chain of thought) sharply improves its math and logic answers.Google · Jan 202222.1k
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeekMath 7B improves mathematical reasoning through enhanced data pre-training and Group Relative Policy Optimization, achieving high scores on MATH benchmark without external tools.DeepSeek · Feb 2024 · Unverified8,994
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningReinforcement learning taught a model long step-by-step reasoning on par with OpenAI o1.DeepSeek · Jan 20255,694
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceDeepSeek-Coder-V2, a Mixture-of-Experts language model, excels in code-specific tasks by enhancing coding and mathematical reasoning capabilities while expanding language support and context length.DeepSeek · Jun 2024 · Unverified502
MolmoAct: Action Reasoning Models that can Reason in SpaceAction Reasoning Models (ARMs) integrate perception, planning, and control to enable adaptable and explainable robotic behavior, achieving superior performance across various tasks and settings.Ai2 · Aug 2025 · Unverified185
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningThinkAct, a dual-system framework, uses reinforced visual latent planning to enable few-shot adaptation, long-horizon planning, and self-correction in embodied AI tasks by bridging high-level reasoning with low-level action execution.NVIDIA · Jul 2025 · Unverified172
About this paper
Authors
Jie Wu, Haoling Li, Xin Zhang and 7 more
arXiv
2601.06953 · PDF
Citations
1, 0 influential · Semantic Scholar
Upvotes
48 · Hugging Face
Code
github.com/JieWu02/X-Coder
Lab
Microsoft · on Companies · on Acquisitions · on Quarterly · on Paydays · on Releases · on TechConf

Changes

What changed
Influential citationsfirst count: 0Sep 25, 2026
Citationsfirst count: 1Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.