Skip to content
Papers.

Action100M: A Large-scale Video Action Dataset research paper by Meta, 2026

Meta · Jan 15, 2026 · Retrieval and data · 12 citations · 33 upvotes · unverified

Read on arXiv

What it shows

Action100M is a large-scale video action dataset constructed from internet instructional videos using automated pipelines with V-JEPA embeddings and GPT-based reasoning for structured annotations.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Meta

All 34
PaperCitations
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataWearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.Agents and evaluation · Sep 2026 · Unverified0
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesControlled experiments on when text, image understanding and image generation help or compete when trained together.Multimodal and robotics · Aug 20262
HumanCLAW: Can Vision-Language Models Act Through a Body?HumanCLAW decouples high-level vision-language decisions from low-level motor execution to evaluate embodied action intelligence, revealing that current vision-language models lack embodied self-awareness.Multimodal and robotics · Jul 2026 · Unverified2
TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsTUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.Agents and evaluation · Jun 2026 · Unverified2
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved ReasoningProcess-driven image generation decomposes synthesis into iterative steps involving textual planning, visual drafting, textual reflection, and visual refinement, with step-wise supervision ensuring consistency and interpretability.Reasoning · Apr 2026 · Unverified5
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningV-JEPA 2.1 is a self-supervised model that learns dense visual representations for images and videos through a combination of dense predictive loss, deep self-supervision, multi-modal tokenizers, and effective scaling.Multimodal and robotics · Mar 2026 · Unverified83
ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuningResearchers address imbalance in routing weights of Mixture-of-LoRAs models by proposing Reinforcement Routing (ReMix), which uses non-learnable weights and reinforcement learning techniques to improve model expressiveness and performance.Inference and efficiency · Mar 2026 · Unverified3
Beyond Language Modeling: An Exploration of Multimodal PretrainingControlled multimodal pretraining experiments reveal key insights about unified visual representations, data complementarity, world modeling emergence, and efficient scaling through mixture-of-experts architectures.Multimodal and robotics · Mar 2026 · Unverified32
Topic
PaperCitations
From Local to Global: A Graph RAG Approach to Query-Focused SummarizationGraphRAG builds a knowledge graph of a document set so a model can answer questions about the whole collection.Microsoft · Apr 20242,198
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic DataGenerating extensive Lean 4 proof data from mathematical competition problems improved DeepSeekMath 7B's theorem-proving capabilities over GPT-4 and other methods.DeepSeek · May 2024 · Unverified260
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and RankingThe Qwen3-VL-Embedding and Qwen3-VL-Reranker models form an end-to-end multimodal search pipeline, leveraging multi-stage training and cross-attention mechanisms to achieve high-precision retrieval across diverse modalities.Alibaba (Qwen) · Jan 2026 · Unverified231
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree SearchDeepSeek-Prover-V1.5 improves theorem proving by optimizing training and inference, utilizing reinforcement learning, and proposing RMaxTS for diverse proof paths, achieving state-of-the-art results on miniF2F and ProofNet benchmarks.DeepSeek · Aug 2024 · Unverified207
Learning to Discover at Test TimeTest-time training enables AI systems to discover optimal solutions for specific scientific problems through continual learning focused on individual challenges rather than generalization.Stanford University · Jan 2026 · Unverified82
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding ModelsOpen text embedding models from 22M to 334M parameters that led MTEB retrieval for their size at release.Snowflake · May 202480
About this paper
Authors
Delong Chen, Tejaswi Kasarla, Yejin Bang and 6 more
arXiv
2601.10592 · PDF
Venue
arXiv.org
Citations
12, 2 influential · Semantic Scholar
Upvotes
33 · Hugging Face
Code
github.com/facebookresearch/Action100M
Lab
Meta · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf

Changes

What changed
Influential citationsfirst count: 2Sep 25, 2026
Citationsfirst count: 12Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.