Skip to content
Papers.

GPT-4o System Card research paper by OpenAI, 2024

OpenAI · Oct 25, 2024 · Multimodal and robotics · 4,980 citations · 88 upvotes · unverified

Read on arXiv

What it shows

GPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.

UnverifiedHugging Face's summary; not yet checked by hand.

More from OpenAI

All 8
PaperCitations
Reasoning Models Struggle to Control their Chains of ThoughtChain-of-thought controllability measures how effectively models can be constrained to follow reasoning steps, with findings showing significantly lower controllability in reasoning versus output generation, and varying impacts from model size, training methods, and task complexity.Reasoning · Mar 2026 · Unverified21
Why Language Models HallucinateModels hallucinate because training and benchmarks reward confident guessing over saying they do not know.Alignment and safety · Sep 2025331
OpenAI o1 System CardThe o1 model series, trained with reinforcement learning and chain of thought, enhances safety and robustness by reasoning about policies, leading to superior performance on risk benchmarks while highlighting the need for robust alignment and risk management.Alignment and safety · Dec 2024 · Unverified0
GPT-4 Technical ReportA multimodal model that passes a simulated bar exam with a score around the top 10% of test takers.Foundation models · Mar 202327.3k
Training language models to follow instructions with human feedbackInstructGPT: fine-tuning on human feedback made a 1.3B model preferred over the 175B GPT-3.Alignment and safety · Mar 202224.3k
Language Models are Few-Shot LearnersGPT-3, a 175B-parameter model, does new tasks from a few examples in the prompt, with no fine-tuning.Foundation models · May 202063.7k
Scaling Laws for Neural Language ModelsLanguage model loss falls as a smooth power law as model size, data and compute grow.Training and scaling · Jan 20209,065
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
LongLive: Real-time Interactive Long Video GenerationLongLive is a frame-level autoregressive framework for real-time and interactive long video generation, addressing efficiency and quality challenges through causal attention, KV-recache, streaming long tuning, and short window attention.NVIDIA · Sep 2025 · Unverified227
About this paper
Authors
OpenAI, Aaron Hurst, Adam Lerer and 416 more
arXiv
2410.21276 · PDF
Citations
4,980, 993 influential · Semantic Scholar
Upvotes
88 · Hugging Face
Lab
OpenAI · on Companies · on Acquisitions · on Paydays · on Releases · on TechConf

Changes

What changed
Influential citationsfirst count: 993Sep 25, 2026
Citationsfirst count: 4,980Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.