Skip to content
Papers.

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding research paper by Google, 2018

Google · Oct 11, 2018 · Foundation models · 121,314 citations · 33 upvotes

Read on arXiv

What it shows

Pre-training a Transformer to fill in masked words, then fine-tuning it, beat task-specific models across language understanding tests.

Summarised by hand from the abstract.

More from Google

All 4
Topic
About this paper
Authors
Jacob Devlin, Ming-Wei Chang, Kenton Lee and 1 more
arXiv
1810.04805 · PDF
Venue
North American Chapter of the Association for Computational Linguistics
Citations
121,314, 22,971 influential · Semantic Scholar
Upvotes
33 · Hugging Face
Lab
Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases

Changes

What changed
New paperAdded to the listSep 24, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.