BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding research paper by Google, 2018
Google · Oct 11, 2018 · Foundation models · 121,314 citations · 33 upvotes
What it shows
Pre-training a Transformer to fill in masked words, then fine-tuning it, beat task-specific models across language understanding tests.
Summarised by hand from the abstract.
More from Google
All 4Other foundation models papers
TopicAbout this paper
- Authors
- Jacob Devlin, Ming-Wei Chang, Kenton Lee and 1 more
- arXiv
- 1810.04805 · PDF
- Venue
- North American Chapter of the Association for Computational Linguistics
- Citations
- 121,314, 22,971 influential · Semantic Scholar
- Upvotes
- 33 · Hugging Face
- Lab
- Google · on Companies · on Acquisitions · on Paydays · on TechConf · on Releases
Changes
| What changed | |
|---|---|
| Sep 24, 2026 | New paperAdded to the listSep 24, 2026 |
Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.