BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation research paper by Salesforce, 2022
Salesforce · Jan 28, 2022 · Multimodal and robotics · 5 upvotes 4 years ago
What it shows
BLIP, a Vision-Language Pre-training framework, improves performance across both understanding and generation tasks by bootstrapping captions from noisy web data, achieving state-of-the-art results on image-text retrieval, image captioning, and VQA, and showing strong generalization to video-language tasks.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other multimodal and robotics papers
TopicAbout this paper
- Authors
- JunnanLi, Dongxu Li, Caiming Xiong and 1 more
- arXiv
- 2201.12086 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 5 · Hugging Face
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |