BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models research paper by Salesforce, 2023
Salesforce · Jan 30, 2023 · Multimodal and robotics · 3 upvotes 3 years ago
What it shows
BLIP-2 efficiently pre-trains vision-language models using off-the-shelf frozen encoders and decoders, achieving state-of-the-art performance with fewer parameters.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other multimodal and robotics papers
TopicAbout this paper
- Authors
- Junnan Li, Dongxu Li, Silvio Savarese and 1 more
- arXiv
- 2301.12597 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 3 · Hugging Face
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |