Skip to content

Adapting Vision-Language Models for E-commerce Understanding at Scale research paper by eBay, 2026

eBay · Feb 12, 2026 · Multimodal and robotics · 0 citations · 13 upvotes · unverified 7 months ago

Read on arXiv

What it shows

General-purpose Vision-Language Models can be effectively adapted for e-commerce applications through targeted techniques that enhance product understanding while maintaining broad multimodal capabilities.

By Matteo Nulli, Vladimir Orshulevich, Tala Bazazo and 9 more · arXiv 2602.11733 · PDF

UnverifiedHugging Face's summary; not yet checked by hand.

More from eBay

All 2
Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 20233 years ago15.9k
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsBLIP-2 efficiently pre-trains vision-language models using off-the-shelf frozen encoders and decoders, achieving state-of-the-art performance with fewer parameters.Salesforce · Jan 20233 years ago9,360
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationBLIP, a Vision-Language Pre-training framework, improves performance across both understanding and generation tasks by bootstrapping captions from noisy web data, achieving state-of-the-art results on image-text retrieval, image captioning, and VQA, and showing strong generalization to video-language tasks.Salesforce · Jan 20224 years ago7,410
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified1 year ago5,071
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified10 months ago1,105
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified2 years ago896
About this paper
Authors
Matteo Nulli, Vladimir Orshulevich, Tala Bazazo and 9 more
arXiv
2602.11733 · PDF
Venue
Conference of the European Chapter of the Association for Computational Linguistics
Citations
0, 0 influential · Semantic Scholar
Upvotes
13 · Hugging Face
Lab
eBay · on Companies · on Quarterly

Changes

What changed
Influential citationsfirst count: 0Oct 5, 2026today
Citationsfirst count: 0Oct 5, 2026today
New paperFound by the weekly scan, unverifiedOct 5, 2026today

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.