Evidence-Backed Video Question Answering research paper by Salesforce, 2026
Salesforce · Jul 13, 2026 · Multimodal and robotics · 4 upvotes 2 months ago
What it shows
Evidence-Backed Video Question Answering requires models to provide answers with precise spatio-temporal segmentation evidence, revealing gaps between reasoning and visual grounding that are improved by large-scale instruction tuning.
Hugging Face's summary; not yet checked by hand.
More from Salesforce
All 33Other multimodal and robotics papers
TopicAbout this paper
- Authors
- Shijie Wang, Honglu Zhou, Ziyang Wang and 5 more
- arXiv
- 2607.11862 · PDF
- Citations
- Not counted yet · Semantic Scholar
- Upvotes
- 4 · Hugging Face
- Lab
- Salesforce · on Companies · on Acquisitions · on Quarterly · on Paydays · on TechConf
Changes
| What changed | |
|---|---|
| Sep 26, 2026today | New paperAdded to the listSep 26, 2026today |