Skip to content
Papers.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation research paper by Character.AI, 2025

Character.AI · Sep 30, 2025 · Multimodal and robotics · 74 citations · 37 upvotes · unverified

Read on arXiv

What it shows

Ovi is a unified audio-video generation model using twin-DiT modules with blockwise cross-modal fusion, enabling natural synchronization and high-quality multimodal outputs.

UnverifiedHugging Face's summary; not yet checked by hand.

Topic
PaperCitations
Segment AnythingA promptable model and a dataset of over a billion masks that cut out any object in any image.Meta · Apr 202315.7k
GPT-4o System CardGPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.OpenAI · Oct 2024 · Unverified4,980
SAM 3: Segment Anything with ConceptsSegment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.Meta · Nov 2025 · Unverified999
DeepSeek-VL: Towards Real-World Vision-Language UnderstandingDeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.DeepSeek · Mar 2024 · Unverified889
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationJanus, an autoregressive framework with separate visual encoding pathways within a unified transformer architecture, enhances performance in unified multimodal understanding and generation.DeepSeek · Oct 2024 · Unverified481
SAM 3D: 3Dfy Anything in ImagesSAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.Meta · Nov 2025 · Unverified256
About this paper
Authors
Chetwin Low, Weimin Wang, Calder Katyal
arXiv
2510.01284 · PDF
Venue
arXiv.org
Citations
74, 31 influential · Semantic Scholar
Upvotes
37 · Hugging Face
Code
github.com/character-ai/Ovi
Lab
Character.AI · on Companies · on Acquisitions

Changes

What changed
Influential citationsfirst count: 31Sep 25, 2026
Citationsfirst count: 74Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.