Skip to content
Papers.

Toto: Time Series Optimized Transformer for Observability research paper by Datadog, 2024

Datadog · Jul 10, 2024 · Architectures · 43 citations · 33 upvotes · unverified

Read on arXiv

What it shows

Toto, a Time Series Optimized Transformer for Observability, achieves state-of-the-art performance in observability and general-purpose forecasting using a vast dataset of time series data.

UnverifiedHugging Face's summary; not yet checked by hand.

More from Datadog

All 3
Topic
PaperCitations
Attention Is All You NeedIntroduced the Transformer, the attention-only architecture behind nearly every large language model since.Google · Jun 2017194k
Mamba: Linear-Time Sequence Modeling with Selective State SpacesA selective state space model that scales linearly with sequence length and matches Transformers on language.Carnegie Mellon University · Dec 20239,203
VGGT: Visual Geometry Grounded TransformerVGGT, a feed-forward neural network, efficiently infers multiple 3D attributes from single or multiple views, outperforming alternatives and enhancing downstream tasks without post-processing.Meta · Mar 2025 · Unverified1,798
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsThe DeepSeekMoE architecture improves expert specialization in Mixture-of-Experts models by segmenting experts and isolating shared ones, achieving better performance and computational efficiency compared to GShard and other models.DeepSeek · Jan 2024 · Unverified1,123
Hymba: A Hybrid-head Architecture for Small Language ModelsHymba, a family of small language models with a hybrid-head architecture combining transformer attention and state space models, achieves state-of-the-art performance with improved efficiency and reduced cache size.NVIDIA · Nov 2024 · Unverified115
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLMOmniVinci, an open-source omni-modal LLM, enhances cross-modal understanding and performance across audio, vision, and robotics applications with innovative architecture and efficient data curation.NVIDIA · Oct 2025 · Unverified58
About this paper
Authors
Ben Cohen, Emaad Khwaja, Kan Wang and 4 more
arXiv
2407.07874 · PDF
Venue
arXiv.org
Citations
43, 9 influential · Semantic Scholar
Upvotes
33 · Hugging Face
Lab
Datadog · on Companies · on Acquisitions · on Quarterly · on Paydays · on Releases · on TechConf

Changes

What changed
Influential citationsfirst count: 9Sep 25, 2026
Citationsfirst count: 43Sep 25, 2026
New paperFound by the weekly scan, unverifiedSep 25, 2026

Sources: each lab's own papers and arXiv, with citation and upvote counts from Semantic Scholar and Hugging Face. One-line summaries are for orientation, not a substitute for the paper. Logos via logo.dev; trademarks belong to their owners.

New papers by email

Monday afternoons, only in weeks with new papers from the labs.

Double opt-in. Unsubscribe any time.