Google TimesFM-3: Native Multivariate Forecasting with Alternating Attention
Google's TimesFM-3 introduces native multivariate time series forecasting with alternating causal temporal and full variate attention, non-autoregressive decoding via contiguous patch masking, and support for historical and future-known covariates, achieving top rankings on Gift-Eval, FEV-Bench, and TIME benchmarks without sacrificing univariate performance.
TimesFM-3 (330M parameters, 1 trillion+ time points pre-trained) is a native multivariate zero-shot time series foundation model. A single forward pass outputs nine quantiles (10th to 90th) for the entire prediction horizon, ranking first among pre-trained foundation models on Gift-Eval, FEV-Bench, and TIME for both point and probabilistic forecasting.
Background: From Univariate to Native Multivariate
For two years, the time series foundation model (TSFM) landscape split into two camps: Google's TimesFM championed univariate forecasting — using only the target series' history and relying on pre-training breadth for zero-shot ability — while competitors like Chronos and Toto pushed multivariate and covariate capabilities that real business scenarios demand. In August 2026, Google released TimesFM-3, which natively supports three scenarios: joint multivariate forecasting, historical covariates, and future-known covariates, and switches decoding from autoregressive to a single forward pass.
Method: Alternating Attention + Masked Decoding
01 Alternating Attention: Strict Causal in Time, Fully Connected Across Variables
TimesFM-3 retains a decoder-only Transformer backbone. Data points are split into patches of length 32, each series normalized independently. The key design for multivariate handling splits attention into two types that alternate in the main stack:
Causal Temporal Attention: Tokens attend along the time dimension with strict causality, seeing only their own series' past. Cross-series future information cannot leak — a baseline requirement for zero-shot settings.
Full Variate Attention: Tokens attend across the series dimension; any time step can view all other series in the dataset, modeling cross-series correlations (e.g., how one brand's promotion affects another's sales).
02 Three Input Types, Two Token Constructions
The framework accepts three input categories:
Multiple Targets: Simultaneously forecast multiple related series.
Historical Covariates: Features known only in the past (e.g., historical foot traffic).
Future-Known Covariates (past-future/dynamic covariates): Features known for the future (e.g., scheduled promotions, weather forecasts).
The first two use a single patch per token. Future-known covariates employ a lookahead strategy: each token concatenates the current patch with the future patch, allowing the model to "legally" peek at known future signals while temporal attention remains causal. This is the most elegant design for handling covariates.
03 Non-Autoregressive Decoding: One Forward Pass Fills the Entire Horizon
TimesFM-3 uses Contiguous Patch Masking for decoding: after the observed context, masked placeholder tokens for the future horizon are appended. Target series and historical covariates are masked in the horizon (future values unknown), while future-known covariates remain visible (holidays, schedules known). Through alternating attention layers, all masked patches are filled simultaneously in a single forward pass — no iterative per-patch loop, no error accumulation. Each target series at each horizon step outputs nine quantiles (10th–90th), delivering point and probabilistic forecasts in one model.
Experiments: Two Modes More Informative Than Rankings
01 Three Benchmarks, Dual Top Rankings
On Gift-Eval, FEV-Bench, and TIME, TimesFM-3 ranks first among pre-trained foundation models for both point and probabilistic forecasting, compared against Chronos-2, Toto 2.0 series, and TimesFM-2.5, evaluated by cross-task average rank (lower is better).
02 Notable Ablation Design
Each benchmark chart shows two TimesFM-3 entries, a design more informative than the ranks themselves:
Univariate Mode: No covariates, no cross-series information; each target series processed independently. Even in this "self-handicapped" mode, TimesFM-3 matches or surpasses competitors' full-mode performance, proving the pre-training quality was not sacrificed for multivariate capability.
Full Multivariate Mode: All multivariate capabilities enabled, further boosting performance to the best average rank.
In other words, multivariate gains are purely additive, not a trade-off. Many models regress on univariate performance when adding new capabilities; TimesFM-3 does not.
03 Promotion Example: Value of Future-Known Covariates
The paper illustrates with an ice-cream sales forecast. A standard univariate model (red line) is blind to scheduled promotion days, predicting a flat trend. TimesFM-3 receives the promotion calendar as a future-known covariate (blue line), learns the historical promotion–sales uplift relationship, and applies it to future promotion days, predicting ~20% sales lift per promotion day — a problem previously solved only by post-hoc manual adjustments.
Conclusion
TimesFM's pivot signals the end of univariate as a standalone camp. The four major TSFM players (TimesFM, Chronos, Moirai, Toto) have all moved to multivariate and/or covariates; competition now centers on "how native is the multivariate implementation." TimesFM-3's answer — alternating attention and lookahead tokens — provides a direct reference for followers.
Engineering-wise, non-autoregressive decoding is the most practical contribution: a single forward pass for the full horizon yields orders-of-magnitude latency improvement for low-latency scenarios (real-time prediction services, BigQuery inline forecasting). BigQuery AI.FORECAST integration is underway (currently TimesFM-2.5 for univariate tasks); Google's typical path — paper → open source → cloud product — continues. Model and code are open (HuggingFace + GitHub), ready to use.
Boundaries must be clear: the blog claims "top-ranked pre-trained foundation model," excluding models fine-tuned on specific datasets. Training data is only described as "real + synthetic, over 1 trillion time points"; the composition of multivariate corpora (how many real multivariate series, how target-covariate pairs were constructed) is undisclosed — replicate and evaluate with caution. At 330M parameters, TimesFM-3 is mid-sized for TSFMs, pursuing "architecture innovation for capability" rather than parameter scaling.
Recommendation: run the HuggingFace version directly. If your scenario has any known future signals (rosters, promotion calendars, weather forecasts), TimesFM-3's past-future covariate interface is among the most convenient implementations available.
References:
Paper:
https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/Code:
https://github.com/google-research/timesfm/releases/tag/v3.0.0Model:
https://huggingface.co/google/timesfm-3.0-pytorchSigned-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Party THU
Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
