Minkey Chang

Data Scientist

All posts

Could multivariate time series have their own representations?

For images and text, people mostly agree on what a good representation should do. For multivariate time series it is less obvious. You might want something low-dimensional that carries time structure, helps forecasting, and, if you care about causality, does not change meaning under every reparameterisation. Dynamic factor models and deep forecasters both compress the panel, but they aim at different goals, and identifiability is where the two stories part ways. This post is about iVDFM, the identifiable variational dynamic factor model from a paper I wrote with J.-Y. Kim. The short version: on synthetic data with temporal structure it recovers the true factors better than the baselines, interventions behave as the theory predicts until the dynamics couple factors, and forecasting quality stays competitive. The rest is what that does and does not mean.


Where the friction is

A model can learn an embedding that predicts well while the axes of that embedding stay arbitrary: rotate or warp the latent space and the forecast barely changes. That is fine if all you want is next-step error. It is awkward if you want to talk about factors as stable objects across runs, or to read a move along one coordinate as a shock with a fixed meaning, which is the concern behind causal representation learning.

The useful way to make this precise is the size of the residual ambiguity group. Classical Gaussian dynamic factor models and linear state-space models are identified only up to the full general linear group GL(r)\mathrm{GL}(r) on the factors: any invertible rotation of the latent space leaves the observation law unchanged. Identifiable VAEs shrink that group, in a static setting, to the much smaller class T\mathcal{T} of permutations and component-wise affine maps, by conditioning the latent prior on observed auxiliary variables. The hard part is time: how to carry T\mathcal{T} through stochastic dynamics without reintroducing a free rotation at every step.


How iVDFM is put together

The idea is to put identifiability on the innovations ηt\boldsymbol{\eta}_t, the shocks that drive the system, rather than on a loosely defined state, and to use dynamics simple enough that whatever is identified at the innovation level can still be read off in the factors ft\boldsymbol{f}_t. The model has three parts: an encoder and prior for innovations, diagonal linear dynamics that propagate them to factors, and a decoder from factors to observations.

Three-panel iVDFM diagram. (a) Extraction: encoder maps observations and auxiliary context to a variational distribution over innovations, with a matched conditional prior. (b) Propagation: innovations drive factors through diagonal linear dynamics. (c) Forecasting and generation: factors are decoded to observations or forecasts.
iVDFM in three steps. (a) Extraction: observations and auxiliary context map to innovations through a matched encoder/prior. (b) Propagation: innovations drive factors via diagonal time-varying linear dynamics. (c) Forecasting and generation: factors are decoded to reconstructions or forecasts.

Innovations get a conditional exponential-family prior that depends on auxiliary variables ut\mathbf{u}_t (calendar, covariates) and a regime embedding et\mathbf{e}_t, itself a deterministic soft mixture of learnable embeddings weighted by a small network on ut\mathbf{u}_t. Keeping (ut,et)(\mathbf{u}_t, \mathbf{e}_t) a deterministic function of observed context is what keeps the iVAE-style identifiability argument valid. Under enough variation in the prior's natural parameters, the components of ηt\boldsymbol{\eta}_t are identifiable up to T\mathcal{T}. Gaussian innovations defeat the argument in practice, for the same reason independent component analysis needs non-Gaussian sources: the Gaussian family is closed under rotation, so coordinates can mix back together. The implementation uses non-Gaussian innovations such as Laplace.

Dynamics are linear and diagonal: ft+1=Aˉtft+Bˉtηt\boldsymbol{f}_{t+1} = \bar{A}_t \boldsymbol{f}_t + \bar{B}_t \boldsymbol{\eta}_t, with both matrices formed from the same regime weights and each diagonal. Every factor coordinate then depends only on its own history and its own innovation, so the ambiguity class T\mathcal{T} transfers unchanged from innovations to factor trajectories. This is a small propagation lemma, and it fails as soon as the dynamics mix factors or the innovations are Gaussian. Higher-order autoregressive dependence fits the same picture through the usual companion-form stacking.

Observations come from yt=g(ft)+εt\boldsymbol{y}_t = g(\boldsymbol{f}_t) + \boldsymbol{\varepsilon}_t with an injective MLP decoder and fixed-scale noise. Training is standard variational inference: infer innovations, roll the dynamics forward, and maximise the ELBO along the Markov structure.


What I looked at

Partial identifiability is supposed to buy three concrete things: faithful factor recovery, well-defined interventions, and no collapse in forecast quality. The experiments are organised around those three.

Synthetic factor recovery. On a dynamic data-generating process (AR dynamics driven by innovations, T=200T{=}200, N=20N{=}20, r=5r{=}5, ten seeds), iVDFM gave the highest mean correlation with the true factors (MCC about 0.65) and trace R2R^2 (about 0.82) against DDFM, iVAE, VAE and DFM. On a static iVAE-style process, a plain VAE was the strongest baseline and iVDFM only matched iVAE. The gap between the two settings is where the identification signal lives: temporal innovation structure is what iVDFM exploits, so the dynamic case is the regime where the mechanism is active.

Side-by-side scatterplots of recovered versus ground-truth factors on the dynamic DGP for iVDFM, DDFM, iVAE, VAE, and DFM, with per-factor correlations annotated.
Factor recovery on the dynamic DGP: recovered vs. ground-truth factors after MCC matching, with iVDFM at the top of the row.

Synthetic interventions. On synthetic structural causal models, dodo-interventions on innovation components give model-implied impulse responses that can be compared with the truth on error, sign accuracy and correlation. On the base linear and the regime-switching models, errors stay bounded and sign and correlation are workable. On the chain model, whose dynamics violate the diagonality assumption, impulse response error roughly doubles and sign and correlation fidelity drop. That is the failure mode the propagation lemma predicts: cross-factor interactions break the carry-through of T\mathcal{T}, and chain-like structure points to richer transition models as the natural relaxation.

Forecasting. On ETTh1/2, ETTm1/2 and Weather at horizons 96 to 720, probabilistic scores (CRPS and standardised MSE) put iVDFM in the same neighbourhood as iTransformer, TimeMixer, TimeXer and DDFM: competitive on CRPS, not always best on MSE. The lesson is narrow. The constraints used to obtain partial identifiability do not destroy distributional forecast quality on these benchmarks, and they do not turn iVDFM into a forecasting specialist either.


A small qualitative case study

Two real panels make the claim that loadings stay portable across runs concrete.

iVDFM factor analysis on the exchange-rate panel: loadings of two recovered factors on AUD, JPY, CAD, CNY, NZD against USD, plus the corresponding factor trajectories over time.
Two-factor iVDFM on a real panel: loadings on individual series and the corresponding factor trajectories.

On daily exchange rates for AUD, JPY, CAD, CNY and NZD against USD, two factors separate a broad market-wide USD component, loading 0.85 to 0.94 on the freely floating rates and weakly on CNY, from a secondary re-pricing direction with moderate negative loadings on the same currencies. On weekly influenza-like-illness surveillance, the same setup yields one coordinate that moves with shared epidemic burden, correlating strongly with both ILI rates and outpatient visits, and a second, weaker coordinate for utilisation-linked residual variation. The operational point of partial identifiability is that under T\mathcal{T} the burden axis stays attached to the same series across retraining, so dashboards and downstream rules do not have to re-learn which coordinate means what after each refit.


When it is worth it

If all you need is a good number on one benchmark horizon, a forecast-first model is simpler to ship. iVDFM is for settings where you also care whether the latent axes are more than a rotating embedding, and where you may later want to connect latents to regimes or shocks in factor-model language. That pays off only when the auxiliaries move the innovation prior enough to identify, when non-Gaussian innovations are acceptable, and when diagonal dynamics with fixed-scale observation noise are not badly wrong. None of that is automatic, and the chain result shows what happens when the last assumption fails: couple the factors and the residual ambiguity grows back.

The broader point is that multivariate series can be modelled with shocks and states in mind, not only with next-row prediction. Partial identifiability up to T\mathcal{T} is the one thing iVDFM adds over a Gaussian factor model or a generic VAE, and the experiments above are an attempt to measure what that one thing buys.


References

  1. iVDFM — Chang, M., & Kim, J.-Y. (2026). Conditionally Identifiable Latent Representation for Multivariate Time Series with Structural Dynamics. ICLR 2026 FinAI Workshop, long paper, and ProbML 2026 proceedings. arXiv:2603.22886.

  2. Dynamic factor models — Stock, J. H., & Watson, M. W. (2002). Macroeconomic Forecasting Using Diffusion Indexes. Journal of Business & Economic Statistics, 20(2), 147–162.

  3. iVAE / identifiable latents — Khemakhem, I., Kingma, D., Monti, R., & Hyvärinen, A. (2020). Variational Autoencoders and Nonlinear ICA: A Unifying Framework. AISTATS.

  4. ICA / non-Gaussianity — Hyvärinen, A., Karhunen, J., & Oja, E. (2001). Independent Component Analysis. Wiley.

  5. Deep dynamic factors — Andreini, P., Izzo, C., & Ricco, G. (2020). Deep Dynamic Factor Models. Working paper.

  6. Causal representation — Schölkopf, B., et al. (2021). Toward Causal Representation Learning. Proceedings of the IEEE, 109(5), 612–634.

Other posts