Could multivariate time series have their own representations?
For images and text, people mostly agree on what a good representation should do. For multivariate time series it is less obvious. You might want something low-dimensional that carries time structure, helps forecasting, and, if you care about causality, does not change meaning under every reparameterisation. Dynamic factor models and deep forecasters both compress the panel, but they aim at different goals, and identifiability is where the two stories part ways. This post is about iVDFM, the identifiable variational dynamic factor model from a paper I wrote with J.-Y. Kim. The short version: on synthetic data with temporal structure it recovers the true factors better than the baselines, interventions behave as the theory predicts until the dynamics couple factors, and forecasting quality stays competitive. The rest is what that does and does not mean.
Where the friction is
A model can learn an embedding that predicts well while the axes of that embedding stay arbitrary: rotate or warp the latent space and the forecast barely changes. That is fine if all you want is next-step error. It is awkward if you want to talk about factors as stable objects across runs, or to read a move along one coordinate as a shock with a fixed meaning, which is the concern behind causal representation learning.
The useful way to make this precise is the size of the residual ambiguity group. Classical Gaussian dynamic factor models and linear state-space models are identified only up to the full general linear group on the factors: any invertible rotation of the latent space leaves the observation law unchanged. Identifiable VAEs shrink that group, in a static setting, to the much smaller class of permutations and component-wise affine maps, by conditioning the latent prior on observed auxiliary variables. The hard part is time: how to carry through stochastic dynamics without reintroducing a free rotation at every step.
How iVDFM is put together
The idea is to put identifiability on the innovations , the shocks that drive the system, rather than on a loosely defined state, and to use dynamics simple enough that whatever is identified at the innovation level can still be read off in the factors . The model has three parts: an encoder and prior for innovations, diagonal linear dynamics that propagate them to factors, and a decoder from factors to observations.

Innovations get a conditional exponential-family prior that depends on auxiliary variables (calendar, covariates) and a regime embedding , itself a deterministic soft mixture of learnable embeddings weighted by a small network on . Keeping a deterministic function of observed context is what keeps the iVAE-style identifiability argument valid. Under enough variation in the prior's natural parameters, the components of are identifiable up to . Gaussian innovations defeat the argument in practice, for the same reason independent component analysis needs non-Gaussian sources: the Gaussian family is closed under rotation, so coordinates can mix back together. The implementation uses non-Gaussian innovations such as Laplace.
Dynamics are linear and diagonal: , with both matrices formed from the same regime weights and each diagonal. Every factor coordinate then depends only on its own history and its own innovation, so the ambiguity class transfers unchanged from innovations to factor trajectories. This is a small propagation lemma, and it fails as soon as the dynamics mix factors or the innovations are Gaussian. Higher-order autoregressive dependence fits the same picture through the usual companion-form stacking.
Observations come from with an injective MLP decoder and fixed-scale noise. Training is standard variational inference: infer innovations, roll the dynamics forward, and maximise the ELBO along the Markov structure.
What I looked at
Partial identifiability is supposed to buy three concrete things: faithful factor recovery, well-defined interventions, and no collapse in forecast quality. The experiments are organised around those three.
Synthetic factor recovery. On a dynamic data-generating process (AR dynamics driven by innovations, , , , ten seeds), iVDFM gave the highest mean correlation with the true factors (MCC about 0.65) and trace (about 0.82) against DDFM, iVAE, VAE and DFM. On a static iVAE-style process, a plain VAE was the strongest baseline and iVDFM only matched iVAE. The gap between the two settings is where the identification signal lives: temporal innovation structure is what iVDFM exploits, so the dynamic case is the regime where the mechanism is active.

Synthetic interventions. On synthetic structural causal models, -interventions on innovation components give model-implied impulse responses that can be compared with the truth on error, sign accuracy and correlation. On the base linear and the regime-switching models, errors stay bounded and sign and correlation are workable. On the chain model, whose dynamics violate the diagonality assumption, impulse response error roughly doubles and sign and correlation fidelity drop. That is the failure mode the propagation lemma predicts: cross-factor interactions break the carry-through of , and chain-like structure points to richer transition models as the natural relaxation.
Forecasting. On ETTh1/2, ETTm1/2 and Weather at horizons 96 to 720, probabilistic scores (CRPS and standardised MSE) put iVDFM in the same neighbourhood as iTransformer, TimeMixer, TimeXer and DDFM: competitive on CRPS, not always best on MSE. The lesson is narrow. The constraints used to obtain partial identifiability do not destroy distributional forecast quality on these benchmarks, and they do not turn iVDFM into a forecasting specialist either.
A small qualitative case study
Two real panels make the claim that loadings stay portable across runs concrete.

On daily exchange rates for AUD, JPY, CAD, CNY and NZD against USD, two factors separate a broad market-wide USD component, loading 0.85 to 0.94 on the freely floating rates and weakly on CNY, from a secondary re-pricing direction with moderate negative loadings on the same currencies. On weekly influenza-like-illness surveillance, the same setup yields one coordinate that moves with shared epidemic burden, correlating strongly with both ILI rates and outpatient visits, and a second, weaker coordinate for utilisation-linked residual variation. The operational point of partial identifiability is that under the burden axis stays attached to the same series across retraining, so dashboards and downstream rules do not have to re-learn which coordinate means what after each refit.
When it is worth it
If all you need is a good number on one benchmark horizon, a forecast-first model is simpler to ship. iVDFM is for settings where you also care whether the latent axes are more than a rotating embedding, and where you may later want to connect latents to regimes or shocks in factor-model language. That pays off only when the auxiliaries move the innovation prior enough to identify, when non-Gaussian innovations are acceptable, and when diagonal dynamics with fixed-scale observation noise are not badly wrong. None of that is automatic, and the chain result shows what happens when the last assumption fails: couple the factors and the residual ambiguity grows back.
The broader point is that multivariate series can be modelled with shocks and states in mind, not only with next-row prediction. Partial identifiability up to is the one thing iVDFM adds over a Gaussian factor model or a generic VAE, and the experiments above are an attempt to measure what that one thing buys.
References
-
iVDFM — Chang, M., & Kim, J.-Y. (2026). Conditionally Identifiable Latent Representation for Multivariate Time Series with Structural Dynamics. ICLR 2026 FinAI Workshop, long paper, and ProbML 2026 proceedings. arXiv:2603.22886.
-
Dynamic factor models — Stock, J. H., & Watson, M. W. (2002). Macroeconomic Forecasting Using Diffusion Indexes. Journal of Business & Economic Statistics, 20(2), 147–162.
-
iVAE / identifiable latents — Khemakhem, I., Kingma, D., Monti, R., & Hyvärinen, A. (2020). Variational Autoencoders and Nonlinear ICA: A Unifying Framework. AISTATS.
-
ICA / non-Gaussianity — Hyvärinen, A., Karhunen, J., & Oja, E. (2001). Independent Component Analysis. Wiley.
-
Deep dynamic factors — Andreini, P., Izzo, C., & Ricco, G. (2020). Deep Dynamic Factor Models. Working paper.
-
Causal representation — Schölkopf, B., et al. (2021). Toward Causal Representation Learning. Proceedings of the IEEE, 109(5), 612–634.
Other posts
- Driving through underground rocks
Steering a drill bit through rock you cannot see: why a good particle filter still drifts, and how simulation fixed the drift.
- Can we really get alpha from market data?
The efficient market view, the micro alpha counter-argument, and why a weak signal only becomes a position once you know its uncertainty.
- What works for forecasting macro economic series with deep learning?
Korean output and investment nowcasting with seven deep models: what the data allows, which families worked, and why it depends on the target.
- Can we make a more risk-aware portfolio agent from utility theory?
Epstein–Zin recursive utility inside actor–critic RL: the Bellman backup that changes, and what it did on Korean ETF splits.
- Classifying bird sounds in the field
Sound as a picture: what a spectrogram is, why the mel scale matches how we hear, and how a network finds a bird call in the image.
- Creating and Evaluating Synthetic Tabular Data
Sequential synthesis for credit bureau data, and three checks: pMSE distinguishability, confidence interval overlap, and attribute disclosure.