Creating and Evaluating Synthetic Tabular Data
Synthetic tabular data is a table generated to look like a real one without containing any real rows, so that it can be shared, tested on, or analysed where the original cannot. Two things have to hold at once: the synthetic table must support the same analyses as the real one, and it must not let anyone learn about a real person. This post explains one generation method, sequential synthesis, and three checks that tell you whether the output is both useful and safe. I used all of them in a capstone project with NICE, a Korean credit bureau, on consumer credit tables.
Sequential synthesis
Sequential synthesis builds the table one column at a time. Pick an order, fit a model for the first column and sample it, then fit a model for the second column conditional on the first and sample that, and so on. Each model is small and matched to the column type, a regression for a numeric column and a multinomial for a categorical one, exactly as the R package synthpop does. The order encodes the dependencies you want preserved, such as age before income.
The appeal is practical. Every step is a model you can inspect, it runs on a CPU, and it scales to large tables, which is what credit bureau data looks like. The cost is that chained simple models miss complex joint structure that a deep generator such as CTGAN or a diffusion model like TabSyn might capture. For this project the trade was worth it, and I contributed the implementation as the syn_seq plugin in Synthcity so it could be reused and compared against those generators on the same footing.
Check 1: Can a classifier tell real from synthetic? (pMSE)
Stack the real and synthetic rows, label each with where it came from, and train a classifier to predict the label. If the two tables have the same distribution, the classifier cannot beat the base rate: every row gets a propensity score close to , the share of synthetic rows. The propensity score mean squared error measures how far it gets:
A value near zero means the synthetic rows are indistinguishable from the real ones in whatever feature space the classifier can see. This is a utility measure, not a privacy one. A large pMSE says the generator has drifted from the real distribution, and anything estimated from the synthetic data will drift with it.

Check 2: Do the statistics you care about agree? (Confidence interval overlap)
Analysts do not use the whole table. They use a handful of estimates from it, such as a rate by segment or a regression coefficient. For each such estimate, compute a confidence interval from the real data and one from the synthetic data, and measure how much the two intervals overlap as a fraction of their combined span. Averaged over the estimates you care about, this says whether the synthetic data would lead an analyst to the same conclusions. It is a specific utility measure, complementary to pMSE, which is a global one.

Check 3: Can the synthetic data be used against real people? (Attribute disclosure)
Removing names does not make a table anonymous. A few ordinary columns together, such as age, postcode and gender, act as a quasi-identifier that picks out individuals. With synthetic data the risk is indirect: an attacker who knows a real person's quasi-identifiers could train a model on the synthetic table to predict a sensitive column, then apply it to that person. The check does exactly that. Fit a model on the synthetic data with the quasi-identifiers as inputs and the sensitive column as the target, predict the sensitive column for the real rows, and score it with or accuracy. A high score means the synthetic table teaches an attacker something true about real individuals, and the generation has to be made coarser for that column.

Tools
The R package synthpop implements sequential synthesis together with the pMSE and confidence interval utility measures. In Python, Synthcity has the syn_seq plugin alongside the deep generators, so the same evaluation can be run against several methods on one table.
References
-
Synthpop — Nowok, B., Raab, G. M., & Dibben, C. (2016). synthpop: Bespoke creation of synthetic data in R. Journal of Statistical Software, 74(11), 1–26. doi:10.18637/jss.v074.i11.
-
pMSE and CI overlap — Snoke, J., Raab, G. M., Nowok, B., Dibben, C., & Slavkovic, A. (2018). General and specific utility measures for synthetic data. Journal of the Royal Statistical Society: Series A, 181(3), 663–688.
-
CTGAN / TVAE — Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. NeurIPS, 32. arXiv:1907.00503.
-
TabSyn — Zhang, H., Zhang, J., Srinivasan, B., et al. (2024). Mixed-type tabular data synthesis with score-based diffusion in latent space. ICLR. arXiv:2310.09656.
-
Synthcity — Qian, Z., Cebere, B.-C., & van der Schaar, M. (2023). Synthcity: facilitating innovative use cases of synthetic data. arXiv:2301.07573.
Other posts
- Driving through underground rocks
Steering a drill bit through rock you cannot see: why a good particle filter still drifts, and how simulation fixed the drift.
- Can we really get alpha from market data?
The efficient market view, the micro alpha counter-argument, and why a weak signal only becomes a position once you know its uncertainty.
- What works for forecasting macro economic series with deep learning?
Korean output and investment nowcasting with seven deep models: what the data allows, which families worked, and why it depends on the target.
- Could multivariate time series have their own representations?
Why forecast embeddings are not factors, and how identifiable innovations with diagonal dynamics recover them without losing forecast quality.
- Can we make a more risk-aware portfolio agent from utility theory?
Epstein–Zin recursive utility inside actor–critic RL: the Bellman backup that changes, and what it did on Korean ETF splits.
- Classifying bird sounds in the field
Sound as a picture: what a spectrogram is, why the mel scale matches how we hear, and how a network finds a bird call in the image.