Minkey Chang

Data Scientist

All posts

Creating and Evaluating Synthetic Tabular Data

Synthetic tabular data is a table generated to look like a real one without containing any real rows, so that it can be shared, tested on, or analysed where the original cannot. Two things have to hold at once: the synthetic table must support the same analyses as the real one, and it must not let anyone learn about a real person. This post explains one generation method, sequential synthesis, and three checks that tell you whether the output is both useful and safe. I used all of them in a capstone project with NICE, a Korean credit bureau, on consumer credit tables.


Sequential synthesis

Sequential synthesis builds the table one column at a time. Pick an order, fit a model for the first column and sample it, then fit a model for the second column conditional on the first and sample that, and so on. Each model is small and matched to the column type, a regression for a numeric column and a multinomial for a categorical one, exactly as the R package synthpop does. The order encodes the dependencies you want preserved, such as age before income.

The appeal is practical. Every step is a model you can inspect, it runs on a CPU, and it scales to large tables, which is what credit bureau data looks like. The cost is that chained simple models miss complex joint structure that a deep generator such as CTGAN or a diffusion model like TabSyn might capture. For this project the trade was worth it, and I contributed the implementation as the syn_seq plugin in Synthcity so it could be reused and compared against those generators on the same footing.


Check 1: Can a classifier tell real from synthetic? (pMSE)

Stack the real and synthetic rows, label each with where it came from, and train a classifier to predict the label. If the two tables have the same distribution, the classifier cannot beat the base rate: every row gets a propensity score p^i\hat p_i close to cc, the share of synthetic rows. The propensity score mean squared error measures how far it gets:

pMSE=1N∑i=1N(p^i−c)2\mathrm{pMSE} = \frac{1}{N}\sum_{i=1}^{N} \big(\hat p_i - c\big)^2

A value near zero means the synthetic rows are indistinguishable from the real ones in whatever feature space the classifier can see. This is a utility measure, not a privacy one. A large pMSE says the generator has drifted from the real distribution, and anything estimated from the synthetic data will drift with it.

Diagram of the pMSE check: real and synthetic rows are stacked and labeled, a classifier is trained on the label, and its predicted probabilities are compared with the synthetic share.
The pMSE check. Real and synthetic rows are stacked with a source label, a classifier is fit to the label, and the spread of its predicted probabilities around the synthetic share c is the score.

Check 2: Do the statistics you care about agree? (Confidence interval overlap)

Analysts do not use the whole table. They use a handful of estimates from it, such as a rate by segment or a regression coefficient. For each such estimate, compute a confidence interval from the real data and one from the synthetic data, and measure how much the two intervals overlap as a fraction of their combined span. Averaged over the estimates you care about, this says whether the synthetic data would lead an analyst to the same conclusions. It is a specific utility measure, complementary to pMSE, which is a global one.

Two confidence intervals, one from real data and one from synthetic data, with the overlapping region marked and the overlap formula.
Confidence interval overlap for one statistic: the shared stretch of the real and synthetic intervals, divided by their combined span, summed over the statistics of interest.

Check 3: Can the synthetic data be used against real people? (Attribute disclosure)

Removing names does not make a table anonymous. A few ordinary columns together, such as age, postcode and gender, act as a quasi-identifier that picks out individuals. With synthetic data the risk is indirect: an attacker who knows a real person's quasi-identifiers could train a model on the synthetic table to predict a sensitive column, then apply it to that person. The check does exactly that. Fit a model on the synthetic data with the quasi-identifiers as inputs and the sensitive column as the target, predict the sensitive column for the real rows, and score it with R2R^2 or accuracy. A high score means the synthetic table teaches an attacker something true about real individuals, and the generation has to be made coarser for that column.

Three-step diagram: fit a model on synthetic data, predict the sensitive variable for the original rows, and calculate R-squared or accuracy.
The attribute disclosure check: a model fit on synthetic covariates predicts a sensitive variable for the real rows, and its R-squared or accuracy is the risk score.

Tools

The R package synthpop implements sequential synthesis together with the pMSE and confidence interval utility measures. In Python, Synthcity has the syn_seq plugin alongside the deep generators, so the same evaluation can be run against several methods on one table.


References

  1. Synthpop — Nowok, B., Raab, G. M., & Dibben, C. (2016). synthpop: Bespoke creation of synthetic data in R. Journal of Statistical Software, 74(11), 1–26. doi:10.18637/jss.v074.i11.

  2. pMSE and CI overlap — Snoke, J., Raab, G. M., Nowok, B., Dibben, C., & Slavkovic, A. (2018). General and specific utility measures for synthetic data. Journal of the Royal Statistical Society: Series A, 181(3), 663–688.

  3. CTGAN / TVAE — Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. NeurIPS, 32. arXiv:1907.00503.

  4. TabSyn — Zhang, H., Zhang, J., Srinivasan, B., et al. (2024). Mixed-type tabular data synthesis with score-based diffusion in latent space. ICLR. arXiv:2310.09656.

  5. Synthcity — Qian, Z., Cebere, B.-C., & van der Schaar, M. (2023). Synthcity: facilitating innovative use cases of synthetic data. arXiv:2301.07573.

Other posts