Concepts

Synthetic data

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Synthetic data is artificial data built from seed data so some patterns of that seed can be reused.

1 · What it is

Synthetic data is artificial data built from real starting examples, called seed data. Some statistical patterns of that seed remain in the new file. The process does not promise every pattern the seed held.

An establishment is one physical place where a business operates. The Longitudinal Business Database is a confidential research product developed at the Census Bureau. It includes only employer establishments. The Synthetic Longitudinal Business Database keeps that file’s structure and its broad statistical properties. The original content is not reused. The public file has 21 million establishment records for 1976 through 2000, covering all sectors. A record can include a synthetic three-digit industry code, the first year, the last year, annual payroll, and annual employment. Multi-unit status can be included too. That status marks a workplace in a firm with two or more workplaces. Geography and firm-level information are left out.

phi-1 is a code model with 1.3 billion parameters, trained with only 7 billion tokens. GPT-3.5 wrote a set of Python textbooks, under 1 billion tokens, as teaching text mixed with code. A smaller set holds under 180 million tokens of exercises and solutions. GPT-3.5 wrote that exercise set as well. Those exercises aim the model at finishing a function from written instructions. Filtered code from the web is still part of the training mix. The reported scores include pass@1 on HumanEval and 55.5% pass@1 on MBPP.

On May 18, 2016, with the experiment still underway, a preliminary analysis found that synthetic data could replace the original data for data-science work.

2 · Why it exists

A public stand-in is how a confidential business file can be released.

Confidential sourceThe Longitudinal Business Database is a confidential research product. The Business Register is restricted data, and it is the source for that database.
Broad patternsProtection changes each establishment and still tries to keep the large-scale distributions intact.
Weak code lessonsCommon code collections are a weak way to teach a model to reason and plan.
3 · How it works

Information on establishments is synthesized.

  1. 1 · fitModels are fit on sensitive information in the collected data.
  2. 2 · drawReplacement values are simulated from those models.
  3. 3 · checkThe new values are tested for disclosure risk.
  4. 4 · releaseThe public file keeps the structure and the statistical properties, and the original content stays out.

Simulated values are what replace the secret ones. The disclosure test sits after that draw.

4 · Where it's used
WhoWhat they askWhat it works with
Census economist“Will this hiring pattern still show up in the confidential file?”A public establishment file built as a stand-in for the secret one
Hospital privacy lead“Can research staff practice on records that are not our patients?”A table sampled from a model fit to the real columns
Code-model trainer“Where do the textbook exercises in this training run come from?”Lessons written by an earlier model, mixed with filtered web code
Methods class“What is left out when the public business file is built?”Geography and firm identity that the public file does not carry
5 · What it solves, and what it doesn't
solves
  • Artificial data can keep some statistical characteristics of the seed it was built from.
  • A synthesizer trained on a real table can then create synthetic rows on demand.
  • Sampling from the model creates synthetic data.
  • Generated textbooks and exercises can be part of the mix that trains a code model.
doesn't solve
  • Results from the public business file are not guaranteed to match the confidential data unless they are validated.
  • A differentially private synthetic dataset is a narrower product, produced by a mechanism that satisfies differential privacy.
  • The Census Bureau is still seeking a differentially private version of this business file, so the current public file is not that product.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialsynthetic data generation - Glossary | CSRC, NIST · read 28 Sept 2026
  2. officialSynthetic Longitudinal Business Database (SynLBD), U.S. Census Bureau · read 28 Sept 2026
  3. officialDefinitions, U.S. Census Bureau · read 28 Sept 2026
  4. officialMethodology, U.S. Census Bureau · read 28 Sept 2026
  5. officialdifferentially private synthetic dataset - Glossary | CSRC, NIST · read 28 Sept 2026
  6. paperTextbooks Are All You Need, Gunasekar et al., Microsoft Research · read 28 Sept 2026
  7. docsWelcome to the SDV!, DataCebo · read 28 Sept 2026
  8. paperThe Synthetic Data Vault : generative modeling for relational databases, Patki, Massachusetts Institute of Technology · read 28 Sept 2026