SparkyData

Articles

Short analyses built on the synthetic population. Every article ships with the dataset it used and the code that reproduces its numbers, and lists the sources of any external figures it quotes.

methodologywhite paper

A Population Stitched from Its Statistics

Everything published about a population is a slice — a census microdata sample that knows income but not wealth, a national wealth survey that knows balance sheets but not the metro, a housing survey that knows rent regulation and little else — and no two slices were cut from the same block, so the questions that live in their intersections cannot be answered by joining tables. This paper sets out how SparkyData treats every published statistic as a constraint and satisfies them all at once with a synthetic population: whole households are sampled from real disclosure-protected microdata so the demographic joints are inherited rather than modelled; the sample is calibrated to published household and person margins simultaneously by constrained optimisation; attributes the microdata lacks — wealth, spending, rent regulation, health — are layered on conditional on what is already fixed and renormalised to their own published marginals; a repair-based solver drives record, aggregate and trajectory constraints to a fixed point; and the cross-section is extended into a multi-decade open-cohort history anchored to recorded annual series with provenance on every year. Validation is the same constraint system run again with a higher bar: every target held within a band that combines published margin of error with the sample's own error, two-way joints checked cell by cell, and held-out statistics used as a genuine out-of-sample test. The result is a dataset for questions the slices cannot answer, with every value labelled inherited, modelled or assumed, and no real person in it.

How articles work

  • Reproducible. The numbers in an article are computed by its code from its companion dataset — not typed in.
  • Sourced. External figures quoted for comparison link to where they came from.
  • Separately anonymised. Identifiers in each article's dataset are re-randomised, so datasets from different articles cannot be joined.