About SparkyData
One place to look at what the data actually says about New York City, past the headlines, the averages, and the people with something to sell.
The trouble with averages
Most of what gets written about a city is a handful of averages and a story to hang them on. The story is usually somebody's, and the averages hide more than they show. A median rent says nothing about who pays it or what they have left. A citywide commute time says nothing about which neighbourhoods carry the long ones. The interesting part is almost always in the intricacies: which households, which years, which boroughs, and how much the answer moves when you change the cut.
This site exists to sift through the noise of news, commentary and interested parties, and to understand not just the generalisations but what sits underneath them. The aim is clarity, not a conclusion.
What an article does instead
Rather than offer one narrative, each article aims for a general view built on the best available public data, then hands you the controls. The charts are interactive, so you can filter to the facets that interest you rather than the angle the article happened to pick. Where the data runs out, the article says so, in the text and on the chart.
Every article also ships a detailed write-up of its methods, the code that computes each figure, and the data it was computed from, so that everything on the page can be reproduced and checked. If a number here disagrees with its source, the source wins; please say so. Every article carries a version number, and anything that changes after publication is logged on the corrections page.
Why New York first
New York is a highly dynamic and unusually well-documented metropolis. The Census Bureau, the city and the state publish more about it than about almost anywhere else, which makes it the right place to build the methods and the tooling. The site starts there. It may later wander into other questions and other places, following wherever the data and curiosity lead.
How the work is done
A lot of the research here is done with the assistance of a combination of AI agents. They pull and harmonise the datasets, draft and run the analysis code, check figures against their sources and help draft the text. A project of this scope would not be possible for one person otherwise. What they do not do is decide what is true.
There is a human in the loop at every stage. Every article is read, checked and proofread by an actual expert in the field before it is published: the numbers are reproduced against the official tables, the joins and assumptions are reviewed, and the writing is edited by hand. The point of using these tools is to manage a project of this size, not to add to the slop. The aim is to provide tools for understanding the data, and the responsibility for everything published here is a person's, not a model's.
Who is behind it
SparkyData is written and maintained by one person: a data scientist with an Ivy League education at both undergraduate and doctoral level and more than ten years in the financial industry. The site is anonymous for now, to keep the focus on getting the models and the procedures right rather than on a byline.
The questions here are ones I have a genuine interest in, and I would be looking into them regardless. Publishing the investigations is a way of sharing them with whoever else finds them interesting. Suggestions, corrections and questions are welcome at sparkytech.dev@gmail.com.