Synthetic data do not simply reproduce the world. By determining which cases and relations survive generation, they can turn the majority into normality and make the infrequent disappear.
In the 26 August Radar we pointed from Sweden to a question: what happens when we delegate part of our capacity to represent the world to models? The workshop The Politics of Models – Bias, Representation & Transparency in Generative Models closed yesterday in Rånäs. One line of research directly connected to its organisers now makes that question much more concrete.
What happens when nine countries disappear from the data as a population is synthesised?
In On Exactitude in Science, Borges imagined an empire whose cartography became so precise that it eventually produced a map the same size as the territory itself. The paradox remains useful because every representation must select, reduce and omit in order to function as a representation at all. Generative models introduce an almost inverse tension. They do not necessarily fail by trying to contain everything, but because in synthesising the world they can turn a statistical distribution into a criterion of relevance and decide, without ever stating the decision in those terms, which differences remain and which become dispensable. Between Borges’s impossible map and the synthetic dataset lies the same question about the distance between the world and its models, and about who gets to decide what may be lost in that distance.
The politics comes before the answer
The WASP and WASP-HS workshop brought technical research and social sciences together for three days at Rånäs Slott, with a maximum of twenty places. The call did not present generative models merely as tools capable of producing text, images, sound, video or data. It described them as “new forms of knowledge representation”.
That distinction matters.
If a model participates in the representation of knowledge, politics does not begin only when it generates a problematic answer. It is also present in earlier decisions: which ontologies are used, how categories are constructed, which sources enter the data, how those data are gathered, and which relations between elements become recognisable to the system.
WASP-HS organised the workshop around four themes: representation and bias, power and control, transparency and accountability, and societal alignment. On social media, one of the organisers, Md Fahim Sikder, presented the meeting as a space for conversation across WASP’s technical perspectives and the social-science perspectives of WASP-HS.
The public page describes the agenda and research questions; we do not treat it here as a record of conclusions. Instead, to understand the stakes, we follow a line of research directly connected to the workshop.
When nine countries stop appearing
Francis Lee, one of the workshop organisers, is part of the WASP-HS research environment Synthetic Data: Facts, Representations, and Transparency. The group studies synthetic data as computational versions of real information and asks what counts as truth, what such data represent, how they can be audited and what happens to bias when an artificial version of the world begins to circulate as data.
An experiment by Ericka Johnson and Saghi Hajisharif makes the problem unusually tangible.
In The intersectional hallucinations of synthetic data, they repeatedly synthesised a dataset they call the 1990 US Adult Census Data using different generative models. One of their first questions was whether cases at the edges of the distributions would survive.
Not all of them did.
The original data contained people born in 40 countries. In one synthetic dataset, 31 remained. The countries that disappeared were precisely those with very little representation in the original dataset.
This is not the familiar hallucination of a chatbot inventing a name, date or reference. Something quieter has happened: the system can produce a dataset that still looks statistically plausible while part of the original diversity has ceased to exist within it.
The later paper by Francis Lee, Saghi Hajisharif and Ericka Johnson, The ontological politics of synthetic data: Normalities, outliers, and intersectional hallucinations, states the consequence clearly. Synthetic generation tends to “highlight majority elements as the ‘normal’ and minimize minority elements”.
The majority can therefore become normality not because someone explicitly declares it so, but because the statistical operation reproduces it more readily while less frequent cases are attenuated or erased.
A society is not a collection of columns
The second finding matters even more from an anthropological perspective.
The researchers found that a synthetic version could preserve some individual distributions reasonably well whilst transforming relations between variables.
In the original data there was one record simultaneously classified with marital status husband and sex female. In one synthetic version there were 259 records with that combination. The same dataset also produced 333 records labelled simultaneously husband/wife and single, whilst other combinations that did exist in the original data disappeared.
Johnson and Hajisharif call such alterations intersectional hallucinations. They use intersectional fidelity for relations between variables that the synthetic dataset preserves sufficiently well.
The distinction is fundamental.
Age, gender, occupation, income, origin or family situation can all become columns, but a society is not a collection of independent columns. Those categories acquire meaning within historical, institutional and social relations.
An audit that checks only whether each column retains similar percentages can therefore miss the main problem. A system may preserve the parts reasonably well and still alter the world they form when related to one another.
The question is no longer only who is represented.
We also need to ask which relations survive representation.
To be synthetic, something has to change
A paradox appears here.
If a synthetic dataset reproduced every datum and every relation in the original exactly, it would cease to fulfil some of the purposes for which synthesis is used. Part of its usefulness lies precisely in introducing differences that can, depending on the case, protect privacy, enable circulation or augment insufficient datasets.
Intersectional hallucinations are therefore not simply an accident that can be removed in full. Johnson and Hajisharif argue that they are part of what makes synthetic data synthetic.
This shifts the question away from an impossible search for a perfect copy and towards a contextual decision.
Which relations must be preserved for a particular use?
Which deviations are acceptable?
Which infrequent cases contain indispensable information?
Which difference is noise, and which represents a social experience that should not disappear?
In their experiments, the researchers were able to control some edge cases and alter the generation process. But to do so they first had to identify them and decide that they deserved attention.
That is where politics appears.
Mathematics can help preserve a relation. It cannot decide by itself which relation deserves to be preserved.
From fidelity to trust
Ericka Johnson recently condensed this line of work on LinkedIn into three words: “Words = worlds.” In that post, she noted that the concepts of intersectional fidelity and intersectional hallucinations had begun to travel into medical research.
That matters because, in health, a computational representation can feed into decisions with material consequences.
In The Lancet Digital Health, Arman Koul, Deborah Duran and Tina Hernandez-Boussard use another useful concept: synthetic trust. They describe unwarranted confidence in models trained on artificial data when those data are assumed to preserve clinical validity or demographic realities.
Their warning broadens the problem.
More data do not necessarily mean more knowledge. A dataset can grow through synthesis and produce an appearance of abundance whilst retaining, amplifying or reorganising the limitations of the material from which it was generated.
The question is therefore not whether data are ‘real’ or ‘artificial’ in the abstract. It is which relations they preserve, for which uses they can be considered valid, and where they cease to represent what we claim to be studying.
What — and who — counts as real
The conversation continues almost immediately.
On 8 and 9 September, at EASST 2026 in Kraków, Charlotte Högberg, Stefano Canali and Francis Lee convene the panel Synthetic data and representation: The politics of AI generated computational practices.
Its contributions range across medicine, urban infrastructures, facial recognition, historical archives, bodies and social-science methodology. Its explicit question is how synthetic data reorganise the politics of representation and “what — and who — counts as real”.
One accepted contribution takes the argument a step further. In The politics of plausibility: Synthesis and generativity beyond excess, Chiara Carboni proposes thinking about synthesis as a politics of plausibility.
The shift is suggestive.
The power of a generative system may not lie only in asserting what is probable. It may also operate by reducing the space of the possible to what it has learned to produce as plausible and, through repeated production, reshaping our own sense of plausibility.
We are no longer speaking only about representing an existing world.
We are speaking about machines that can generate populations, fill absences, produce archives, simulate bodies, reconstruct pasts and project futures.
What disappears matters too
Debates about algorithmic bias have helped make visible who is poorly represented by a system. Synthetic data require us to add another dimension.
Sometimes the problem is not that a minority is stereotyped.
It may simply stop appearing.
And at other times every category remains present, but the relations that connected them have been reorganised.
This is why the warning built into the Rånäs workshop matters. Ontologies, categories, data sources and collection practices are not technical decisions situated before politics. They are among the places where politics begins to materialise.
The political question of a model does not begin when it produces an offensive answer or a stereotyped image.
It begins much earlier, when a particular version of the world becomes computable.
And it leaves a question worth carrying into the next stage of this conversation:
What world begins to exist when a machine learns what deserves to be represented?
Sources and references
- WASP-HS. The Politics of Models – Bias, Representation & Transparency in Generative Models, 25–27 August 2026.
- WASP. The Politics of Models: Bias, Representation and Transparency in Generative Models.
- Md Fahim Sikder. LinkedIn post on the workshop and its WASP/WASP-HS interdisciplinary framing.
- WASP-HS. Synthetic Data: Facts, Representations, and Transparency.
- Johnson, E. & Hajisharif, S. The intersectional hallucinations of synthetic data. AI & Society 40, 1575–1577. DOI 10.1007/s00146-024-02017-8.
- Lee, F., Hajisharif, S. & Johnson, E. The ontological politics of synthetic data: Normalities, outliers, and intersectional hallucinations. Big Data & Society 12(2). DOI 10.1177/20539517251318289.
- Ericka Johnson. “Words = worlds”. LinkedIn.
- Koul, A., Duran, D. & Hernandez-Boussard, T. Synthetic data, synthetic trust: navigating data challenges in the digital revolution. The Lancet Digital Health 7(11), 100924. DOI 10.1016/j.landig.2025.100924.
- EASST 2026. Synthetic data and representation: The politics of AI generated computational practices, 8–9 September 2026.
- Borges, Jorge Luis. On Exactitude in Science, in El hacedor. Critical reference at The Borges Center.
- Carboni, C. The politics of plausibility: Synthesis and generativity beyond excess. EASST 2026.