An Essay on Data and Judgement · 2026

Abundant and Uncurated

Why synthetic data scales human judgement — and why it cannot become its source
Anuj Sadani
Data has never been more plentiful or less trustworthy. The interesting question is not whether machines can generate more of it. They can. The question is what happens to a system that learns only from itself — and what, precisely, a human is still for.
MODEL LOOP converges on what it already knows EXOGENOUS SIGNAL HUMAN JUDGEMENT the only unbounded source
A verified essay. Every factual claim in this text is bound to a quoted passage in a source captured to disk and checked by hash. 45 claims in the ledger · 43 cited here · 2 rejected in verification and reported as such · 22 distinct works · 57 quoted bindings, zero mismatches.

Short on time? The argument in brief — 3–5 min read

  Get the full PDF on Ko-fi

Copyright, permissions, and how this was made

Overture

The Argument in Five Minutes

Everyone agrees that data is abundant and that good data is scarce. The usual explanation is that curation is slow, expensive and unglamorous, and that better tooling will eventually fix it. That explanation is comfortable and it is wrong in an important way.

Start with the received wisdom, because it collapses quickly. The popular case for keeping humans in the loop is that human judgement is simply better. It is not. In a peer-reviewed comparison, a language model beat crowd workers on annotation accuracy by roughly twenty-five percentage points, at under a third of a cent per annotation — about thirty times cheaper than the human alternative.22 Strong model judges agree with human preferences more than eighty per cent of the time, which is the rate at which humans agree with each other.7 Any argument for the human that rests on comparative accuracy has already lost.

The argument that survives

Two mechanisms, neither about accuracy, hold the line.

The first is informational. A synthetic loop that verifies itself with a model converges on what that verifier already knows — the retraining process drives the estimate toward the verifier’s own knowledge centre, and gains plateau or reverse unless the verifier is perfectly reliable.16 A closed loop of models redistributes information. It does not create any. The human is not in the loop as a quality check. The human is in the loop as the only unbounded source of new information.

The second is structural. Evaluation criteria cannot be fully specified before a human has looked at outputs.8 The dependency is circular: you need criteria in order to grade, and grading is how you discover the criteria. Every result showing machines matching humans measures performance against a fixed rubric — and the rubric is precisely the artefact that cannot be fixed in advance. A better model makes the applying cheaper. It does not remove the step that comes first.

Why curation stays expensive

Curation resists automation for reasons that better tooling does not touch. Its failures are invisible at the time of the work: data cascades are opaque and delayed, with poor indicators and metrics,13 which means there is no contemporaneous quality signal to automate against. And curation is normative rather than clerical — annotation instructions reproduce and normalise the worldviews of whoever commissioned them.18 It encodes a position on what the right answer is. That is not a transcription task, and it never becomes one.

The verdict
Machines should absorb the application of judgement, where they are demonstrably accurate and radically cheaper. Humans own two things that cannot be delegated: the formation of the criteria, and the supply of signal from outside the system. The thesis — and the counter-evidence is real

The honest counter-case is not hypothetical. One system reached state-of-the-art reasoning performance trained by self-play with no human-curated data at all.17 That result is addressed head-on in Chapter III rather than managed. Its margin is 0.3 percentage points, and it works because a code executor hands it free ground truth — a luxury that exists for arithmetic and does not exist for tone.

Contents

What’s inside

  1. Overture: The Argument in Five Minutesthe whole case, compressed
  2. IThe Collapse That Wasn’ta regime, not a fact about synthetic data
  3. IIThe Verifier’s Knowledge Centrewhy a model cannot bootstrap itself
  4. IIIWhere Synthetic Earns Its Placeand the strongest case against this essay
  5. IVThe Price of Judgementwhy curation resists automation
  6. VThe Rubric Problemwhat machines can and cannot take over
  7. VIA Scarcity of Permissionthe data is not running out the way you think
  8. VIIThe Division of Labourthe verdict, and what would break it
  9. AAppendix: What Did Not Survivethe method, and two rejected claims
  10. Notes & Sourcestwenty-two works, tiered
Audio editionListen to this book10 chapters · 33 min

I

Chapter One

The Collapse That Wasn’t

In July 2024, Nature published a result that gave the field a nightmare with a name. Indiscriminately training generative models on data produced by other models causes model collapse: the models progressively forget the true underlying data distribution, even when that distribution is not itself changing.1 The characteristic symptom is elegant and grim — the tails of the original content distribution disappear.1 The rare things go first.

The finding travelled fast, and it travelled stripped of its conditions. It became a general claim about synthetic data, which is not what the paper demonstrates.

Replace versus accumulate

Later work qualified the result heavily. When each generation’s real data is replaced by synthetic data, test error does rise with every model-fitting iteration.3 But when synthetic data accumulates alongside the original real data, collapse is avoided.3 Same phenomenon, opposite outcome, and the only difference is what happens to the old data.

REPLACE ACCUMULATE real synth 1 synth 2 synth 3 error ↑ proportion of real data hits zero after the first iteration real real + synth 1 real + synth 1–2 real + synth 1–3 bounded real data is diluted but never removed — collapse is avoided
Fig. 1 — The two regimes. Only one of them describes a real training pipeline.

A position paper pushes the criticism further on two fronts. First, the literature is arguing about eight different definitions of degradation rather than one.4 Second, and more damagingly, the replace paradigm drives the proportion of real data to zero immediately after the first iteration4 — a condition that describes no training pipeline anyone actually operates.

A tier inversion worth noticing

These are not contradictory findings. They are findings about different setups, and the later work simply ran an arm the earlier work did not.3,4 But something about how they are shelved deserves comment, because it will mislead a careful reader who is being careful in the usual way.

The peer-reviewed paper carries the less qualified claim. The preprints carry the more careful one.1,3,4 A reader who weights sources by venue — the ordinary, sensible heuristic — gets this backwards. In a field moving this fast, recency can outrank review status, and citing only the peer-reviewed framing produces a more alarming account than the literature supports.

One housekeeping note, because a heavily cited paper carrying a correction invites suspicion. The Nature paper does carry a 2025 Author Correction, amending text in its Theoretical intuition section.2 It is a symbol substitution. It is not a retraction and it does not disturb the findings.

Model collapse is real.
It is also a description of a laboratory condition that nobody trains in.

II

Chapter Two

The Verifier’s Knowledge Centre

Underneath the argument about regimes sits a mechanism that generalises better than either result, and it is the load-bearing wall of this essay. It also very nearly went unnoticed, for reasons worth admitting.

Continual improvement requires that data be curated in a way that injects signal exogenous to the system that produced the original data.11 Without such curation, performance can plateau or collapse after many iterations.11

Read that carefully. It specifies exogenous signal. It does not specify human signal. And at first pass, the distinction appears to dissolve the case for the human entirely: a verifier “whether a human or a better model” is sufficient to prevent collapse.16 On that reading, a stronger model substitutes cleanly for a person, and the whole hand-wringing about human data is a transitional worry.

Two sentences later

It does not substitute, and the qualification is decisive. Where a model rather than a human supplies the verification, early gains plateau and may reverse unless that verifier is perfectly reliable.16 More fundamentally, verifier-guided retraining ultimately drives the parameter estimate toward the verifier’s own knowledge centre.16

That is the whole argument in one clause. A model verifier does not add information to the system. It transfers information the verifier already holds, and then the process terminates. You can distil a stronger model into a weaker one indefinitely and never exceed the stronger one. The ceiling is not a training problem. It is an information-theoretic one.

verifier’s knowledge model verifier — converges inward model human signal — arrives from outside
Fig. 2 — A model verifier bounds the system at its own knowledge. Nothing inside the loop can exceed it.

Where this claim is weak

This is the most important argument in the essay and the least corroborated. The knowledge-centre result rests on a single preprint whose long-run analysis is theoretical — demonstrated on linear regression, a variational autoencoder on MNIST, and a 135-million-parameter language model.16 It has not been shown at frontier scale.

If that source fails to replicate, this chapter weakens from a formal result to a plausibility argument. Readers should know which brick the wall is resting on.

A system that learns only from itself
is not learning. It is settling.

III

Chapter Three

Where Synthetic Earns Its Place

None of this makes synthetic data a mistake. It makes it a tool with conditions, and the conditions are more interesting than the tool. Synthetic data improves models conditionally rather than generally, and the binding condition is the strength of the verification applied to it — not the volume generated.10,14

Filtering is not automatically virtuous

The practitioner consensus holds that a judge filter is the critical step, and that an unfiltered synthetic dataset is worse than a smaller filtered one. Half of that is supported. Filtering enhances learning only when the underlying verification is strong enough to supply a high-resolution correctness signal.10

The corollary is routinely omitted, and it is the useful half: naive or rigid verification filtering has been observed to produce no improvement at all.10 Weak filtering is theatre. And over-strict filtering is actively harmful, because it strips out the diversity that measurably correlates with both pre-training and supervised fine-tuning performance.21

There is a related finding that looks like good news and is really a relocation of the problem. The quality of the seed data fed to a generator matters more to downstream performance than whether the generated knowledge is novel.14 Which is to say: synthetic data inherits its quality from curated data. The curation did not disappear. It moved upstream, where it is easier not to look at.

One place synthetic genuinely holds up

Synthetic test data has been found as effective as hand-crafted data both for refining evaluation criteria and for aligning with user expectations.9 This is a real result and it deserves stating precisely, because it is routinely over-read. It concerns the refinement of judge criteria, not the quality of training data, and it rests on a preprint with twenty-four participants.

The companion figure gets quoted more often than the finding: eighty-three per cent of those participants preferred the synthetic-generation tool to creating or selecting test cases by hand.9 That measures perceived workload and tool preference. It is evidence that people would rather not hand-write test cases — a genuine finding about the economics of curation effort, and not evidence that synthetic data is good.

The strongest case against this essay

Now the part that should worry me. One system, trained by self-play with no human-curated data whatsoever, outperformed models trained on expert-curated human data in the coding category.17 It is a spotlight result at a top venue. If a model can bootstrap state-of-the-art reasoning from nothing, the thesis of this essay is in serious trouble.

Three considerations bound it. None of them dismisses it.

The margin

0.3 absolute percentage points, on one category.17 A near-tie, not a rout — and the abstract-level framing obscures this completely.

The scope

It works because a code executor supplies validation and verifiable feedback.17 Free ground truth exists for code and mathematics. It does not exist for tone, safety or appropriateness.

The missing experiment

A domain with no formal correctness oracle, where self-generated data improves a model across many generations with no human signal at any point. That experiment has not been run.

Confidence: low

This chapter is held at low confidence, set by its weakest link rather than its average. The seed-quality finding is vendor-authored by a company selling data curation, whose commercial interest points the same direction as its result.14

And the strongest filtering evidence comes from the code domain, where a compiler supplies ground truth for free10 — the setting most favourable to automated verification and least representative of the judgements that are actually hard.

IV

Chapter Four

The Price of Judgement

Human curation is costly for structural reasons rather than temporary ones. This matters because the two diagnoses lead to opposite strategies. If curation is expensive because tooling is immature, you wait. If it is expensive because of what it fundamentally is, you budget for it permanently.

The shape of the cost

1M+ / week
Binary model-generation comparisons collected on a weekly basis in one frontier alignment pipeline.19
92%
Of AI practitioners surveyed (n=53) reported experiencing at least one data cascade.13
cents
Per task, for data workers producing training data — typically without employment protections.18

Volume is the least interesting half

The scale is genuinely large. One frontier alignment pipeline collected human preference data on a weekly basis, comprising over one million binary model-generation comparisons.19 That is not a dataset assembled once and amortised. It is a standing operation with a payroll.

But volume is the part that tooling can attack, and focusing on it is what makes people predict that curation costs are about to fall. Two other findings explain why they do not.

Failure is invisible until late

Data cascades — compounding downstream failures originating in data problems — are opaque and delayed, with poor indicators and metrics for detecting them.13 Ninety-two per cent of the practitioners surveyed reported experiencing at least one.13

This is the load-bearing explanation for why curation resists automation, and it is a chicken-and-egg problem rather than an engineering gap. Automated quality control needs a quality signal to optimise against. A process whose defects surface only far downstream does not emit one at the time the work is being done. You cannot cheaply automate against a signal that does not yet exist.

Curation is normative, not clerical

The second reason is less discussed and more fundamental. Annotation instruction documents reproduce and normalise the worldviews of the parties commissioning the data.18 Curation is not transcription that a cheaper worker could perform equally well. It encodes a position on what the right answer is.

This is why quality standards are expensive to write and contested to apply. It also explains something the industry prefers not to name: the cost has largely been displaced rather than removed. Data workers producing training data are paid as little as a few cents per task and typically lack the social protections attached to employment.18 The field’s own account of itself is that data is the most under-valued and de-glamorised aspect of AI practice.13

Curation looks like a clerical cost because we have arranged for someone else, somewhere else, to absorb it. On the economics of not looking

A gap in this chapter

No peer-reviewed cost-per-label figure for expert curation was found. The available comparison uses crowd workers on classification tasks as the human baseline22 — not the expert judgement this chapter is about. This is the thinnest evidence in the essay, and it is thin in the direction that would most strengthen the argument if filled.

The wage figures come from fieldwork in Venezuela and Argentina18 and should not be generalised to expert annotation in high-income markets.

V

Chapter Five

The Rubric Problem

Here is where the sources genuinely disagree, so the disagreement goes first and the reconciliation second. Anything else would be advocacy.

The case against the human

It is strong. Model judges match human-human agreement rates, exceeding eighty per cent.7 Models beat crowd workers on annotation accuracy by roughly twenty-five percentage points.22 And they do it at under a third of a cent per annotation, about thirty times cheaper than the human alternative.22

What automation demonstrably wins

>80%
Judge–human agreement, matching the rate at which humans agree with each other.7
+25pp
Annotation accuracy over crowd workers, across four datasets of tweets and news articles.22
30×
Cheaper per annotation than the crowd-work baseline.22

The case for the human

It rests on one finding, and the finding is structural rather than comparative. It is impossible to completely determine evaluation criteria prior to human judging of model outputs.8 The circularity has a name — criteria drift: users need criteria in order to grade outputs, and grading outputs is what allows them to define the criteria.8

The obvious objection is to grade first and specify afterwards. It was tested, and it does not hold: participants who graded before specifying criteria still refined those criteria on further grading.8 The loop does not terminate by reordering it.

The reconciliation

The two bodies of evidence are not in conflict, because they measure different things. Every result favouring automation measures performance against a fixed rubric. And the rubric is precisely the artefact that cannot be fixed in advance.8

Cannot be delegated
  • Criteria formation — discovering what should be measured at all, which requires looking at outputs first.8
  • Exogenous signal — supplying information the system does not already contain.11,16
Should be delegated
  • Criteria application — grading at volume against a settled rubric, where machines match human agreement rates.7
  • First-pass labelling — where accuracy is higher and cost is thirtyfold lower.22

This distinction is worth holding onto because it is a claim about logic, not about capability. It does not say models are not yet good enough at judging. It says the specification of what counts as good cannot be completed before the examining happens. A more capable model makes the applying cheaper; it does not remove the step that precedes it. The conclusion would fall if someone demonstrated stable criteria formation without human grading — and that is the experiment to watch for.

A note on citation hygiene

The paper reporting agreement above eighty per cent7 is the same paper documenting that model judges exhibit position bias, verbosity bias, self-enhancement bias and limited reasoning ability.7 The headline figure and the failure taxonomy come from one source. They should always travel together, and in practice the first travels alone.

VI

Chapter Six

A Scarcity of Permission

The premise underneath all of this is that high-quality human data is running out. It is worth checking, because the claim blends two different phenomena that call for opposite responses.

The forecast. If current trends continue, models will be trained on datasets roughly equal in size to the stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained.5 This is a projection rather than an observation, and its authors are an organisation whose visibility depends on scaling forecasts being newsworthy.5,6 It has not happened yet.

The measurement. More than twenty-eight per cent of the most actively maintained, critical sources in the C4 corpus have become fully restricted from use through robots.txt.15 This is not a forecast. It is an audit of something that already occurred.

Two different shortages

2026–2032
Projected window in which training-set size meets the stock of public human text.5 A forecast.
28%+
Of the most actively maintained critical C4 sources, now fully restricted via robots.txt.15 A measurement.

These should not be blended, because only one of them is a data problem. Exhaustion is a supply constraint that better curation could in principle address — more efficient use of what exists. Consent withdrawal is a governance phenomenon that curation effort cannot touch at all, because the data still exists in full and is simply no longer permitted.

Which reframes the essay’s own premise. A significant part of the scarcity in “abundant but uncurated” is not scarcity of text. It is scarcity of permission15 — and the response to that is legal and commercial, not technical.

The web did not run out of words.
It started saying no.

VII

Chapter Seven

The Division of Labour

The proposition that data is abundant while high-quality curated data is scarce is correct. The reasons are more structural than the usual account of effort and process suggests, and the prescription that follows is narrower than “keep humans in the loop.”

Curation is expensive because its failures are undetectable at the time of the work13 and because it encodes contested judgement rather than recording facts.18 Neither of those is fixed by better tooling, which is why the cost has stayed stubborn while everything around it got cheaper.

The verdict
Machines should absorb the application of judgement at scale, where they are demonstrably accurate and radically cheaper. Humans should own two things that cannot be delegated: the formation of the criteria, because criteria cannot be specified before outputs are examined; and the supply of exogenous signal, because a loop verified by models converges on what those models already know. Held at moderate confidence — and contested

Two independent mechanisms converge on this, which is the main reason to believe it. A synthetic loop guided by a model verifier converges on that verifier’s knowledge centre,16 so a closed system of models redistributes information without creating any. And evaluation criteria cannot be fully specified before outputs are examined,8 so the definition of quality cannot be handed over in advance. The first is informational; the second is logical. They do not depend on each other, and they point the same way.

What this changes in practice

Budget

Treat curation as a permanent operating cost, not a capital project that finishes. The weekly cadence of frontier preference collection19 is the shape to plan for.

Pipeline

Put humans where criteria are formed, not where labels are produced. Machines are better and cheaper at application;7,22 they cannot do formation at all.

Synthetic data

Generate freely, but assume the gain is bounded by verification strength.10 Unfiltered volume is not an asset, and weak filtering is not verification.

Access

Treat the consent withdrawal as a first-order constraint.15 No amount of curation skill recovers data you are not permitted to use.

What would break this

The conclusion is held at moderate confidence and it is formally contested. The strongest counter-evidence — state-of-the-art reasoning learned with no human data — is real.17 Its scope condition, a code executor providing ground truth,17 is exactly the condition that does not hold in the domains where curation is hardest. But that is a bound, not a refutation, and the bound could be loosened by better automated verifiers in domains that currently have none.

The falsifying experiment is specific: a domain with no formal correctness oracle, in which self-generated data improves a model across many generations with no human signal at any point. If that lands, this essay is wrong, and the honest thing is to say so in advance rather than after.

The machine can tell you whether the answer fits the rule.
It cannot tell you the rule was worth having.

A

Appendix

What Did Not Survive

This essay was assembled under a verification discipline, and the discipline is worth describing because it changed the argument twice and rejected two claims outright.

The method

Every source was captured to disk before being cited, and hashed. Every factual claim was bound to an exact quoted passage inside one of those captured files, and the binding was checked mechanically — not against the live URL, and not against memory of the page. Claims that could not be bound were marked failed and kept in the ledger rather than quietly dropped.

Forty-five claims were written; forty-three passed and are cited here. Fifty-seven bindings were checked against twenty-seven captured snapshots, with zero mismatches. Twenty-two distinct works underpin the essay: nine peer-reviewed, thirteen preprints or conference papers, no press coverage, and — importantly — no systematic review with a registered protocol. Nothing here carries meta-analytic support.

Two claims were rejected

Rejected

“Golden datasets built from production traces reflect real usage in a way synthetic test cases cannot.” Widely repeated in practitioner writing. The nearest captured source runs the other way, noting that real-world data is noisy and unbalanced and may overlook rare but critical edge cases.9

Rejected

“A layered LLM-annotator, judge, synthetic gap-fill and human-review pipeline is now standard practice.” “Standard” is an empirical claim about industry adoption. It would need a practitioner survey, and none exists in this corpus.

Two findings emerged only from reading past the abstract

Both changed the argument. The paper stating that a verifier “whether a human or a better model” prevents collapse reads, at abstract level, as though the two are interchangeable. Two sentences later it states that gains plateau or reverse unless the verifier is perfectly reliable, and that retraining converges on the verifier’s knowledge centre.16 That inversion became Chapter II.

And the self-play result that beats expert-curated human data does so by 0.3 absolute percentage points.17 The headline survives; the decisiveness does not. A summary that stopped at the abstract would have reported a rout.

Known limits

Single-source

The knowledge-centre bound — the formal basis of Chapter II — rests on one preprint at small scale.16

Domain skew

The strongest verification evidence is from code, where ground truth is free.10 Generalisation to tone and safety is directional, not established.

Out of scope

Copyright and licensing; privacy regulation; non-text modalities; data-market economics. The common claim that synthetic data wins in health and finance was not assessed and should not be inferred from this text.

Unretrieved

Aroyo and Welty (2015), the standard treatment of annotator disagreement as signal rather than noise, was paywalled and could not be obtained. It is not cited, and Chapter IV is weaker for its absence.

Unknowable from public sources: what frontier labs currently spend on human data, and the present ratio of synthetic to human data in frontier training mixes. The closest public figure19 is self-reported and three years old.

Appendix

Notes & Sources

Twenty-two works. Tier labels follow an academic source standard: T2 peer-reviewed primary studies and proceedings; T3 preprints, conference papers and lab technical reports, which is where the frontier of this particular literature lives. All sources were captured to disk and accessed 2026-08-16. Citation numbers in the text map to this list.

  1. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. — AI models collapse when trained on recursively generated data. Nature, 24 Jul 2024. nature.com [T2]
  2. Shumailov, I., et al. — Author Correction: AI models collapse when trained on recursively generated data. Nature 640, E6, 21 Mar 2025. nature.com [T2 — correction; symbol substitution, not a retraction]
  3. Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., et al. — Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413, Apr 2024. arxiv.org [T3]
  4. Schaeffer, R., et al. — Position: Model Collapse Does Not Mean What You Think. arXiv:2503.03150, Mar 2025. arxiv.org [T3]
  5. Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., & Hobbhahn, M. — Position: Will we run out of data? Limits of LLM scaling based on human-generated data. ICML 2024 / arXiv:2211.04325. proceedings.mlr.press [T2 — Epoch AI authors]
  6. Epoch AI — Will we run out of data to train large language models? Jun 2024. epoch.ai [T3 — publisher’s own summary of 5]
  7. Zheng, L., Chiang, W.-L., Sheng, Y., et al. — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets & Benchmarks. arxiv.org [T2]
  8. Shankar, S., Zamfirescu-Pereira, J.D., Hartmann, B., Parameswaran, A.G., & Arawjo, I. — Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. ACM UIST 2024 / arXiv:2404.12272. arxiv.org [T2]
  9. Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges. arXiv:2511.04478, Nov 2025. arxiv.org [T3 — preprint, N=24 user study]
  10. Verification Limits Code LLM Training. arXiv:2509.20837, Sep 2025. arxiv.org [T3 — code domain only]
  11. Escaping Collapse: The Strength of Weak Data for Large Language Model Training. arXiv:2502.08924, Feb 2025. arxiv.org [T3]
  12. Lee, H., Phatale, S., Mansoor, H., et al. — RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. arXiv:2309.00267. arxiv.org [T3 — Google Research; disconfirming source]
  13. Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. — “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. ACM CHI 2021. dl.acm.org [T2 — n=53 practitioners; Google Research authors]
  14. BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining. arXiv:2508.10975, Aug 2025. arxiv.org [T3 — DatologyAI; commercial interest in the finding]
  15. Longpre, S., et al. (Data Provenance Initiative, MIT) — Consent in Crisis: The Rapid Decline of the AI Data Commons. NeurIPS 2024 D&B / arXiv:2407.14933. arxiv.org [T2 — audit of 14,000 domains]
  16. Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence. arXiv:2510.16657, Oct 2025. arxiv.org [T3 — single-source basis for Ch. II]
  17. Zhao, A., et al. (LeapLab, Tsinghua) — Absolute Zero: Reinforced Self-play Reasoning with Zero Data. NeurIPS 2025 spotlight / arXiv:2505.03335. arxiv.org [T3 — strongest disconfirming result]
  18. Miceli, M., & Posada, J. — The Data-Production Dispositif. Proc. ACM Human-Computer Interaction, CSCW 2022 / arXiv:2205.11963. arxiv.org [T2 — 210 instruction documents, 55 interviews]
  19. Touvron, H., et al. (Meta AI) — Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288, Jul 2023. arxiv.org [T3 — self-reported lab technical report]
  20. Long, L., Wang, R., Liu, R., et al. — On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey. arXiv:2406.15126, Jun 2024. arxiv.org [T3 — narrative survey, not a systematic review]
  21. On the Diversity of Synthetic Data and its Impact on Training Large Language Models. arXiv:2410.15226, Oct 2024. arxiv.org [T3]
  22. Gilardi, F., Alizadeh, M., & Kubli, M. — ChatGPT outperforms crowd workers for text-annotation tasks. PNAS 120(30), 2023 / arXiv:2303.15056. arxiv.org [T2 — disconfirming source; crowd-worker baseline]

On the essay’s own limits. The verification workspace — the claims ledger with every quoted binding, the source ledger with URLs and content hashes, the synthesis, and the gate’s verdict — is published alongside this essay. The captured snapshots themselves are held back, being third-party copyrighted full text; the hashes let anyone re-capture from the original URLs and confirm the match. Three of the seven conclusions are formally mixed-stance, meaning the underlying sources disagree; each is presented above with the disagreement rather than a chosen side. A claim-auditing pass using an independent reviewer that never sees the narrative was not run; load-bearing claims were audited by re-reading surrounding passages instead, which is a weaker check and is recorded as one.

© 2026 Anuj Sadani. This essay is licensed under CC BY-NC-ND 4.0 — share it with credit, not commercially, and without redistributing modified versions. Quoted passages from the works listed above remain the property of their publishers and are reproduced here for citation and verification. Source ledgers and the typeset edition: github.com/asadani/abundant-and-uncurated.

About the author

Anuj Sadani builds AI systems and the teams that build them. He spent a decade at NVIDIA, has worked inside AI-first organizations, and has led multicultural engineering pods across Europe — sixteen-plus years, most of them on the AI era's unglamorous half: turning hype into systems that ship, and systems that ship into outcomes that hold up in production. That work has fed into industry recognition from Gartner and Everest Group, and an Innovator of the Year nod.

He is the author of The Clean Vibe Coder: A Code of Conduct for Programmers in the Age of AI Agents and Borrow the Line. Own the Move., and writes about engineering, leadership, and what stays human when the tools get good. He still believes the best technology is the kind that makes the people around it braver.

Writing and work Essays, books, and what he is building. anujsadani.in
This book, and the evidence Essays, the source ledgers, and the rest of the shelf. tech.anujsadani.in
Abundant and Uncurated — a verified inquiry by Anuj Sadani, 2026.
Repository, claim ledger, and verification workspace: github.com/asadani/abundant-and-uncurated
Built on 45 claims bound to exact quotations; two failed verification and are disclosed rather than dropped.
Text licensed CC BY-NC-ND 4.0. Quoted passages remain the property of their publishers and are reproduced for citation and verification.