All Papers
https://lifearchitect.ai
ai
llm
paper

What’s in GPT-5

Alan D. Thompson27 July 2026
Download PDF
What’s in GPT-5

ABSTRACT

What’s in GPT-5? A Deep Dive into the Data Powering Next-Generation AI

As large language models grow more powerful, one thing becomes increasingly scarce: transparency. OpenAI, like many frontier AI labs, has become notably quiet about the specific data used to train its flagship models. But that hasn't stopped researchers from piecing together the puzzle.

In August 2024, Alan D. Thompson of LifeArchitect.ai released a comprehensive 27-page report titled "What's in GPT-5?" — a meticulous investigation into the datasets likely used to train OpenAI's next-generation model. The report has been reviewed by all major AI labs and intergovernmental organizations, and was even cited in the 2024 G7 AI document provided to world leaders including President Biden, Prime Minister Starmer, President Macron, and others

The investigation reveals that GPT-5’s raw pre-training corpus approaches approximately 500 trillion tokens (2 petabytes), condensing to over 70 trillion tokens (281 terabytes) after filtering—representing a dramatic scale increase over its predecessors. The dataset draws from over 27 distinct sources, categorized across web data (Common Crawl, Reddit), commercial news partnerships (News Corp, Associated Press, Financial Times), books (Books1–3), academic literature (Wikipedia, Consensus NLP), legal documents (FreeLaw), code repositories (GitHub, Stack Exchange), and multilingual resources. Critically, the analysis documents a paradigm shift toward synthetic data generation, tracing its evolution from Microsoft’s 1-billion-token experiments in late 2023 to Hugging Face’s 25-billion-token Cosmopedia pipeline in early 2024. For GPT-5, synthetic content—generated across 110 curated topics and audience types—is projected to constitute a substantial fraction of the final training mix. The findings highlight a broader industry trend of diminishing transparency, while quantitatively establishing the foundational data infrastructure required for next-generation frontier models. This work provides essential empirical grounding for assessing GPT-5’s emergent capabilities, biases, and the strategic implications of synthetic data as a compounding asset in AI development.

Tags:
chatgpt 5