OpenAI Allegedly Used Paywalled O'Reilly Books for AI Training

Apr 2, 2025
A recent paper by the AI Disclosures Project suggests that OpenAI's GPT-4o model was trained on paywalled O'Reilly Media books without a licensing agreement.
OpenAI Allegedly Used Paywalled O'Reilly Books for AI Training

OpenAI has been accused of training its AI models on copyrighted content without permission, with a new paper from the AI Disclosures Project suggesting that the company used paywalled books from O'Reilly Media for its GPT-4o model. The paper, authored by Tim O'Reilly, Ilan Strauss, and Sruly Rosenblat, indicates that GPT-4o shows strong recognition of non-public O'Reilly book content compared to earlier models like GPT-3.5 Turbo.

The research employed a method known as DE-COP, which detects copyrighted content in language models' training data. This method revealed that GPT-4o likely has prior knowledge of many non-public O'Reilly books published before its training cutoff date. The findings highlight the need for increased transparency in AI model training data sources.

The AI Disclosures Project, co-founded by Tim O'Reilly and Ilan Strauss, aims to address the societal impacts of AI's commercialization by advocating for better corporate transparency. The paper's findings suggest that OpenAI, despite having some licensing agreements, may have used unlicensed paywalled content to enhance its AI models.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like AI Policy Brief or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

AI Policy Brief

Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.