OpenAI Allegedly Used Paywalled O'Reilly Books for AI Training
OpenAI has been accused of training its AI models on copyrighted content without permission, with a new paper from the AI Disclosures Project suggesting that the company used paywalled books from O'Reilly Media for its GPT-4o model. The paper, authored by Tim O'Reilly, Ilan Strauss, and Sruly Rosenblat, indicates that GPT-4o shows strong recognition of non-public O'Reilly book content compared to earlier models like GPT-3.5 Turbo.
The research employed a method known as DE-COP, which detects copyrighted content in language models' training data. This method revealed that GPT-4o likely has prior knowledge of many non-public O'Reilly books published before its training cutoff date. The findings highlight the need for increased transparency in AI model training data sources.
The AI Disclosures Project, co-founded by Tim O'Reilly and Ilan Strauss, aims to address the societal impacts of AI's commercialization by advocating for better corporate transparency. The paper's findings suggest that OpenAI, despite having some licensing agreements, may have used unlicensed paywalled content to enhance its AI models.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from Regulation
Oct 2 California Keeps AI Terminology Despite Federal SI Order Oct 2 Oversight Board Calls for Independent AI Governance Oct 2 MPTF Coalition Seeks Platform Liability for AI Scams Oct 2 Trump Consulted Grok Before Maduro Capture, Report Says Oct 1 RFK Jr. says AI can give patients better medical second opinionsAI Policy Brief
Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
OpenAI Pauses Latest Model Training After Agent Incidents
OpenAI Publishes Model Misalignment Reporting Framework
Researchers Used Anthropic Tool to Access OpenAI Employee Account
OpenAI apologizes for access to Australian government systems
OpenAI Links Model Reasoning Extraction Campaign to Moonshot AI
Daily AI Brief: the AI news that matters, in your inbox.