AI & TechArtificial IntelligenceBigTech CompaniesDigital PublishingNewswireTechnologyWhat's Buzzing

Oxford, OpenAI Partner for Bodleian Text AI Training

▼ Summary

– Oxford University allowed OpenAI to use scanned texts from the Bodleian Library for AI training, a deal initially presented as a scanning initiative.
– By June 2025, the library had provided 125,000 scans of old PhD theses, prompting internal staff concerns about reputational harm and energy usage.
– Oxford defended the arrangement by stating the data was out of copyright, non-exclusive, and that the AI training aspect was disclosed rather than hidden.
– The agreement highlights a broader industry trend where AI firms acquire physical books for high-quality training data to avoid internet-generated ‘slop’.
– Unlike competitors who destroy books during scanning, Oxford’s deal keeps the rare volumes intact while making them accessible to scholars.

Oxford University has entered into a controversial agreement with OpenAI, granting the artificial intelligence giant access to historical documents from its Bodleian Library for the purpose of training machine learning models. This partnership, which came to light through internal documents reviewed by The Guardian, reveals that scanned texts were integrated directly into OpenAI’s dataset. The story was initially reported by journalists Ethan Penny and Dan Milmo on Saturday.

While the university announced the collaboration publicly in March 2025, the initial framing focused on digitization efforts rather than AI development. At that time, Oxford stated that OpenAI’s technology would facilitate the scanning of rare manuscripts, thereby increasing accessibility for students and scholars. The institution did not disclose at that stage that these materials would also serve as fuel for AI algorithms. By June 2025, reports indicated that the Bodleian had provided OpenAI with 125,000 scans. These digital copies encompass doctoral dissertations from European and American universities dating back to the 19th and 20th centuries.

Internal records obtained via a freedom of information request highlight significant concerns among library staff. Meeting notes reveal anxiety regarding potential reputational damage to the university and the substantial energy consumption associated with AI operations. Despite these reservations, Oxford defended the arrangement. A spokesperson emphasized that the volume of data was limited, the materials were out of copyright, and the license was not exclusive to OpenAI. Furthermore, the library retains ownership rights and plans to make the scans available online in the coming months. The spokesperson clarified that the dual use of the data was transparent, noting that while digitization was the primary objective, staff had been clear about the secondary intent to train AI models.

“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” an OpenAI spokesperson told the Guardian.

This collaboration positions Oxford as the sole United Kingdom representative within OpenAI’s NextGenAI group. Other participating institutions include Boston Public Library, Caltech, MIT, and the University of Michigan. The deal emerges against a backdrop of intensifying competition among AI firms for high-quality, human-generated content. As the internet becomes saturated with AI-generated text, companies are increasingly purchasing physical books to ensure their training data remains free of synthetic “slop.” This trend has sparked backlash from secondhand booksellers, particularly after reports surfaced in August that some vendors were destroying rare volumes to scan them for AI purposes. In contrast, the Bodleian’s agreement ensures that the physical books remain intact.

(Source: The Next Web)

Topics

ai training data 95% academic partnerships 85% Ethical Concerns 80% library digitization 75% book preservation 70%
Show More