The Guardian reported on September 26 that historical texts digitised through Oxford’s Bodleian Libraries partnership with OpenAI were used to populate an OpenAI training set, citing internal university documents. That is a different use from the one a reader would take from Oxford’s March 2025 announcement: making library material searchable and accessible to students and researchers. The distinction matters because a public-access digitisation project and a commercial model-training dataset can create different benefits, obligations and questions for the institution.
The reported training use should be read with its provenance intact. NextWith.ai has not reviewed the internal documents or the contract, and Oxford disputes the suggestion that this aspect of the project was hidden. The Guardian’s account quotes an Oxford spokesperson saying digitisation was the university’s primary interest but that staff had been open about contributing training data. OpenAI told the newspaper that it wanted its models to preserve historical knowledge and reflect a broader range of cultures and perspectives. Neither statement specifies the precise model versions, the full list of works or the technical way the data was used.
What Oxford disclosed publicly
Oxford’s March 4, 2025 announcement described a five-year collaboration with OpenAI and a pilot to digitise public-domain material held by the Bodleian. It said the resulting collections, previously unavailable online, would become searchable and accessible to students and researchers worldwide. The announcement identified 3,500 global dissertations dated 1498 to 1884 as one collection to digitise. It did not spell out model-training use in that announcement. That omission does not by itself establish that a contractual term was secret; it does explain why the new reporting changes the public picture of the partnership.
The Bodleian’s project page describes a pilot that began in February 2025, funded by OpenAI. It frames the work as an investigation into scanning capacity, metadata, transcription and search. The library says around 125,000 images from its Global Dissertations collection had been captured for digitisation, along with hundreds of thousands of catalogue-card images. The Guardian separately reports that 125,000 dissertation images had been shared with OpenAI by June 2025. Those are related but distinct statements: the library page establishes capture, while the reported transfer rests on the newspaper’s document-based account. Neither figure tells us how many pages entered a completed model-training run.
Why the distinction matters
For researchers, digitising hard-to-reach material can be valuable even if no AI model is trained on it. Better scans, searchable text and improved catalogues can let scholars inspect sources without traveling to Oxford. The library’s project page describes tests of optical and handwriting recognition, metadata extraction and human review by specialist librarians. These are concrete library services. Training a general AI model is a separate downstream use, with different questions about attribution, reproducibility, access and who can verify what was included.
The Guardian says minutes obtained through a freedom-of-information request record staff concerns about reputational risk and the environmental implications of partnering with an energy-intensive technology. That is evidence of internal debate, not proof that the project breached a policy or that its climate impact has been measured here. It also does not erase a potential public benefit from digitisation. A sound assessment needs both sides of the arrangement: what becomes openly available to readers, and what use OpenAI gains from the scans.
Oxford told the newspaper that the digitised material was modest in scale and out of copyright, that OpenAI’s use was non-exclusive, and that the Bodleian retained rights to the scans. The university said it would begin publishing the materials openly online within months. Those are consequential commitments, but the article does not supply a dated release inventory or a public training-data manifest. Until those exist, readers cannot independently check which works have been made available or which were used to train a model. That is a disclosure gap, not a claim that the partnership is unlawful.
The practical test is therefore measurable. A library can show the public what was digitised, when the promised scans and transcriptions are available, and what permission a partner received to use them in AI training. Such a record would let scholars assess the access benefit and the training exchange separately, rather than relying on a single headline about preservation or a single concern about data use.
When Oxford publishes the promised Bodleian scans, check the public collection inventory and release date against the reported training use; ask the university to disclose the training-use terms if those remain absent.