Skip to content

openai

Oxford's Bodleian Library texts ended up training OpenAI

Promtime

OpenAI has scanned a rare set of 10,000 16th-century broadside ballads at Oxford's Bodleian Library. These are song lyrics and musical notes that once circulated on Tudor street corners, and according to internal documents reported by The Guardian, material digitised by OpenAI has gone to "populate the OpenAI training set". When Oxford announced the partnership in March 2025, it described the project as digitisation for students and researchers and did not say the texts would train OpenAI's models.

At a glance

  • Oxford's own minutes, obtained through a freedom of information request, show staff and members of the Bodleian governance committee worrying about reputational risk and the university's environmental commitments.
  • By June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including 19th- and 20th-century PhD theses from European and American universities.
  • Oxford says the material is modest in scale, out of copyright and not exclusive to OpenAI, and that the Bodleian keeps the rights and will publish the scans openly online within months.

If you missed the original deal, Oxford's own announcement described it as a five-year collaboration. It gave students and faculty research grant funding, enterprise-level security and AI tools, plus access to OpenAI models including o1 and 4o. In the same announcement, Oxford presented the library side as a pilot to digitise public domain materials from the Bodleian, including 3,500 global dissertations from 1498 to 1884, so that researchers worldwide could search them.

Oxford's minutes record staff concerns over reputational risk and energy use

The paper trail comes from Oxford's own meeting minutes, which were released under a freedom of information request. They record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI. They also record worries about how a deal built on an energy-intensive technology fits with the university's environmental commitments.

OpenAI frames the project as preservation. A spokesperson said the company was "proud" to ensure "the AI models of today preserve the world's historical knowledge for the future". They added that with more than a billion people using this technology in everyday life, it matters that it reflects different cultures, histories and perspectives.

Oxford rejects the suggestion that it hid the machine-learning element from the public and students. A university spokesperson said digitisation was Oxford's primary interest, but that staff had been open about the project also contributing training data.

By June 2025, 125,000 dissertation images had been shared with OpenAI

By June 2025, the Bodleian had shared 125,000 images scanned from historical dissertations with OpenAI. They include PhD theses written at European and American universities in the 19th and 20th centuries. The 10,000 Tudor broadside ballads were scanned as well.

Staff have also discussed digitising 18th-century Irish state papers, the private letters of the Irish novelist Marie Edgeworth and Dorothy Hodgkin's penicillin notebooks. The reporting describes these as discussed, not as already scanned.

The contract also opens the door to mass digitisation of the Bodleian's whole collection, which runs to 23 million items, and the minutes mention an "Ask the Bod" chatbot. Oxford says the amount of text being digitised is "modest in scale" and covers only out-of-copyright material.

Oxford is the only UK member of OpenAI's NextGenAI project

Oxford is not alone. Under a project called NextGenAI, OpenAI has struck similar agreements with US research libraries including Boston Public Library, Caltech, MIT and the University of Michigan. Oxford is the project's only UK member.

The search for text goes beyond libraries. Booksellers report a run of orders for obscure titles, such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Secondhand bookshop owners speculate that because these books are unlikely to exist online in digitised form, they count as fresh data for the next generation of models.

Many of those books are destroyed. Anthropic has spent tens of millions of dollars buying books, slicing off their spines to scan the pages and then having them pulped, though it says it does not buy and destroy rare and antiquarian books. 404 Media hid a tracking device in a secondhand book order and followed it to an Amazon facility in the US, where books were also taken apart and scanned. Under the Oxford deal, the Bodleian's collections stay intact.

AI-generated text is filling the web, so developers are turning to old books

Language models learn by taking in huge amounts of text and picking up its patterns, which is how they learn to write complete sentences. IBM describes training as starting with billions or trillions of words from books, articles, websites and code. That text is cleaned and split into tokens, the small chunks of text a model actually works with.

According to IBM, the first stage is self-supervised. Nobody labels a correct answer, and the model finds structure in unlabelled text on its own. Underneath are transformer networks with billions of parameters, whose self-attention mechanism lets them focus on different tokens at different moments.

That is why the source of the text matters. Scraped websites are increasingly full of AI-generated material, which makes them less useful for training, so developers have turned to physical, often historical, book collections. Imagine learning to cook from a recipe box that keeps filling up with copies of your own recipes. A Tudor ballad or a 19th-century thesis is a card no model wrote.

The reporting leaves the scale and the terms open. Oxford calls the material modest in scale without giving a total, and nobody has said whether money changes hands or what OpenAI's training rights cover beyond being non-exclusive. In our view, the weak spot is the March 2025 announcement: Oxford says staff had been open about the training use, yet the public pitch talked only about digitisation for students and researchers.

When the Bodleian scans go public. Oxford says the library will start publishing the digitised material openly online in the next few months, as it does with work from other digitisation partnerships. That should let outsiders see what was scanned. No date has been given for mass digitisation of the 23 million items, and so far the "Ask the Bod" chatbot has only come up in meeting minutes.

Related stories

  1. OpenAI's 2019 note: "We trained GPT-3 on pirated stuff"
  2. Microsoft disclaims its own director's "astonishing theft"
  3. Anthropic planned a $2T IPO until doom fears landed
  4. SoftBank turns to junk bonds to fund its next OpenAI check
  5. FT puts OpenAI's cash burn at almost $280B through 2030
  6. DeepSeek and six rivals make a tenth of OpenAI and Anthropic

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.