AI ‘gold rush’ for chatbot training data could run out of human-written text

Advertisement

Advertise with us

Artificial intelligence systems like ChatGPT could soon run out of what keeps making them smarter — the tens of trillions of words people have written and shared online.

Read this article for free:


or

Already have an account? Log in here »

To continue reading, please subscribe:

Subscribe and receive a limited-edition Free Press branded hat or tote.

Digital Subscription

One year of digital access for only $205*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles

*First annual payment billed as $205.00 + GST for one year. This annual subscription will automatically renew at $233.00 + GST every 52 weeks (10% off the regular annual price of $259.35). Offer available to new and qualified returning subscribers only. Cancel any time.

To continue reading, please subscribe:

Add Free Press access to your Brandon Sun subscription for only an additional

$1 for the first 4 weeks*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles
Start now

*Your next Brandon Sun subscription payment will increase by $1.00 and you will be charged $17.95 plus GST for four weeks. After four weeks, your payment will increase to $24.95 plus GST every four weeks.

Hey there, time traveller!
This article was published 06/06/2024 (821 days ago), so information in it may no longer be current.

Artificial intelligence systems like ChatGPT could soon run out of what keeps making them smarter — the tens of trillions of words people have written and shared online.

A new study released Thursday by research group Epoch AI projects that tech companies will exhaust the supply of publicly available training data for AI language models by roughly the turn of the decade — sometime between 2026 and 2032.

Comparing it to a “literal gold rush” that depletes finite natural resources, Tamay Besiroglu, an author of the study, said the AI field might face challenges in maintaining its current pace of progress once it drains the reserves of human-generated writing.

Artificial intelligence systems like ChatGPT are gobbling ever-larger collections of human writings they need to get smarter. (AP Digital Embed)
Artificial intelligence systems like ChatGPT are gobbling ever-larger collections of human writings they need to get smarter. (AP Digital Embed)

In the short term, tech companies like ChatGPT-maker OpenAI and Google are racing to secure and sometimes pay for high-quality data sources to train their AI large language models – for instance, by signing deals to tap into the steady flow of sentences coming out of Reddit forums and news media outlets.

In the longer term, there won’t be enough new blogs, news articles and social media commentary to sustain the current trajectory of AI development, putting pressure on companies to tap into sensitive data now considered private — such as emails or text messages — or relying on less-reliable “synthetic data” spit out by the chatbots themselves.

“There is a serious bottleneck here,” Besiroglu said. “If you start hitting those constraints about how much data you have, then you can’t really scale up your models efficiently anymore. And scaling up models has been probably the most important way of expanding their capabilities and improving the quality of their output.”

The researchers first made their projections two years ago — shortly before ChatGPT’s debut — in a working paper that forecast a more imminent 2026 cutoff of high-quality text data. Much has changed since then, including new techniques that enabled AI researchers to make better use of the data they already have and sometimes “overtrain” on the same sources multiple times.

But there are limits, and after further research, Epoch now foresees running out of public text data sometime in the next two to eight years.

The team’s latest study is peer-reviewed and due to be presented at this summer’s International Conference on Machine Learning in Vienna, Austria. Epoch is a nonprofit institute hosted by San Francisco-based Rethink Priorities and funded by proponents of effective altruism — a philanthropic movement that has poured money into mitigating AI’s worst-case risks.

Besiroglu said AI researchers realized more than a decade ago that aggressively expanding two key ingredients — computing power and vast stores of internet data — could significantly improve the performance of AI systems.

The amount of text data fed into AI language models has been growing about 2.5 times per year, while computing has grown about 4 times per year, according to the Epoch study. Facebook parent company Meta Platforms recently claimed the largest version of their upcoming Llama 3 model — which has not yet been released — has been trained on up to 15 trillion tokens, each of which can represent a piece of a word.

But how much it’s worth worrying about the data bottleneck is debatable.

“I think it’s important to keep in mind that we don’t necessarily need to train larger and larger models,” said Nicolas Papernot, an assistant professor of computer engineering at the University of Toronto and researcher at the nonprofit Vector Institute for Artificial Intelligence.

Papernot, who was not involved in the Epoch study, said building more skilled AI systems can also come from training models that are more specialized for specific tasks. But he has concerns about training generative AI systems on the same outputs they’re producing, leading to degraded performance known as “model collapse.”

Training on AI-generated data is “like what happens when you photocopy a piece of paper and then you photocopy the photocopy. You lose some of the information,” Papernot said. Not only that, but Papernot’s research has also found it can further encode the mistakes, bias and unfairness that’s already baked into the information ecosystem.

If real human-crafted sentences remain a critical AI data source, those who are stewards of the most sought-after troves — websites like Reddit and Wikipedia, as well as news and book publishers — have been forced to think hard about how they’re being used.

“Maybe you don’t lop off the tops of every mountain,” jokes Selena Deckelmann, chief product and technology officer at the Wikimedia Foundation, which runs Wikipedia. “It’s an interesting problem right now that we’re having natural resource conversations about human-created data. I shouldn’t laugh about it, but I do find it kind of amazing.”

While some have sought to close off their data from AI training — often after it’s already been taken without compensation — Wikipedia has placed few restrictions on how AI companies use its volunteer-written entries. Still, Deckelmann said she hopes there continue to be incentives for people to keep contributing, especially as a flood of cheap and automatically generated “garbage content” starts polluting the internet.

AI companies should be “concerned about how human-generated content continues to exist and continues to be accessible,” she said.

From the perspective of AI developers, Epoch’s study says paying millions of humans to generate the text that AI models will need “is unlikely to be an economical way” to drive better technical performance.

As OpenAI begins work on training the next generation of its GPT large language models, CEO Sam Altman told the audience at a United Nations event last month that the company has already experimented with “generating lots of synthetic data” for training.

“I think what you need is high-quality data. There is low-quality synthetic data. There’s low-quality human data,” Altman said. But he also expressed reservations about relying too heavily on synthetic data over other technical methods to improve AI models.

“There’d be something very strange if the best way to train a model was to just generate, like, a quadrillion tokens of synthetic data and feed that back in,” Altman said. “Somehow that seems inefficient.”

——————

The Associated Press and OpenAI have a licensing and technology agreement that allows OpenAI access to part of AP’s text archives.

Report Error Submit a Tip

More Stories

Neighbours relieved after fire strikes Quickie Mart

Morgan Modjeski 4 minute read Preview

Neighbours relieved after fire strikes Quickie Mart

Morgan Modjeski 4 minute read Wednesday, Sep. 2, 2026

As Janina McNaughton watched the Quickie Mart on River Avenue go up in flames on Tuesday, she felt relieved.

“The first thing I see is flames, like a fireball of flames inside,” McNaughton said about watching the blaze from her bedroom across the street.

On Wednesday, the bright green building, at 292 River Ave., was boarded up and charred on the outside.

Crews from the Winnipeg Fire Paramedic Service responded at 8:13 p.m. Tuesday to battle the blaze, said the city in a news release. It was under control within the hour.

Read
Wednesday, Sep. 2, 2026

Frequency of private plasma donations called into question

Marsha McLeod and Rylee Gerrard 9 minute read Preview

Frequency of private plasma donations called into question

Marsha McLeod and Rylee Gerrard 9 minute read Updated: 8:07 AM CDT

Two weeks after the death of a second plasma donor in Winnipeg at a privately run collection site, provincial and territorial health officials compiled a list of urgent questions for Health Canada.

A health department official in Nova Scotia, which leads the Provincial Territorial Blood Liaison Committee, wrote in mid-February to the federal regulator that their members had expressed “a great deal of curiosity tempered with concern” and forwarded two dozen questions, according to documents obtained by the Free Press.

Under a section titled “MANITOBA ESTABLISHMENT INCIDENTS,” the provinces wanted to know: “What requirements are in place to ensure Grifols is ensuring collections frequency (does) not exceed the limit under the Health Canada licence?” and “Is Health Canada able to identify whether Grifols Canada has a donor system across establishments? (e.g., can identify when a donor last donated and where, similar to their operations in the United States).”

“Does the situation of two high-severity incidents in one Province or at establishments run by a relatively new entity to Canada, within a short period of time, initiate any additional protocols?”

Read
Updated: 8:07 AM CDT

A life's story: translator, poet, radio host, actor loved life to the fullest

Tiago Resko 7 minute read Preview

A life's story: translator, poet, radio host, actor loved life to the fullest

Tiago Resko 7 minute read 3:00 AM CDT

Born with an innate sense of curiosity, Charles Leblanc always wanted to understand the world around him. To him, words and theatre weren’t just pastimes. They were means of understanding the human experience and helping others see different perspectives.

The poet, translator, actor and radio host dedicated his life to this practice, founding the Winnipeg International Writers’ Festival, establishing French community radio station En’Vol — where he hosted a show for 25 years — and acting as a core member of the Association of Translators, Terminologists and Interpreters of Manitoba.

Leblanc’s contributions and dedication to Manitoba’s arts and literary scenes made an impact that continues to be felt by the community.

He died June 26 at the age of 75.

Read
3:00 AM CDT

Agriculture sector wary of ‘strong headwind’ from south

Laura Rance-Unger 5 minute read 2:01 AM CDT

Industry analysts talked in circles on a recent webinar as they discussed the implications for agriculture of Canada’s deteriorating relationship with its largest trading partner.

On one hand, panellists for the hastily convened Canadian AgriFood Policy Institute event acknowledged the futility of negotiating on quicksand: a deal made today with the United States might be sucked into the muck tomorrow.

But on the other, webinar panellists speculated with varying degrees of optimism over what it will take to get negotiators back to bargaining.

This disconnect isn’t unique to this cluster of analysts. However, it underscores the difficulty of coming to terms with the seismic shift in America’s attitudes towards Canada.

Smith’s pretzel logic confuses

Editorial 4 minute read Preview

Smith’s pretzel logic confuses

Editorial 4 minute read 2:01 AM CDT

Alberta Premier Danielle Smith had both sides of her mouth working overtime on Aug. 30, when she spoke at an event in Calgary to commemorate Alberta Day, the annual celebration of the province’s entry into Confederation in 1905 (officially Sept. 1, it was celebrated on the weekend).

Smith’s comments — and the things she did not say — revealed a lot about the political tightrope the premier is trying to walk over the challenges currently faced by her province. At the top of that list of challenges is a persistent campaign to separate Alberta from the aforementioned Confederation. A campaign that Smith appears to both support and oppose, depending on which day of the week it is.

Her Alberta Day remarks were an excellent case in point.

Smith has decided that on Oct. 19, Albertans will vote on a series of non-binding propositions in a province referendum. One of those questions will ask Albertans whether they want to hold a referendum on separation, thus making this fall’s vote a “referendum on whether to hold a referendum,” as some observers have dubbed it.

Read
2:01 AM CDT

Politicians who harp on Klein’s history with Nygard do so at their own peril

Tom Brodbeck 5 minute read Preview

Politicians who harp on Klein’s history with Nygard do so at their own peril

Tom Brodbeck 5 minute read Yesterday at 1:13 PM CDT

It was probably the easiest prediction to make in Winnipeg politics this week: the moment Kevin Klein officially entered the mayoral race, someone would bring up Peter Nygard.

Sure enough, they did.

And it was led by Klein’s chief political opponent in this race, Mayor Scott Gillingham, who called Klein a “fixer” for the convicted sex offender.

None of this is particularly surprising.

Read
Yesterday at 1:13 PM CDT