WEATHER ALERT

Chatbots sometimes make things up. Is AI’s hallucination problem fixable?

Advertisement

Advertise with us

Spend enough time with ChatGPT and other artificial intelligence chatbots and it doesn't take long for them to spout falsehoods.

Read this article for free:

or

Already have an account? Log in here »

To continue reading, please subscribe:

Subscribe and receive a limited-edition Free Press branded hat or tote.

Digital Subscription

One year of digital access for only $205*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles

*First annual payment billed as $205.00 + GST for one year. This annual subscription will automatically renew at $233.00 + GST every 52 weeks (10% off the regular annual price of $259.35). Offer available to new and qualified returning subscribers only. Cancel any time.

To continue reading, please subscribe:

Add Free Press access to your Brandon Sun subscription for only an additional

$1 for the first 4 weeks*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles
Start now

*Your next Brandon Sun subscription payment will increase by $1.00 and you will be charged $17.95 plus GST for four weeks. After four weeks, your payment will increase to $24.95 plus GST every four weeks.

Hey there, time traveller!
This article was published 01/08/2023 (1091 days ago), so information in it may no longer be current.

Spend enough time with ChatGPT and other artificial intelligence chatbots and it doesn’t take long for them to spout falsehoods.

Described as hallucination, confabulation or just plain making things up, it’s now a problem for every business, organization and high school student trying to get a generative AI system to compose documents and get work done. Some are using it on tasks with the potential for high-stakes consequences, from psychotherapy to researching and writing legal briefs.

“I don’t think that there’s any model today that doesn’t suffer from some hallucination,” said Daniela Amodei, co-founder and president of Anthropic, maker of the chatbot Claude 2.

FILE - Text from the ChatGPT page of the OpenAI website is shown in this photo, in New York, Feb. 2, 2023. Anthropic, ChatGPT-maker OpenAI and other major developers of AI systems known as large language models say they're hard at work to make them more truthful. (AP Photo/Richard Drew, File)
FILE - Text from the ChatGPT page of the OpenAI website is shown in this photo, in New York, Feb. 2, 2023. Anthropic, ChatGPT-maker OpenAI and other major developers of AI systems known as large language models say they're hard at work to make them more truthful. (AP Photo/Richard Drew, File)

“They’re really just sort of designed to predict the next word,” Amodei said. “And so there will be some rate at which the model does that inaccurately.”

Anthropic, ChatGPT-maker OpenAI and other major developers of AI systems known as large language models say they’re working to make them more truthful.

How long that will take — and whether they will ever be good enough to, say, safely dole out medical advice — remains to be seen.

“This isn’t fixable,” said Emily Bender, a linguistics professor and director of the University of Washington’s Computational Linguistics Laboratory. “It’s inherent in the mismatch between the technology and the proposed use cases.”

A lot is riding on the reliability of generative AI technology. The McKinsey Global Institute projects it will add the equivalent of $2.6 trillion to $4.4 trillion to the global economy. Chatbots are only one part of that frenzy, which also includes technology that can generate new images, video, music and computer code. Nearly all of the tools include some language component.

Google is already pitching a news-writing AI product to news organizations, for which accuracy is paramount. The Associated Press is also exploring use of the technology as part of a partnership with OpenAI, which is paying to use part of AP’s text archive to improve its AI systems.

In partnership with India’s hotel management institutes, computer scientist Ganesh Bagler has been working for years to get AI systems, including a ChatGPT precursor, to invent recipes for South Asian cuisines, such as novel versions of rice-based biryani. A single “hallucinated” ingredient could be the difference between a tasty and inedible meal.

When Sam Altman, the CEO of OpenAI, visited India in June, the professor at the Indraprastha Institute of Information Technology Delhi had some pointed questions.

“I guess hallucinations in ChatGPT are still acceptable, but when a recipe comes out hallucinating, it becomes a serious problem,” Bagler said, standing up in a crowded campus auditorium to address Altman on the New Delhi stop of the U.S. tech executive’s world tour.

“What’s your take on it?” Bagler eventually asked.

Altman expressed optimism, if not an outright commitment.

“I think we will get the hallucination problem to a much, much better place,” Altman said. “I think it will take us a year and a half, two years. Something like that. But at that point we won’t still talk about these. There’s a balance between creativity and perfect accuracy, and the model will need to learn when you want one or the other.”

But for some experts who have studied the technology, such as University of Washington linguist Bender, those improvements won’t be enough.

Bender describes a language model as a system for “modeling the likelihood of different strings of word forms,” given some written data it’s been trained upon.

It’s how spell checkers are able to detect when you’ve typed the wrong word. It also helps power automatic translation and transcription services, “smoothing the output to look more like typical text in the target language,” Bender said. Many people rely on a version of this technology whenever they use the “autocomplete” feature when composing text messages or emails.

The latest crop of chatbots such as ChatGPT, Claude 2 or Google’s Bard try to take that to the next level, by generating entire new passages of text, but Bender said they’re still just repeatedly selecting the most plausible next word in a string.

File - OpenAI CEO Sam Altman speaks in Abu Dhabi, United Arab Emirates, Tuesday, June 6, 2023. Anthropic, ChatGPT- maker OpenAI and other major developers of AI systems known as large language models say they're hard at work to make them more truthful. (AP Photo/Jon Gambrell, File)
File - OpenAI CEO Sam Altman speaks in Abu Dhabi, United Arab Emirates, Tuesday, June 6, 2023. Anthropic, ChatGPT- maker OpenAI and other major developers of AI systems known as large language models say they're hard at work to make them more truthful. (AP Photo/Jon Gambrell, File)

When used to generate text, language models “are designed to make things up. That’s all they do,” Bender said. They are good at mimicking forms of writing, such as legal contracts, television scripts or sonnets.

“But since they only ever make things up, when the text they have extruded happens to be interpretable as something we deem correct, that is by chance,” Bender said. “Even if they can be tuned to be right more of the time, they will still have failure modes — and likely the failures will be in the cases where it’s harder for a person reading the text to notice, because they are more obscure.”

Those errors are not a huge problem for the marketing firms that have been turning to Jasper AI for help writing pitches, said the company’s president, Shane Orlick.

“Hallucinations are actually an added bonus,” Orlick said. “We have customers all the time that tell us how it came up with ideas — how Jasper created takes on stories or angles that they would have never thought of themselves.”

The Texas-based startup works with partners like OpenAI, Anthropic, Google or Facebook parent Meta to offer its customers a smorgasbord of AI language models tailored to their needs. For someone concerned about accuracy, it might offer up Anthropic’s model, while someone concerned with the security of their proprietary source data might get a different model, Orlick said.

Orlick said he knows hallucinations won’t be easily fixed. He’s counting on companies like Google, which he says must have a “really high standard of factual content” for its search engine, to put a lot of energy and resources into solutions.

“I think they have to fix this problem,” Orlick said. “They’ve got to address this. So I don’t know if it’s ever going to be perfect, but it’ll probably just continue to get better and better over time.”

Techno-optimists, including Microsoft co-founder Bill Gates, have been forecasting a rosy outlook.

“I’m optimistic that, over time, AI models can be taught to distinguish fact from fiction,” Gates said in a July blog post detailing his thoughts on AI’s societal risks.

He cited a 2022 paper from OpenAI as an example of “promising work on this front.” More recently, researchers at the Swiss Federal Institute of Technology in Zurich said they developed a method to detect some, but not all, of ChatGPT’s hallucinated content and remove it automatically.

But even Altman, as he markets the products for a variety of uses, doesn’t count on the models to be truthful when he’s looking for information.

“I probably trust the answers that come out of ChatGPT the least of anybody on Earth,” Altman told the crowd at Bagler’s university, to laughter.

Report Error Submit a Tip

More Stories

Trailblazer returns to share success, inspire next generation

Grace Penner 6 minute read Preview

Trailblazer returns to share success, inspire next generation

Grace Penner 6 minute read Saturday, Jul. 25, 2026

Starstruck kids and fans zigzagged throughout the rink, anxiously waiting to get an autograph from Kati Tabin and to see their role model on the same side of the glass.

Tabin, Montreal Victoire defender and Team Canada silver medallist, on Saturday came home to the rink where it all began, the Oakbank Community Club. She brought the PWHL championship trophy, the Walter Cup, won in her third season with the club, as well as her 2026 Milan Cortina Olympic medal.

The arena was packed to the brim with fans from all over coming to see Tabin — the local player who made history in women’s hockey. Hundreds of enthusiasts of all ages lined up with their Victoire or Team Canada merchandise, some waiting over two hours to see their idol.

Eight-year-old Loxley Cierzan came with her mother, Ainsley, and other family members to meet the one and only Tabin. Cierzan has loved seeing other girls grace the ice and play the competitive sport she adores. When the PWHL takeover stopped in Winnipeg in March, she not only got to watch the game but also a practice beforehand where Tabin signed her Victoire hat.

Read
Saturday, Jul. 25, 2026

A Life's Story: Beloved matriarch, kindergarten teacher lived a musical life

Gabrielle Piché 6 minute read Preview

A Life's Story: Beloved matriarch, kindergarten teacher lived a musical life

Gabrielle Piché 6 minute read Saturday, Jul. 25, 2026

Phyllis Dana faced a choice: leave Winnipeg to study opera at a prestigious Philadelphia school, or marry the man she loved.

She chose the latter.

Now, her Juno-winning granddaughter cherishes a Happy Birthday voicemail sung by Phyllis. Three generations of Dana women — Phyllis, her daughter, and granddaughter — would sing in Winnipeg’s prominent Jewish spaces. And hundreds of Manitobans know the matriarch as their kindergarten teacher.

“My mom said, ‘A sheet of music won’t keep you warm at night,’” Karen Dana recalled.

Read
Saturday, Jul. 25, 2026

Brown bombs as Bombers blasted

Taylor Allen 7 minute read Preview

Brown bombs as Bombers blasted

Taylor Allen 7 minute read Saturday, Jul. 25, 2026

Winnipeg Blue Bombers quarterback Dru Brown took Friday night's ‘Christmas in July’ theme a little too seriously.

He wasn't carrying around a red sack, but he did gift-wrap four interceptions to the Calgary Stampeders in what ended up being a 52-30 loss for the Blue and Gold in front of a sold-out crowd.

The result dropped the Bombers, now 1-3 at home this season, to 4-3. Calgary improved to 3-4 and now have a chance to win the tiebreaker when the two sides meet for a final time on Oct. 10 at Princess Auto Stadium.

“Yeah, not nearly good enough...You certainly don't want to go get into a track meet against a good offence like that,” said Bombers head coach Mike O’Shea.

Read
Saturday, Jul. 25, 2026

Sultans one win away from MJBL three-peat

Cassidy Dankochik 3 minute read Preview

Sultans one win away from MJBL three-peat

Cassidy Dankochik 3 minute read Yesterday at 6:31 PM CDT

A small delay due to lightning didn’t slow down the Carillon Sultans as they stormed into Koskie Field and defeated the Elmwood Giants 12-0 Sunday to take a 2-0 series lead in the Manitoba Junior Baseball League final.

Across the first two games of the best-of-five series, Carillon has a 27-5 run advantage after a similar offensive output during Game 1 in Steinbach on Friday. The Giants and Sultans have played in the finals the past three seasons, with Carillon winning both previous years.

The Sultans offence was quick to strike, plating eight runs in the top of the second inning. Anders Meilleur cleared the bases with a three-run double to start the barrage, with Dayton Christensen crushing a three-run home run to left-centre to cap off the inning.

“If we get on a roll in an inning, it’s contagious,” head coach Don Meilleur said. “We don’t hurt teams by getting one or two runs an inning, we hurt teams because we put up a crooked number most of the time and today was the perfect example… Anybody on that bench and anybody in the lineup can hurt another team. We just always find ways to get somebody to be a hero.”

Read
Yesterday at 6:31 PM CDT

Daycare grapples with five-year-old’s outbursts, choking incidents

Maggie Macintosh 4 minute read Preview

Daycare grapples with five-year-old’s outbursts, choking incidents

Maggie Macintosh 4 minute read Friday, Jul. 24, 2026

A Windsor Park child-care facility is weighing a five-year-old’s right to care versus the safety of other children who have been choked.

Little People’s Place expelled a pre-schooler in the spring in response to his aggressive behaviour toward other children and staff.

The child’s family has appealed to the province for him to be reinstated.

The case is raising concerns about inclusion support resources and measures to protect everyone at the Cottonwood Road daycare.

Read
Friday, Jul. 24, 2026

Puzzles Palace

1 minute read Monday, Jul. 13, 2026

To solve our puzzles, please subscribe with this special offer: |