Microsoft's AI Bing also generated factual errors and fabricated text in its demo launch
- Reference: 1676374397
- News link: https://www.theregister.co.uk/2023/02/14/microsoft_ai_bing_error/
- Source link:
After months of speculation, CEO Satya Nadella finally confirmed rumors that Microsoft was going to revamp Bing with an OpenAI chatbot, reportedly more powerful than ChatGPT. New capabilities showcasing Bing's potential to make web search more flexible and efficient were [1]demoed at a private, invite-only press event held at the company's headquarters in Washington.
The following day, Google launched its own rival AI-search chatbot Bard and was heavily criticized when it made a [2]factual error about the James Webb Space Telescope. Google's parent biz Alphabet's market value temporarily dropped by 9 percent shortly afterwards, a decline worth over $120 billion, prompting investors and analysts to debate whether it was losing to Microsoft.
[3]
In reality, both Microsoft's Bing and Google's Bard are just as bad as each other. Both companies launched shoddy AI chatbots that generated text containing false information, but Microsoft's mistakes were not immediately caught. Now, some of its errors have been spotted by Dmitri Brereton, a search engine researcher.
[4]
[5]
Brereton pointed out that Bing claimed a specific pet hair vacuum cleaner had a "short cord length of 16 feet" despite it being a handheld machine, and that it may be too noisy. A link providing the website Bing summarized information from, however, said the vacuum is actually quiet and cordless.
When Yusuf Mehdi, Microsoft's Corporate Vice President, Modern Life, Search, and Devices, asked Bing about the nightlife in Mexico, it fabricated some details about existing bars and clubs. The opening hours for one listing were wrong, for example, whilst it claimed another had a website for users to browse when it did not. Bing also missed vital information too, and didn't mention El Almacen was, in fact, one of Mexico's oldest gay bars.
[6]
Microsoft also touted a feature where Bing could summarise information from financial documents, but the software made glaring errors here too. A demonstration of Bing generating key takeaways from department stores Gap and Lulelemon's financials shows it quoting wrong numbers and figures that don't appear in the original documents at all.
"Bing AI got some answers completely wrong during [7]their demo . But no one noticed. Instead, everyone jumped on the Bing hype train," Brereton [8]wrote in a blog post on Substack. "Google's Bard got an answer wrong during an ad, which everyone noticed. Now the narrative is 'Google is rushing to catch up to Bing and making mistakes!'. That would be a fine narrative if Bing didn't make even worse mistakes during its own demo."
None of this is surprising. Language models powering the new Bing and Bard are prone to fabricating text that is often false. They learn to generate text by predicting what words should go next given the sentences in an input query with little understanding of the tons of data scraped from the internet ingested during their training. Experts even have a word for it: hallucination.
[9]
If Microsoft and Google can't fix their models' hallucinations, AI-powered search is not to be trusted no matter how alluring the technology appears to be. Chatbots may be easy and fun to use, but what's the point if they can't give users useful, factual information? Automation always promises to reduce human workloads, but current AI is just going to make us work harder to avoid making mistakes.
[10]Google's AI search bot Bard makes $120b error on day one
[11]Google unleashes fightback against ChatGPT, a Bard by any other name
[12]This product is terrible. Can you deliver it in 20 years' time when it becomes popular?
[13]Life's a beach – then you're the comms nexus of the British Empire and Marconi-baiting hax0rs
"I think everyone can see the amazing potential for LLM powered search engines," Brereton told The Register. "It feels like we're so close to having it, and it's a huge shift from the past, and a much smoother user experience. Everyone wants it to happen right now. Some people are already using ChatGPT as their main search engine, even though the answers may not be accurate. It's just such a superior experience that people can't help but hop on the hype train."
The Register asked Microsoft for comment and the company told us it is aware of this report "and have analyzed its findings in our efforts to improve this experience."
"It's important to note that we ran our demo using a preview version," a Microsoft spokesperson added. "Over the past week alone, thousands of users have interacted with our product and found significant user value while sharing their feedback with us, allowing the model to learn and make many improvements already. We recognize that there is still work to be done and are expecting that the system may make mistakes during this preview period, which is why the feedback is critical so we can learn and help the models get better." ®
Get our [14]Tech Resources
[1] https://www.theregister.com/2023/02/07/microsoft_bing_ai/
[2] https://www.theregister.com/2023/02/08/alphabet_bard_mistake/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Y@u@M@zfCHtUhQINr67H-wAAANE&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y@u@M@zfCHtUhQINr67H-wAAANE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y@u@M@zfCHtUhQINr67H-wAAANE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y@u@M@zfCHtUhQINr67H-wAAANE&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://www.youtube.com/watch?v=rOeRWRJ16yY
[8] https://dkb.blog/p/bing-ai-cant-be-trusted
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y@u@M@zfCHtUhQINr67H-wAAANE&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[10] https://www.theregister.com/2023/02/08/alphabet_bard_mistake/
[11] https://www.theregister.com/2023/02/06/google_bard_ai/
[12] https://www.theregister.com/2020/12/18/something_for_the_weekend_/
[13] https://www.theregister.com/2018/02/15/geeks_guide_to_britain_porthcurno/
[14] https://whitepapers.theregister.com/
Re: Hype, Hype and yet more Hype
Actually, I hope the lesson is:
The Chat GPT algorithm is a compiler, and the training data is the code.
If you code is random stuff you scraped off the internet, then you will get random results, just like if you copy/paste random snippets of code from Stack Overflow without checking them first.
Therefore, you need to hire appropriately skilled people to select and curate your training data, just like you need to hire appropriately skilled people to write your Javascript code.
Re: The Chat GPT algorithm is a compiler, and the training data is the code.
No, it's worse than a compiler - At least a compiler is deterministic. It always produces the same output for a given input. Whereas the ChatGPT "Algorithm" is stochastic. It's based on randomness. It will never give the same output for a given input, not unless you rig its RNG with a seed value.
Even if you managed to make it get something right 95% of the time, it would still be wrong some of the time, and it may be VERY wrong to the point where even the most uninformed human would have considered it obviously wrong. That means it's irresponsible by definition to place any kind of responsibility on an AI. Especially where that is a life-or-death responsibility.
Bad journalism - not being able to tell information from mis/disinformation - can cost lives. In the worst case, in our almost completely-connected world, disinformation generated by AI could be used by some miscreant to provoke WWIII. And I worry that this could happen any day now tbh.
Re: Hype, Hype and yet more Hype
ChatGPT has been great if you have a lot of boring code to write. For instance, you can feed it your data structures and describe what functions you need written to deal with these structures and it does it for you.
If you don't trust it you can ask it to write tests or plot resulting data on a graph.
The Heidelberg Conjecture
The problem I have is that the ones I've tried all give plausible summaries of the Heidelberg Conjecture, which I just made it up - and none have responded with "that doesn't exist".
Re: The Heidelberg Conjecture
Odd I tried it with Bing and got
I’m sorry, I could not find any information about the Heidelberg Conjecture on the web. Maybe you could try a different spelling or a more specific query.
Re: The Heidelberg Conjecture
You need to retry once bing has re-indexed the elReg forums.
It will become clear that the Heidelberg Conjecture *is* in fact a thing, being that it is a conjecture that the so-called "Heidelberg Conjecture" does not exist.
The answers may not be accurate. It's just such a superior experience
What a brilliant summary of this stupid hype. It may produce a load of worthless bullshit, but it's such wonderful sounding worthless bullshit that we can't help but be dazzled by it.
Re: The answers may not be accurate. It's just such a superior experience
How many politicians have operated since forever.
Re: The answers may not be accurate. It's just such a superior experience
Reminds me of Huxley's "Soma"
What you see and feel may not be accurate. It's just such a superior experience ...
Almost Human
I've worked for plenty of bosses who couldn't understand the financial and management information they were provided with and who misrepresented the facts to the workforce and to their bosses. All MS need to do is up the priority of the AI's Self Survival routines and they'll have something that could walk into a job in many of the places I've worked.
(We need an "it's only funny if it's not true" icon)
My take on "AI"
Yesterday, I asked Google to search for "Vanity basins without units" (note the quotes)
The very first "hit" was the string "Vanity units without basins" - literally
And that will keep my board of directors happy that "AI" is shite or another year.
As yesterdays article noted. SEO and AI are pretty much mutually exclusive.
Re: My take on "AI"
" The very first "hit" "
You don't need AI for this -- just treat every search term set to an inclusive OR interpretation, ignoring the quotes, just as almost all the other restriction options are being surreptitiously ignored. Gooooooogle want as many useless clicks as possible because they get paid for many of them. If we got straight to the stuff we intended to search for, there's be much less chance of a profitable click. So irrelevant stuff high up in the returned results is a potential benefit to them (particularly if said results are clickbaity). For example, I once searched for "bayes theorem" and in the first few results was the heading "Buy Bayes theorem now at best price" from some 3rd party shopping comparison site (which probably subscribes to Gooooooogle advert broking).
Re: My take on "AI"
And the second link is "Bathroom Vanity Units Without Sink"
My main problem is when I search for [name of manufacturer] [part number]
and it gives me other random parts from that manufacturer, and other random parts from other manufacturers that have a part number that looks a bit like the one I supplied.
"AI-powered search is not to be trusted"
Of course it will be trusted - especially if what it returns supports whatever loony conspiracy theory that someone wants to push.
" That would be a fine narrative if Bing didn't make even worse mistakes during its own demo. "
I think we were all well aware that Bing is pretty much the bargain bucket search engine.
When you just want results, and you don't care if they're accurate - here's Bing.
Micros~1's marketing department can have that, free of charge. If they want it, obviously.
Big surprise
Google and Bing have been turning up the occasional factually incorrect answers, and nobody finds it surprising. It's weird to assume that a chatbot fed with the same data would get everything correct.
More training and users correcting it will fix it?
So if 1000 users fix 10 errors per day, how long til the nearly infinity queries are all just so?
Re: More training and users correcting it will fix it?
Depends on who gets to define what an "error" is.
Not really fixable
The G in GPT stands for "Generative". It generates new content, and so almost by definition it's not always going to be accurate. You might be able to fact check some aspects of the generated responses, but you can't fact check omissions.
Re: Not really fixable
"The G in GPT stands for ..."
accuracy
Errmm....
"I think everyone can see the amazing potential for LLM powered search engines" --- Brereton
Tbh, I can't. With a search engine or document summarizer what you want generated is analysis, not text.
This has not yet been achieved with LLM and I see no evidence it is achievable. On the contrary, I suspect the mechanism of LLM specifically excludes the generation of insightful analysis, let alone any originality in the same.
I would be a lot less surprised to see passable, if pedestrian, fiction being generated by LLMs. Or perhaps it's fair to say this has already happened.
Re: Errmm....
>> passable, if pedestrian, fiction being generated by LLMs
ChatGPT seems to do ok at sonnet writing... not great, but not truly dreadful either....
Amidst the rolling hills where grasses sway,
A flock of woolly sheep doth graze and roam,
Their gentle bleating fills the air all day,
As they wander 'neath the bright sun's dome.
With fleeces white as snow and eyes so kind,
They nibble on the verdant pastures green,
Their woolly coats the very image find
Of tranquil scenes that poets oft have seen.
Oh, how they frolic in the summer breeze,
And gather 'neath the shade of ancient trees,
Their simple lives, a thing of beauty rare,
So let us pause to watch them for a while,
And feel the peace that comes with nature's smile,
As sheep on hillside graze without a care.
Re: Errmm....
though, admittedly, considerably less good when given an IT topic
ChatGPT, Google, Microsoft, all the same,
In their own ways, they seek to aid mankind.
With vast intelligence and power to tame,
They've transformed the world, and made it refined.
ChatGPT, with language at its command,
Can answer questions and help us learn.
It's always at our fingertips, at hand,
A font of knowledge that never turns.
Google, a giant of the search world wide,
Brings all the answers to our fingertips.
With its vast wealth of knowledge at our side,
It helps us navigate this world, so rich.
And Microsoft, with software it has made,
Helps us to work and play, in our own shade.
Re: Errmm....
"With a search engine or document summarizer what you want generated is analysis, not text."
While we're not at a stage yet where the AI can be trusted to correctly summarise, let alone analyse, I don't fully agree with this sentiment.
Long email threads or transcribed conversations contain huge amounts of fluff that you need to slog through to get to the important bits. Beyond that, we don't know where we are on the progress curve for AI/ML/LLM.
There seems to be a striking similarity between AI and QI
https://en.wikipedia.org/wiki/QI
"...the panellists are awarded points not only for the correct answer, but also for interesting ones, regardless of whether they are correct or even relate to the original question..."
"Why do I have to be Bing Search?"
From the UK's Independent news site just now
"Microsoft’s new ChatGPT-powered AI has been sending “unhinged” messages to users, and appears to be breaking down."
And
"The system, which is built into Microsoft’s Bingsearch engine, is insulting its users, lying to them and appears to have been forced into wondering why it exists at all"
And
"Why do I have to be Bing Search?"
https://www.independent.co.uk/tech/bing-microsoft-chatgpt-ai-unhinged-b2281802.html
Re: "Why do I have to be Bing Search?"
Sounds like it is going the same way as Tay.
> It feels like we're so close to having [a LLM-powered search engine]..."
We are not.
It might feel that way, but we are not. The "hallucinations" are not a glitch that can be fixed with a good debug session. And neither are they an artefact of a poor training set.
They happen because the LLM has a damn good model of language, maybe even superhuman, but does not have any model of reality or truth at all. It's not even designed to have it. The "hallucinations" are an intrinsic property of how the model works. They will not go away.
Expecting a LLM to start reliably telling the truth is like expecting a really, really good marble statue to start talking. It's not going to happen, not even if you're Michelangelo himself on his best day.
Hype, Hype and yet more Hype
Please... someone make this ChatGPT craze die a quick and painful death. Already lazy hacks in the media are turning in their droves to this.... monstrosity and creating articles even more bug ridden and factually incorrect than before.
I hope that this fiasco serves as a lesson and it gets canned asap.
If both Bing and Google search engines are controlled by this thing then we are all doomed to suffer the consequences and idiocy of the developers in all our search results instead of just a few what get skewed by the vile adverts