News: 1708687995

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Are you ready to back up your AI chatbot's promises? You'd better be

(2024/02/23)


Opinion I keep hearing about businesses that want to fire their call center employees and front-line staffers as fast as possible and replace them with AI. They're upfront about it.

Meta CEO Mark Zuckerberg recently said the company behind [1]Facebook was laying off employees "so we can invest in these long-term, ambitious visions around AI." That may be a really dumb move. Just ask Air Canada.

Air Canada recently found out the hard way that [2]when your AI chatbot makes a promise to a customer, the company has to make good on it. Whoops!

[3]

In Air Canada's case, a virtual assistant told Jake Moffatt he could get a bereavement discount on his already purchased Vancouver to Toronto flight because of his grandmother's death. The total cost of the trip without the discount: CA$1,630.36. Cost with the discount: $760. The difference may be petty cash to an international airline, but it's real money to ordinary people.

[4]

[5]

The virtual assistant told him that if he purchased a normal-price ticket, he would have up to 90 days to claim back a bereavement discount. A real-live Air Canada rep confirmed he could get the bereavement discount.

When Moffatt later submitted his refund claim with the necessary documentation, Air Canada refused to pay out. That did not work out well for the company.

[6]

Moffatt took the business to small claims court, claiming Air Canada was negligent and had misrepresented its policy. Air Canada replied, in effect, that " [7]The chatbot is a separate legal entity that is responsible for its own actions. "

I don't think so!

The court agreed. "This is a remarkable submission. While a chatbot has an interactive component, it is still just a part of Air Canada's website. It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot."

[8]

The money quote for other businesses to pay attention to going forward with their AI plans is: "I find Air Canada did not take reasonable care to ensure its chatbot was accurate."

This is one case, and the damages were minute. Air Canada was ordered to pay Moffatt back the refund he was owed. Yet businesses need to know that they are as responsible for their AI chatbots being accurate as they are for their flesh-and-blood employees. It's that simple.

And, guess what? AI LLMs often aren't right. They're not even close. According to a study by non-profits AI Forensics and AlgorithmWatch, [9]a third of Microsoft Copilot's answers contained factual errors . That's a lot of potential lawsuits!

As Avivah Litan, a Gartner distinguished vice president analyst focused on AI, said, if you let your AI chatbots be your front-line of customer service, [10]your company "will end up spending more on legal fees and fines than they earn from productivity gains."

Attorney Steven A. Schwartz knows all about that. He relied on ChatGPT to find prior cases to support his case. And, Chat GPT found prior cases right enough. There was only one little problem. [11]Six of the cases he cited didn't exist. US District Judge P. Kevin Castel was not amused. The judge [12]fined him $5,000 , but it could have been much worse. Anyone making a similar mistake in the future is unlikely to face such leniency.

Accuracy alone isn't the only problem. Prejudices baked into your Large Language Models (LLMs) can also bite you. The [13]iTutorGroup can tell you all about that. This [14]company lost a $365,000 lawsuit to the US Equal Employment Opportunity Commission (EEOC) because AI-powered recruiting software automatically rejected female applicants aged 55 and older and male applicants aged 60 and older.

[15]Microsoft extends Copilot in Windows for Insiders

[16]Prompt engineering is a task best left to AI models

[17]ChatGPT starts spouting nonsense in 'unexpected responses' shocker

[18]Mitchell Baker logs off for good as CEO of Firefox maker Mozilla

To date, the [19]biggest mistake caused by relying on AI was the American residential real estate company Zillow's real estate pricing blunder.

In November 2021, Zillow wound down its Zillow Offers program. This AI program advised the company on making cash offers for homes that would then be renovated and flipped. However, with a median error rate of 1.9 percent and error rates as high as 6.9 percent, the company lost serious money. How much? Try a $304 million inventory write-down in one quarter alone. Oh, and Zillow laid off 25 percent of its workforce.

I'm not a Luddite, but the simple truth is AI is not yet trustworthy enough for business. It's a useful tool, but it's no replacement for workers, whether they're professionals or help desk staffers. In a few years, it will be a different story. Today, you're just asking for trouble if you rely on AI to improve your bottom line. ®

Get our [20]Tech Resources



[1] https://www.nytimes.com/2024/02/05/technology/why-is-big-tech-still-cutting-jobs.html

[2] https://www.theregister.com/2024/02/15/air_canada_chatbot_fine/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZdjPNxFYancHB1hCMqrKoAAAAI8&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZdjPNxFYancHB1hCMqrKoAAAAI8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZdjPNxFYancHB1hCMqrKoAAAAI8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZdjPNxFYancHB1hCMqrKoAAAAI8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://decisions.civilresolutionbc.ca/crt/crtd/en/item/525448/index.do

[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZdjPNxFYancHB1hCMqrKoAAAAI8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[9] https://aiforensics.org/work/bing-chat-elections

[10] https://www.computerworld.com/article/3713100/air-canada-chatbot-error-underscores-ais-enterprise-liability-danger.html

[11] https://www.theregister.com/2023/06/12/a_federal_judge_is_considering/

[12] https://www.theregister.com/2023/06/22/lawyers_fake_cases/

[13] https://www.teachaway.com/schools/itutorgroup

[14] https://www.akingump.com/en/insights/blogs/ag-data-dive/eeoc-settles-over-recruiting-software-in-possible-first-ever-ai-related-case

[15] https://www.theregister.com/2024/02/22/microsoft_extends_copilot_in_windows/

[16] https://www.theregister.com/2024/02/22/prompt_engineering_ai_models/

[17] https://www.theregister.com/2024/02/21/chatgpt_bug/

[18] https://www.theregister.com/2024/02/09/mozilla_ceo_mitchell_baker_departs/

[19] https://edition.cnn.com/2021/11/09/tech/zillow-ibuying-home-zestimate/index.html

[20] https://whitepapers.theregister.com/



AI LLMs often aren't right. They're not even close.

Anonymous Coward

Worse than that. They actively make shit up.

Re: AI LLMs often aren't right. They're not even close.

amanfromMars 1

Worse than that. They actively make shit up. .... Anonymous Coward

So virtually human and being modelled on the politically incorrect and serially incompetent, AC. ........ See Yes, Prime Minister ..... but it does require balls other than jugglers’ below

Re: AI LLMs often aren't right. They're not even close.

Lurko

"Worse than that. They actively make shit up."

Yes, but for many years I was a customer of a well known UK cable company, and the cheaply offshored customer service agents routinely made shit up. A visit to their customer help forum shows they still do. So if you've got poorly trained humans making it up as they go along, often with poor language skills, then AI can at least improve on the language skills, and make shit up more cheaply. What's not to like for PHBs?

Yes, Prime Minister ..... but it does require balls other than jugglers'

amanfromMars 1

Air Canada discovered the hard way that when your AI chatbot makes a commitment, your company will be on the hook for it

Are we then to reasonably expect political parties be on the hook for promises and commitments to goals they and their leading cheer-leading chatboxes/ministers and constituent members came nowhere near to fulfilling?

Yeah, why not? That seems perfectly fair and not all crooked for a rigged game.

Re: Yes, Prime Minister ..... but it does require balls other than jugglers'

Tron

Political promises should come with stats, a timescale and be legally binding. Fail and you should be excluded from office, fined and imprisoned.

At which point the lying hypocrites will all switch from 'commitments' to 'aspirations'.

They have been doing this a long time and citizens don't get any less gullible. Brexit proved that beyond reasonable doubt. The most you can do is erase the current lot from power at the next election and be screwed over by different politicians for a bit. They don't do any of it for us and they are 'all in it together'. So insulate yourself from them as best you can.

Computers fail when they try to be human. AI is unreliable. The mugs will throw money at it the way they did at the metaverse. We get to suffer from the failures and sometimes to laugh at it. Then politicians step in, tap them for free money in fines and then take control of it all.

GoneFission

>Air Canada replied, in effect, that "The chatbot is a separate legal entity that is responsible for its own actions."

Imagine if this comes up again in a higher court and the ruling sides with the company. This would result in them cashing in on all of the cost savings of replacing humans with barely purpose-functional LLMs and none of the burden of associated risks and liabilities.

Separation

elsergiovolador

Does it mean you can train your LLM to say whatever you need it to say, connect it to e.g. Twatter and if confronted say not me guv, it's LLM, sue that not me?

I'm not a Luddite, but

Rafael #872397

... we need a word for the complete opposite of a Luddite, for people and companies that jump on unproved tech concepts and start planning around it before seeing whether it is going to work or be useful, sometimes just because others are thinking about maybe doing it.

Tech bro, bandwagon jumper, "visionary", early adopter* and technophile* just aren't enough.

Musketeer?

*ChatGPT suggestions.

Re: I'm not a Luddite, but

StewartWhite

How about TechnoTwat?

Re: I'm not a Luddite, but @Rafael #872397

amanfromMars 1

Future Builder Pioneer? ...... Wild Wacky Westernised Cowboy?...... Exotic Erotic Eastern Imperialist? ......... Brave Heart? ....... Bold Leader?

Re: I'm not a Luddite, but

David-M

Since Luddite is likely named after Mr. Lud, we'd maybe be looking for a word like Altmanite or AirCanadite...

"In a few years, it will be a different story."

John H Woods

I'm not sure I have seen any real evidence for that. Is the training going to get better? Are the LLMs going to get so much better the training can be the same? Or is some new form of AI chatbot that isn't really an LLM going to appear? Absent any of that, I'm sceptical.

Re: "In a few years, it will be a different story."

Anonymous Coward

"Is the training going to get better?" - Well google just did a deal to use Reddit as training material, so there goes that hope.

found out the hard way

heyrick

I disagree. The hard way was the torturous logic they tried in order to make their chatbot somehow not responsible for what it said.

Disclaimer

Gordon861

Does this mean that all business chatbot/AIs will now have a disclaimer on the page saying that any answers should be confirmed and are not binding?

And will these disclaimers be legally enforceable if the bot does give out false information.

cornetman

Someone needs to come up with a way of pairing some kind of "fact script" with the part of Chat GPT that can hold a conversation.

Think of a real person sitting with a company handbook on company rules and current promotional offers.

Relying on ChatGPT to not only hold a conversation with context but also generate the factual basis for its responses is never going to be reliable.

Helcat

Can see a potential problem here:

ChatGPT was asked what the VAT paid was on £500 where VAT was 20%. It came back with an answer of £100

okay...

so it was asked if it had the calculations correct: It admitted it did not. It then gave the corrected formula...£500 * 20 / 1+(20/100).

So 10,000/ 1.2.

oops... tax paid: £8,333.33

Now, if you're claiming your tax back...

(for clarity, the actual tax paid was £83.34 as tax is always rounded up :p )

Quality Control, n.:
The process of testing one out of every 1,000 units coming off
a production line to make sure that at least one out of 100 works.