News: 1687343165

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

We just don't get enough time, contractor tasked with fact-checking Google Bard tells us

(2023/06/21)


Workers tasked with improving the output of Google's Bard chatbot say they've been told to focus on working fast at the expense of quality. Bard sometimes generates inaccurate information simply because there isn't enough time for these fact checkers to verify the software's output, one of those workers told The Register .

Large language models like Bard learn what words to generate next from a given prompt by ingesting mountains of text from various sources – like the web, books, and papers. But this information is complex, and sentence-predicting AI chatbots cannot tell fact from fiction. They just try their best to emulate us humans from our own work.

Hoping to make large language models like Bard more accurate, [1]crowdsource workers are hired to assess the accuracy of the bot's responses; that feedback is then passed back into the pipeline so that future answers from the bot are of a higher quality. Google and others put humans in the loop to bump up the apparent abilities of the trained models.

[2]

Ed Stackhouse – a long-time contractor hired by data services provider Appen, working on behalf of Google to improve [3]Bard – claims workers aren't given adequate time to analyze the accuracy of Bard's outputs.

[4]

[5]

They have to read an input prompt and Bard's responses, search the internet for the relevant information, and write up notes commenting on the quality of the text. "You can be given just two minutes for something that would actually take 15 minutes to verify," he told us. That doesn't bode well for improving the chatbot.

An example could be looking at a blurb generated by Bard describing a particular company. "You would have to check that a business was started at such and such date, that it manufactured such and such project, that the CEO is such and such," he said. There are multiple facts to check, and often not enough time to verify them thoroughly.

[6]AI is going to eat itself: Experiment shows people training bots are using bots

[7]Euro Parliament green lights its AI safety, privacy law

[8]Out with the old, in with the new – Accenture declares AI is 'mature and delivers value'

[9]Google warns its own employees: Do not use code generated by Bard

Stackhouse is part of a group of contract workers raising the alarm over how their working conditions can make Bard inaccurate and potentially harmful. "Bard could be asked 'can you tell me the side effects of a certain prescription?' and I would have to go through and verify each one [Bard listed]. What if I get one wrong?" he asked. "Every prompt and answer we see in our environment is one that could go out to customers – to end users."

It's not just medical issues – other topics can be risky, too. Bard spewing incorrect information on politicians, for example, could sway people's opinions on elections and undermine democracy.

[10]

Stackhouse's concerns aren't far-fetched. OpenAI's ChatGPT notably [11]wrongly accused a mayor in Australia of being found guilty in a financial bribery case dating back to the early 2000s.

If workers like Stackhouse are unable to catch these errors and correct them, AI will continue to spread falsehoods. Chatbots like Bard could fuel a shift in the narrative threads of history or human culture – important truths could be erased over time, he argued. "The biggest danger is that they can mislead and sound so good that people will be convinced that AI is correct."

Appen contractors are penalized if they don't complete tasks within an allotted time, and attempts to persuade managers to give them more time to assess Bard's responses haven't been successful. Stackhouse is one of a group of six workers who said they were fired for speaking out, and have filed an unfair labor practice complaint with America's labor watchdog – the National Labor Relations Board – the Washington Post [12]first reported .

[13]

The workers accuse Appen and Google of unlawful termination and interfering with their efforts to unionize. They were reportedly told they were axed on account of business conditions. Stackhouse said he found this hard to believe, since Appen had previously sent emails to workers stating that there was "a significant spike in jobs available" for Project Yukon – a program aimed at evaluating text for search engines, which includes Bard.

Appen was offering contractors additional $81 on top of base pay for working 27 hours per week. Workers are reportedly normally limited to working 26 hours per week for up to $14.50 per hour. The company has active job postings looking for Search Engine Evaluations specifically to work on Project Yukon. Appen did not respond to The Register 's questions.

The group also tried to reach out to Google, and contacted senior vice president Prabahkar Raghavan – who leads the tech behemoth's search business – and were ignored.

Courtenay Mencini, a spokesperson from Google, did not address the workers' concerns that Bard could be harmful. "As we've shared, Appen is responsible for the working conditions of their employees – including pay, benefits, employment changes, and the tasks they're assigned. We, of course, respect the right of these workers to join a union or participate in organizing activity, but it's a matter between the workers and their employer, Appen," she told us in a statement.

Stackhouse, however, said: "It's their product. If they want a flawed product, that's on them." ®

Get our [14]Tech Resources



[1] https://www.theregister.com/2023/06/16/crowd_workers_bots_ai_training/

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZJMeoAr1ElK92yHRnUYOSAAAAJQ&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://bard.google.com/

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJMeoAr1ElK92yHRnUYOSAAAAJQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJMeoAr1ElK92yHRnUYOSAAAAJQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://www.theregister.com/2023/06/16/crowd_workers_bots_ai_training/

[7] https://www.theregister.com/2023/06/15/european_parliament_ai_act/

[8] https://www.theregister.com/2023/06/14/accenture_doubles_ai_workforce/

[9] https://www.theregister.com/2023/06/19/even_google_warns_its_own/

[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZJMeoAr1ElK92yHRnUYOSAAAAJQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[11] https://www.reuters.com/technology/australian-mayor-readies-worlds-first-defamation-lawsuit-over-chatgpt-content-2023-04-05/

[12] https://www.washingtonpost.com/technology/2023/06/14/google-ai-bard-raters-chatbot-accuracy/

[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZJMeoAr1ElK92yHRnUYOSAAAAJQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[14] https://whitepapers.theregister.com/



If they want a flawed product, that's on them

xanadu42

Terminator...

Re: If they want a flawed product, that's on them

ThatOne

They don't want a flawed product, they want a cheap one.

Reminder: Profit = Money earned - Money spent...

Subverted use case??

jmch

Google's origin and raison d'etre as a search engine seems to have been turned upside down. It used to be that if I type "what are the side effects of drug X?", or "who is Mr Y?", I would get back multiple links to source material that I can easily find. If necessary I can quickly cross-check multiple sources. Inserting an LLM in between is of zero utility if I anyway have to crosscheck the output with a different source. Even worse if the LLM is unable to direct me to the source material (which it can't ever do because of how it works).

Bard, Chat-GPT etc etc are also pre-trained, meaning they are immediately out of date (therefore useless on current or recent events), and require gigantic amounts of processing power to deliver a search result that can be generated much more easily by a search engine's indexed search. While LLMs could be useful for generating (bland, grammatically correct but possibly inaccurate) sections of text, they are pretty useless as a search engine replacement.

Re: Subverted use case??

bo111

I find value in generated summaries before digging into specific URLs

Re: Subverted use case??

FrogsAndChips

Bard, Chat-GPT etc etc are also pre-trained, meaning they are immediately out of date

That's incorrect in the case of Bard, which can access the internet to search for information and isn't limited to its training data.

TheMaskedMan

"While LLMs could be useful for generating (bland, grammatically correct but possibly inaccurate) sections of text, they are pretty useless as a search engine replacement."

This! chatGPT is great for turning out outlines that you might then edit - much easier than writing it all from scratch, sometimes it comes up with angles I haven't thought of, even. In its own way it is very useful.

But it isn't a search engine, and I can't see the point in trying to use it as one.

It was great for a while

Will Godfrey

But the Internet is being destroyed by the likes of Google, Facebook, et al.

Just on search alone, I've noticed a steady decline in relevance of results to my queries.

Re: It was great for a while

ThatOne

Unfortunately "relevance" has long stopped being relevant (yes, ironic, I know). "Profitability" is the one and only goal now.

Re: It was great for a while

Anonymous Coward

One of the quality issues is due to PageRank itself. Early days people used to link to quality pages. Then SEO happened and shit hit the fan with scammers occupying the top results. Also search engines consider most links as likes, which is incorrect. For example some news sites link to disinformation sites just for reference, but this pushes them up in the search results. There should be a "dislike" link type (a href). But this would be too complicated.

Plausible nonsense

Dom 3

Or as someone I know put it: it's not "give me an answer to this question" it's "give me something that looks like an answer to this question".

Fun can be had asking for recipes involving random ingredients.

Maybe they should just

Robert Carnegie

Employ and train reliable professional human researchers to BE "Google Bard", without the so-called "AI".

When a place gets crowded enough to require ID's, social collapse is not
far away. It is time to go elsewhere. The best thing about space travel
is that it made it possible to go elsewhere.
-- R. A. Heinlein, "Time Enough For Love"