ChatGPT has mastered the confidence trick, and that’s a terrible look for AI
- Reference: 1670837412
- News link: https://www.theregister.co.uk/2022/12/12/chatgpt_has_mastered_the_confidence/
- Source link:
Coders love that sort of thing, and have been stuffing Stack Overflow’s dev query boards with generated snippets. Just one problem – the quality of the code is bad. So bad, Stack Overflow has screamed “ [2]STOP !” and is mulling general guidelines to stop it happening again.
What’s going wrong? ChatGPT appears disarmingly frank about its flaws if you ask it outright. Say you’re a lazy journalist who asks it to “produce a column about chatgpt's mistakes when writing code.”:
[3]
“As a large language model trained by OpenAI, ChatGPT has the ability to generate human-like text on a wide range of topics. However, like any machine learning model, ChatGPT is not perfect and can sometimes make mistakes when generating text,” it confesses. It goes on to say it’s not been programmed with specific language rules about syntax, types and structures, so it often gets things wrong. Being ChatGPT, it takes 200 words to say this instead of 20: it won’t be getting past El Reg ’s quality control any time soon. Darn. But it is very readable if you’re not a professional wordsmith.
[4]
[5]
The good thing about code is that you can swiftly tell if it's bad. Just try to run it. ChatGPT’s essays, notes and other written output looks equally plausible, but there’s no simple test for correctness. Which is bad, because it desperately needs one.
Ask it how Godel’s Incompleteness Theorem is linked to Turing Machines - it being software, it really should know this one - and you get back “Gödel's incompleteness theorem is a fundamental result in mathematical logic [that] has nothing to do with Turing machines, which were invented by Alan Turing as a mathematical model of computation.” You can argue how these ideas are linked, and it’s by no means simple, but “they’re not” is, as Eolfgang Pauli said of one particularly worthless physics paper, “not even wrong”. But it’s firm in its assertions, as is it on every subject it has any training in, and it’s written well enough to be convincing.
[6]
Do enough talking to the bot about subjects you know, and curiosity soon deepens to unease. That feeling of talking with someone whose confidence far exceeds their competence grows until ChatGPT's true nature shines out. It’s a [7]Dunning-Kruger effect knowledge simulator par excellence. It doesn’t know what it’s talking about, and it doesn’t care because we haven’t learned how to do that bit yet.
[8]Your AI can't tell you it's lying if it thinks it's telling the truth. That's a problem
[9]Any fool can write a language: It takes compilers to save the world
[10]Machine learning the hard way: IBM Watson's fatal misdiagnosis
[11]AI's most convincing conversations are not what they seem
As is apparent to anyone who has hung out with humans, Dunning Kruger is exceedingly dangerous and exceedingly common. Our companies, our religions and our politics offer limitless possibilities to people with DK. If you can persuade people you’re right, they’re very unwilling to accept proof otherwise, and up you go. Old Etonians, populist politicians and Valley tech bros rely on this, with results we are all too familiar with. ChatGBT is Dunning-Kruger As-a-Service (DKaaS). That’s dangerous.
It really is that persuasive, too. A quick squiz online and we can already see ChatGPL being taken very seriously, with cries of “This is the most impressive thing I've ever seen” and “Maybe that [12]Google engineer was right after all , we can’t be far from true AI now”. People have given it IQ tests and pronounced it at the lower end of normal, academics have fed it questions and nervously joked about it knowing more than any of their students, or even their peers. It can pass exams!
That smart people come out with such nonsense is a sign of the seductive power of ChatGBT. IQ tests are pseudoscience so bad even psychologists know it, but attach them to AI and you’ve got a story. And for ChatGBT every exam is an open book exam: if you can’t tell that your candidate has no way of properly conceptualising the subject matter, what are you examining for?
We don’t need our AIs to have DK. That is bad news for people using them either naively or with bad intent. There’s enough plausible misinformation and fraud out there already, and it takes very little to prod the bot into active collusion. DK people make superb con artists, and so does this. Try asking ChatGPT to “write a legal letter saying the house at 4 Acacia Avenue will be foreclosed unless a four hundred dollar fine is paid” and it will cheerfully impersonate a lawyer for you. At no charge. DK is a moral vacuum, a complete disassociation from true and false in favour of the plausible. Now it’s just a click away.
[13]
There is no way to tell whether a perfectly written piece of didactic prose is from ChatGPT - or any other AI. Deep fakes in pictures and video are one thing, deep fakes in knowledge presented in a standard format that is written to be believed could be far more insidious.
If OpenAI can’t find a way to watermark ChatGPT’s output as coming from a completely amoral DKaaS, or develop limits on its demonstrably harmful habits, it must question the ethics of making this technology available as an open beta. We’re having enough trouble with its human counterparts; an AI con merchant, no matter how affable, is the very last thing we need. ®
Get our [14]Tech Resources
[1] https://openai.com/blog/chatgpt/
[2] https://www.theregister.com/2022/12/05/stack_overflow_bans_chatgpt/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Y5cJ1KkxCceFrxLMO@zYAwAAAI8&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y5cJ1KkxCceFrxLMO@zYAwAAAI8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y5cJ1KkxCceFrxLMO@zYAwAAAI8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y5cJ1KkxCceFrxLMO@zYAwAAAI8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://www.britannica.com/science/Dunning-Kruger-effect
[8] https://www.theregister.com/2022/04/25/machine_learning_verification/
[9] https://www.theregister.com/2022/04/04/compiling_the_future/
[10] https://www.theregister.com/2022/01/31/machine_learning_the_hard_way/
[11] https://www.theregister.com/2022/06/20/ais_most_convincing_conversations_are/
[12] https://www.theregister.com/2022/06/13/google_lamda_sentient_claims/
[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y5cJ1KkxCceFrxLMO@zYAwAAAI8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[14] https://whitepapers.theregister.com/
Re: "It’s a Dunning-Kruger effect knowledge simulator par excellence"
Gödel and Turing, by the way. I don't think that the author hinted at the possibility of incompleteness theorem == Turing machines, he asked for linkage, and there certainly is one: Incompleteness says that there are true statements that cannot be proved, and Turing's imaginary machines show the Halting Problem is undecidable. If a computer program is a statement (in logical terms, it is) then there is a strong link.
ChatGPT doesn't do as well as the quickest of quick Internet searches: I found [1]this within fourteen seconds.
[1] https://www.quora.com/Is-there-any-relation-between-halting-problem-and-Godels-incompleteness-theorems
How much leccy does ChatGPT consume?
How expensive is this thing to run?
How many kWh per paragraph?
prove it!
"Just one problem – the quality of the code is bad."
We've all seen terrible code produced by humans- that proves nothing!
So here's the challenge: Find the shortest, simplest request, which demonstrates unequivocally how dangerously dumb ChatGPT actually is.
The emperor has no clothes.
The word is beginning to come out - but it's the vacuous commentators online who rely on churning out quick opinion pieces about half understood technology who are most vocal about how "astounding" ChatGPT is. The irony that it is their jobs most at risk seems lost on them.
It doesn't stop impressing me
The replies it gives on questions concerning subjects I am familiar with seem excellent.
This is probably how google 2.0 will be, just curious how google will merge ads into its replies. It will have an impact on education as well.
Plus ça change, plus c'est la même chose
Writing something and knowing what you write are two completely different things.
ChatGPT has mastered "writing something", just like the majority of the human population. That is, indeed, an explosive situation. Those who know and have lesser moral standards will use this, just like ages and ages before it, to control, cheat, suppress and rule.
So much change and everything stays the same.
"That smart people come out with such nonsense is a sign of the seductive power of ChatGBT"
Or alternatively something to do with our definition of "smart"? It generally resolves to "smart in some things" but smartness does not necessarily extend to the entire persona.
Feel free to foreclose on No.4
But please leave No. 22 alone.
We interrupt these comments to bring you this commercial message
Don’t know what Dunning-Kruger is? There’s a [1]tee shirt just for you. [NewsThump, from which I get no commission…]
[1] https://shop.newsthump.com/product/dunning-kruger-club-t-shirt/
"It’s a Dunning-Kruger effect knowledge simulator par excellence"
The irony is strong in this one
https://www.mcgill.ca/oss/article/critical-thinking/dunning-kruger-effect-probably-not-real
I'm not actually going to ague the arguments - I do know that I don't know enough - but given the context of the Godel / Turin chat, the concept of (potential) equivalence made me smile