ChatGPT's odds of getting code questions correct are worse than a coin flip
- Reference: 1691446211
- News link: https://www.theregister.co.uk/2023/08/07/chatgpt_stack_overflow_ai/
- Source link:
The Purdue team analyzed ChatGPT’s answers to 517 Stack Overflow questions to assess the correctness, consistency, comprehensiveness, and conciseness of ChatGPT’s answers. The US academics also conducted linguistic and sentiment analysis of the answers, and questioned a dozen volunteer participants on the results generated by the model.
"Our analysis shows that 52 percent of ChatGPT answers are incorrect and 77 percent are verbose," the team's paper concluded. "Nonetheless, ChatGPT answers are still preferred 39.34 percent of the time due to their comprehensiveness and well-articulated language style." Among the set of preferred ChatGPT answers, 77 percent were wrong.
[1]
OpenAI on the ChatGPT website acknowledges its software "may produce inaccurate information about people, places, or facts." We've asked the lab if it has any comment about the Purdue study.
Only when the error in the ChatGPT answer is obvious, users can identify the error
The [2]pre-print paper is titled, "Who Answers It Better? An In-Depth Analysis of ChatGPT and Stack Overflow Answers to Software Engineering Questions." It was written by researchers Samia Kabir, David Udo-Imeh, Bonan Kou, and assistant professor Tianyi Zhang.
"During our study, we observed that only when the error in the ChatGPT answer is obvious, users can identify the error," their paper stated. "However, when the error is not readily verifiable or requires external IDE or documentation, users often fail to identify the incorrectness or underestimate the degree of error in the answer."
[3]
[4]
Even when the answer has a glaring error, the paper stated, two out of the 12 participants still marked the response preferred. The paper attributes this to ChatGPT's pleasant, authoritative style.
"From semi-structured interviews, it is apparent that polite language, articulated and text-book style answers, comprehensiveness, and affiliation in answers make completely wrong answers seem correct," the paper explained.
They do say always be polite...
"The cases where participants preferred incorrect and verbose ChatGPT's answers over Stack Overflow's answers were due to several reasons, as reported by the participants," Samia Kabir, a doctoral student at Purdue and one of the paper's authors, told The Register .
"One of the main reasons was how detailed ChatGPT’s answers are. In many cases, participants did not mind the length if they are getting useful information from lengthy and detailed answers. Also, positive sentiments and politeness of the answers were the other two reasons.
[5]
"Participants ignored the incorrectness when they found ChatGPT’s answer to be insightful. The way ChatGPT confidently conveys insightful information (even when the information is incorrect) gains user trust, which causes them to prefer the incorrect answer."
Kabir said the user study is intended to complement the in-depth manual and large-scale linguistic analysis of ChatGPT answers.
"Nevertheless, it would always be beneficial to have a bigger sample size," she said. "We also welcome other researchers to reproduce our study – our dataset is publicly available to foster future research."
[6]
The authors observe that ChatGPT answers contain more "drives attributes" – language that suggests accomplishment or achievement – but doesn't describe risks as frequently as Stack Overflow posts.
"On many occasions we observed ChatGPT inserting words and phrases such as 'of course I can help you', 'this will certainly fix it', etc," the paper stated.
[7]Cybercrooks are telling ChatGPT to create malicious code
[8]ChatGPT has mastered the confidence trick, and that's a terrible look for AI
[9]OpenAI predicts biz can break a billion in revs by 2024
[10]Stack Overflow bans ChatGPT as 'substantially harmful' for coding issues
Among other findings, the authors found ChatGPT is more likely to make conceptual errors than factual ones. "Many answers are incorrect due to ChatGPT’s incapability to understand the underlying context of the question being asked," the paper found.
The authors' linguistic analysis of ChatGPT answers and Stack Overflow answers suggests the bot's responses are "more formal, express more analytic thinking, showcase more efforts towards achieving goals, and exhibit less negative emotion." And their sentiment analysis concluded ChatGPT answers express "more positive sentiments" than Stack Overflow answers.
Kabir said, "From our findings and observation from this research, we would suggest that Stack Overflow may want to incorporate effective methods to detect toxicity and negative sentiments in comments and answers in order to improve sentiment and politeness.
"We also think that Stack Overflow may want to improve the discoverability of their answers to help in finding useful answers. Additionally, Stack Overflow may want to provide more specific guidelines to help answerers structure their answers, eg: in a step-by-step, detail-oriented manner."
Stack Overflow versus an overflowing stack
There's some positive news here for Stack Overflow, which in 2018 was called out for being the source of incorrect code snippets in about [11]15 percent of 1.3 million Android apps. In the study 60 percent of respondents found the (presumably) human-authored answers to be more correct, concise and useful.
Nonetheless, Stack Overflow's use seems to have declined, though the amount is disputed. It appears traffic has been down six percent every month since January 2022 and was down 13.9 percent in March, according to an [12]April report from SimilarWeb that suggested usage of ChatGPT may be contributing to the decline.
Community members from Stack Exchange, the network of Q&A sites that includes Stack Overflow, have apparently come to [13]a similar conclusion , based on a drop in new question activity, new answers being posted to the site, and in new user registrations.
Stack Overflow, [14]under new ownership since 2021, disagreed with SimilarWeb's assessment in an email to The Register .
A spokesperson said the biz in May 2022 recategorized its analytics cookie from a "Strictly Necessary" to a "Performance" cookie and, in September 2022 shifted to Google Analytics version 4, both of which affect traffic reporting and comparisons over time.
Friendly AI chatbots will be designing bioweapons for criminals 'within years' [15]READ MORE
"Although we have seen a small decline in traffic, in no way is it what the graph is showing," the company spokesperson told us. "This year, overall, we're seeing an average of ~5 percent less traffic compared to 2022.
"That said, Stack Overflow’s traffic, along with traffic to many other sites, has been impacted by the surge of interest in ChatGPT over the last few months. In April of this year, we saw an above average traffic decrease (~14 percent), which we can likely attribute to developers trying GPT-4 after it was released in March. Our traffic also changes based on search algorithms, which have a big influence on how our content is discovered."
Asked about the study's findings, Stack Overflow's spokesperson said no one at the outfit had time to explore the report.
"We know there is no shortage of ways how developers can leverage AI, however from our own findings, there is one core deterrent in its adoption – trust in the accuracy of AI-generated content," the rep said.
"Stack Overflow’s annual Developer Survey of 90,000 coders recently found that 77 percent of developers are favorable of AI tools, but only 42 percent trust the accuracy of those tools. [16]OverflowAI developed with community at the core and with a focus on the accuracy of data and AI-generated content.
"With OverflowAI, we are offering the ability to check, validate, attribute and confirm accuracy and trustworthiness across the Stack Overflow community and its more than 58 million questions and answers." ®
Get our [17]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZNG95tjLf8oFKjHHFkK4jwAAAUw&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://arxiv.org/abs/2308.02312
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZNG95tjLf8oFKjHHFkK4jwAAAUw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZNG95tjLf8oFKjHHFkK4jwAAAUw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZNG95tjLf8oFKjHHFkK4jwAAAUw&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZNG95tjLf8oFKjHHFkK4jwAAAUw&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[7] https://www.theregister.com/2023/01/06/chatgpt_cybercriminals_malicious_code/
[8] https://www.theregister.com/2022/12/12/chatgpt_has_mastered_the_confidence/
[9] https://www.theregister.com/2022/12/19/in_brief_ai/
[10] https://www.theregister.com/2022/12/05/stack_overflow_bans_chatgpt/
[11] https://ieeexplore.ieee.org/document/7958574
[12] https://www.similarweb.com/blog/insights/ai-news/stack-overflow-chatgpt/
[13] https://meta.stackexchange.com/questions/387278/did-stack-exchanges-traffic-go-down-since-chatgpt
[14] https://www.theregister.com/2021/06/02/stack_overflow_prosus/
[15] https://www.theregister.com/2023/07/28/ai_senate_bioweapon/
[16] https://stackoverflow.blog/2023/07/27/announcing-overflowai/
[17] https://whitepapers.theregister.com/
Re: We tried it
Sorry, bit of ambiguity there.
By "try it" did you mean ChatGPT or Stack Overflow?
Of course, your conclusion can apply to both equally well.
Soooo....
ChatGPT spouts bullshit and is usually wrong. Maybe promote it to a management position?
Re: Soooo....
Aiming to low. Take out the racism filters an run it for president.
Re: Soooo....
Yes bullshit but good looking bullshit!
Useful as a guide, not the end-all-be-all
I've typically gone to ChatGPT for some more obscure technical problems that I struggle finding meaningful answers to on Google (nowadays, it seems like Google gives a few pages of mostly unique answers, then it just starts repeating itself). I've asked things like how a particular daemon config needs to be written to accomplish X when the documentation doesn't give you enough details or maybe just a real quick script that I don't feel like writing, like a batch file to loop through a list of subdirectories and create a separate ZIP file of each one (I do more bash, not batch).
In most cases, I'd say the answer ChatGPT provides is at least mostly correct. Usually at a minimum, it leads me in the right direction to solving the problem. I think therein lies the difficulty it will have at being the big job replacement tool for many technical-type roles. If you don't understand the nuances of the concept you are dealing with, you probably can't figure out how to fix little things that are wrong. Your best bet is just asking again and seeing it can fix it. However, that often leads you down a rabbit hole of frustration.
For example, I was asking questions about using some PowerShell commands to do something. It kept giving me commands (which were valid) with parameters that were not. I had to keep correcting it by saying, "Command X doesn't support the Y parameter". It would apologize, then continue to give answers that simply did not work. I had similar results when asking it how to do some routing/firewall config for a switch. It kept giving me directives that didn't exist for my model or firmware version, even when I told it what I had.
That's why I see ChatGPT as simply another helpful tool, but not something that you should expect will give you exactly what you want or need.
Re: Useful as a guide, not the end-all-be-all
"I do more bash, not batch"
Ever considered just installing bash and leveraging what you do know instead of what sounds like a lot of time going round in a loop piddling about with ChatGPT?
Participants ignored the incorrectness
> when they found ChatGPT’s answer to be insightful. The way ChatGPT confidently conveys insightful information (even when the information is incorrect) gains user trust, which causes them to prefer the incorrect answer.
Really stressing the key word here: INCORRECT.
This isn't exactly a new behaviour from SO participants, which has been going downhill for a while.
If anything good is going to come from this work, the very best outcome would be to give SO a damn good wake up call: your traffic is going down because you are as crap as ChatGPT but at least the LLM is polite about it! Not that management can actually do anything to change SO answers (and, more importantly wrt politeness the comments) for the better, is there?
I know it’s fashionable to hate on SO, but I think it’s a great resource. Particularly the debates in the comments under the top answers. Even if my question isn’t directly answered it almost always sends me looking in the right direction.
(Of course you have to understand the answer and test that it actually works, yadda yadda.)
MO
" The way ChatGPT confidently conveys insightful information (even when the information is incorrect) gains user trust, which causes them to prefer the incorrect answer."
Isn't that pretty much the way a con artist operates too?
We tried it
And then tried a subset of an infinite bunch of monkeys hammering at an infinite number of keyboards.
The monkeys made better code