ChatGPT can't pass these medical exams – yet
- Reference: 1684913346
- News link: https://www.theregister.co.uk/2023/05/24/chatgpt_gastroenterology_exams/
- Source link:
A study led by physicians at the Feinstein Institutes for Medical Research tested both variants of ChatGPT – powered by OpenAI's older GPT-3.5 model and the latest GPT-4 system. The academic team copy and pasted the multiple choice questions taken from the 2021 and 2022 American College of Gastroenterology (ACG) Self-Assessment Tests into the bot, and analyzed the software's responses.
Interestingly, the less advanced version based on GPT-3.5 answered 65.1 percent of the 455 questions correctly while the more powerful GPT-4 scored 62.4 percent. How that happened is hard to explain as OpenAI is secretive about the way it trains its models. Its spokespeople told us, at least, both models were trained on data dated as recent as September 2021.
[1]
In any case, neither result was good enough to reach the 70 percent threshold to pass the exams.
[2]
[3]
Arvind Trindade, an associate professor at The Feinstein Institutes for Medical Research and senior author of the study [4]published in the American Journal of Gastroenterology , told The Register .
"Although the score is not far away from passing or obtaining a 70 percent, I would argue that for medical advice or medical education, the score should be over 95."
[5]
"I don't think a patient would be comfortable with a doctor that only knows 70 percent of his or her medical field. If we demand this high standard for our doctors, we should demand this high standard from medical chatbots," he added.
[6]Professor freezes student grades after ChatGPT claimed AI wrote their papers
[7]Stanford sends 'hallucinating' Alpaca AI model out to pasture over safety, cost
[8]ChatGPT talks its way through Wharton MBA, medical exams
[9]ChatGPT has mastered the confidence trick, and that's a terrible look for AI
The American College of Gastroenterology trains physicians, and its tests are used as practice for official exams. To become a board-certified gastroenterologist, doctors need to pass the American Board of Internal Medicine Gastroenterology examination. That takes knowledge and study – not just gut feeling.
ChatGPT generates responses by predicting the next word in a given sentence. AI learns common patterns in its training data to figure out what word should go next, and is partially effective at recalling information. Although the technology has improved rapidly, it's not perfect and is often prone to hallucinating false facts – especially if it's being quizzed on niche subjects that may not be present in its training data.
"ChatGPT's basic function is to predict the next word in a string of text to produce an expected response based on available information, regardless of whether such a response is factually correct or not. It does not have any intrinsic understanding of a topic or issue," the paper explains.
Trindade told us that it's possible that the gastroenterology-related information on webpages used to train the software is not accurate, and that the best resources like medical journals or databases should be used.
[10]
These resources, however, are not readily available and can be locked up behind paywalls. In that case, ChatGPT may not have been sufficiently exposed to the expert knowledge.
"The results are only applicable to ChatGPT – other chatbots need to be validated. The crux of the issue is where these chatbots are obtaining the information. In its current form ChatGPT should not be used for medical advice or medical education," Trindade concluded. ®
Get our [11]Tech Resources
[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZG3gShwrTZ7UTjqK6Ts9zgAAANg&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZG3gShwrTZ7UTjqK6Ts9zgAAANg&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZG3gShwrTZ7UTjqK6Ts9zgAAANg&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[4] https://journals.lww.com/ajg/Abstract/9900/ChatGPT_Fails_the_Multiple_Choice_American_College.751.aspx
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZG3gShwrTZ7UTjqK6Ts9zgAAANg&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2023/05/17/university_chatgpt_grades/
[7] https://www.theregister.com/2023/03/21/stanford_ai_alpaca_taken_offline/
[8] https://www.theregister.com/2023/01/24/chatgpt_exam_study/
[9] https://www.theregister.com/2022/12/12/chatgpt_has_mastered_the_confidence/
[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZG3gShwrTZ7UTjqK6Ts9zgAAANg&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[11] https://whitepapers.theregister.com/
Re: multiple guess
That would probably be 25%. Then, using some great chatGPT math, we can conclude that the real success rate of the chat bot is between 87.4% and 90.1%.
Therefore, the doctors are wrong! It is great medical advice to ask chatGPT. You'll get, again using great chatGPT math, between 9.9% and 12.6% correct answers from your robotic overlord. Nothing can possibly go wrong, chatGPT told me so.
If I were allowed unlimited time and free access to the Internet, I suspect I could pass this exam (or almost any written test), even though I know nothing about gastroenterology.
Passing grade
> "I don't think a patient would be comfortable with a doctor that only knows 70 percent of his or her medical field. If we demand this high standard for our doctors, we should demand this high standard from medical chatbots," he added.
Then why is 70% the passing grade?
As the joke goes: "What do you call the person who graduates medical school with the worst grades in their year?" "Doctor."
Re: Passing grade - mirror, indicate before passing
Yes, this does seem to be an example of the double standards (bias against?) employed in many fields.
Many articles that talk about autonomous vehicles give the impression that nothing less than perfection is acceptable. Yet the standard of AI driving is already (measured as accidents per 100,000 mile / km) better than the average for police drivers.
This seem to be a confidence issue rather than one about actual abilities. Maybe what's needed is some blind comparisons: doctors' diagnoses vs. machine,and see where the truth lies.
OpenAI is secretive about the way it trains its models
And herein lies the real problem: do you want your medical advice coming from a secret source whose origin cannot be revealed?
The word for something that gives the appearance of improbable success in circumstances where you're not allowed to look behind the curtain is "magic".
That's the kind of medicine it's taken generations to get away from.
Re: OpenAI is secretive about the way it trains its models
> do you want your medical advice coming from a secret source whose origin cannot be revealed?
Isn't that what is commonly called "experience"?
Some years ago a large computer consultancy set about automating a lot of processes. I has an external person following me around for a week, taking note of how I worked and in particular how I debugged operational issues. Each time I solved a problem, this person would ask me the process I had used to reach a solution. Most times the reply would come down "I have 25 years of experience, it looked like something I'd seen before".
Which wasn't very helpful - though admittedly I felt no obligation to be helpful ;) -, but was the truth.
I have a major problem with this ChatGPT rush
.. which is also why I *seriously* question Microsoft sticking it in anything it can lay its hands on in the hope that that will somehow liven up their sales.
To me, this feels like we may be heading for a computer version of [1]thalidomide .
There, a 'wonderdrug' was found in the 1950s to be useful for all sorts of purposes and its use kept spreading under various different names (which made it harder to tie the problems together). It took 5 years for the appaling truth to be traced back to the drug. Nowadays, safe use cases have been found (in some cases even spectacular), but not before there was a full generation of fairly dramatic human disasters and decades of dealing with them.
I see four problems here:
1 - sticking it everywhere without having fully evaluated consequences (some of that is simply because we don't know them yet);
2 - the utter lack of accountability of software companies for the aforementioned consequences;
3 - it's now a buzzword, which means the critical thinking dial is turned *way* down;
4 - it's Microsoft. For those who know its history, I don't need to say more (and if you don't, look it up).
Even if your risk management has arrived at the conclusion that you should avoid Microsoft products doesn't mean you may not be exposed to the consequences of others using it, in volume.
So no, I am not filled with enthusiasm for the new toy being put into places where it can cause massive harm without someone being able to say 'no' to the whole idea.
[1] https://www.sciencemuseum.org.uk/objects-and-stories/medicine/thalidomide
I think there is a problem with the test
Several years ago, a machine learning based facial gaydar was tested on photographs with the faces blurred out. Hiding the faces did not reduce the accuracy. The software must have been basing its decision on something other than the face, like clothing, background, lighting or the composition of the photo.
As paywalled gastroenterology text books and medical journals are not part of the training data there could be unexpected reasons why ChatGPT is scoring better than a random number generator. My guess is there is some kind of pattern to the the phrasing of many of the wrong answers in the multiple choice tests that ChatGPT can exploit to improve its score. A more useful test would be to feed it transcripts of patient consultations and compare ChaptGPT's proposed treatments with the recommendations of gastroenterologists with a high success rate.
multiple guess
60ish percent -- hmm. How high is the "tick boxes at random" score for this test?