Professor freezes student grades after ChatGPT claimed AI wrote their papers
- Reference: 1684364903
- News link: https://www.theregister.co.uk/2023/05/17/university_chatgpt_grades/
- Source link:
As detailed in a [1]now-viral Reddit thread this week, Jared Mumm, a coordinator at the American university's department of agricultural sciences and natural resources, informed students he used ChatGPT to assess whether their submitted assignments were human-written or produced by computer.
We're told OpenAI's bot labeled at least some of the submitted work as machine crafted, leading to grades being withheld pending an investigation. Students caught up in the row hit back, saying their essays were indeed written by them. As a result of the probe, diplomas were temporarily withheld for those graduating. It's [2]understood about half the class had their diplomas put on hold.
[3]
Specifically, Mumm said he ran his seniors' final three essays through ChatGPT twice, and if the bot said both times for each piece that it wrote the work, he would flunk that paper.
[4]
[5]
"I will be giving everyone in this course an X," he reportedly told his class, and apparently told several students: "I'm not grading AI s***."
The University of Texas A&M-Commerce confirmed the X grade means incomplete, and was a temporary measure while the affair was investigated. Several students have now been cleared of any cheating, we note, while some others opted to submit fresh essays to be graded. At least one pupil so far has admitted using ChatGPT to complete assignments.
[6]
"A&M-Commerce confirms that no students failed the class or were barred from graduating because of this issue," the institution [7]said in a statement. "University officials are investigating the incident and developing policies to address the use or misuse of AI technology in the classroom."
"They are also working to adopt AI detection tools and other resources to manage the intersection of AI technology and higher education. The use of AI in coursework is a rapidly changing issue that confronts all learning institutions. ChatGPT," it continued.
A representative from the university declined to comment further. The Register has asked Mumm for comment.
[8]
One person familiar with the brouhaha at the uni told us: "So far it seems the situation is mostly resolved: the school admitted to students that the grades should not have been withheld in the first place. It was completely out of protocol and an inappropriate use of ChatGPT. They haven’t addressed the foul language in accusations yet."
The kerfuffle highlights whether or not educators should use software to detect AI-produced content within submitted coursework. ChatGPT is not the greatest tool to use to classify machine-generated text; it cannot even accurately determine whether someone used it to write an essay. Basically, it shouldn't be used this way, to detect text output by ChatGPT or some other model.
Other types of software specifically built to detect text generated by AI models are often not reliable, either, as is becoming increasingly apparent.
[9]OpenAI claims GPT-4 will beat 90% of you in an exam
[10]China cracks down on AI-generated news anchors
[11]Top AI execs tell US Senate: Please, please pour that regulation down on us
[12]Will LLMs take your job? Only if you let them
A pre-publication study [13]suggested it will be impossible to discern AI-written text as models improve. Vinu Sankar Sadasivan, a PhD student at the University of Maryland, and the first author of [14]that paper , told us the chances of detecting AI-generated text using the best detectors is no better than flipping a coin.
"Generative AI text models are trained using human text data with the objective of making their output resemble that of humans," Sadasivan said.
"Some of these AI models even memorize human text and output them in some instances without citing the actual text source. As these large language models improve over time to mimic humans, the best possible detector would achieve only an accuracy of nearly 50 percent.
"This is because the probability distribution of text output from human and AI models can nearly be the same for a sufficiently advanced [large language model], making detection hard. Hence, we theoretically show that the task of reliable text detection is impossible in practice."
The paper also showed that such software can be easily tricked into classifying AI text as human, if users make a few quick edits to paraphrase the outputs of a large language model. Sadasivan says universities and schools should not be using these detectors to check for plagiarism since they're unreliable.
"We should not use these detectors to make the final verdict. Borrowing words from my advisor, Prof Soheil Feizi: 'I think we need to learn to live with the fact that we may never be able to reliably say if a text is written by a human or an AI'," he said. ®
Get our [15]Tech Resources
[1] https://www.reddit.com/r/ChatGPT/comments/13isibz/texas_am_commerce_professor_fails_entire_class_of/
[2] https://www.reddit.com/r/ChatGPT/comments/13isibz/comment/jkeqnam/?utm_source=reddit&utm_medium=web2x&context=3
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZGWi5PRE6oh7lcZ-iBmFxwAAAIg&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZGWi5PRE6oh7lcZ-iBmFxwAAAIg&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZGWi5PRE6oh7lcZ-iBmFxwAAAIg&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZGWi5PRE6oh7lcZ-iBmFxwAAAIg&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[7] https://www.tamuc.edu/news/texas-am-university-commerce-addresses-concerns-about-chatgpt-in-ag-classroom/
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZGWi5PRE6oh7lcZ-iBmFxwAAAIg&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://www.theregister.com/2023/03/14/openai_gpt4_ai/
[10] https://www.theregister.com/2023/05/16/china_crackdown_on_ai_generated_news/
[11] https://www.theregister.com/2023/05/17/ai_oversight_hearing/
[12] https://www.theregister.com/2023/05/15/will_llms_take_your_job/
[13] https://www.theregister.com/2023/03/21/detecting_ai_generated_text/
[14] https://arxiv.org/abs/2303.11156
[15] https://whitepapers.theregister.com/
Why not go back to oral exams for the Finals? Sure they take longer but it makes way harder to cheat.
Yes, surely the best way to tell if a student wrote the paper is to make them answer questions about the topic and argument.
And if they use an LLM anyways, and produce a convincing paper, which they then study thoroughly that they might be able to pass the exam, I'd still say they've learned what they were supposed
"I'm not grading AI s***."
The odd things is that "AI s***" is probably a lot easier to read and mark than the corresponding "student s***" that they would have otherwise got.
As a past teacher/lecturer I'd be inclined to shut up and mark it, and be glad that I am receiving something more intelligible.
Self Referential
The fundamental problem this instructor has is that he's using software to tell him which students are using software to write their papers. The operation of the detection software is as opaque as the software that's used to write the papers so the whole thing becomes a festering mess of confusion.
I'm so glad I'm not a student these days. Learning isn't learning so much as a regimented exercise in swallowing and regurgitating facts as fast as possible, where adherence to schedules and work formatting carries as much weight in grading as actual content. It almost demands cheating of the student just as a way of getting through the course (....especially as courses and course work aren't coordinated, each subject and instructor has their own fiefdom and if assignments and deadlines clash or overlap then that's not their problem).