News: 1710853213

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

AI researchers have started reviewing their peers using AI assistance

(2024/03/19)


Academics focused on artificial intelligence have taken to using generative AI to help them review the machine learning work of peers.

A group of researchers from Stanford University, NEC Labs America, and UC Santa Barbara recently analyzed the peer reviews of papers submitted to leading AI conferences, including ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023.

The authors – Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A McFarland, and James Y Zou – reported their findings in [1]a paper titled "Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews."

[2]

They undertook the study based on the public interest in, and discussion of, large language models that dominated technical discourse last year.

The authors found a small but consistent increase in apparent LLM usage for reviews submitted three days or less before the deadline

The difficulty of distinguishing between human- and machine-written text and the reported rise in [3]AI news websites led the authors to conclude that there's an urgent need to develop ways to evaluate real-world data sets that contain some indeterminate amount of AI-authored content.

Sometimes AI authorship stands out – as in a [4]paper from Radiology Case Reports entitled "Successful management of an Iatrogenic portal vein and hepatic artery injury in a 4-month-old female patient: A case report and literature review."

[5]

[6]

This jumbled passage is a bit of a giveaway: "In summary, the management of bilateral iatrogenic I'm very sorry, but I don't have access to real-time information or patient-specific data, as I am an AI language model."

But the distinction isn't always obvious, and past attempts to develop an automated way to sort human-written text from robo-prose have not gone well. OpenAI, for example [7]introduced an AI Text Classifier for that purpose in January 2023, only to shutter it six months later " [8]due to its low rate of accuracy ."

[9]

Nonetheless, Liang et al contend that focusing on the use of adjectives in a text – rather than trying to assess entire documents, paragraphs, or sentences – leads to more reliable results.

The authors took two sets of data, or corpora – one written by humans and the other one written by machines. And they used these two bodies of text to evaluate the evaluations – the peer reviews of conference AI papers – for the frequency of specific adjectives.

[10]Grok-1 chatbot code released – open source or open Pandora's box?

[11]Top LLMs struggle to make accurate legal arguments

[12]In the rush to build AI apps, please, please don't leave security behind

[13]Google launches Gemini AI systems, claims it's beating OpenAI and others - mostly

"[A]ll of our calculations depend only on the adjectives contained in each document," they explained. "We found this vocabulary choice to exhibit greater stability than using other parts of speech such as adverbs, verbs, nouns, or all possible tokens."

It turns out LLMs tend to employ adjectives like "commendable," "innovative," and "comprehensive" more frequently than human authors. And such statistical differences in word usage have allowed the boffins to identify reviews of papers where LLM assistance is deemed likely.

[14]

Word cloud of top 100 adjectives in LLM feedback, with font size indicating frequency (click to enlarge)

"Our results suggest that between 6.5 percent and 16.9 percent of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates," the authors argued, noting that reviews of work in the scientific journal Nature do not exhibit signs of mechanized assistance.

Several factors appear to be correlated with greater LLM usage. One is an approaching deadline: The authors found a small but consistent increase in apparent LLM usage for reviews submitted three days or less before the deadline.

The researchers emphasized that their intention was not to pass judgment on the use of AI writing assistance, nor to claim that any of the papers they evaluated were written completely by an AI model. But they argued the scientific community needs to be more transparent about the use of LLMs.

[15]

And they contended that such practices potentially deprive those whose work is being reviewed of diverse feedback from experts. What's more, AI feedback risks a homogenization effect that skews toward AI model biases and away from meaningful insight. ®

Get our [16]Tech Resources



[1] https://arxiv.org/abs/2403.07183

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZfnEp2W47fMNOW@9pnQglAAAAAs&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://www.theregister.com/2020/06/09/microsoft_ai_thirlwall/

[4] https://www.sciencedirect.com/science/article/pii/S1930043324001298

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZfnEp2W47fMNOW@9pnQglAAAAAs&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZfnEp2W47fMNOW@9pnQglAAAAAs&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[7] https://www.theregister.com/2023/01/31/openai_tool_chatgpt_detection/

[8] https://openai.com/blog/new-ai-classifier-for-indicating-ai-written-text

[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZfnEp2W47fMNOW@9pnQglAAAAAs&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[10] https://www.theregister.com/2024/03/18/grok_chatbot_code_released/

[11] https://www.theregister.com/2024/01/10/top_large_language_models_struggle/

[12] https://www.theregister.com/2024/03/17/ai_supply_chain/

[13] https://www.theregister.com/2023/12/06/google_gemini_ai/

[14] https://regmedia.co.uk/2024/03/18/llm_word_cloud.jpg

[15] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZfnEp2W47fMNOW@9pnQglAAAAAs&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[16] https://whitepapers.theregister.com/



How embarassing for Elsevier

Anonymous Coward

Proof that nobody actually read the final document before submitting it, and then nobody read the document before publishing it? And eight doctors putting their names to this, I'll bet they feel proud.

Sheesh.

LLM are evidently trained on sales, marketing and PR drivel

Lurko

If one is to judge by that word cloud. Which will come as no surprise to anyone.

Blackjack

Is stuff like this why that burned out street lightt in front of my house has not been replaced in years, congrats you are good enough for government work!

It's AI all the way down

yetanotheraoc

AI says: Commendable, innovative, and comprehensive ... and would have been more so had there been even more AI input.

Putting on my QA hat, next testing steps:

1. Create some obviously bogus (to human eyes) paper. Typical AI output would be perfect for this, especially if AI creates the footnotes.

2. Review paper with AI "assistance" (e.g. how it's really done, AI review with human polishing). N.B. Technically this qualifies as a peer review. Expect: ( commendable + innovative + comprehensive )

3. Review review with AI "assistance". Expect: ( commendable + innovative + comprehensive )^2

Ultimately ...

Mike 137

I'm waiting for when 'AI' generated content gets reviewed by 'AI', generating reports that get used as training data for 'AI', leading to the the entire closed loop system disappearing up its own portal without any human having to write or read anything ever again.

Who raised you people?

Omnipresent

Teaching the AI about it's self seems like a good way to destroy the world. Why is it that people with book smarts have zero common sense? Do you not ever "say it out loud"? Do you just not care about anything anymore?

Can we sue the ever loving crap out of AI for grabbing our identities, information, and work yet?

To program is to be.