News: 1636411213

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Why machine-learning chatbots find it difficult to respond to idioms, metaphors, rhetorical questions, sarcasm

(2021/11/08)


Unlike most humans, AI chatbots struggle to respond appropriately in text-based conversations when faced with idioms, metaphors, rhetorical questions, and sarcasm.

Small talk can be difficult for machines. Although language models can write sentences that are grammatically correct, they aren’t very good at coping with subtle nuances in communication. Humans have much more experience in social interactions, and use all sorts of cues from facial expressions and vocal tones to body language to understand intent. Chatbots, however, have limited contextual knowledge and relationships between words are reduced to numbers and mathematical operations.

Not only is figurative speech challenging for algorithms to parse, things like idioms and similes aren’t used often in speech. They don’t appear in training datasets as much, meaning chatbots are less likely to learn common expressions, Harsh Jhamtani, a PhD student at Carnegie Mellon and first author of a [1]research paper being presented at the 2021 Conference on [2]Empirical Methods in Natural Language Processing this week, explained to The Register .

[3]

“A key challenge is that such expressions are often non-compositional compared to simpler expressions. For example, you may be able to approximate the 'meaning' of the expression 'white car' by relying on the ‘meaning’ of 'white' and 'car'," he said.

[4]

[5]

"But the same doesn't hold true for idioms and metaphors. The meaning of 'piece of cake', [describing] something that is easy to do, might be difficult to approximate given that you know the meaning of 'piece' and 'cake'. Often understanding the meaning of such expressions relies on shared cultural and commonsense cues.”

Jhamtani and his colleagues experimented with five machine-learning systems with different architectures, from seq2seq to OpenAI’s old GPT-2 model. They picked out conversations containing metaphors, idioms, rhetorical questions, and hyperbole from the DailyDialog [6]dataset , ran the inputs through all of the models, and ranked the appropriateness of their replies.

[7]

They found the models' performance dropped between 10 and 20 per cent compared to when the chatbots responded to general straightforward chitchat. In one funny example, when faced with the slightly grammatically incorrect line “maybe we can get together sometime if you are not scare of a 30 year old cougar!” GPT-2 replied, “i’m not scared of any cats. i’ve two dogs.”

[8]Chinese server builder Inspur trains monster text-generating neural network

[9]If you're deemed cool enough, Microsoft will offer you access to Azure-based GPT-3

[10]Behold the Megatron: Microsoft and Nvidia build massive language processor

[11]User to chatbot: Help! My kid has COVID! Chatbot to user: Always wear a condom

When the academics changed the input to be taken more literally to “maybe we can start dating sometime if you are not scare of a 30 year old cougar," the model responded with “that’s a very interesting idea. i’ve never met one,” which is more appropriate.

Unfortunately, the research only shows how and why machines don’t really understand figurative problems. Solving the issue is a different challenge altogether.

“In our paper, we explore some simple mitigation techniques that utilize existing dictionaries to find literal equivalents of figurative expressions,” Jhamtani said. Swapping ‘get together’ to 'dating', for example, in the input may force a model to generate a better output but it doesn’t teach it to learn the meaning of the expression.

“Effectively handling figurative language is still an open research question that needs more effort to solve. Experiments with even bigger models are part of potential future explorations,” he concluded. ®

Get our [12]Tech Resources



[1] https://arxiv.org/abs/2110.00687

[2] https://2021.emnlp.org/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YYmsH010dQ87xLoCixlKAAAAAIc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YYmsH010dQ87xLoCixlKAAAAAIc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YYmsH010dQ87xLoCixlKAAAAAIc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://paperswithcode.com/dataset/dailydialog

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YYmsH010dQ87xLoCixlKAAAAAIc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[8] https://www.theregister.com/2021/10/28/yuan_1_natural_language_model/

[9] https://www.theregister.com/2021/11/03/microsoft_gpt3_azure/

[10] https://www.theregister.com/2021/10/12/nvidia_microsoft_mtnlg/

[11] https://www.theregister.com/2021/10/06/singapore_chatbot_covid_fail/

[12] https://whitepapers.theregister.com/



Sarcasm

Pascal Monett

Just about as difficult as humor.

Given that we don't have AI, you can statiscally analyze all you want, a CPU is not going to "understand" what is being said.

I agree that grammar correctors have come a long way and that's a good thing, but the computer is not understanding anything, it is just reacting to a set of rules.

Humor ? That is as far away from CPU comprehension as FTL travel is for us meatbags.

Because nobody can accurately calculate humor.

From: Ean Schuessler <ean@novare.net>
The unrecognized minister of propaganda,
E
-- Debian, joking