News: 1649151665

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Google talks up its 540-billion-parameter text-generating AI system

(2022/04/05)


Though AI models will continue to get increasingly more powerful the larger they become, performance improvements from scale have not yet plateaued, according to researchers at Google.

But while neural networks have grown, are they really any smarter? Companies are making larger and larger machine-learning systems, though they still suffer from the same weaknesses: they all generate toxic, biased, and inaccurate text. Experts have argued against making language models larger, comparing them to " [1]stochastic parrots ;" they don't understand language and simply regurgitate patterns in the training data.

They can spit out racist remarks, produce misinformation, or [2]memorise personal identifiable information . The safety and ethical risks involved in building such systems increases as they grow in size, prompting academics to argue against scaling up. Some believe more time and effort should be spent inventing new algorithms that are smaller and less computationally intensive, instead of just making existing architectures larger.

[3]

The latest 540-billion parameter transformer-based system built by researchers at Google, however, shows the performance of language models can still improve with size.

[4]

[5]

"We evaluated [Pathways Language Model] (PaLM) on hundreds of language understanding and generation tasks, and found that it achieves state-of-the-art few-shot performance across most tasks, by significant margins in many cases," Sharan Narang and Aakanksha Chowdhery, software engineers at Google Research, [6]said .

PaLM was better at a wide range of tasks, from question-answering, reading comprehension to common sense reasoning, than OpenAI's GPT-3, Nvidia and Microsoft's Megatron-Turing NLG, and DeepMind's Chinchilla and Gopher language models, they claimed.

[7]

PaLM is bigger, and contains more parameters than all of these models. It can also generate code, and seems to perform comparably to OpenAI's Codex 12B model despite being trained on less Python code, according to results published in a recent paper [8][PDF] ;

PaLM excels in another area: training efficiency. It was trained on 6,144 chips across two Cloud TPU v4 Pods, the company's largest training system configuration to date. A total of 2.56x10 FLOPs, equivalent to 29,600 petaFLOPs per day were performed during the process.

[9]AI beats top players at Bridge in two-day tournament

[10]How Google hopes to build more efficient, multi-capability AI systems

[11]DARPA to build life-saving AI models that think like medics

[12]DeepMind 'grossly inadequate' at tackling sexual harassment, says former staffer

"The goal is always to optimize the parallelism strategy, model architecture, compiler implementation together to maximize the FLOPs utilization, but the theoretical maximum throughput may not be achievable on any system," Chowdhery told The Register .

"Essentially when the accelerator chips (TPU or GPU) are not being used for matrix multiplication operations, this metric counts it as less than the theoretical maximum utilization."

Some computation is wasted transferring data from the memory, and passing it back and forth between neighboring chips. PaLM achieves training efficiency of 57.8 per cent hardware FLOPs utilization, making it more efficient than other language models. The researchers reckon there are still more performance gains to be realized by training language models on higher quality text or more data on top of training efficiency.

Some things don't change

Despite PaLM's capabilities, it still generates offensive and untruthful text and reflects biases in its training data. For example, it is more likely to associate Muslims with violence or terrorism stereotypes. Like other language models, PaLM was trained on text scraped from the internet. In fact, 50 percent of its training data come from conversations on social media websites.

"Our analysis reveals that our training data, and consequently PaLM, do reflect various social stereotypes and toxicity associations around identity terms," the team admitted in the paper. "Removing these associations, however, is non-trivial; for instance, filtering off content that is deemed toxic by an automated tool may disproportionately exclude content about or authored by marginalized subgroups in the training data."

[13]

PaLM's capabilities and limitations are partly due to it memorizing snippets of its training data. It has a memorization rate of 40 percent for examples that appear more than 500 times in the datatset, compared with 0.75 percent for an example that appears once. Memorization is double-edged sword; it's useful for recalling facts in information, but it also makes the system more likely to learn prejudices too.

Still, the researchers claim PaLM "shows breakthrough capabilities on numerous very difficult tasks". It is able to explain jokes, or perform multi-step arithmetic problems, and repair broken code. "Further understanding of risks and benefits of these models is a topic of ongoing research, together with developing scalable solutions that can put guardrails against malicious uses of language models," Narang and Chowdhery said.

PaLM is being used for research purposes. Google researchers developed the model as a proof of concept to scale up a language model using its [14]Pathways architecture . The goal is to experiment with the new technique to build a single AI system that can generalize across thousands or millions of tasks and is trained on different types of data, one day. ®

Get our [15]Tech Resources



[1] https://dl.acm.org/doi/10.1145/3442188.3445922

[2] https://www.theregister.com/2021/03/18/openai_gpt3_data/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2YkwTUX-Q9XHQNqz3S2TAhgAAAJQ&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YkwTUX-Q9XHQNqz3S2TAhgAAAJQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YkwTUX-Q9XHQNqz3S2TAhgAAAJQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://ai.googleblog.com/2022/04/pathways-language-model-palm-scaling-to.html

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44YkwTUX-Q9XHQNqz3S2TAhgAAAJQ&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[8] https://storage.googleapis.com/pathways-language-model/PaLM-paper.pdf

[9] https://www.theregister.com/2022/04/03/in_brief_ai/

[10] https://www.theregister.com/2022/03/31/google_tpu_ai/

[11] https://www.theregister.com/2022/03/30/darpa_ai_medicine/

[12] https://www.theregister.com/2022/03/31/deepmind_harassment_open_letter/

[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33YkwTUX-Q9XHQNqz3S2TAhgAAAJQ&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[14] https://www.theregister.com/2022/03/31/google_tpu_ai/

[15] https://whitepapers.theregister.com/



Slightly obvious flaw

b0llchit

I found a significant error in the model:

In fact, 50 percent of its training data come from conversations on social media websites.

How many hardware guys does it take to change a light bulb?

"Well the diagnostics say it's fine buddy, so it's a software problem."