Filing NeMo: Nvidia's AI framework hit with copyright lawsuit
(2024/03/11)
- Reference: 1710182709
- News link: https://www.theregister.co.uk/2024/03/11/authors_file_lawsuit_to_torpedo/
- Source link:
Nvidia is the latest tech giant to face allegations that it used copyrighted works to train AI models without obtaining the permission of the authors.
A proposed class action [1]lawsuit [PDF] filed against the GPU supremo in San Francisco on Friday March 8 claims the company used copyrighted material to train large language models in the Megatron library for its [2]NeMo generative AI framework .
The complaint was filed by three authors, Abdi Nazemian, Brian Keene, and Stewart O'Nan, who claim that books they wrote were among the material used to train the Megatron LLMs.
[3]
From the court filing, it appears that Nvidia is not accused of overtly copying the work of the authors itself, but instead using a dataset to train the Megatron models that was known to contain a number of unlicensed copyrighted works.
[4]
[5]
The lawsuit refers specifically to models that Nvidia released in September 2022, namely NeMo Megatron-GPT 1.3B, NeMo Megatron-GPT 5B, NeMo Megatron-GPT 20B, and NeMo Megatron-T5 3B.
These are hosted on the website operated by AI outfit [6]Hugging Face , along with information about each model, including its training dataset. In this case, the information states that the models were trained on "The Pile" dataset prepared by EleutherAI.
[7]
The Pile is described as "an 800GB Dataset of Diverse Text for Language Modeling," and one of its constituent parts is a collection of books called Books3, which contains the contents of about 196,640 books, including those created by the three authors.
According to the court filing, the Books3 dataset was available separately on Hugging Face until October 2023, when it was removed because it "is defunct and no longer accessible due to reported copyright infringement."
The authors want the case to proceed as a class action, with themselves serving as class representatives, and are asking for a jury trial and for damages for the alleged violations of their copyrights.
[8]
In a statement sent to The Register , an Nvidia spokesperson said: "We respect the rights of all content creators and believe we created NeMo in full compliance with copyright law."
[9]Don't feel left out, chip designers. Nvidia's made a chatbot assistant for you, too
[10]Dell pumps out reference designs, plumps services, to bring AI on-prem
[11]New York Times sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT
[12]OpenAI urges court to throw out authors' claims in AI copyright battle
This isn't the first case of an AI company being sued over accusations of copyright infringement regarding the data used to train AI models. In December last year, The New York Times launched a [13]case against Microsoft and OpenAI over claims the pair had used its articles without permission to build ChatGPT and similar models.
That case was perhaps made more interesting by OpenAI's assertion in January that it would be [14]"impossible" to build top-tier neural networks that meet today's needs without using people's copyrighted works.
Meanwhile, Nvidia is still priming the AI pump with the announcement of a new professional certification in generative AI to help developers to establish technical credibility in this area.
Set to become available to coincide with the Santa Clara-based giant's GTC event later this month, the [15]professional certification program will offer two associate-level generative AI accreditations, focusing on proficiency in large language models and multimodal workflow skills. ®
Get our [16]Tech Resources
[1] https://regmedia.co.uk/2024/03/11/nvidia_comp2.pdf
[2] https://www.theregister.com/2023/06/27/nvidia_snowflake_custom_ai/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2023/08/24/hugging_face_big_tech_investment/
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://www.theregister.com/2023/11/01/nvidia_ai_chatbot/
[10] https://www.theregister.com/2023/08/01/dell_infrastructure_services_ai/
[11] https://www.theregister.com/2023/12/27/the_new_york_times_files/
[12] https://www.theregister.com/2023/08/31/openai_class_action_fair_use/
[13] https://www.theregister.com/2024/03/05/ms_openai_vs_nyt/?td=keepreading
[14] https://www.theregister.com/2024/01/08/midjourney_openai_copyright/
[15] https://www.nvidia.com/en-us/learn/certification/generative-ai-llm-associate/
[16] https://whitepapers.theregister.com/
A proposed class action [1]lawsuit [PDF] filed against the GPU supremo in San Francisco on Friday March 8 claims the company used copyrighted material to train large language models in the Megatron library for its [2]NeMo generative AI framework .
The complaint was filed by three authors, Abdi Nazemian, Brian Keene, and Stewart O'Nan, who claim that books they wrote were among the material used to train the Megatron LLMs.
[3]
From the court filing, it appears that Nvidia is not accused of overtly copying the work of the authors itself, but instead using a dataset to train the Megatron models that was known to contain a number of unlicensed copyrighted works.
[4]
[5]
The lawsuit refers specifically to models that Nvidia released in September 2022, namely NeMo Megatron-GPT 1.3B, NeMo Megatron-GPT 5B, NeMo Megatron-GPT 20B, and NeMo Megatron-T5 3B.
These are hosted on the website operated by AI outfit [6]Hugging Face , along with information about each model, including its training dataset. In this case, the information states that the models were trained on "The Pile" dataset prepared by EleutherAI.
[7]
The Pile is described as "an 800GB Dataset of Diverse Text for Language Modeling," and one of its constituent parts is a collection of books called Books3, which contains the contents of about 196,640 books, including those created by the three authors.
According to the court filing, the Books3 dataset was available separately on Hugging Face until October 2023, when it was removed because it "is defunct and no longer accessible due to reported copyright infringement."
The authors want the case to proceed as a class action, with themselves serving as class representatives, and are asking for a jury trial and for damages for the alleged violations of their copyrights.
[8]
In a statement sent to The Register , an Nvidia spokesperson said: "We respect the rights of all content creators and believe we created NeMo in full compliance with copyright law."
[9]Don't feel left out, chip designers. Nvidia's made a chatbot assistant for you, too
[10]Dell pumps out reference designs, plumps services, to bring AI on-prem
[11]New York Times sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT
[12]OpenAI urges court to throw out authors' claims in AI copyright battle
This isn't the first case of an AI company being sued over accusations of copyright infringement regarding the data used to train AI models. In December last year, The New York Times launched a [13]case against Microsoft and OpenAI over claims the pair had used its articles without permission to build ChatGPT and similar models.
That case was perhaps made more interesting by OpenAI's assertion in January that it would be [14]"impossible" to build top-tier neural networks that meet today's needs without using people's copyrighted works.
Meanwhile, Nvidia is still priming the AI pump with the announcement of a new professional certification in generative AI to help developers to establish technical credibility in this area.
Set to become available to coincide with the Santa Clara-based giant's GTC event later this month, the [15]professional certification program will offer two associate-level generative AI accreditations, focusing on proficiency in large language models and multimodal workflow skills. ®
Get our [16]Tech Resources
[1] https://regmedia.co.uk/2024/03/11/nvidia_comp2.pdf
[2] https://www.theregister.com/2023/06/27/nvidia_snowflake_custom_ai/
[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://www.theregister.com/2023/08/24/hugging_face_big_tech_investment/
[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ze@NCCk7BS90zuXJZOJJIAAAARc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[9] https://www.theregister.com/2023/11/01/nvidia_ai_chatbot/
[10] https://www.theregister.com/2023/08/01/dell_infrastructure_services_ai/
[11] https://www.theregister.com/2023/12/27/the_new_york_times_files/
[12] https://www.theregister.com/2023/08/31/openai_class_action_fair_use/
[13] https://www.theregister.com/2024/03/05/ms_openai_vs_nyt/?td=keepreading
[14] https://www.theregister.com/2024/01/08/midjourney_openai_copyright/
[15] https://www.nvidia.com/en-us/learn/certification/generative-ai-llm-associate/
[16] https://whitepapers.theregister.com/
Re: neural networks that meet today's needs
Anonymous Coward
Nvidea made $31 billion profit in the 12 months to October 2023 and returned $10.4 billion dollars to shareholders in the last quarter of 2023. Yet just like all the other mega-rich corporations involved in AI currently, they chose to steal data from creatives, which will make them even more billions, rather than pay them a fraction of their quarterly profits.
Just another day in the enshitification of the whole planet by immoral and psychopathic billionaires and toxic corporations.
neural networks that meet today's needs
So we want search engines that make up the answers if they don't find them, image generators that can't count arms, writing systems with added blandness?
Amazing. I never knew.