News: 1653507551

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Despite 'key' partnership with AWS, Meta taps up Microsoft Azure for AI work

(2022/05/25)


Meta’s AI business unit set up shop in Microsoft Azure this week and announced a strategic partnership it says will advance PyTorch development on the public cloud.

The deal

[1]PDF

will see Mark Zuckerberg’s umbrella company deploy machine-learning workloads on thousands of Nvidia GPUs running in Azure. While a win for Microsoft, the partnership calls in to question just how strong Meta’s commitment to Amazon Web Services (AWS) really is.

Back in those long-gone days [2]of December , Meta named AWS as its “key long-term strategic cloud provider." As part of that, Meta promised that if it bought any companies that used AWS, it would continue to support their use of Amazon's cloud, rather than force them off into its own private datacenters. The pact also included a vow to expand Meta’s consumption of Amazon’s cloud-based compute, storage, database, and security services.

[3]

The AWS-Meta team up also included a collaboration to optimize workloads using the PyTorch machine learning framework — which Meta, then Facebook, released in 2016 — for deployment in the cloud provider’s Elastic Compute Cloud and SageMaker services.

[4]Meta releases code for massive language model to AI researchers

[5]Facebook opens political ad data vaults to researchers

[6]Meta won't migrate future acquisitions out of AWS

[7]Zuckerberg sued for alleged role in Cambridge Analytica data-slurp scandal

It appears Meta is more than happy to play the field, though, deploying workloads wherever it pleases. Guess that's what they call multi-cloud; it also demonstrates the difference between "key" and "exclusive."

The announcement this week revealed Meta began deploying workloads on Azure’s Nvidia A100-accelerated instances to train AI models in 2021. Meta now plans to expand deployments on Azure to a dedicated cluster consisting of 5,400 of Nvidia’s 80GB A100 GPUs to accelerate AI research and development on “cutting-edge ML training workloads” for its AI business unit.

[8]

[9]

In fact, the social media giant says it trained its 175 billion parameter [10]OPT-175B natural-language processing transformer model, released in March, in Azure.

“With Azure’s compute power and 1.6TB/s of interconnect bandwidth per VM, we are able to accelerate our ever-growing training demands to better accommodate larger and more innovative AI models,” Meta’s VP of AI Jerome Pesenti said in a statement.

[11]

While neither Microsoft nor Meta provided specifics as to how exactly the massive GPU cluster will be used moving forward, there’s a fair chance that, much like the social media giant’s earlier AWS collab, it’ll involve PyTorch.

In addition to the infrastructure deal, Meta said it would collaborate with Microsoft to “scale PyTorch adoption on Azure.”

“We’re happy to work with Microsoft in extending our experience to their customers using PyTorch in their journey from research to production,” Pesenti said.

[12]

Later this year, Microsoft plans to roll out PyTorch development accelerators it says will make it easier to deploy the framework on Azure.

The Register reached out to Meta for further comment; we’ll let you know if we hear back. ®

Get our [13]Tech Resources



[1] https://news.microsoft.com/wp-content/uploads/prod/sites/636/2022/05/Meta-MSFT-announcement-FINAL.pdf

[2] https://www.theregister.com/2021/12/02/meta_aws_alliance/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Yo6nBhxYCtDUlruYq6ovAgAAAEk&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://www.theregister.com/2022/05/04/meta_releases_code_for_175billion/

[5] https://www.theregister.com/2022/05/24/facebook_political_ad_targeting_data/

[6] https://www.theregister.com/2021/12/02/meta_aws_alliance/

[7] https://www.theregister.com/2022/05/23/zuckerberg_sued_for_his_role/

[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Yo6nBhxYCtDUlruYq6ovAgAAAEk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Yo6nBhxYCtDUlruYq6ovAgAAAEk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[10] https://www.theregister.com/2022/05/04/meta_releases_code_for_175billion/

[11] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Yo6nBhxYCtDUlruYq6ovAgAAAEk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[12] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/paasiaas&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Yo6nBhxYCtDUlruYq6ovAgAAAEk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[13] https://whitepapers.theregister.com/



Different strokes for different data

HildyJ

Since I don't run a mega meta multinational, I've never researched it, but it makes sense that each cloudy provider has its strengths and weaknesses (although finding them goes beyond cloudy and into foggy or even pea soup territory).

But I suspect that Suckerberg's primary concern is making sure he has an appendage (let's say feet in AWS and Azure, arms in Google and IBM) because that lets him bully the clouds into discounts and perks with the threat of shifting his business among them.

Audience: What will become of Linux when the Hurd is ready?
Eric Youngdale: Err... is Richard Stallman here?
-- From the Linux conference in spring '95, Berlin