News: 1657575083

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Meta's AI-based Wikipedia successor 'may be the next big break in NLP'

(2022/07/11)


Meta has open-sourced a machine-learning resource that could one day supplant Wikipedia as the world's biggest publicly available knowledge-verification database.

Dubbed [1]Sphere , it can be used to perform knowledge-intensive natural language processing, or KI-NLP, we're told. In practical terms, that means it can be used to answer complicated questions using natural language, and find sources for claims.

A given example of its use is asking Sphere, "Who is Joëlle Sambi Nzeba?" Wikipedia doesn't have an entry for her, but Sphere said she was "born in Belgium and grew up partly in Kinshasa (Congo). She currently lives in Brussels. She is a writer and slammer, alongside her activism in a feminist movement," and links to a website where it got that information about her work.

[2]

Wikipedia has pretty much served as the corpus of record, Meta's eggheads wrote [3]in a paper discussing the design of Sphere, claiming the volunteer-maintained uber-wiki is "accurate, well-structured, and small enough to use easily in testing environments."

The tech and social impact of AI's powerful, emerging 'foundation models' [4]READ MORE

Seeking to build something bigger and better than Wikipedia, though, Meta pulled together content from all over the web to form a "universal, uncurated and unstructured knowledge source for multiple KI-NLP tasks at once." The result is Sphere, which is more or less a mountain of processed data that can be queried using a bunch of machine-learning tools.

The team adds that Sphere "can match and outperform baselines grounded in Wikipedia" on some tasks using the [5]KILT AI benchmark. That is to say, Sphere performs better than AI systems built on Wikipedia's content.

[6]

[7]

The primary aim of Sphere was to see what impact replacing Wikipedia, as a source, had on the performance of knowledge-intensive systems, and while the team did report that Sphere had some issues, its performance indicates that, at the very least, it can add value to KI-NLP tasks beyond what Wikipedia corpora can offer.

The researchers behind Sphere claim their work marks "the first time a general purpose search index improves language models on common sense tasks."

[8]

Sphere isn't the only AI platform Meta has released on GitHub: last week it released [9]NLLB-200 , the first translation AI to pass the 200 language threshold, or so the Facebook parent claimed. Like Sphere, NLLB-200 has been put to use at Wikipedia; the former system for automatically checking citations in edited articles, and the latter to improve translation of pages into less commonly spoken languages.

When transitioning to a web corpus, we no longer have the certainty that any document is good, truthful or unique

Sphere goes beyond similar web corpora in terms of scale, consisting of 906 million passages and 134 million documents. The next largest in terms of passages/documents is the [10]Internet Augmented Dialog generator, which pulls data from 250 million passages and 109 million documents.

But the internet contains no controls for quality or accuracy, which the researchers admit is a key problem for actually deploying this thing. "Using Wikipedia as the knowledge source allows researchers to assume the high quality of the corpus documents. When transitioning to a web corpus, we no longer have the certainty that any document is good, truthful or unique," the researchers wrote.

Sphere's creators think iterative efforts should focus on assessing quality of the data it retrieves, detecting false claims and contradictions, determining how to prioritize trustworthy sources, and when to decide not to answer a question because of a lack of information. You know, making it actually useful.

If it can successfully turn Sphere into a white-box AI with reliable and trustworthy information, Meta said, Sphere "may be the next big break in NLP." ®

Get our [11]Tech Resources



[1] https://github.com/facebookresearch/sphere

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Ysydiz3ztl12NTtWxhsL9wAAAAs&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://arxiv.org/abs/2112.09924

[4] https://www.theregister.com/2021/08/23/percy_liang_qa/

[5] https://ai.facebook.com/tools/kilt/

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ysydiz3ztl12NTtWxhsL9wAAAAs&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Ysydiz3ztl12NTtWxhsL9wAAAAs&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[8] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Ysydiz3ztl12NTtWxhsL9wAAAAs&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[9] https://www.theregister.com/2022/07/06/facebooks_new_translation_ai_breaks/

[10] http://arxiv.org/abs/2107.07566

[11] https://whitepapers.theregister.com/



What could possible go wrong?

The Man Who Fell To Earth

So they think Wikipedia is accurate, eh? Explains a lot.

Let's all ask it about [1]Scotland ...

[1] https://www.theregister.com/2020/08/26/scots_wikipedia_fake/

Another tool for Suckerberg

HildyJ

I assume we'll see this in Farcebook any day now. Unlike the scientists, they don't care about accuracy. They want eyeballs and engagement for ads. Controversy over accuracy, if anything, increases eyeballs and engagement.

Too bad this wasn't from a more reputable NLP research source.

We are preparing to think about contemplating preliminary work on plans to
develop a schedule for producing the 10th Edition of the Unix Programmers
Manual.
-- Andrew Hume