News: 1716441607

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Read AI about it... OpenAI does deal with News Corp

(2024/05/23)


OpenAI and News Corp on Wednesday announced a partnership that will bring the publisher's output to the super-lab's models, marking yet another in a series of data content deals for the industry.

The tie-up means that "OpenAI has permission to display content from News Corp mastheads in response to user questions and to enhance its products." What that possibly means: The models can be trained on News Corp articles, answer queries using that info, and cite those sources.

OpenAI hasn't always sought permission. For example, it was recently accused of [1]failing to respect celebrity Scarlett Johansson's refusal to license her voice to the biz for speech synthesis. OpenAI said in response it didn't do anything wrong, and [2]insisted again on Wednesday it did not use the movie star's voice, and that it had hired an actress for the speech synthesis role before it tried to tap up Johansson.

[3]

Nor has OpenAI fully disclosed the data used to train its latest models.

[4]

[5]

But given the growing number of [6]AI-related lawsuits , asking permission rather than seeking forgiveness looks like the wiser option. Non-consensual use of copyrighted content for AI training, arguably rampant throughout the machine learning world lately, isn't likely to go unnoticed anymore.

Other recent agreements reached by OpenAI include content access arrangements with [7]Reddit and [8]Stack Overflow . And rival Google struck a similar deal [9]with Reddit for its user-generated comments.

[10]

The latest partnership – rumored to land News Corp $250 million over the next five years – will give OpenAI access to current and past content from various publications today owned by the corporation Rupert Murdoch founded, including The Wall Street Journal, Barron’s, MarketWatch, Investor’s Business Daily, FN, and the New York Post; The Times, The Sunday Times and The Sun; The Australian, news.com.au, The Courier Mail, The Advertiser, and Herald Sun; and others. It doesn't cover content from News Corp's other businesses.

The two companies in separate but similar [11]press [12]releases said the ultimate goal of the pact is to help people "make informed choices based on reliable information and news sources."

AI models often come with a caution that they can't be relied upon, because they may "hallucinate" – make things up. The OpenAI's ChatGPT web page, for example, contains the disclaimer, "ChatGPT can make mistakes. Check important info."

[13]

Sam Altman, CEO of OpenAI, characterized the deal as "a proud moment for journalism and technology."

"Together, we are setting the foundation for a future where AI deeply respects, enhances, and upholds the standards of world-class journalism," he said.

Neither OpenAI nor News Corp made it clear how their data deal will uphold the standards of world-class journalism, which tends to favor disclosure of information rather than withholding it through the imposition of a [14]non-disclosure, non-disparagement agreement . OpenAI has since denied that it threatened staff with loss of equity if they broke an unusually restrictive NDA, and Sam Altman said he knew nothing of such threats, although [15]it seems plenty of management did.

[16]Top AI players pledge to pull the plug on models that present intolerable risk

[17]Microsoft AI Studio opens for business, with a nod to safety

[18]If you find Microsoft's Copilot offerings overwhelming, it's no wonder: There are 130-plus of them now

[19]AI might be coming for your job, but Sam Altman can't go on dinner dates anymore

The impact of AI on journalism is a complex topic, covered in depth in a [20]recent report from the Tow Center for Digital Journalism at Columbia University's Graduate School of Journalism.

The report notes, "The growing use of AI in news work tilts the balance of power toward technology companies, raising concerns about 'rent' extraction and potential threats to publishers’ autonomy business models, particularly those reliant on search-driven traffic."

Ron Bodkin, co-founder and CEO of Theoriq, a decentralized AI firm, told The Register that a compensation model is necessary for those making the material that sustains AI.

It's essential that content creators and journalists are fairly compensated for their work

"We believe it's essential that content creators and journalists are fairly compensated for their work, especially as AI agents and chatbots increasingly consume and summarize content," he said. "This compensation is critical to maintaining high-quality journalism and supporting the creative industry."

Bodkin however said that current trends favor the largest tech companies through exclusive content access deals.

"This dynamic often results in only the largest publishers receiving funds for their content, further consolidating power and resources," he explained. "The worst case is licensing deals are exclusive, leading to walled gardens that restrict access to a monopolist."

Bodkin argues that AI access to data needs to be democratized using distributed technologies, such as micropayments on a blockchain.

"This approach would ensure that journalists and other data providers are fairly compensated while also allowing startups and nonprofit researchers equitable access to data," he said. "Such a system would be based on the value realized from the data, promoting a more inclusive and competitive environment in AI development."

The Register has asked OpenAI and News Corp to confirm whether "enhance its products" should be interpreted to mean that OpenAI intends to use News Corp material to train its models. We also asked whether News Corp journalists will be compensated for having their work used thus.

We've not heard back. ®

Get our [21]Tech Resources



[1] https://www.theregister.com/2024/05/21/scarlett_johansson_openai_accusation/

[2] https://www.washingtonpost.com/technology/2024/05/22/openai-scarlett-johansson-chatgpt-ai-voice/

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Zk8T5Bcu22yZfvU05E0tFAAAAEk&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Zk8T5Bcu22yZfvU05E0tFAAAAEk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Zk8T5Bcu22yZfvU05E0tFAAAAEk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://blogs.gwu.edu/law-eti/ai-litigation-database/

[7] https://openai.com/index/openai-and-reddit-partnership/

[8] https://openai.com/index/api-partnership-with-stack-overflow/

[9] https://blog.google/inside-google/company-announcements/expanded-reddit-partnership/

[10] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Zk8T5Bcu22yZfvU05E0tFAAAAEk&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[11] https://openai.com/index/news-corp-and-openai-sign-landmark-multi-year-global-partnership/

[12] https://newscorp.com/2024/05/22/news-corp-and-openai-sign-landmark-multi-year-global-partnership/

[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Zk8T5Bcu22yZfvU05E0tFAAAAEk&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[14] https://x.com/sama/status/1791936857594581428

[15] https://www.vox.com/future-perfect/351132/openai-vested-equity-nda-sam-altman-documents-employees

[16] https://www.theregister.com/2024/05/22/ai_safety_seoul_declaration_signed/

[17] https://www.theregister.com/2024/05/21/microsoft_ai_studio_opens_for/

[18] https://www.theregister.com/2024/05/21/microsoft_extends_reach_of_copilot/

[19] https://www.theregister.com/2024/05/20/openai_safety_culture/

[20] https://www.cjr.org/tow_center_reports/artificial-intelligence-in-the-news.php

[21] https://whitepapers.theregister.com/



RIP ChatGPT

Sorry that handle is already taken.

The Register has asked OpenAI and News Corp to confirm whether "enhance its products" should be interpreted to mean that OpenAI intends to use News Corp material to train its models. I appreciate that this might have been the joke but... I can't wait to read their response. That's not what "enhance" means.

Selling your soul...

chuckufarley

...has never been so cheap.

Re: Selling your soul...

Anonymous Coward

So cheap! $50 million a year is a bargain. Faust will be laughing in his fiery grave at the fast one the big D has pulled this time, as he awaits Murdoch arrival whereon Faust can taunt at Murdoch for eternity, singing, "He didn't even get the wishes". Who says the Devil is all bad!?

AI is rampant in the news business. Nationals use LLMs for almost all subbing now (if not officially then by staff). It is really quick and good at headlines, subheads, captions, etc. All the stuff that takes time. Now, it's just mostly proofing its output and its quality is top-notch. Subbing desks can lose 50% of their staff now. Also, same for the Picture desk and Legal in the next 6 months. Those staff that remain have perfected prompt engineering. That is the most desirable skill in Production now.

This is great for OpenAI cause they now have a chunk of validated and verified data which can be labelled as such. Lots of pluses for that and OpenAI will most likely earn 20 times what they paid. This should also give OpenAI access to archives that aren't online.

There will be a TImesAI. The Reg should push for a similar deal to survive.

Re: Selling your soul...

Dan 55

Ah, Murdoch has a soul to sell does he now?

Nice words

Pascal Monett

" Together, we are setting the foundation for a future where AI deeply respects, enhances, and upholds the standards of world-class journalism "

Love the idea. Now all we need is to find some world-class journalism.

News outlets today are owned by billionnaires who shape the news and the way it is broadcast. Gone are the days when the news was actual news, with facts and nothing but. Today it's all about shaping the message. Facts are carefully selected. Not everything is said, and some things are said without proper sourcing (which doesn't keep those things from being repeated).

Re: Nice words

A Non e-mouse

I remember at secondary school we had a couple of lessons specifically about the media and how they can easily put a filter on the story to project their own bias.

A Non e-mouse

For those outside the UK, Newscorp own two of the slimiest newspapers in the UK: News Of The World and The Sun.

The News Of The World had to shut down after they were found to have hacked the voicemail of a dead teenager. (This caused distress and confusion as the investigation wrongly thought she was still alive)

The Sun is infamous for being the first British newspaper to publish daily pictures of topless women (Somehow this is journalism) and also blamed the victims of the [1]Hillsborough disaster for their own deaths. (It was later proved that Police incompetence was to blame)

There are far more incidents of these publications having no morals.

So if OpenAI are using Newcorp as a source for their training data, it tells you how much you should trust the output of OpenAI: Not at all.

[1] https://en.wikipedia.org/wiki/Hillsborough_disaster

Bendacious

Not to forget two weeks ago The Sun labelled a 14 year old boy, murdered in the street by a maniac, "sword lad". Amazingly crass "world-class journalism"

xanadu42

Does this mean that ChatGPT will soon be spewing extremist content?

There are benefits to exclusive deals

SVD_NL

For example, if i decide on using an LLM, i can choose to use one that isn't trained on The Sun.

Mastheads?

tfewster

AKA the banner? Or article headlines? Or the article contents?

I get that newspapers are usually desirable inputs, as they have to stick fairly close to the truth or face the consequences [YMMV*]. Unlike Reddit, where any yahoo can just make shit up with impunity.

[World War 2 Bomber Found on Moon, Sunday Sport, 14 August 1988]

they can't be relied upon, because they may "hallucinate" – make things up

Howard Sway

Well, it's a good tie-up, because history has shown that The Sun also sometimes couldn't be relied on and made things up - which is why they've been sued so many times for libel. As well as having to pay damages when they lost, they then often had to publish a retraction in later issues. I doubt the AI training will link these together, so it may ingest some old stories that turned out to be libellous, and then regurgitate them as "facts". At this point, as money has changed hands for the libellous content, and it has in effect been republished, OpenAI could find itself on the end of some serious lawsuits.

Cite your sources

Bendacious

"The models can be trained on News Corp articles, answer queries using that info, and cite those sources."

How do they cite those sources? This goes against everything I've read about LLMs. I don't think the models know what the source is. They are set to actively avoid regurgitating training data and I don't think they know when they are doing that. The sources are usually thousands or millions of pieces of different text, stolen from all over the internet.

Huh?