News: 1709598669

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Anthropic releases Claude 3 and claims it's better than ChatGPT and Gemini

(2024/03/05)


AI startup Anthropic has released Claude 3, the latest iteration of its large language model, which it claims is more powerful than OpenAI's GPT-4.

Announced on Monday, Claude 3 comes in three different sizes: [1]Opus, Sonnet, and Haiku [badly formatted PDF]. Opus is the most powerful of the three and is available to developers and users via Anthropic's API and Claude Pro subscription. Sonnet can be accessed by developers through an API and currently powers Anthropic's free web chatbot. The smallest model, Haiku, isn't available just yet.

In academic benchmark tests – assessing LLMs' ability to retain common knowledge, solve math problems, generate code, and show reasoning skills – Opus scored higher than OpenAI's GPT-4 and Google’s Gemini Ultra, Anthropic reports. The developer went so far as to boast that Opus "exhibits near-human levels of comprehension and fluency on complex tasks, leading the frontier of general intelligence."

[2]

Meanwhile, Sonnet and Haiku are more powerful than OpenAI's previous GPT-3.5 model, but less capable than Google's Gemini Ultra and Pro models.

[3]

[4]

Anthropic explained that the context window – the amount of input it can process at once – will be 200K tokens at first but is capable of going up to a million tokens.

Opus is pricey, and designed for users looking to use AI for tasks that require top levels of data comprehension and generation – like scientific research or analyzing long, complex reports. It costs $15 to process an input prompt stretching to a million tokens, and $75 to generate a million tokens for output. By way of comparison, OpenAI charges between $10 and $30 for processing and generating a million tokens on its GPT-4 Turbo model.

[5]

Sonnet is aimed at mainstream enterprise users that need a capable yet fast model that can do things like search and retrieve information, write marketing copy, or generate code. It has been optimized for large-scale deployments and costs $3 and $15 to handle a million tokens at input and output, respectively. Haiku will be even cheaper, costing $0.25, and $1.25 to process and generate a million tokens. It should be useful for things like content moderation, language translation, or customer service.

[6]More and more LLMs in biz products, but who'll take responsibility for their output?

[7]Top LLMs struggle to make accurate legal arguments

[8]FTC drills into Amazon, Microsoft, Google over billions pledged to OpenAI, Anthropic

[9]How 'sleeper agent' AI assistants can sabotage your code without you realizing

Amazon announced it will host Anthropic's Claude 3 models on its Bedrock cloud platform: Sonnet today, and Opus and Haiku sometime soon. It's a similar story for Google Cloud's Vertex AI Model Garden: Sonnet is available today in private preview, with API access to all three models arriving soon.

Claude 3 is also less cautious than its predecessor. Claude 2.1 would often refuse to comply with prompts that weren't necessarily harmful – like requests to write a fictional story. The developer's announcement [10]assured users : "We've made meaningful progress in this area: Opus, Sonnet, and Haiku are significantly less likely to refuse to answer prompts that border on the system's guardrails than previous generations of models."

Large language models' surprise emergent behavior written off as 'a mirage' [11]READ MORE

The biggest issue that plagues LLMs, however, is their tendency to generate inaccurate information or straight-up make things up with such confidence that users may well believe it. The errors – referred to as hallucinations – make it difficult to trust the output of AI software let alone give computers more autonomy in tasks.

Anthropic promised Opus offers a "twofold improvement" in accuracy compared to Claude 2.1, and will introduce a feature that will cite sources in the outputs generated by its latest models for users to inspect. That's similar to say, Google Gemini, which also says where it got its info from in some of its answers to prompts.

"We do not believe that model intelligence is anywhere near its limits, and we plan to release frequent updates to the Claude 3 model family over the next few months. We're also excited to release a series of features to enhance our models' capabilities, particularly for enterprise use cases and large-scale deployments," Anthropic's announcement concluded.

Interestingly, Anthropic has chosen to not make Claude 3 a multi-modal system. Although it can process images, it cannot produce them and cannot handle audio or video inputs, unlike ChatGPT or Gemini. ®

Get our [12]Tech Resources



[1] https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Zeam9xh0vpRzNWX4mDSpCgAAAU8&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Zeam9xh0vpRzNWX4mDSpCgAAAU8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Zeam9xh0vpRzNWX4mDSpCgAAAU8&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Zeam9xh0vpRzNWX4mDSpCgAAAU8&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://www.theregister.com/2023/09/28/llms_business_risks/

[7] https://www.theregister.com/2024/01/10/top_large_language_models_struggle/

[8] https://www.theregister.com/2024/01/25/ftc_ai_inquiry/

[9] https://www.theregister.com/2024/01/16/poisoned_ai_models/

[10] https://www.anthropic.com/news/claude-3-family

[11] https://www.theregister.com/2023/05/16/large_language_models_behavior/

[12] https://whitepapers.theregister.com/



No true AI

Throatwarbler Mangrove

I've had it up to here with people hailing large language models as the epitome of artificial intelligence. Let me make this clear: these models are far from being true AI, and it's time we stop pretending they are!

First off, what do these models really do? They generate text based on patterns and information they've learned during training. It's like saying a parrot is a genius because it can mimic human speech. Just because these models can regurgitate information and produce seemingly coherent sentences doesn't mean they truly understand the content. True AI should comprehend and reason, not just repeat like a parrot.

Moreover, large language models lack genuine understanding of context. They can string together words and phrases, but do they grasp the nuances, emotions, or cultural intricacies embedded in the language? No! True AI should possess the ability to empathize, understand context, and interpret information beyond mere pattern recognition.

Let's not forget the issue of biased outputs. These models are trained on massive datasets, often reflecting the inherent biases present in the data. As a result, they perpetuate and even amplify existing societal biases. True AI should be capable of recognizing and rectifying such biases, not reinforcing them.

Don't even get me started on the lack of common sense! These models can spew out absurd and nonsensical responses, showcasing their fundamental inability to apply common sense reasoning. True AI should be able to navigate the world with logic and sound judgment, not leave us scratching our heads at their bizarre outputs.

In conclusion, folks, it's high time we stopped throwing around the term "AI" when referring to these large language models. They are sophisticated tools with impressive capabilities, but they fall far short of what true artificial intelligence should embody. Let's demand more from the AI community and set higher standards instead of settling for the illusion of intelligence!

End of rant.

Re: No true AI

HuBo

Cool rant (and, on top of that, the LLMs don't even do spell-checking!)! In a recent interview of [1]Sampo Pyysalo , he seems (maybe contradictorily) to echo both these thoughts and their opposite. If I read it well, it seems that your perspective (of separate linguistic/comm and cognitive/mind subsystems) is also that of [2]Noam Chomsky (good intellectual company!).

[1] https://www.eetimes.eu/how-university-of-turku-researcher-sampo-pyysalo-trains-finnish-llms/

[2] https://iep.utm.edu/chomsky-philosophy/

Re: No true AI

claimed

I dunno, if I tell you I’ve got an artificial leg, what do you expect? A true leg should do (blah blah), these artificial legs are just not there yet….

Re: No true AI

Throatwarbler Mangrove

And just so everyone knows, my post above was created by ChatGPT with the prompt, "Create an angry forum post complaining that large language models are not true artificial intelligence."

Setting the bar low :)

Bebu

《exhibits near-human levels of comprehension》

Often not a big ask. Once had a dog (begian) that had more intelligence than most of surrounding population.

《fluency on complex tasks》

Presumably on linguistic tasks? Don't need much fluency with arc welding only with the flux I imagine.

《leading the frontier of general intelligence.》

General Intelligence - the head of military intelligence (the definitive oxymoron.)

In a couple of years all this will be last year's tulips. :)

I was thinking what LLM might actually be useful for - the obvious one to me was handwriting recognition and probably speech recognition. They are pretty reasonable now but I wonder whether that is the case with non latin scripts and specialist domains like mathematics - would be convenient to be able to enter a p.d.e. like Schroedinger's equations by hand. :)

Comparing software engineering to classical engineering assumes that software
has the ability to wear out. Software typically behaves, or it does not. It
either works, or it does not. Software generally does not degrade, abrade,
stretch, twist, or ablate. To treat it as a physical entity, therefore, is
misapplication of our engineering skills. Classical engineering deals with
the characteristics of hardware; software engineering should deal with the
characteristics of *software*, and not with hardware or management.
-- Dan Klein