Keeping Knowledge Free for Over a Decade

Claude 3 surpasses GPT-4 on Chatbot Arena for the first time

posted onMarch 28, 2024

by l33tdawg

Arstechnica

Credit: Arstechnica

On Tuesday, Anthropic's Claude 3 Opus large language model (LLM) surpassed OpenAI's GPT-4 (which powers ChatGPT) for the first time on Chatbot Arena, a popular crowdsourced leaderboard used by AI researchers to gauge the relative capabilities of AI language models. "The king is dead," tweeted software developer Nick Dobos in a post comparing GPT-4 Turbo and Claude 3 Opus that has been making the rounds on social media. "RIP GPT-4."

Since GPT-4 was included in Chatbot Arena around May 10, 2023 (the leaderboard launched May 3 of that year), variations of GPT-4 have consistently been on the top of the chart until now, so its defeat in the Arena is a notable moment in the relatively short history of AI language models. One of Anthropic's smaller models, Haiku, has also been turning heads with its performance on the leaderboard.

"For the first time, the best available models—Opus for advanced tasks, Haiku for cost and efficiency—are from a vendor that isn't OpenAI," independent AI researcher Simon Willison told Ars Technica. "That's reassuring—we all benefit from a diversity of top vendors in this space. But GPT-4 is over a year old at this point, and it took that year for anyone else to catch up."

Source

Tags

Artificial Intelligence

previous article

You May Also Like

Recent News

Friday, November 29th

North Korean hackers posing as IT workers steal over $1B in cyberattack

OpenAI is at war with its own Sora video testers following brief public leak

Found on VirusTotal: The world’s first UEFI bootkit for Linux

Tuesday, November 19th

CISA Director Jen Easterly, in Place Since 2021, to Step Down

Korea extradites Russian, Vietnamese suspects linked to $16M ransomware scheme

WhatsApp: NSO Group Operates Pegasus Spyware for Customers

Friday, November 8th

North Korean hackers target cryptocurrency with malware

Man sick of crashes sues Intel for allegedly hiding CPU defects

Law enforcement operation takes down 22,000 malicious IP addresses worldwide

Friday, November 1st

Youth of today say passwords are old news, passkeys are the future

Chinese attackers accessed Canadian government networks – for five years

OpenAI launches ChatGPT with Search, taking Google head-on

Here’s the paper no one read before declaring the demise of modern cryptography

Not just ChatGPT anymore: Perplexity and Anthropic’s Claude get desktop apps

Tuesday, July 9th

China's APT40 gang is ready to attack vulns within hours or days of public release

AI-Powered Super Soldiers Are More Than Just a Pipe Dream

The president ordered a board to probe a massive Russian cyberattack. It never did.

Massive car dealer ransom attack is mostly over after 2 weeks of work-arounds

Wednesday, July 3rd

“RegreSSHion” vulnerability in OpenSSH gives attackers root on Linux

Two of the German military’s new spy satellites appear to have failed in orbit

Friday, June 28th

Indonesian Airports, Data Centres Hit By Worst Cyberattack in Years

Cisco Talos warns of wider security implications following Snowflake breach

Thursday, June 27th

I Wore Meta Ray-Bans in Montreal to Test Their AI Translation Skills. It Did Not Go Well

Researchers upend AI status quo by eliminating matrix multiplication in LLMs

YouTube tries convincing record labels to license music for AI song generator

Critical MOVEit vulnerability puts huge swaths of the Internet at severe risk

Julian Assange to plead guilty but is going home after long extradition fight

Thursday, June 13th

Hackers show off jailbroken checkm8-vulnerable iPad and Apple TV running iPadOS 18 & tvOS 18 respectively

One of the major sellers of detailed driver behavioral data is shutting down

Turkish student creates custom AI device for cheating university exam, gets arrested

Wednesday, June 12th

Chinese-Made Biometric Access System Has 24 Vulnerabilities

Apple and OpenAI currently have the most misunderstood partnership in tech

Adobe to update vague AI terms after users threaten to cancel subscriptions

China state hackers infected 20,000 Fortinet VPNs

Tuesday, June 11th

Russia Is Targeting Germany With Fake Information as Europe Votes