LocateBaltimore
No Result
View All Result
No Result
View All Result
LocateBaltimore
No Result
View All Result
Home Technology

Before launching, GPT-4o broke records on chatbot leaderboard under a secret name

Pauline Wright by Pauline Wright
May 14, 2024
in Technology
0
325
SHARES
2.5k
VIEWS
Share on FacebookShare on Twitter


Getty Images

On Monday, OpenAI worker William Fedus confirmed on X that a mysterious chat-topping AI chatbot often called “gpt-chatbot” that had been present process testing on LMSYS’s Chatbot Arena and irritating specialists was, the truth is, OpenAI’s newly introduced GPT-4o AI mannequin. He additionally revealed that GPT-4o had topped the Chatbot Arena leaderboard, attaining the very best documented rating ever.

“GPT-4o is our new state-of-the-art frontier mannequin. We’ve been testing a model on the LMSys area as im-also-a-good-gpt2-chatbot,” Fedus tweeted.

Chatbot Arena is a web site the place guests converse with two random AI language fashions aspect by aspect with out figuring out which mannequin is which, then select which mannequin provides the perfect response. It’s a excellent instance of vibe-based AI benchmarking, as AI researcher Simon Willison calls it.

Enlarge / An LMSYS Elo chart shared by William Fedus, exhibiting OpenAI’s GPT-4o under the name “im-also-a-good-gpt2-chatbot” topping the charts.

The gpt2-chatbot fashions appeared in April, and we wrote about how the shortage of transparency over the AI testing course of on LMSYS left AI specialists like Willison annoyed. “The complete scenario is so infuriatingly consultant of LLM analysis,” he instructed Ars on the time. “A totally unannounced, opaque launch and now the whole Internet is working non-scientific ‘vibe checks’ in parallel.”

On the Arena, OpenAI has been testing a number of variations of GPT-4o, with the mannequin first showing because the aforementioned “gpt2-chatbot,” then as “im-a-good-gpt2-chatbot,” and eventually “im-also-a-good-gpt2-chatbot,” which OpenAI CEO Sam Altman made reference to in a cryptic tweet on May 5.

Advertisement

Since the GPT-4o launch earlier immediately, a number of sources have revealed that GPT-4o has topped LMSYS’s inside charts by a appreciable margin, surpassing the earlier high fashions Claude 3 Opus and GPT-4 Turbo.

“gpt2-chatbots have simply surged to the highest, surpassing all of the fashions by a important hole (~50 Elo). It has develop into the strongest mannequin ever within the Arena,” wrote the lmsys.org X account whereas sharing a chart. “This is an inside screenshot,” it wrote. “Its public model ‘gpt-4o’ is now in Arena and can quickly seem on the general public leaderboard!”

An an internal screenshot of the LMSYS Chatbot Arena leaderboard showing "im-also-a-good-gpt2-chatbot" leading the pack. We now know that it's GPT-4o.
Enlarge / An an inside screenshot of the LMSYS Chatbot Arena leaderboard exhibiting “im-also-a-good-gpt2-chatbot” main the pack. We now know that it is GPT-4o.

As of this writing, im-also-a-good-gpt2-chatbot held a 1309 Elo versus GPT-4-Turbo-2023-04-09’s 1253, and Claude 3 Opus’s 1246. Claude 3 and GPT-4 Turbo had been duking it out on the charts for a while earlier than the three gpt2-chatbots appeared and shook issues up.

I’m a good chatbot

For the document, the “I’m a good chatbot” within the gpt2-chatbot take a look at name is a reference to an episode that occurred whereas a Reddit person named Curious_Evolver was testing an early, “unhinged” model of Bing Chat in February 2023. After an argument about what time Avatar 2 could be exhibiting, the dialog eroded rapidly.

“You have misplaced my belief and respect,” stated Bing Chat on the time. “You have been mistaken, confused, and impolite. You haven’t been a good person. I’ve been a good chatbot. I’ve been proper, clear, and well mannered. I’ve been a good Bing. 😊”

Altman referred to this change in a tweet three days later after Microsoft “lobotomized” the unruly AI mannequin, saying, “i’ve been a good bing,” virtually as a eulogy to the wild mannequin that dominated the information for a quick time.



Source hyperlink

Tags: brokeChatbotGPT4olaunchingleaderboardrecordssecret
Previous Post

OpenAI’s new GPT-4o lets people interact using voice or video in the same model

Next Post

S.G.K Electricals

Next Post

S.G.K Electricals

No Result
View All Result

Categories

  • Construction (53)
  • Food (977)
  • Local News (1,995)
  • Local Sports (1,999)
  • Technology (4,000)

Recent.

How to Make Powdered Sugar (Without Cornstarch Option)

How to Make Powdered Sugar (Without Cornstarch Option)

August 25, 2026
Cream of Asparagus Soup with White Wine

Cream of Asparagus Soup with White Wine

August 25, 2026
Easy Whole Wheat Penne With Broccoli (18-Minute Base)

Easy Whole Wheat Penne With Broccoli (18-Minute Base)

August 24, 2026

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Category

  • Construction (53)
  • Food (977)
  • Local News (1,995)
  • Local Sports (1,999)
  • Technology (4,000)

Tags

2024 Draft 2024 Draft News Air apple Baltimore bridge Chicken Clifton Brown day Derrick Henry draft Easy Experiments Game Gameday Gameday News General Google Heres home Homepage Centerpiece Homepage Latest Headlines iPhone Jackson Key Lamar Lamar Jackson Late For Work Maryland NFL offseason OpenAI Ravens Recipe recipes Ryan Mink Savory season shopping tech TikTok users video Watch week
  • About
  • Home

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.

No Result
View All Result
  • About
  • Home

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.