Grok-3 Beats DeepSeek-R1 at Reasoning, is as Capable as OpenAI’s o1 Pro: Karpathy

2 months ago 20
  • Published on February 18, 2025
  • In AI News

Andrej Karpathy, founder of Eureka Labs, and a former researcher at OpenAI was given early access to Grok-3. 

Andrej Karpathy

xAI, the AI model maker headed by Elon Musk, unveiled its latest family of models, the Grok-3. 

According to benchmarks, the Grok-3 outperforms several competing models and is also the first to score over 1400 on Chatbot Arena, a platform for comparing and evaluating AI models. 

Grok-3 also offers reasoning (Think) capabilities and a deep research feature called DeepSearch. 

Andrej Karpathy, founder of Eureka Labs, who was also once a part of OpenAI and Tesla, was given early access to Grok-3. 

He shared a post on X detailing his experience. He revealed that the model performed well on complex tasks, such as creating a hex grid for the popular board game Settlers of Catan. 

“Few models get this right reliably. The top OpenAI thinking models (e.g. o1-pro, at $200/month) get it too, but all of DeepSeek-R1, Gemini 2.0 Flash Thinking, and Claude do not,” he said. 

Karpathy also uploaded OpenAI’s GPT-2 technical paper to estimate the number of flops required to train the model. He revealed that while Grok-3 and GPT-4o failed at this task, Grok-3, with thinking (reasoning), solved it ‘great’, and even OpenAI’s o1 Pro failed at the task. 

“The impression overall I got here is that this is somewhere around o1-pro capability, and ahead of DeepSeek-R1, though, of course, we need actual, real evaluations to look at,” he added. 

Karpathy also tested Grok-3’s DeepSearch capabilities, which he found comparable to Perplexity’s deep research but not yet at the level of that offered by OpenAI. He found that the model was hallucinating URLs that do not exist and reporting incorrect facts without providing citations. 

“When I asked it to create a report on the major LLM labs and their amount of total funding and estimate of employee count, it listed 12 major labs but not itself (xAI),” he added. 

After using the model for around 2 hours, he concluded by saying, “Grok 3 + thinking feels somewhere around the state of the art territory of OpenAI’s strongest models (o1-pro, $200/month), and slightly better than DeepSeek-R1 and Gemini 2.0 Flash Thinking.” 

Others like Lex Fridman, who also received early access to the model, said, “My mind is blown, very impressive model,” in a post on X.

Picture of Supreeth Koundinya

Supreeth Koundinya

Supreeth is an engineering graduate who is curious about the world of artificial intelligence and loves to write stories on how it is solving problems and shaping the future of humanity.

Association of Data Scientists

GenAI Corporate Training Programs

India's Biggest Women in Tech Summit

Mar 20 and 21, 2025 | 📍 J N Tata Auditorium, Bengaluru

Download the easiest way to
stay informed

Subscribe to The Belamy: Our Weekly Newsletter

Biggest AI stories, delivered to your inbox every week.

Rising 2025 | DE&I in Tech & AI

Mar 20 and 21, 2025 | 📍 J N Tata Auditorium, Bengaluru

AI Startups Conference.
April 25, 2025 | 📍 Hotel Radisson Blue, Bangalore, India

Data Engineering Summit 2025

15-16 May, 2025 | 📍 Taj Yeshwantpur, Bengaluru, India

MachineCon GCC Summit 2025

19-20th June 2025 | 📍 ITC Grand, Goa

17-19 September, 2025 | 📍KTPO, Whitefield, Bangalore, India

India's Biggest Developers Summit Nimhans Convention Center, Bangalore

discord icon

Our Discord Community for AI Ecosystem.

Read Entire Article