Grok-3 Beats DeepSeek-R1 at Reasoning, is as Capable as OpenAI’s o1 Pro: Karpathy

2 months ago 20

Published on February 18, 2025
In AI News

Andrej Karpathy, founder of Eureka Labs, and a former researcher at OpenAI was given early access to Grok-3.

Andrej Karpathy

xAI, the AI model maker headed by Elon Musk, unveiled its latest family of models, the Grok-3.

According to benchmarks, the Grok-3 outperforms several competing models and is also the first to score over 1400 on Chatbot Arena, a platform for comparing and evaluating AI models.

Grok-3 also offers reasoning (Think) capabilities and a deep research feature called DeepSearch.

Andrej Karpathy, founder of Eureka Labs, who was also once a part of OpenAI and Tesla, was given early access to Grok-3.

He shared a post on X detailing his experience. He revealed that the model performed well on complex tasks, such as creating a hex grid for the popular board game Settlers of Catan.

“Few models get this right reliably. The top OpenAI thinking models (e.g. o1-pro, at $200/month) get it too, but all of DeepSeek-R1, Gemini 2.0 Flash Thinking, and Claude do not,” he said.

Karpathy also uploaded OpenAI’s GPT-2 technical paper to estimate the number of flops required to train the model. He revealed that while Grok-3 and GPT-4o failed at this task, Grok-3, with thinking (reasoning), solved it ‘great’, and even OpenAI’s o1 Pro failed at the task.

“The impression overall I got here is that this is somewhere around o1-pro capability, and ahead of DeepSeek-R1, though, of course, we need actual, real evaluations to look at,” he added.

Karpathy also tested Grok-3’s DeepSearch capabilities, which he found comparable to Perplexity’s deep research but not yet at the level of that offered by OpenAI. He found that the model was hallucinating URLs that do not exist and reporting incorrect facts without providing citations.

“When I asked it to create a report on the major LLM labs and their amount of total funding and estimate of employee count, it listed 12 major labs but not itself (xAI),” he added.

After using the model for around 2 hours, he concluded by saying, “Grok 3 + thinking feels somewhere around the state of the art territory of OpenAI’s strongest models (o1-pro, $200/month), and slightly better than DeepSeek-R1 and Gemini 2.0 Flash Thinking.”

Others like Lex Fridman, who also received early access to the model, said, “My mind is blown, very impressive model,” in a post on X.

Supreeth Koundinya

Supreeth is an engineering graduate who is curious about the world of artificial intelligence and loves to write stories on how it is solving problems and shaping the future of humanity.