Sign in

OpenAI's deep research can complete 26% of ‘Humanity’s Last Exam': What is it and what does it mean?

The recently-launched deep research AI agent by OpenAI is more than a quarter of the way to surpassing boundaries of human knowledge. Find out how

Updated on: Feb 12, 2025, 16:59:59 IST
By | Edited by
Share
Share via
  • facebook
  • twitter
  • linkedin
  • whatsapp
Copy link
  • copy link

Artificial intelligence critics argue that the technology may soon outsmart humans, which may lead to a ‘Terminator ’-style situation for humanity. For some, AI is already on its way to turning this future into a reality.

OpenAI's DeepResearch may soon become more intelligent than humans. Find out how (Reuters)
OpenAI's DeepResearch may soon become more intelligent than humans. Find out how (Reuters)

The deep research AI model, launched by ChatGPT-maker OpenAI earlier this month, has shown an over two-fold increase in performance over the next-best AI model in one of the world's toughest exams for large language models (LLMs) - Humanity's Last Exam.

What is the exam about?

Humanity's Last Exam is a recently released exam for AI models, also called large language models, like ChatGPT, Grok-2 and deep research. It is used to judge the performance of the AI model against a preset list of characteristics.

According to the people behind the exam, it was created as AI models are already scoring 90% accuracy on existing tests. This means that in a way, the scale to measure their performance is falling short. Thus, a larger scale was created in the form of Humanity's Last Exam.

The exam consists of 2,700 challenging questions, most of which are released to the public, over a hundred subjects.

Did OpenAI prove its dominance?

The Sam Altman-led company's AI models performed with varied accuracy in the exam. The company's model which performed with the least accuracy was 4o, which displayed an accuracy of 3.1% with a calibration error of 92.3%.

OpenAI's o1 model scored 8.8% accuracy and 92.8% calibration error while o3-mini (medium) and o3-mini (high) mediums scored 11.1% and 14% accuracy levels with 91.5% and 92.8% calibration error rates respectively.

OpenAI's newest model, deep research, scored a staggering 26.6% accuracy in Humanity's Last Exam. This is over two-fold more accuracy than the next-best performer, which is OpenAI's o3-mini (high) model.

How did other models fare?

According to the exam's website, which was last updated on February 11, Elon Musk's ambitious AI model Grok-2 scored a meagre 3.9% accuracy with a 90.8% calibration error. Another competitor, Anthropic's Claude 3.5 Sonnet model scored 4.8% accuracy and 88.5% calibration error.

Google's Gemini Thinking scored 7.2% accuracy and 90.6% calibration error. Chinese firm DeepSeek's R1 model, which caused a global stock rout for technology firms after it was launched last month, scored higher than all other competitors but couldn't outperform even the o3-mini (medium) model of OpenAI. It scored 8.6% accuracy with 81.4% calibration error.

What does deep research's score mean?

The performance showcased by OpenAI's deep research shows that the model can answer a wide range of analytical, subjective and objective questions with more accuracy than any of its competitors. It also means that the model is more capable of delivering well-rounded answers than other AI models.

This is likely because the model was released primarily to aid people in researching any topic of their choice without the hassle. According to its creators, deep research can conduct multi-step research on the internet for complex tasks in tens of minutes. The same task would otherwise take humans many hours.

  • HT News Desk
    ABOUT THE AUTHOR
    HT News Desk

    Follow the latest breaking news, major developments and agenda-setting stories from India and around the world with the newsdesk at Hindustan Times. Operating round the clock, the desk brings together experienced editors, reporters and correspondents to deliver fast, accurate and contextual reporting across subjects that influence public policy, governance, business, society and international affairs. The HT News Desk covers politics, elections, government policies, the economy, business and markets, science and technology, the environment, law and order, infrastructure, education, climate issues and geopolitics, while closely tracking developments across states, institutions and global capitals. The team also leads coverage of major breaking news events, policy announcements, court proceedings, natural disasters, public emergencies and significant international developments. Reports published by the newsdesk are based on information gathered from reporters on the ground, official statements, government agencies, court records, regulatory filings, recognised institutions and other authoritative sources. Stories undergo editorial scrutiny and verification processes to ensure accuracy, fairness and relevance, and are updated as events evolve and additional information becomes available. Whether covering a key political decision in New Delhi, an economic policy shift affecting millions, a landmark court ruling or a major global event, the HT News Desk aims to provide readers with reliable, fact-based journalism that delivers not only the latest developments but also the context and analysis needed to understand their wider implications.Read More

Stay updated with the latest Technology News, gadget launches, app updates, artificial intelligence and digital trends. Find reviews, comparisons and useful insights from the world of tech.