OpenAI published a table on Sept. 5, 2025. In it, the o4-mini model answered 99 percent of the questions on the SimpleQA evaluation and got 75 percent of all questions wrong1. A year and 24 days later, President Donald Trump signed Executive Order 14434, which directs federal agencies to write "Super Intelligence" in place of "Artificial Intelligence"2. Six questions follow from those two documents. The AILately.com Desk answers them below and traces every figure to a page it opened.
What is the scariest thing about AI?
Hallucinations are the problem OpenAI's research post says "remains stubbornly hard to fully solve." The post defines them in one sentence: "By this we mean instances where a model confidently generates an answer that isn't true."1. Its example involves one of its own authors. Asked for the title of Adam Tauman Kalai's PhD dissertation, a widely used chatbot gave three different answers. All three were wrong1.
On SimpleQA, the evaluation the post uses as its example, o4-mini abstained on 1 percent of questions. It answered 24 percent correctly and 75 percent wrongly1. OpenAI's newer gpt-5-thinking-mini abstained on 52 percent, answered 22 percent correctly and answered 26 percent wrongly1. The older model scored two points higher on accuracy and made nearly three times as many errors.
Chart 1
OpenAI's o4-mini got 75% of SimpleQA questions wrong, against 26% for gpt-5-thinking-mini, which abstained on 52%
SimpleQA, share of questions by outcome, percent, from OpenAI's table of Sept. 5, 2025
Each row is one outcome: the blue dot is o4-mini, the magenta dot is gpt-5-thinking-mini, and the rod between them is the gap.
Source: OpenAI [1]. Chart by The AILately.com Desk.
The numbers behind this chart
| Item | OpenAI o4-mini | gpt-5-thinking-mini |
|---|---|---|
| Abstained | 1% | 52% |
| Answered correctly | 24% | 22% |
| Answered wrongly | 75% | 26% |
OpenAI attributes the gap to grading. Its post continues: "Nonetheless, accuracy-only scoreboards dominate leaderboards and model cards, motivating developers to build models that guess rather than hold back."1. The post summarizes a paper by Adam Tauman Kalai of OpenAI, Ofir Nachum, Santosh Vempala and Edwin Zhang. Its abstract says "language models are optimized to be good test-takers, and guessing when uncertain improves test performance," and adds that hallucinations "persist even in state-of-the-art systems and undermine trust."3.
Sam Altman, OpenAI's chief executive, said on an episode of the OpenAI Podcast in June 2025: "People have a very high degree of trust in ChatGPT, which is interesting because AI hallucinates. It should be the tech that you don't trust that much."45. Read as strategy, the sentence assigns the remedy to the listener. The adjustment Altman names is the user's trust. OpenAI's post supplies the model-side half: "Hallucinations remain a fundamental challenge for all large language models, but we are working hard to further reduce them."1.
Dario Amodei, Anthropic's chief executive, took the other side of the trust question on May 22, 2025. He spoke at Code with Claude, Anthropic's first developer event, in San Francisco: "It really depends how you measure it, but I suspect that AI models probably hallucinate less than humans, but they hallucinate in more surprising ways," Amodei said in response to a TechCrunch question15. TechCrunch added that the claim is hard to verify, because most hallucination benchmarks compare models with each other and leave humans out15. Earlier that month, TechCrunch noted, a lawyer representing Anthropic had apologized in court after Claude created citations with wrong names and titles15.
Workers describe the same pattern. Hao-Ping Lee of Carnegie Mellon University and six researchers at Microsoft Research surveyed 319 knowledge workers. The workers shared 936 examples of generative AI at work, and the team found that "higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking."6. Those figures rest on self-reports, so the finding describes how workers characterize their own habits.
What is super intelligence, and how is it different from AI?
Executive Order 14434 states the federal position in its first section: "As these capabilities continue to improve, they increasingly represent not merely artificial intelligence, but a new era of Super Intelligence."2. President Donald Trump signed it on Sept. 29, 2026. Section 3 defines the new term by pointing at existing law. Super Intelligence and SI mean the technologies and systems covered by the statutory definition of artificial intelligence in 15 U.S.C. 9401(3)2. For now, the two terms cover the same systems.
Congress wrote that statute in 2021. It describes a machine-based system that makes predictions, recommendations or decisions for a given set of human-defined objectives7. The order gives the White House science adviser, Michael Kratsios, 60 days to propose legislative language that may change the definition. Sixty days from Sept. 29 is Nov. 28, 202628.
The research definition predates the order by 28 years. Nick Bostrom of Oxford University wrote in a paper first published in 1998: "By a 'superintelligence' we mean an intellect that is much smarter than the best human brains in practically every field, including scientific creativity, general wisdom and social skills."9. Bostrom's standard is a comparison with the best human brains in practically every field. The order's standard is a citation of the statute.
Ana-Maria Stanciuc, editor-in-chief of The Next Web, compared the order with the federal renamings of a gulf and a lake8. Both went through the Board on Geographic Names, and she concluded: "A technology has no such board; its name belongs to the labs, universities, courts, and regulators that use it, and none of them work from the White House style guide."8.
How do I keep AI from doing my thinking?
The Carnegie Mellon and Microsoft survey finds where critical thinking goes once workers use AI: "GenAI shifts the nature of critical thinking toward information verification, response integration, and task stewardship."6. Each of the three activities maps to one habit:
- Verification: check each factual claim against a primary source before using it. Ask the model to mark each claim it is unsure of, a behavior the OpenAI post says its Model Spec endorses: "it is better to indicate uncertainty or ask for clarification than provide confident information that may be incorrect."1.
- Integration: write your own conclusion first, then ask the model to argue against it and name its weakest points.
- Stewardship: keep the sign-off on your own work, because the next section shows courts assigning it to people.
Who answers for a wrong AI answer?
On June 22, 2023, federal Judge P. Kevin Castel in New York sanctioned two lawyers and their firm $5,000. The firm, Levidow, Levidow & Oberman, had cited six cases that ChatGPT invented1011. Castel's order distinguished the tool from the duty: "Technological advances are commonplace and there is nothing inherently improper about using a reliable artificial intelligence tool for assistance," Castel wrote. "But existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings."10.
Castel found bad faith. It came after the cases were questioned: the lawyers kept standing by them, and he cited "shifting and contradictory explanations" from attorney Steven A. Schwartz10. Levidow, Levidow & Oberman said it would comply, disputed the bad-faith finding and weighed an appeal10.
Which jobs does AI change by 2030?
The World Economic Forum's Future of Jobs Report 2025 draws on more than 1,000 employers. Together they employ more than 14 million workers across 22 industry clusters and 55 economies12. Through 2030, the report projects job creation and destruction equal to 22 percent of today's jobs. New roles equal to 14 percent of current employment, or 170 million jobs, are offset by displacement of 8 percent, or 92 million jobs. Net growth comes to 7 percent, or 78 million jobs12.
Chart 2
The WEF expects 170 million jobs created and 92 million displaced by 2030, a net gain of 78 million
Projected job creation and displacement, 2025 to 2030, millions of jobs
The first column is jobs created, the magenta column is jobs displaced, and the last column is the net gain once both are counted.
Source: World Economic Forum [12]. Chart by The AILately.com Desk.
The numbers behind this chart
| Item | Value |
|---|---|
| Jobs created | 170 million |
| Jobs displaced | −92 million |
| Net growth | 78 million |
AI is one driver of that total among several. Among technology trends, AI and information processing technologies drew the highest share: 86 percent of employers expect them to transform their business by 203012. Clerical and secretarial workers see the largest decline in absolute numbers, including cashiers and ticket clerks and administrative assistants and executive secretaries. The fastest-declining roles include postal service clerks, bank tellers and data entry clerks12. Technology-related roles grow fastest in percentage terms, including Big Data Specialists, Fintech Engineers and AI and Machine Learning Specialists. Frontline roles add the most jobs in absolute terms: farmworkers, delivery drivers, construction workers, salespersons and food processing workers12.
Read as planning data, the net figure of 78 million combines two groups12. Clerical workers appear among the report's largest declines. Technology, care and frontline roles appear among its largest gains12.
What did Stephen Hawking warn about AI?
Stephen Hawking, the theoretical physicist, told the BBC in December 2014: "The development of full artificial intelligence could spell the end of the human race."13. The question came from a revamp of the speech technology Hawking used, which contains a basic form of AI. According to the BBC, he called such primitive forms already very useful while fearing the consequences of creating something that can match or surpass humans13.
He described what such a system would do: "It would take off on its own, and re-design itself at an ever increasing rate," Hawking said. "Humans, who are limited by slow biological evolution, couldn't compete, and would be superseded."13. Rollo Carpenter, creator of the chatbot Cleverbot, answered in the same article: "I believe we will remain in charge of the technology for a decently long time and the potential of it to solve many of the world problems will be realised."13.
Almost two years later, at the Oct. 19, 2016 launch of the Leverhulme Centre for the Future of Intelligence in Cambridge, Hawking named both outcomes: "In short, the rise of powerful AI will be either the best, or the worst thing, ever to happen to humanity. We do not yet know which."14.
Each side of the trust question needs a different fact. Amodei's reading needs a comparison of human and model error on the same questions15. OpenAI's table reports model error alone: 75 percent for o4-mini and 26 percent for gpt-5-thinking-mini on SimpleQA1. Altman's reading needs users who trust the answers anyway, and the Carnegie Mellon and Microsoft survey ties confidence in AI to less critical thinking6. Hawking's reading needs full artificial intelligence, which Carpenter expected within a few decades in the same BBC article13.
By the numbers
- Error rate: 75 percent for o4-mini on SimpleQA, 26 percent for gpt-5-thinking-mini1.
- Abstentions: gpt-5-thinking-mini declined 52 percent of SimpleQA questions, o4-mini 1 percent1.
- Survey: 319 knowledge workers shared 936 examples of generative AI at work6.
- Sanction: $5,000 on two lawyers and their firm for six cases ChatGPT invented, June 22, 20231011.
- Jobs: the World Economic Forum expects 170 million created and 92 million displaced by 2030, a net gain of 78 million12.
- Deadline: Executive Order 14434 gives the science adviser 60 days, to Nov. 28, 2026, to propose a definition of Super Intelligence2.
What to watch
- OpenAI's next evaluation table shows whether abstention rises and error rates fall on SimpleQA, the two measures in the Sept. 5, 2025 post1.
- The science adviser's legislative proposal shows whether Super Intelligence receives a definition that departs from the statute's definition of AI27.
- A follow-up study of worker behavior would test the link between confidence and critical thinking that the Carnegie Mellon and Microsoft survey drew from self-reports6.
- Head-to-head scoring of people and models on identical questions would check Amodei's estimate15.
Sources
- OpenAI, "Why language models hallucinate," OpenAI, Sept. 5, 2025, https://openai.com/index/why-language-models-hallucinate/
- Donald J. Trump, "Inaugurating the Era of Super Intelligence (Executive Order 14434)," The White House, Sept. 29, 2026, https://www.whitehouse.gov/presidential-actions/2026/09/inaugurating-the-era-of-super-intelligence/
- Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala and Edwin Zhang, "Why Language Models Hallucinate," arXiv, Sept. 4, 2025, https://arxiv.org/abs/2509.04664
- Celeste Martin, "Expert warns 'we should not rely on AI too much' as hallucinations persist," Eyewitness News, June 29, 2025, https://www.ewn.co.za/2025/06/29/expert-warns-we-should-not-rely-on-ai-too-much-as-hallucinations-persist
- Vuyile Madwantsi, "ChatGPT's CEO on AI trust: a surprising confession you need to hear," Independent Online, June 27, 2025, https://iol.co.za/lifestyle/2025-06-27-chatgpts-ceo-on-ai-trust-a-surprising-confession-you-need-to-hear/
- Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks and Nicholas Wilson, "The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers," Microsoft Research, April 2025, https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/
- U.S. Congress, "15 U.S. Code Section 9401: Definitions," Legal Information Institute, Jan. 1, 2021, https://www.law.cornell.edu/uscode/text/15/9401
- Ana-Maria Stanciuc, "First a gulf, then a lake, now a technology. Trump orders US agencies to call AI 'super intelligence'," The Next Web, Sept. 30, 2026, https://thenextweb.com/news/trump-super-intelligence-executive-order-signed
- Nick Bostrom, "How Long Before Superintelligence?," nickbostrom.com, 1998, https://nickbostrom.com/superintelligence
- Larry Neumeister, "Lawyers submitted bogus case law created by ChatGPT. A judge fined them $5,000," Associated Press, June 22, 2023, https://www.clickondetroit.com/tech/2023/06/22/lawyers-submitted-bogus-case-law-created-by-chatgpt-a-judge-fined-them-5000/
- Meghan L. Newcomer, "Judge Castel Sanctions Lawyers Who Submitted Fake Cases Generated by ChatGPT," Steptoe, June 28, 2023, https://www.steptoe.com/en/news-publications/sdny-blog/judge-castel-sanctions-lawyers-who-submitted-fake-cases-generated-by-chatgpt.html
- World Economic Forum, "Future of Jobs Report 2025," World Economic Forum, January 2025, https://www3.weforum.org/docs/WEF_Future_of_Jobs_Report_2025.pdf
- Rory Cellan-Jones, "Stephen Hawking warns artificial intelligence could end mankind," BBC News, Dec. 2, 2014, https://www.bbc.com/news/technology-30290540
- University of Cambridge, "'The best or worst thing to happen to humanity' - Stephen Hawking launches Centre for the Future of Intelligence," University of Cambridge, Oct. 19, 2016, https://www.cam.ac.uk/research/news/the-best-or-worst-thing-to-happen-to-humanity-stephen-hawking-launches-centre-for-the-future-of
- Maxwell Zeff, "Anthropic CEO claims AI models hallucinate less than humans," TechCrunch, May 22, 2025, https://techcrunch.com/2025/05/22/anthropic-ceo-claims-ai-models-hallucinate-less-than-humans/





