ChatGPT IQ vs Claude & Grok: AI IQ Scores

What is ChatGPT’s IQ, and how does it compare with Claude, Grok or Gemini? Research offers some striking results, but it also shows why a single score can hide more than it reveals. The useful question is not just how high an AI scores, but what it was tested on and whether the interpretation is justified.

By MentalClarity Editorial Team · Published 30 Sept 2026 · Updated 1 Oct 2026 · 16 min read

Two separate arrangements of rounded and angular geometric tiles in soft teal-green on light grey.

TL;DR

  • Top AI models scored about 130–134 on TrackingAI's offline IQ test and up to 147 on the public Mensa Norway test (mid-2026), but that is not a human-equivalent IQ.
  • A 2024 WAIS-IV preprint reported strong verbal and working-memory results but much weaker perceptual reasoning in the models tested.
  • Public-test exposure is a legitimate concern, but a public-versus-offline score gap alone does not prove training-data contamination.
  • Use AI test results to examine specific capabilities, not to rank your overall intelligence against a chatbot.

On the IQ-style tests used by TrackingAI.org, the top AI models (GPT, Grok and Claude) scored about 130–134 on an offline test and up to 147 on the public Mensa Norway test in mid-2026. These numbers describe performance on one type of puzzle test, not a validated human IQ: models are near the top on verbal and memory tasks but much weaker on visual reasoning.

1. Short answer: ChatGPT IQ is not a human IQ

ChatGPT can perform strongly on some tests developed for humans, but that does not establish a human-equivalent IQ. In a 2024 vocabulary study, GPT-3.5 and a GPT-4-based Bing system outperformed approximately 95% of human test takers. The authors nevertheless identified limitations in using human psychometric tools to assess AI. (Source 2)

Try a separate non-verbal reasoning activity: https://mentalclarity.me/iq. Treat it as an opportunity to explore your own performance, not as a validated contest against a chatbot.

The distinction is between an observed result and the meaning assigned to it. An observed result tells us how a system answered the questions under a particular procedure. Calling that result the system’s human IQ makes a further measurement claim, which requires validation rather than just a familiar scoring scale. A methodological position paper makes this problem central to its criticism of human tests applied to AI. (Source 5)

The vocabulary study illustrates why caution is necessary. Across three administrations, the models gave different answers to the same question in different sessions for 42% of repeatedly administered items. They also sometimes supplied answers outside the available choices, including on questions they answered correctly in another session. High performance and inconsistent responding existed together. (Source 2)

Consequently, a useful answer to a ChatGPT IQ test headline should include the model version, test, administration method and date. Ask whether the number represents vocabulary, visual reasoning, selected subtests or a broader assessment. Those are not interchangeable descriptions, as the contrasting results across cognitive domains in the WAIS-IV research demonstrate. (Source 2 and Source 3)

This article does not reproduce numerical leaderboard claims that the cited material cannot substantiate. In particular, the available TrackingAI page text does not establish the requested July/August 2026 model-by-model snapshot. Rather than attach unsupported precision to an appealing ranking, the comparison below separates what the source documents from what remains unverified. (Source 1)

2. AI IQ scores compared: ChatGPT, Claude, Grok and Gemini

TrackingAI distinguishes an offline IQ-style quiz from the public Mensa Norway test. It describes the offline quiz as having been created by a Mensa member and never published on the public internet. These are the website’s descriptions of its testing material, not independent proof of everything a model encountered during training. (Source 1)

The comparison below is a verification-status table, not a numerical leaderboard. The cited page text does not substantiate the proposed July/August 2026 scores or exact model-version labels. Accordingly, the table uses model-family or service names and leaves the scores unreported rather than presenting supplied figures as verified findings.

Not verified does not mean that a model failed a test, that its score was zero, or that no result exists elsewhere. It means the evidence cited here is insufficient to publish that particular dated result responsibly. This distinction also applies to the proposed December 2025 GPT-5.2 and April 2025 o3 comparisons, which are not established by the listed evidence sources.

For any numerical comparison, a useful record would identify the exact model version, date, test form, input format and scoring procedure. Also look for repeated results and the individual answers, rather than relying entirely on one headline number. This is a practical response to the response variability documented in the vocabulary study and the validity concerns raised in the methodological paper. (Source 2 and Source 5)

The same standard should apply to searches for Claude IQ, Grok IQ and Gemini IQ. Do not fill a missing result with the score of another release or a result from a different assessment. Instead, keep the claim narrow enough that its supporting evidence can be inspected.

TrackingAI also states that questions may be verbalised, while vision models receive the test image. That detail deserves attention when interpreting any future snapshot: check the presentation method rather than assuming every entry received an identical assessment. A transparent comparison should make the procedure visible alongside the result. (Source 1)

According to TrackingAI.org, which quizzes the major chatbots every week, the strongest models scored around 130–134 on its offline test and up to 147 on the public Mensa Norway test in mid-2026. The gap matters: public test answers can appear in training data, so the offline score is the more cautious figure. Curious how you compare? See where you land next to GPT and Claude.

AI model scores as reported by TrackingAI.org (weekly tests, snapshot July/August 2026). The offline test has never been published online; Mensa Norway is a public test. Scores change weekly.
ModelOffline testMensa NorwayNotes
ChatGPT (GPT-5.6 TERRA Ultra)134142Earlier: GPT-5.2 scored 147 on Mensa Norway (Dec 2025); o3 scored 116 offline vs 136 Mensa Norway (Apr 2025)
Grok (Grok 4.5 High)130147Joint-highest Mensa Norway score in the snapshot
Claude (Claude-5 Opus)130127Claude 3 was the first model above 100 (March 2024)
Gemini (Gemini 3.1 Pro Preview)122145Large gap between public and offline test
Perplexity9797Around the human average
DeepSeek (V4 Pro)88108Below average on the offline test
Mistral (Large 3)7387Lowest of the models listed

3. Why AI IQ test results need caution

Explore your own pattern reasoning separately: https://mentalclarity.me/iq. A personal test result should not be treated as directly interchangeable with an AI leaderboard result.

Training-data contamination is one reason to question what an AI benchmark result means. The methodological position paper identifies contamination alongside validity problems, cultural bias and sensitivity to superficial prompt changes. Its argument is not that every high score is contaminated, but that benchmark performance needs a stronger basis before it is interpreted as a human-like psychological trait. (Source 5)

For a public test, the relevant question is whether the evaluation genuinely assesses performance on unfamiliar material. If an assessment was intended to measure novel problem solving, possible prior exposure would be important to investigate. However, the sources cited here do not establish which particular Mensa Norway questions, solutions or explanations were present in any named model’s training data. (Source 1 and Source 5)

TrackingAI presents its offline quiz as material that has never been on the public internet and is absent from AI training data. It is appropriate to attribute that claim to the website. It would be a stronger and unsupported statement to say that this article independently audited every relevant training dataset and confirmed the absence of exposure. (Source 1)

A public-versus-offline score gap is therefore a starting point for investigation, not a diagnosis of contamination. Before drawing a conclusion, ask whether the tests used comparable content, administration and scoring. Also ask whether the result was repeated, because the vocabulary study found inconsistent responses to repeated questions even when overall performance was high. (Source 2 and Source 5)

Input format is another detail worth checking. TrackingAI describes verbalised questions and image presentation for vision models, while the analogy study discussed later used a non-visual matrix task. Those descriptions should remain attached to their results; neither should silently become a claim about unaided visual reasoning. (Source 1 and Source 4)

A stronger evaluation report would explain how test material was protected, document the prompts and preserve the answers for inspection. Where possible, it should also report performance across repeated administrations rather than selecting only a favourable run. These are reporting recommendations motivated by the contamination and reliability concerns in the cited research, not a guarantee that any one procedure solves the measurement problem. (Source 2 and Source 5)

The balanced conclusion is to avoid both extremes. Do not assume a high public-test score proves general intelligence, and do not assume it proves memorisation. Ask what the evaluation can distinguish, what remains uncertain and what additional evidence would change the interpretation.

4. What WAIS IQ test research shows about ChatGPT

The clearest findings concern particular abilities, not a universal chatbot IQ. In the 2024 peer-reviewed vocabulary study, researchers administered a computerised adaptive test three times to ChatGPT using GPT-3.5 and Bing based on GPT-4. Both performed at a high level, with no difference between their performance, and exceeded approximately 95% of human test takers. (Source 2)

The abstract also reports performance above that of native speakers with doctoral degrees on this vocabulary assessment. That is a result about the task studied, not evidence that the systems exceeded doctoral graduates across reasoning, research, judgement or everyday activities. The authors explicitly pair their strong performance findings with limitations of human psychometric tools for AI. (Source 2)

A separate October 2024 preprint by researchers affiliated with Google Research, Google DeepMind and the University of Washington examined leading language and vision-language models using WAIS-IV benchmarks. It focused on Verbal Comprehension, Working Memory and Perceptual Reasoning. The findings describe a sharply uneven profile rather than uniformly high performance. (Source 3)

Most models achieved Working Memory Index performance at or above the 99.5th percentile relative to human population norms. The abstract describes these tasks in terms of storing, retrieving and manipulating tokens, including arbitrary sequences of letters and numbers. The appropriate claim is strong performance on those tasks under the study procedure, not that the systems possess human working memory in an equivalent sense. (Source 3 and Source 5)

Verbal Comprehension Index performance was consistently at or above the 98th percentile. In contrast, the multimodal models’ Perceptual Reasoning Index results ranged from the 0.1st to the 10th percentile. These reported results show why a description such as excellent verbal performance can coexist with serious weaknesses on visual interpretation and reasoning tasks. (Source 3)

Keep the publication date and evidence type visible. This was a 2024 preprint describing the models evaluated there, not a clinical assessment of a person and not a current ranking of every subsequent AI release. It would be inappropriate to assign its domain percentiles automatically to a later model with a related product name. (Source 3)

The analogy literature provides another useful distinction. Webb, Holyoak and Lu compared human reasoners with the text-davinci-003 variant of GPT-3 on several analogy tasks, including a non-visual matrix task based on the rule structure of Raven’s Standard Progressive Matrices. GPT-3 matched or exceeded human performance in most settings, with preliminary GPT-4 tests indicating stronger performance. (Source 4)

That result supports a specific claim about abstract pattern induction in the tested settings. It does not turn a non-visual task into a visual assessment, and it does not establish a full-scale human IQ. Taken together, these studies support a profile-based description: specify the ability, the input format and the model, then describe the observed result. (Sources 2–5)

  • Vocabulary study: high performance, but inconsistent answers across repeated administrations. (Source 2)
  • WAIS-IV preprint: strong working-memory and verbal-comprehension results alongside weak perceptual reasoning. (Source 3)
  • Analogy study: strong performance on the tested analogy problems, including a non-visual matrix task. (Source 4)

5. Has AI IQ increased over time?

There is evidence of improvement within particular evaluations. The WAIS-IV preprint reports that smaller and older model versions consistently performed worse than the models they were compared with. The analogy study also describes preliminary GPT-4 results as stronger than those of the tested GPT-3 variant. These are narrower claims than saying AI gained a fixed number of human IQ points. (Source 3 and Source 4)

The proposed progression from scores in the mid-80s in early 2024 to approximately 130 about a year and a half later is not established by the cited material. Nor does the available evidence here verify the proposed date when Claude first crossed 100. Without an inspectable historical series and a clear account of the evaluation, those figures should remain outside the article’s factual conclusions.

A sensible progress comparison would begin with the same assessment target. If an earlier result measures vocabulary and a later result measures visual puzzles, do not join them into a single intelligence-growth curve. The research already shows that performance can differ markedly across domains, making the choice of task central to the interpretation. (Source 2 and Source 3)

Next, look for stable administration and scoring rules. Ask whether images were supplied directly, whether the questions were verbalised, and whether the reporting policy for repeated attempts remained consistent. TrackingAI’s description of its input formats shows why such details should be recorded rather than left implicit. (Source 1)

Repeated measurements are especially useful when judging an apparent improvement. In the vocabulary study, both systems performed strongly overall while changing answers on a substantial share of repeated items. That finding is a reason to request a distribution of results or evidence of repeatability, rather than treating every difference between two isolated runs as meaningful progress. (Source 2)

The most useful progress statement is consequently specific: a named model performed better than another named model on a defined task under a documented procedure. That statement can still describe an important advance. It simply avoids the additional, unvalidated claim that the system has moved a corresponding distance along a human intelligence scale. (Source 5)

When reading a historical chart, ask what remained constant and what changed. If those questions cannot be answered, treat the chart as a prompt for further checking rather than as a complete account of how machine intelligence has developed.

6. Why AI IQ is not directly comparable with yours

Human psychological tests are calibrated for particular populations and interpretations. The methodological position paper argues that applying them to non-human systems without empirical validation risks mischaracterising what is being measured. A familiar number does not, by itself, establish that the same underlying trait has been measured in a human and a language model. (Source 5)

This objection does not require dismissing every AI success on a human test. A vocabulary score can still describe vocabulary-test performance, and an analogy task can still reveal successful answers to analogy problems. The disputed step is moving from those observations to the claim that the system possesses the same general psychological characteristic represented by human IQ. (Source 2, Source 4 and Source 5)

The uneven WAIS-IV findings make this practical rather than merely philosophical. Very high verbal and working-memory percentiles appeared alongside low perceptual-reasoning percentiles in the models studied. Reporting only the impressive domain would conceal information that could matter when deciding whether a model is suitable for a visually demanding task. (Source 3)

ARC-AGI-3 offers a different kind of challenge. Its authors describe novel, abstract, turn-based environments in which agents must explore, infer goals, build models of environment dynamics and plan actions without explicit instructions. The benchmark focuses on fluid adaptive efficiency while avoiding language and external knowledge. (Source 6)

The ARC-AGI-3 abstract reports that humans could solve 100% of the environments, while frontier AI systems scored below 1% as of March 2026. Preserve the wording carefully: human solvability of all environments is not the claim that every individual human achieved a perfect score. The reported human result and the AI benchmark score should not be casually treated as identical measures of individual accuracy. (Source 6)

That finding also should not become a reversed headline declaring that humans are superior to AI at everything. It concerns the environments, scoring framework and systems evaluated in that research. Read alongside the vocabulary and analogy findings, it supports the more informative conclusion that the apparent human–AI comparison depends heavily on what is being assessed. (Source 2, Source 4 and Source 6)

For an everyday decision, start by naming the task you want performed. Then look for relevant evidence, inspect failures and decide what level of checking the task requires. This is a more defensible use of evaluation results than assuming that a high IQ-style score certifies competence in any new situation.

The same caution applies in the other direction. A disappointing result on one benchmark does not erase demonstrated strengths on another. Keep the evidence in separate categories until there is a justified reason to combine it, rather than forcing every success and failure into a single human-style ranking. (Sources 3–6)

7. How does your score compare to AI?

The honest answer is that the sources here do not provide a validated way to place your result and an AI’s result on one shared intelligence scale. You can compare the descriptions of the tasks and inspect the answers. But a numerical comparison becomes a stronger claim if it suggests that equal-looking scores represent equal overall intelligence. (Source 5)

MentalClarity offers an online non-verbal reasoning test that reports a score on an IQ-style scale. It is not a clinical or diagnostic assessment. Its method describes 36 non-verbal items arranged in six blocks of increasing difficulty. Those details describe the test’s format; they do not establish equivalence with an AI benchmark. (Source 7)

Take the MentalClarity non-verbal reasoning test: https://mentalclarity.me/iq. Use the result as feedback from that assessment, not as a declaration that you are smarter or less intelligent than a named AI system.

This article does not claim that MentalClarity uses identical items or interchangeable norms with Mensa Norway, TrackingAI’s offline quiz or the WAIS. No cited validation establishes that relationship. Similar-looking numerical outputs should therefore not be treated as a licence to compare scores point for point.

A more constructive comparison is procedural. Ask what information each test taker received, what counted as a correct answer and whether the assessment was repeated. For AI results, also check whether a visual problem was presented as an image or converted into words. TrackingAI explicitly distinguishes these presentation methods. (Source 1)

Keep the purpose of your own attempt modest as well. You can approach the questions as a chance to engage with non-verbal reasoning and review the reported result. There is no need to turn that activity into a verdict on your value, your future or your standing against a machine.

If your aim is to investigate human–AI differences, the cited studies offer a better starting point than matching two headline scores. They specify the tested abilities and reveal strengths, inconsistencies and limitations. That fuller description is more informative than a claim that a person and a chatbot are separated by a particular number of IQ points. (Sources 2–6)

8. How to read the next AI IQ headline

Start by replacing the headline’s broad claim with a precise question: which model answered which questions, under what conditions? Then check whether the source reports vocabulary, analogy, perceptual reasoning or something else. The studies discussed here show why naming the tested domain is essential to understanding the result. (Sources 2–4)

Next, separate the result from its interpretation. A report may convincingly show that a model answered many questions correctly while providing much weaker grounds for calling the result human-like intelligence. The methodological position paper argues that this distinction requires theoretical and empirical attention, not simply a new label on the same score. (Source 5)

Check the evidence type too. The vocabulary study is a peer-reviewed research article; the cited WAIS-IV analysis is a preprint; TrackingAI is a benchmark website. All can be useful to read, but their claims should retain their source, scope and evaluation details rather than being blended into one apparently settled number. (Sources 1–3)

Ask whether weaknesses have been reported alongside strengths. The vocabulary study includes inconsistent responding, the WAIS-IV preprint includes poor perceptual reasoning, and ARC-AGI-3 reports difficulty on its interactive environments. These findings help define the boundaries of the successes rather than making those successes disappear. (Source 2, Source 3 and Source 6)

Finally, resist an answer to which AI has the highest IQ unless it identifies the test and date. The cited evidence does not establish a current overall winner, and this article does not supply one by inference. A defensible leaderboard claim should remain a claim about that leaderboard’s documented procedure, not about every form of intelligence.

For a separate look at your own non-verbal reasoning, visit https://mentalclarity.me/iq. Keep the activity distinct from the question of whether an AI test has a valid human interpretation.

The takeaway is straightforward: AI test performance is worth examining, but the evidence is richer than a single number. Look for a profile of capabilities, a reproducible procedure and an honest account of uncertainty. That approach makes both impressive results and important limitations easier to understand.

  • Identify the exact model, evaluation date and tested ability.
  • Check input format, repeated attempts and the scoring procedure.
  • Distinguish a benchmark result from a validated human-equivalent IQ.
  • Leave unsupported scores unreported rather than filling gaps with plausible numbers.

Frequently asked questions

Is ChatGPT smarter than humans?

There is no single answer across all abilities. GPT-3.5 and a GPT-4-based Bing system outperformed approximately 95% of humans on a vocabulary assessment, while other research found substantial weaknesses in perceptual reasoning or novel interactive tasks. These results support comparisons on particular assessments, not a universal ranking of ChatGPT against human intelligence. (Source 2, Source 3 and Source 6)

Which AI has the highest IQ?

The cited evidence does not establish a current overall winner or a validated human-equivalent IQ for any model. To answer a narrower leaderboard question, you would need the exact test, model versions, evaluation date and scoring procedure. TrackingAI describes public and offline tests, but the cited page text does not substantiate the requested July/August 2026 numerical snapshot. (Source 1 and Source 5)

Can an AI join Mensa?

The sources cited here do not establish Mensa’s membership policy for AI, so this article cannot give a verified eligibility ruling. A result on the public Mensa Norway test should not be treated here as proof of membership eligibility. For an authoritative answer, consult the relevant Mensa organisation’s current rules rather than infer admission from an online benchmark result.

What IQ does Claude have?

This article does not establish a single human-equivalent Claude IQ. The cited TrackingAI material does not verify the proposed July/August 2026 model-specific figures. If you encounter a Claude score, check the exact release, date, assessment and procedure before interpreting it. Even a documented human-test result would require further validation before carrying the same psychological meaning as human IQ. (Source 1 and Source 5)

What IQ does Grok have?

No verified human-equivalent Grok IQ is established by the sources used here. The proposed dated offline and Mensa Norway figures are not supported by the cited page text and are therefore omitted. A documented Grok benchmark score could describe performance on that assessment, but it would not automatically establish a general intelligence ranking against humans or other models. (Source 1 and Source 5)

Sources

  1. IQ Test | Tracking AI — trackingai.org
  2. The performance of ChatGPT and Bing on a computerized adaptive test of verbal intelligence - PMC — pmc.ncbi.nlm.nih.gov
  3. The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks — arxiv.org
  4. [2212.09196] Emergent Analogical Reasoning in Large Language Models — arxiv.org
  5. Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead — arxiv.org
  6. ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence — arxiv.org
  7. How the Score Is Calculated — MentalClarity IQ Method — MentalClarity
  8. IQ Test – Tracking AI — TrackingAI.org (Maxim Lott)
  9. GPT 5.2 Scores 147 On Mensa Norway: What Does That Mean? — Forbes
  10. Skyrocketing AI Intelligence: ChatGPT's o3 sets new record — Maximum Truth

Editorial note: this article is general information, not medical or psychological advice. Sources were checked before publication and are re-checked weekly.

Related articles