Guide
What Is AGI?
A plain-English guide to artificial general intelligence: the competing definitions from OpenAI, Google DeepMind, Anthropic and academia, the benchmarks used to measure progress, who says AGI has arrived and who says it has not, the forecasts, and how to read an AGI claim.
Quick answer: AGI (artificial general intelligence) is a hypothetical AI system that can match or exceed human ability across most or all cognitive tasks, rather than excelling at one narrow job. There is no single agreed definition and no agreed test: OpenAI’s Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work”, while Google DeepMind’s “Levels of AGI” paper grades it on a five-level scale and the Center for AI Safety’s 2025 definition scores it against ten areas of human cognition. The question is openly contested: OpenAI President Greg Brockman greeted the 3 September 2026 launch of GPT-6 Astra with “Welcome to the AGI era” (Axios), and Nvidia CEO Jensen Huang posted “AGI has arrived” on 6 September (Yahoo Finance), but the ARC Prize Foundation, whose benchmark OpenAI cited, said it lacked “evidence to call this AGI yet” (The Next Web). The one thing to carry into any AGI headline: whether AGI has “arrived” depends almost entirely on which definition the speaker is using.
This guide lays out the competing definitions, the benchmarks, the forecasts and the arguments, with each claim attributed to whoever made it. It takes no position on when AGI will arrive or whether it already has. For the short definition alongside the rest of the vocabulary, see the AI glossary; for which models are actually strongest today, see best AI models.
What is AGI?
AGI, short for artificial general intelligence, is an AI system with broad, flexible competence: it can learn and perform a wide range of intellectual tasks at or above the level of a capable human, including tasks it was not specifically built for.
The word doing the work is general. Almost every AI system in use before the 2020s was narrow AI: a chess engine, a spam filter or a protein-folding model that is superhuman at one task and useless at everything else. AGI describes the opposite profile, a single system that can switch between writing a contract, debugging software, planning a project and learning a new skill without being rebuilt for each one.
Three terms get used interchangeably and should not be:
- Narrow AI is a system built for, and competent at, a specific task or domain. It exists today in huge numbers.
- AGI is a system broadly competent across most cognitive tasks at roughly human level. Whether any system qualifies is the live dispute this page covers.
- Superintelligence (sometimes ASI, artificial superintelligence) is a system that substantially exceeds the best humans across essentially all cognitive domains. No one claims it exists; several companies now name it as their goal.
Most definitions of AGI are restricted to cognitive work, meaning tasks that can be done with a computer, and leave out physical tasks such as plumbing or surgery. That single choice explains a large share of the disagreement about timelines: a definition that requires robotics sets a much higher bar than one that only counts work done at a keyboard.
The competing definitions of AGI
There are at least eight definitions of AGI in active use by labs, researchers and investors, and they do not agree on what counts. The table sets them side by side.
| Definition | Who uses it | What counts as AGI | Measurable? |
|---|---|---|---|
| Economic output | OpenAI Charter | ”Highly autonomous systems that outperform humans at most economically valuable work” | Partly; “most” and “economically valuable” are not quantified |
| Contractual profit trigger | Microsoft–OpenAI agreement, as reported by The Information in December 2024 (summarised by Simon Willison) | Systems able to generate roughly $100 billion in profits for early investors | Yes, financially, but it measures revenue rather than intelligence; superseded in April 2026 |
| Levels of AGI | Google DeepMind, Morris et al., November 2023 | Performance and generality graded from “Emerging” to “Superhuman” against skilled adults | Designed to be; no official scoring of current models published |
| Exceptional AGI | Google DeepMind, “An Approach to Technical AGI Safety and Security”, April 2025 | A system matching at least the 99th percentile of skilled adults on a wide range of non-physical tasks; the paper calls its development by 2030 “plausible” | In principle |
| Cognitive profile score | Hendrycks et al., “A Definition of AGI”, October 2025 | ”An AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult”, scored across ten equally weighted cognitive domains | Yes; produces a 0–100% score |
| Powerful AI | Anthropic (Dario Amodei, “Machines of Loving Grace”, October 2024; Anthropic submission to the White House OSTP, March 2025, quoted by Redwood Research) | Intellectual capability matching or exceeding Nobel Prize winners across most disciplines; a “country of geniuses in a datacenter” | Partly; described by capabilities rather than a test |
| Human-level AI / world models | Yann LeCun, AMI Labs (Brown Daily Herald, April 2026) | “The ability to accomplish new tasks you’ve never been exposed to and solve new problems without any prior training”, including in the physical world | Not formally |
| Task-by-task | Jensen Huang, Nvidia (August 2026 earnings call, reported by Yahoo Finance) | “For many tasks, we could say that we’ve already achieved AGI” | No; it does not set a threshold |
Two things stand out from the table. First, the definitions measure different things: economic output, percentile performance, a cognitive profile, or contract revenue. A system can satisfy one and fail another at the same time. Second, the strictest academic definitions and the loosest executive ones are years apart in what they imply, which is why a CEO saying “AGI is here” and a researcher saying “AGI is a decade away” can both be internally consistent.
Google DeepMind’s Levels of AGI
The most widely cited attempt to make AGI measurable is the Levels of AGI framework published by Google DeepMind researchers (Meredith Ringel Morris, Shane Legg and colleagues) in November 2023. It rates a system on two axes, performance and generality, and treats AGI as a ladder rather than a finish line.
| Level | Performance threshold (vs skilled adults) | Narrow example given in the paper | General example given in the paper |
|---|---|---|---|
| Level 0: No AI | Not applicable | Calculator software | Human-in-the-loop computing such as Amazon Mechanical Turk |
| Level 1: Emerging | Equal to or somewhat better than an unskilled human | Simple rule-based systems | ChatGPT, Bard and Llama 2 (as of 2023) |
| Level 2: Competent | At least the 50th percentile | Siri, Alexa | Not yet achieved (as of 2023) |
| Level 3: Expert | At least the 90th percentile | Grammarly, image generators such as DALL-E 2 | Not yet achieved (as of 2023) |
| Level 4: Virtuoso | At least the 99th percentile | Deep Blue, AlphaGo | Not yet achieved (as of 2023) |
| Level 5: Superhuman | Outperforms 100% of humans | AlphaFold, AlphaZero | Artificial superintelligence; not yet achieved |
The paper’s own classification put 2023’s chatbots at “Emerging AGI”, Level 1. Google DeepMind has not published an official re-grading of 2026 models on this scale, so any claim that a current model is “Level 3” or “Level 4” is a third party’s interpretation rather than the framework authors’ verdict.
The Center for AI Safety’s AGI score
The most quantitative definition is “A Definition of AGI” (October 2025), led by Dan Hendrycks of the Center for AI Safety with co-authors including Yoshua Bengio and Gary Marcus. It grounds AGI in the Cattell-Horn-Carroll theory of human cognitive abilities, splits cognition into ten domains weighted 10% each (including knowledge, reading and writing, maths, reasoning, working memory, long-term memory storage and retrieval, visual and auditory processing, and speed), and scores a model from 0% to 100%.
The paper scored GPT-4 at 27% and GPT-5 at 57% (agidefinition.safe.ai). The authors’ central finding was that models had a “jagged” profile: strong on knowledge, reading, writing and maths, and near zero on long-term memory storage, meaning the ability to learn something new and retain it across sessions. The project site lists scores for those two models only; later models have no official score.
How would we know? The benchmarks used to measure progress towards AGI
No benchmark measures AGI directly. Each of the tests below measures one slice of general capability, and every one of them has been described by its own creators as insufficient proof of AGI on its own.
| Benchmark | What it measures | Why it is cited in AGI debates | Notable result | Main limitation |
|---|---|---|---|---|
| ARC-AGI-3 (ARC Prize) | Learning the rules of unfamiliar interactive puzzle games with no instructions | Designed to test skill acquisition on novel problems, which humans find easy | GPT-6 Astra: 62.7% on ARC Prize’s standard harness, 99.9% on OpenAI’s provider adapter (ARC Prize, 3 Sep 2026) | Closed, deterministic game worlds; ARC Prize says it “does not represent the complexity and open-endedness of the real world” |
| ARC-AGI-2 | Abstract visual reasoning puzzles | The predecessor test, built to resist memorisation | GPT-5.6 Sol reached 92.5% | Now close to saturated |
| AGI definition score (CAIS) | Breadth across ten human cognitive domains | The only test that outputs an explicit “percent of AGI” figure | GPT-5: 57% (October 2025) | No published scores for 2026 models |
| GDPval (OpenAI) | Real work products in 44 occupations, graded blind by industry professionals | Closest test to OpenAI’s own “economically valuable work” definition | At launch on 25 September 2025, GPT-5-high matched or beat experts on 40.6% of tasks and Claude Opus 4.1 on just under half (YourStory) | Single, well-specified tasks; no iteration, ambiguity or long projects |
| METR time horizon (METR, March 2025) | The length of software task, measured in human working time, a model completes 50% of the time | Tracks autonomy rather than knowledge | METR’s original paper found the horizon doubling roughly every seven months from 2019 to 2024 | Software tasks only; METR has published notes on how sensitive the figures are to modelling assumptions |
| Humanity’s Last Exam | Expert-written, closed-answer questions across academic fields | Probes the edge of expert knowledge | GPT-6 Astra: 57.2% with tools, below Claude Fable 5.1’s 65.0% on OpenAI’s own table | Knowledge and reasoning recall, not autonomy or learning |
Benchmarks keep being beaten faster than their designers expect, which is one side’s evidence that AGI is close. The other side’s point is the same fact read differently: each saturated benchmark turned out not to capture what people meant by general intelligence, so a new one had to be built.
ARC-AGI-3 and the GPT-6 Astra episode
The clearest recent case study of how AGI claims and benchmarks interact is ARC-AGI-3, which launched on 25 March 2026 as 135 interactive game environments that the system must explore without being told the rules or the goal (Pasquale Pillitteri).
| Date | Event | Source |
|---|---|---|
| 25 Mar 2026 | ARC-AGI-3 launches. Human testers solve 100% of environments; the best AI score is 0.37% (Gemini 3.1 Pro) | Pasquale Pillitteri |
| 24 Jul 2026 | Claude Opus 5 scores 30.16% | ARC Prize results |
| 3 Sep 2026 | GPT-6 Astra scores 62.7% on ARC Prize’s standard harness at a cost of $26,098, and 99.9% on OpenAI’s own “provider adapter” harness at $19,817 | ARC Prize |
| 3 Sep 2026 | Greg Brockman: “Welcome to the AGI era” | Axios |
| 3–6 Sep 2026 | ARC Prize states that saturating the benchmark “would not represent proof of achieving AGI”; co-founder Mike Knoop says the foundation lacks “evidence to call this AGI yet” | ARC Prize; The Next Web |
The gap between 62.7% and 99.9% comes from the harness, the software layer that connects a model to the test, not from a different model. ARC Prize also reported that Astra, at maximum reasoning, used fewer actions than the human baseline on 96.0% of levels (ARC Prize). Both readings are defensible: Astra made a very large jump on a test designed to resist exactly that, and the organisation that built the test still declined to call it AGI.
Has AGI arrived? Who says what
Prominent figures disagree openly about whether AGI exists today. The table records each position in the speaker’s words, with the date and source.
| Person | Role | Position | Date | Source |
|---|---|---|---|---|
| Greg Brockman | President, OpenAI | ”Welcome to the AGI era”; asked whether Astra is AGI, “For me personally, I do think we’re there” | 3 Sep 2026 | Axios; iTWire |
| Sam Altman | CEO, OpenAI | OpenAI is “not quite yet” at AGI but expects an internal system he would call AGI by the end of 2026 | Aug 2026 (TIME) | The Decoder |
| Mark Chen | Chief Research Officer, OpenAI | OpenAI is “80% of the way” to AGI | Aug 2026 (TIME) | The Decoder |
| Jensen Huang | CEO, Nvidia | ”AGI has arrived” | 6 Sep 2026 | Yahoo Finance |
| Mike Knoop | Co-founder, ARC Prize Foundation | No “evidence to call this AGI yet” | Sep 2026 | The Next Web |
| Demis Hassabis | Chair, Google DeepMind (chief executive at the time) | AGI “could arrive around 2030”, possibly 2029 “or even sooner” | May 2026 (Axios) | Computerworld |
| Yann LeCun | Founder, AMI Labs; former Meta chief AI scientist | ”If you are interested in human-level AI, don’t work on LLMs” | Apr 2026 | Brown Daily Herald |
| Andrej Karpathy | OpenAI co-founder, independent researcher | Agents that work like a capable employee are roughly a decade away; current models “just don’t work” reliably enough | Oct 2025 (Dwarkesh Podcast) | The Neuron |
Two details matter when reading this table. Altman and Brockman are both OpenAI leaders and did not use the same words about the same model, which shows how loosely “AGI” is applied even inside one company. And Huang’s 6 September post followed his August earnings-call remark that AGI “for many tasks” was already achieved, a task-by-task standard rather than the “most economically valuable work” standard in OpenAI’s Charter.
Independent benchmarking at launch did not show a general jump of the size the AGI language implied: Artificial Analysis measured GPT-6 Astra level with its predecessor GPT-5.6 Sol on its Intelligence Index at launch (61 each), while OpenAI priced it at $10/$50 per million tokens, 2.5 times Sol (The Next Web). Our GPT-6 Astra page and best AI models ranking track how it compares on current index versions.
When will AGI arrive? The forecasts
Forecasts for AGI range from “already here” to “decades away”, and the spread is mostly explained by which definition each forecast resolves on. We report them without endorsing any.
| Forecast | Type | AGI estimate | Definition it resolves on | Source |
|---|---|---|---|---|
| Metaculus, “date of artificial general intelligence” | Crowd forecast | Median January 2033; 25% by 2029 (July 2026) | A demanding multi-part test that includes a robotics component | Metaculus |
| Metaculus, “weakly general AI” | Crowd forecast | Median June 2028 (July 2026) | A lighter test with no robotics | Metaculus |
| Grace et al. survey of 2,778 AI researchers | Expert survey (fielded 2023) | 50% chance of “high-level machine intelligence” by 2047; 10% by 2027 | Unaided machines able to do every task better and more cheaply than humans | arXiv, January 2024 |
| AAAI 2025 Presidential Panel survey of 475 respondents | Expert survey | 76% said scaling up current approaches is “unlikely” or “very unlikely” to produce AGI | Not tied to a date | The Decoder |
| Anthropic | Lab statement | ”Powerful AI systems will emerge in late 2026 or early 2027” | Anthropic’s “powerful AI” definition | March 2025 OSTP submission, quoted by Redwood Research |
| Google DeepMind (Demis Hassabis) | Lab leader | ”Five to 10 years” (January 2026, Davos); “around 2030”, possibly sooner (May 2026) | Human-level AGI across cognitive tasks | Fortune; Computerworld |
| OpenAI (Sam Altman) | Lab leader | An internal system he would classify as AGI by the end of 2026 | OpenAI Charter definition | The Decoder |
Two caveats apply. Crowd forecasts move; the Metaculus figures above are a July 2026 snapshot, taken before the GPT-6 Astra launch. And researcher surveys lag: the largest one was fielded in 2023, before reasoning models and agents existed in their current form.
Why credible people disagree about AGI
The disagreement is not mainly about facts. Most experts agree on what current models score on current benchmarks. They disagree on four things.
1. What the word means
A definition built on economic output, a definition built on human cognitive breadth and a definition that requires physical-world competence produce three different answers about the same model on the same day. The CAIS definition was written specifically because, in its authors’ framing, the lack of a concrete definition lets the goalposts move in both directions.
2. Whether scaling current methods is enough
One camp, which includes the leaders of OpenAI and Anthropic, expects continued scaling of large language models, plus reinforcement learning and agent tooling, to reach AGI. The opposing camp holds that large language models are missing something structural. Yann LeCun argues that systems trained on text lack world models and are “completely helpless when it comes to the physical world” (Brown Daily Herald). Demis Hassabis sits between them, saying in January 2026 that “maybe we need one or two more breakthroughs” (Fortune). The AAAI 2025 survey suggests the academic research community leans sceptical, with 76% doubting that scaling alone is sufficient.
3. Benchmarks versus reliability
Models now beat expert humans on many closed tests while still failing at tasks a new employee would handle, such as remembering last week’s instructions or recovering from an unexpected error in a long project. The CAIS paper found near-zero long-term memory storage in GPT-5 despite a 57% overall score. Andrej Karpathy calls the remaining work the “march of nines”: each step from 90% to 99% to 99.9% reliability takes as much effort as the last (The Neuron). For the practical side of this gap, see what is agentic AI.
4. Incentives
Declaring or denying AGI has had direct financial consequences. Until April 2026, an AGI declaration by OpenAI could change Microsoft’s rights to OpenAI technology (see the next section). AGI language also features in fundraising and hardware sales: Jensen Huang’s “AGI has arrived” post named the roughly 100,000 Nvidia Grace Blackwell chips used to train GPT-6 Astra. None of this makes a claim false, but it is a reason to ask which definition a speaker is using and what follows if they are believed.
AGI and money: the Microsoft–OpenAI AGI clause
The most concrete use of the word AGI in industry was contractual, not scientific. The timeline below is drawn from Simon Willison’s summary of the public record and our OpenAI provider page.
| Date | What changed |
|---|---|
| 2019 | Microsoft invests in OpenAI and licenses its “pre-AGI” technology; rights change once AGI is achieved, as defined by the OpenAI Charter |
| Dec 2024 | The Information reports a previously undisclosed financial definition: AGI is reached when OpenAI’s systems can generate about $100 billion in profits for its earliest investors |
| Oct 2025 | A restructured agreement says any OpenAI AGI declaration “will now be verified by an independent expert panel”; Microsoft’s IP rights run until the panel verifies AGI or through 2030, whichever comes first |
| 27 Apr 2026 | A new agreement makes revenue-share payments to Microsoft continue through 2030 “independent of OpenAI’s technology progress”, which ends the AGI trigger in practice (CNBC) |
The practical result is that an OpenAI declaration of AGI no longer changes the financial terms with its largest partner, according to the April 2026 terms as reported. Microsoft has meanwhile pursued its own goal under different language: in November 2025 Microsoft AI chief Mustafa Suleyman announced an MAI Superintelligence Team aimed at what the company calls “humanist superintelligence” (Fortune). See our Microsoft provider page.
AGI vs narrow AI vs superintelligence
The three categories differ in breadth and in level. The table is the quickest way to keep them apart.
| Narrow AI | AGI | Superintelligence (ASI) | |
|---|---|---|---|
| Breadth | One task or domain | Most or all cognitive tasks | Essentially all cognitive domains |
| Level | Can be superhuman within its domain | Roughly human level, at least that of a skilled or well-educated adult | Substantially beyond the best humans |
| Exists today? | Yes, widely | Contested (see above) | No one claims it does |
| Examples | AlphaFold, a spam filter, a chess engine | Disputed; GPT-6 Astra is the system most often named in 2026 claims | None |
| Who names it as a goal | Not applicable | OpenAI (Charter), Google DeepMind | Microsoft (“humanist superintelligence”), Meta (Meta Superintelligence Labs), Safe Superintelligence Inc. |
Note that today’s frontier models do not sit neatly in any column. A model such as Claude Opus 5.5 or GPT-6 Astra is far more general than classic narrow AI and outperforms most humans on many expert exams, yet still fails at some tasks most adults find routine. The CAIS paper’s word for this is “jagged”.
Why AGI matters beyond the label
Whether or not a model gets called AGI, the capabilities the debate is about already shape policy and safety work. Three examples, each dated:
- Safety frameworks. OpenAI rated GPT-6 Astra at its “critical” cybersecurity threshold, meaning it can find and exploit previously unknown vulnerabilities without explicit human direction, and restricted those capabilities to trusted testers at launch on 3 September 2026 (Axios).
- Evaluation standards. Google DeepMind launched the DeepMind Institute in September 2026, chaired by Demis Hassabis, to publish differing views on AGI; its inaugural essays include a proposal for a US-led standards body to evaluate frontier models (TechCrunch).
- Jobs and the economy. OpenAI’s GDPval was built to measure AI against professional work in 44 occupations because “economically valuable work” is the core of OpenAI’s own AGI definition (YourStory).
In practice, for someone choosing a tool today, the AGI label changes nothing about which model is best for a task. That question is answered by task-specific evidence: see best AI models, best AI for coding and best AI agents.
How to read an AGI claim: a five-question checklist
When a company, executive or headline says AGI has arrived, or is close, five questions separate the substance from the label.
- Which definition? Economic output (OpenAI Charter), percentile performance (Google DeepMind), cognitive breadth (CAIS), or “many tasks”? If none is stated, the claim cannot be checked.
- Which evidence? A named benchmark with a published score is checkable; “it feels different” is not.
- Who ran the test, and how? Vendor-reported scores and independent scores can differ sharply. GPT-6 Astra scored 62.7% on ARC-AGI-3 under ARC Prize’s standard harness and 99.9% under OpenAI’s own harness (ARC Prize).
- What does the benchmark’s creator say? ARC Prize, whose test OpenAI cited, declined to call GPT-6 Astra AGI.
- What follows if the claim is believed? Contracts, fundraising, hardware sales and regulation can all turn on the word. That does not make a claim wrong, but it is context the reader is owed.
Related reading
For the one-line definition of AGI, superintelligence, alignment and 100-plus other terms, see the AI glossary. For how today’s models actually rank, best AI models, plus the model pages for GPT-6 Astra, Claude Opus 5.5 and Claude Opus 5. For the agent capabilities at the centre of the reliability argument, what is agentic AI and best AI agents. For the companies making the claims, our provider pages for OpenAI, Anthropic, Google, Microsoft, Meta and xAI.
Frequently asked questions
What is AGI in simple terms?
AGI, or artificial general intelligence, is an AI system that can do most intellectual tasks as well as a capable human, rather than being good at only one thing. A chess engine is narrow AI because it can only play chess; an AGI could play chess, write a report, fix software and learn a new skill without being rebuilt. There is no single agreed definition, so whether any current system counts is disputed.
What does AGI stand for?
AGI stands for artificial general intelligence. The word “general” distinguishes it from narrow AI, which is built for one task or domain. The term was popularised in the early 2000s by researchers including Shane Legg, now chief AGI scientist at Google DeepMind, and Ben Goertzel.
Has AGI been achieved yet?
There is no consensus that AGI has been achieved. OpenAI President Greg Brockman said “Welcome to the AGI era” at the launch of GPT-6 Astra on 3 September 2026 and Nvidia CEO Jensen Huang posted “AGI has arrived” on 6 September, while OpenAI CEO Sam Altman said OpenAI was “not quite yet” there and the ARC Prize Foundation said it lacked evidence to call GPT-6 Astra AGI. The answer depends on which definition is used.
Is GPT-6 Astra AGI?
GPT-6 Astra, released by OpenAI on 3 September 2026, is the model most often named in 2026 AGI claims, but it is not agreed to be AGI. It scored 62.7% on ARC-AGI-3 under ARC Prize’s standard harness and 99.9% under OpenAI’s own harness, and ARC Prize said saturating its benchmark “would not represent proof of achieving AGI”. Independent testing by Artificial Analysis at launch placed it level with its predecessor GPT-5.6 Sol on general intelligence. See our GPT-6 Astra page.
What is OpenAI’s definition of AGI?
OpenAI’s Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work”. A separate, financial definition in OpenAI’s agreement with Microsoft, reported by The Information in December 2024, tied AGI to systems able to generate about $100 billion in profits; that trigger lost its effect when the April 2026 agreement made Microsoft’s revenue share independent of OpenAI’s technology progress.
What is the difference between AGI and AI?
AI is the broad field of building systems that perform tasks requiring intelligence, and nearly all AI in use is narrow: built for specific tasks. AGI is a specific, disputed milestone within AI where one system is broadly competent across most cognitive tasks at roughly human level. Every AGI would be AI, but almost no AI is AGI.
What is the difference between AGI and superintelligence?
AGI means matching human ability across most cognitive tasks; superintelligence means substantially exceeding the best humans across essentially all of them. No one claims superintelligence exists. Microsoft, Meta and Safe Superintelligence Inc. have each named superintelligence, rather than AGI, as their stated goal.
When will AGI arrive?
Forecasts range from 2026 to the 2040s and depend on the definition used. Sam Altman expects an internal OpenAI system he would call AGI by the end of 2026; Anthropic predicted “powerful AI” in late 2026 or early 2027; Demis Hassabis said in May 2026 that AGI could arrive around 2030; the Metaculus community median for its demanding AGI question was January 2033 in July 2026; and a 2023 survey of 2,778 AI researchers put a 50% chance on high-level machine intelligence by 2047.
How is AGI measured?
No single test measures AGI. The most cited measures are the ARC-AGI benchmarks, which test learning on novel puzzles; the Center for AI Safety’s AGI definition, which scores a model across ten human cognitive domains (GPT-5 scored 57% in October 2025); OpenAI’s GDPval, which grades real work products in 44 occupations; and METR’s time-horizon measure, which tracks how long a task a model can complete autonomously. Each creator says its test alone cannot prove AGI.
Do AI researchers think current models will lead to AGI?
Many are sceptical. In the AAAI 2025 Presidential Panel survey, 76% of 475 respondents said scaling up current AI approaches is “unlikely” or “very unlikely” to produce AGI. Yann LeCun argues that large language models cannot reach human-level intelligence without world models trained on sensory data, while leaders at OpenAI and Anthropic expect continued scaling to get there.
Is AGI dangerous?
AGI-level capability raises risks that the labs themselves treat as serious, including misuse in cyberattacks and systems pursuing goals their developers did not intend. OpenAI rated GPT-6 Astra at its “critical” cybersecurity threshold and restricted those capabilities to trusted testers at launch, and Google DeepMind published a technical AGI safety paper in April 2025. How large the risks are is as contested as the timeline.
What is the ARC-AGI benchmark?
ARC-AGI is a series of benchmarks from the ARC Prize Foundation designed to test whether an AI can learn new skills on problems it has never seen, which humans find easy. ARC-AGI-3, launched on 25 March 2026, uses 135 interactive games with no instructions; the best AI score at launch was 0.37%, Claude Opus 5 reached 30.16% in July 2026, and GPT-6 Astra reached 62.7% on the standard harness in September 2026.