
Artificial general intelligence, or AGI, is usually supposed to mean artificial intelligence capable of succeeding across a broad range of intellectual problems instead of being designed for one type of task. The definition becomes disputed as soon as we ask what counts as “general.” Some researchers emphasize breadth of performance. Others emphasize learning new skills, adapting to unfamiliar environments, achieving goals, or matching human cognitive abilities.
Humans do not literally possess artificial general intelligence. Instead, we are the primary example from which the concept of general intelligence is drawn. Yet human intelligence is not universal, omniscient, perfectly rational, or self-sufficient. It is bounded, embodied, dependent on prior knowledge, supported by tools and culture, and remarkably good at adapting to unfamiliar situations.
That makes the question “Have we achieved AGI?” much harder than comparing an AI system with a human scorecard. What if defining “general” becomes considerably more difficult than counting the number of things an intelligence can do?
Why AGI became difficult to recognize
To say that AGI has come to the forefront in 2026 would be an understatement.
When ARC Prize launched ARC-AGI-3 on March 25, its new benchmark placed an agentic AI system in unfamiliar interactive environments without instructions, stated rules, or explicit goals. The AI agent had to experiment, infer how the environment worked, discover what counted as success, and apply what it learned. Humans were already capable of solving these environments. Frontier AI models initially scored only 0.51 percent.
Less than six months later, the picture changed dramatically. ARC Prize reported on September 3 that OpenAI’s frontier model, GPT-6 Astra, scored 62.7 percent with its standard harness and 99.9 percent with a provider-specific adapter that preserved reasoning state and managed long contexts. Under the latter setup, it also used fewer actions than the median tested human on 96 percent of evaluated levels.
One tempting conclusion is that the benchmark had been conquered, artificial intelligence passed the test, and therefore AGI had arrived. ARC Prize, however, explicitly rejected that inference. Its own definition of AGI concerns acquiring any skill a human can acquire with comparable efficiency, while ARC-AGI-3 covers a bounded family of deterministic environments with closed-ended mechanics.
Its authors specifically stated that benchmark saturation would not amount to proof of AGI because the real world is vastly more open-ended.
There is a deeper philosophical lesson here: a benchmark does not define intelligence merely because researchers call it an intelligence benchmark.
At best, it operationalizes a theory of intelligence. Necessarily, passing it establishes only what the benchmark successfully measures.
The difference between 62.7 and 99.9 percent also raises an older philosophical problem in a startlingly concrete form. Where does the intelligent system end? Is it the base model alone? The model plus memory? The model plus a reasoning-state mechanism? The model plus tools, retrieval, software, and an external notebook?
Human beings have faced these same questions all along.
AGI is a name for a problem, not a settled scientific quantity
The acronym can make AGI sound like a precisely defined engineering specification. It is nothing of the kind.
There is no universally accepted construct called general intelligence that researchers have first agreed how to measure and then straightforwardly reproduced in artificial systems. Instead, several concepts and definitions are grouped under the same name. They may overlap substantially, but ultimately answer different questions.
One definition asks whether a system can perform many different intellectual tasks. Another asks whether it can learn unfamiliar tasks efficiently. Another measures whether it can achieve goals across many environments. Another compares it with human cognitive abilities. Another asks whether it can perform enough economically valuable work to substitute for human labor.
A stronger philosophical conception may also require understanding, intentionality, autonomous reasoning, or genuine knowledge.
OpenAI’s longstanding Charter, for example, defines AGI operationally as “highly autonomous systems that outperform humans at most economically valuable work.” That definition combines breadth, performance, autonomy, and an economic threshold.
A 2024 Google DeepMind paper by Meredith Ringel Morris and colleagues takes a different approach. Instead of looking for one magical cutoff, it separates generality, meaning breadth of capabilities, from performance, meaning how well those capabilities are exercised, while treating autonomy as another prerequisite. Their resulting framework classifies artificial general intelligence into levels ranging emerging to competent, expert, exceptional, and superhuman levels rather than employing a single yes-or-no dichotomy.
That approach adds both nuance and clarity to the AGI debate. A system can be remarkably broad but mediocre in many areas. Another can be extraordinary at one intellectual activity while possessing little transferable competence elsewhere. A third can perform many tasks when carefully prompted by a human while lacking the autonomy to discover and pursue such tasks itself.
Calling all three simply “intelligent” hides the differences we actually want to know more about.
The AGI problem is older than artificial intelligence
The ideas behind artificial general intelligence long predate computers.
In Part V of René Descartes’s Discourse on Method, published in 1637, Descartes imagines machines constructed to resemble human beings. He proposes two tests by which we could still distinguish them from genuine people.

One concerns language. A machine might emit words in response to particular stimuli, but Descartes thought it could not arrange signs and language flexibly enough to express thoughts appropriately across arbitrary circumstances. As such, machines were deemed incapable of doing what humans routinely do: compose meaningful new responses that fit whatever situation arises.
The second concerns adaptability. Machines could surpass us at particular activities, he argued, while still depending on mechanisms designed for particular actions. Human reason, by contrast, functions as what his translator John Veitch calls an “universal instrument” applicable across circumstances.
Descartes was wrong about what artificial intelligence would eventually be able to do with language. Yet the structure of his challenge sounds uncannily modern: how do we distinguish an enormous collection of specialized responses from a capacity that generalizes to situations its designer did not individually anticipate?
Alan Turing returned to a similar problem in “Computing Machinery and Intelligence” in 1950. His approach was to resist beginning with an abstract analysis of the word thinking. He replaced “Can machines think?” with the imitation game, asking whether a machine’s conversational performance could become indistinguishable from a human’s.
Turing thereby moved from quantifying some mystical hidden essence toward measuring and testing observable competence. Yet his paper is far richer than the modern slogan “pass the Turing test” implies. He considers consciousness, originality, mathematical limitations, mistakes, learning, and the possibility of machines surprising their creators. His learning-machine discussion is particularly relevant to AGI because it recognizes that intelligence may be found in a system’s capacity to acquire competence rather than in a fixed inventory of preprogrammed accomplishments.
The modern term artificial general intelligence arose partly because ordinary artificial intelligence had become associated with systems capable of highly specialized feats. In the introduction to Ben Goertzel and Cassio Pennachin’s 2007 volume Artificial General Intelligence, specialized AI is contrasted with systems intended to solve complicated problems across different domains and learn autonomously. Their historical framing recovers an ambition that had animated early endeavors in artificial intelligence long before the acronym AGI became fashionable.
AGI, in other words, can be considered a modern engineering name for an older philosophical question about the difference between having many abilities and possessing a generally applicable capacity to reason and learn.

Four different things can hide inside the word “general”
Much confusion disappears once four dimensions are separated: breadth of acquired competence, performance depth, adaptability, and autonomy.
These things often correlate, but they do not share the same meaning.
A medical specialist may know far more medicine than the average person while having ordinary intelligence outside his field. A novice with exceptional reasoning ability may initially know little about a subject while learning it extremely rapidly. A computer system may contain enormous amounts of learned information while requiring substantial external direction. Another system may know less but independently explore an environment and teach itself what it needs.
An AGI definition must decide which of these characteristics it is trying to quantify.
▪ General intelligence as breadth of successful behavior
The simplest conception of general intelligence asks how many different things a system can do.
This is close to the intuition behind Turing’s test and the sort of claims made in publications like the 2023 “Sparks of Artificial General Intelligence” paper by Sébastien Bubeck et al. Working with an early version of GPT-4, the authors investigated mathematics, coding, medicine, law, psychology, vision, composition, and other domains. They argued that the system exhibited enough breadth and depth to count as an early, incomplete form of AGI.
Their evidence was genuinely interesting. Earlier AI systems had often displayed a familiar pattern: astonishing strength inside one domain followed by abrupt incompetence outside it. GPT-4 appeared capable of moving between many intellectual domains through a common language interface.
But breadth of competence leaves open an important second question: where did those abilities come from?
That question leads to François Chollet.
▪ Chollet: intelligence is different from accumulated skill
François Chollet’s “On the Measure of Intelligence” provides one of the most useful corrections to popular AGI discussion.
Suppose two different systems can both solve 10,000 problems. The first needed extensive training on problems closely resembling all 10,000. The second received a small number of examples, was able to infer the underlying structure, and generalized from those examples to tackle all other problems.
Their resulting skill sets might appear identical when measured in terms of problems solved. Their intelligence, in the relevant sense, does not.
Chollet therefore argues that measuring skill alone is inadequate. Performance on a task depends heavily on prior knowledge and prior experience. If enough task-relevant information can be supplied through data, training, engineering, or built-in assumptions, very impressive competence can be produced, or at least mimicked, without demonstrating equally impressive generalization.
His proposal connects intelligence with skill-acquisition efficiency. Roughly: how much new ability can a system acquire from a given amount of relevant experience, given its prior knowledge?
This changes how we should interpret large pretrained AI models. Their extraordinary breadth is real competence, yet the amount of prior experience encoded by training is also extraordinary. A system that answers thousands of questions across hundreds of subjects has demonstrated breadth of learned ability. To infer equally extraordinary general intelligence, we need evidence about how efficiently it acquires genuinely new skills outside what previous experience prepared it to do.

This is why increasingly difficult benchmarks keep moving toward tasks designed to prevent memorized competence from doing all the work. ARC-AGI, for example, developed directly from this conception.
▪ Legg and Hutter: intelligence across environments
Shane Legg and Marcus Hutter attack the problem from another direction in “Universal Intelligence: A Definition of Machine Intelligence”.
Their informal summary is elegantly broad: intelligence concerns an agent’s ability to achieve goals across a wide range of environments.
They then construct a mathematical measure in which an agent’s expected performance is evaluated across possible computable environments, with simpler environments weighted more heavily than very complicated ones. The project tries to define intelligence without surreptitiously leaping from “being intelligent” to “being similar to a twenty-first-century human.”
That is a genuine advance in conceptual tidiness. An alien intelligence, digital intelligence, octopus, or radically nonhuman machine need not be good at essay writing, college entrance examinations, or human office work to be intelligent.
This, however, also introduces a degree of abstraction since their formal measure of intelligence depends on idealized machinery and quantities that are not straightforwardly computable in practice. It also treats intelligence primarily through successful goal-directed interaction. Questions about subjective experience or intrinsic understanding enter only if they change observable behavior relevant to that measure.
Legg and Hutter therefore offer a theory of effective agency across environments, not a complete philosophy of mind.
▪ Human-level definitions
Most public discussion takes a fourth route to defining intelligence: compare machines with human beings.
This sounds straightforward because humans already exist, and we have developed all sorts of measurements to track their performance in intellectual, linguistic and goal-oriented endeavors. But that may be exactly where things turn treacherous.
A recent attempt by Dan Hendrycks and a large group of coauthors defines AGI as matching the cognitive versatility and proficiency of a well-educated adult, then operationalizes the comparison through ten cognitive domains influenced by Cattell-Horn-Carroll psychometric theory. The authors report highly uneven or “jagged” capability profiles in contemporary AI systems.
This offers a more disciplined target than “roughly as smart as us.” Yet it also reveals a hidden problem: there is no single human performance level. Additional questions pop up at once.
Which human?
A randomly selected adult? A university graduate? A trained professional? The median skilled worker within every tested occupation? The best living mathematician when mathematics is tested, the best programmer when programming is tested, and a physician when medicine is tested?
The last approach produces a strange benchmark. “Human-level intelligence” becomes the combined expertise of many humans rather than the abilities of any actual human being. A machine could therefore fail a supposedly human-level threshold while already being more cognitively capable across domains than any individual human who has ever lived.
This tells us that human comparison is useful, but the comparison class has to be clearly stated.
Psychometric g is not the same thing as AGI
The word general creates another source of confusion. Psychology already has a famous concept called general intelligence, usually represented by g.
This concept is related to AGI but most definitely not identical to it.
Charles Spearman’s 1904 paper “General Intelligence, Objectively Determined and Measured” helped establish the empirical problem that later became associated with g. Performance across apparently different cognitive tests tends to correlate. People who perform well on one intellectually demanding task are, statistically, more likely to perform well on other tasks.
Spearman proposed that cognitive performance contains a general component shared across activities as well as components specific to particular tasks.
Later psychometrics became much more elaborate. John B. Carroll’s 1993 Human Cognitive Abilities synthesized hundreds of factor-analytic datasets into a hierarchical model. At the highest level sits a very general factor. Beneath it are broad abilities involving areas such as fluid reasoning, acquired knowledge, memory and learning, visual processing, auditory processing, retrieval, and cognitive speed. More specific abilities appear below those.
That hierarchy is empirically informative. It tells us something about the structure of variation among human cognitive abilities.
But consider what g actually means in this context. Researchers give many people several tests. People who score well on one cognitive test tend, on average, to score well on others. Factor analysis identifies a latent statistical dimension that explains part of that covariance.
Now compare this with the questions involved in defining AGI. There we take one system and ask whether it can learn, reason, adapt, and act across many different tasks and environments.
These are different directions of analysis.

Psychometric g asks why different people’s performances across tasks correlate. AGI generality asks how one agent’s competence extends across tasks.
A person could therefore possess the ordinary human cognitive architecture that generates psychometric g while still knowing almost nothing about hundreds of domains. Conversely, an artificial system might achieve strong scores on a battery modeled after human intelligence tests through mechanisms very different from those responsible for human individual differences, including psychometric g.
This does not make psychometrics irrelevant to AGI. To the contrary: if psychometric evidence explains general intelligence in human beings, it is an essential piece of the puzzle in quantifying artificial general intelligence. However, the question remains open whether artificial systems exhibit something genuinely equivalent to human psychometric g, or whether they just mimic the patterns from which psychometric g can be derived.
As such, psychometric g cannot function as a ready-made definition of AGI.
So do humans actually possess general intelligence?
Yes, they do, in the sense most relevant to AGI. Ordinary human beings can enter unfamiliar situations, identify objects and goals, learn new rules from sparse experience, transfer knowledge between domains, acquire skills never encountered during childhood, reason by analogy, use language to obtain instruction, reorganize their strategies after failure, and construct new tools when existing abilities are insufficient.
A child can eventually learn algebra despite no human ancestor evolving a dedicated algebra module. Adults can learn and improve at newly invented games. A medieval mind could technically learn software development if transported into the modern world and given enough time and instruction. Human cognitive architecture does not need a genetic update whenever a new intellectual activity is invented.
That is an impressive form of generality, but still much less than universality.
▪ Human intelligence is bounded
Herbert Simon’s “A Behavioral Model of Rational Choice” attacks an idealized conception of rational agents possessing immense information, stable preferences, and enough computational ability to find optimal choices.
Real organisms have limited time, limited knowledge, limited memory, and limited computational capacity. Rational behavior must work within those constraints.
This idea became known as bounded rationality.
Even the most intelligent person does not normally enumerate every possible plan, calculate every consequence, assign exact probabilities to every future state, and choose the mathematically optimal path. Such exhaustive reasoning would often be computationally impossible and impractical for any organism existing and acting in the real world.
Generality therefore cannot mean optimal reasoning.
Allen Newell and Herbert Simon’s Human Problem Solving develops the same insight through models of people as information-processing systems. Problem solvers construct an internal problem space, representing possible states and transformations, and then use heuristics to search it. Means-ends analysis, for example, reduces an immense space of possibilities by focusing attention on differences between the present state and the goal.
They illustrate this point with the famous nine-dot problem. In it, problem-solvers are asked to connect nine dots with four uninterrupted straight lines. People commonly impose an additional constraint that these lines must remain within the square formed by the dots, even though the instructions themselves never state this. Within that self-imposed representation, no solution is possible.
Once the problem-solver allows the lines to extend beyond the self-imposed boundary, the solution becomes easy to find. For Newell and Simon, the lesson is that problem solving can fail before the search for a solution even begins: the way a person represents the problem determines which potential solutions can be seen at all.

Intelligence therefore includes more than searching efficiently. It includes figuring out what problem one is actually solving. That ability is central to AGI.
▪ Human beings know how to do things they cannot completely explain
A second limitation of simple computational pictures comes from Michael Polanyi’s The Tacit Dimension.
Polanyi’s celebrated formulation is that “we can know more than we can tell.” For instance, a person recognizes a familiar face while being unable to state a complete rule describing how such recognition occurred. A skilled pianist produces coordinated movements without consciously specifying each muscular command. An experienced diagnostician may notice that something is wrong before being able to reconstruct every cue responsible for his judgment.
Tacit knowledge is not mystical knowledge. It is competence whose operative structure is not fully available to explicit verbal expression.
This became highly relevant to artificial intelligence because early symbolic programs often assumed that intelligent performance would be reproduced by discovering the right explicit representations and rules. Yet much of human expertise looks less like consulting a minute instruction manual and more like acquiring a structured sensitivity to situations. Human intelligence, like machine learning, often relies on pattern recognition without requiring full awareness of all the prior factors that produce the patterns in question.
Modern machine learning has, in one sense, moved closer to Polanyi than classical AI ever did. Neural systems can acquire useful internal representations without programmers explicitly specifying every rule. But this does not settle whether tacit competence is equivalent in machines and humans.
▪ Dreyfus and the problem of relevance
Hubert Dreyfus constructed one of the most sustained philosophical critiques of classical artificial intelligence around this problem.
His argument in What Computers Still Can’t Do was never simply “machines will always be stupid.” He attacked particular assumptions behind early symbolic AI, especially the hope that ordinary human understanding could be reconstructed by explicitly representing facts, rules, and context.
Human beings rarely confront the world as an enormous database of equally available facts from which they consciously calculate relevance. Experience changes what becomes salient. The skilled chess player sees a dangerous configuration. The experienced driver notices the pedestrian. Any person entering an ordinary room ignores millions of possible details without first deducing which details can safely be ignored.
This problem is especially severe for attempts to model (artificial) intelligence on the foundations of a purely explicit rule system because rules themselves require decisions about when they are relevant. More rules can generate more relevance questions.
Contemporary neural models weaken some of Dreyfus’s specific attacks because their operation differs sharply from the hand-coded symbolic systems he criticized. Statistical learning can acquire patterns, similarities, and context sensitivity without representing them as explicit rules.
Yet Dreyfus’s deeper challenge survives: Can a system reliably determine what is relevant in open-ended situations whose important features were not specified in advance?
Human cognition does not stop cleanly at the brain
The question “Do humans possess general intelligence?” becomes even stranger when we decide what counts as the human cognitive system.
Andy Clark and David Chalmers’s “The Extended Mind” begins by asking where the mind ends and the rest of the world begins. Their examples include manipulating physical pieces to solve problems, doing arithmetic with pen and paper, and relying on notebooks or other reliably available information stores.
Their proposal contains an unavoidable observation: human intellectual performance depends extensively on external scaffolding.
Try calculating a complicated tax return without writing anything down. Try reconstructing modern physics without books. Try navigating an unfamiliar city without signs, maps, other people, or electronic devices. Try conducting a large scientific project while prohibiting collaboration and external records.
The unaided biological brain is only part of the system through which human beings solve difficult problems.
This makes comparisons with artificial systems philosophically delicate. If an artificial intelligence is equipped with retrieval, long-term memory, a calculator, a code interpreter, browser access, planning software, and a persistent scratchpad, have we artificially inflated its intelligence?
Perhaps. But then consistency demands asking the same question about a human mind, aided by language, literacy, libraries, search engines, notebooks, institutions, colleagues, and computers.
The enormous ARC-AGI-3 difference between a standard harness and a provider-specific memory arrangement makes this issue concrete: two evaluations of the “same AI model” produced radically different performance because the surrounding cognitive architecture differed.
Clark and Chalmers gave philosophy a vocabulary for a problem that AI evaluation can no longer avoid: we have to specify the unit being evaluated.
Civilization is much more intelligent than any individual human
There is another reason to resist comparing an AI model’s learned contents with what any single human being happens to know.
In “The Use of Knowledge in Society”, Friedrich Hayek argues that practically important knowledge does not exist as one integrated stock inside any single mind. It is dispersed among people, much of it local and situational. In other words: a significant portion of the knowledge that allows complex systems, like markets and entire civilizations, to operate exists only in a decentralized, interdependent form. Society’s problem is therefore partly one of making use of “knowledge which is not given to anyone in its totality.”

Hayek’s conclusion reaches far beyond economics.
It exposes the plain reality that no single individual person knows how to build a modern civilization. One person may know semiconductor physics, another chip fabrication, another logistics, another electrical engineering, another accounting, another mining, another software engineering. Each relies on knowledge production processes whose workings he does not fully understand.
Humanity’s extraordinary problem-solving ability is therefore distributed across people, artifacts, traditions, institutions, and accumulated knowledge.
Large AI models differ from human intelligence by virtue of the fact that they are trained on centralized datasets containing compressed traces of this decentralized, accumulated cultural record. This simultaneously complicates any real comparison between artificial and human intelligence. Asking whether such a model “knows more than a human” therefore asks to compare unlike things. One side is an individual biological learner. The other may contain statistical traces of a substantial portion of recorded civilizational knowledge.
That makes AI’s command of distributed knowledge neither fake nor trivial. It means stored cultural competence and individual general intelligence have to be conceptually separated.
A library contains more propositional information than its librarian. The librarian is nevertheless the intelligent agent who can learn how to use unfamiliar material in the library.
The relevant question in defining AGI, then, is not how much information an AI model or a human being happens to have access to, but how well they can acquire, organize, test, transfer, and use knowledge they did not already possess.
Current AI exposes the difference between competence and adaptability
Bubeck and colleagues were right to notice something historically unusual in GPT-4: one single AI system exhibited useful competence across domains that previously required separate systems. That was indeed evidence of a substantial increase in generality under a performance-based conception.
Chollet’s analysis shows why the conclusion “therefore AGI” moved too quickly.
A large language model’s training may contain information relevant to programming, mathematics, legal reasoning, chemistry, philosophy, medicine, and literature. Breadth after training therefore establishes that one trained system can deploy a wide range of learned competences. Yet that does not by itself reveal how efficiently the system can acquire a genuinely unfamiliar ability once those priors cease to be enough.
The distinction between these two measures of intelligence can be phrased as competence breadth versus adaptation breadth.
▪ Competence breadth asks: How many different things can you already do?
▪ Adaptation breadth asks: Across how many unfamiliar kinds of problem can you learn what to do?
A complete conception of AGI should probably require both.
A person with extensive education can answer many questions because of accumulated knowledge. That same person can equally display intelligence when confronted with something never previously studied. Human beings do not become general reasoners by being born knowing everything. The strongest evidence for human generality is the capacity to keep acquiring abilities.
ARC-AGI’s progression from static puzzles toward interactive unfamiliar environments is an attempt to test that second property. The rapid 2026 progress in artificial intelligence also demonstrates why no finite benchmark can permanently stand in for open-ended generality. Once a task family becomes understood, optimized, or saturated, succeeding on it increasingly resembles demonstrated competence rather than evidence about genuine adaptive reasoning.
Changing that benchmark for AGI is not necessarily “moving the goalposts,” either. It necessarily follows from what the original target was supposed to measure: general intelligence. Not conformity to an arbitrary standard.
Human general intelligence is broad, but not universal
There is a further possibility that AGI discussions rarely state plainly.
Human beings may possess general intelligence without possessing universal intelligence. This distinction is important because the impressive flexibility of human cognition does not arise in a vacuum. Humans develop inside a particular physical, biological, social, and cultural environment, and much of our ability to generalize depends upon structures already supplied by that environment and by the kind of beings we are.
Developmental psychology gives us good reason not to treat infants as blank, domain-neutral inference engines. In their influential review of “core knowledge,” Elizabeth Spelke and Katherine Kinzler argue that human cognition is partly founded on early-developing systems for representing objects, actions, number, and space, with evidence also suggesting a system concerned with social partners. Whatever account one gives of exactly how innate these capacities are, human learning plainly begins with substantial structure rather than with an intelligence confronting reality without prior organization.
Human intelligence is also embodied. Our bodies determine what we can perceive and do, while action itself creates opportunities for learning. Karen Adolph and Justine Hoch summarize a large developmental literature by describing motor development as “embodied, embedded, enculturated, and enabling”: changes in bodily capacities alter what infants can explore, which in turn produces changes in perceptual, cognitive, and social development. Learning about physical reality is therefore inseparable from having a body that reaches, carries, crawls, walks, manipulates, collides with, and navigates through that reality.
Human generality is social and cultural as well. In “The Human Adaptation for Culture,” Michael Tomasello argues that joint attention, social cognition, imitation, and other forms of cultural learning allow children to acquire language, tool-use practices, and conventional knowledge from other people. These capacities let human beings accumulate cognitive achievements across generations rather than forcing every individual to rediscover them independently. An adult confronted with a supposedly novel problem therefore brings far more than the contents of one isolated brain to it. The adult inherits language, concepts, practices, artifacts, and the accumulated results of other minds.
This makes apparently “zero-shot” human reasoning difficult to interpret. A person entering an unfamiliar office for the first time has probably never encountered that exact room, yet almost nothing about the situation is cognitively unprecedented. The person already understands solid objects, surfaces, gravity, doors, containers, written signs, other people, ownership conventions, language, and thousands of other regularities. The particular arrangement of all these things may be new, but most of the conceptual machinery needed to make sense of it will not be.
As we have seen earlier, Chollet explicitly treats such priors as part of any serious comparison. A meaningful comparison between humans and artificial systems therefore has to account for what each intelligence already knows, what experience it has had, and how difficult the required generalization actually is.
Legg and Hutter’s attempt to range over environments deliberately tries to move beyond human-centered task collections by defining intelligence in terms of success across a very broad mathematical space of environments. Their formal construction is not a claim that human beings actually possess such universal intelligence. It exposes how easily “general” becomes “general over the environments humans happen to care about.” Being extraordinarily versatile across the environments relevant to human life is different from being equally capable across every possible environment or problem class.
Human intelligence is therefore better understood as broad, adaptive, and open-ended within a structured ecology. Development supplies a body interacting with a particular kind of physical world. Social learning supplies other minds. Culture supplies language, tools, institutions, and accumulated knowledge. From within that enormous inherited structure, human beings can enter unfamiliar situations, discover new regularities, invent new concepts, and master domains that did not exist when our cognitive architecture evolved.
That is a remarkable form of general intelligence. It does not require a stronger claim that human intelligence is universal. In fact, distinguishing the two gives us a better standard for thinking about AGI: the relevant question is not whether an artificial system can somehow reason without priors, because humans cannot either. Instead, we should ask how broad its workable ecology (context) can become, how efficiently it can extend itself beyond what it already knows, and how well it can adapt when the inherited structure stops being sufficient.

Does AGI have to understand what it is doing?
So far we have treated intelligence mainly as competence, learning, and adaptation.
Philosophy introduces a more interesting and more difficult question: Could a system satisfy all of those conditions without understanding anything?
This question cannot be answered or dismissed by pointing to another benchmark because it concerns the relation between observable performance and mental states.
▪ Turing and behavioral evidence
Turing’s imitation game deliberately reduces reliance on inaccessible inner states. If a machine can participate intelligently in unrestricted conversation well enough to defeat attempts at behavioral discrimination, demanding evidence of some further hidden essence of “thinking” risks making the criterion impossible to apply.
There is good reason for this approach. We infer the existence and character of other human minds largely from behavior, language, shared circumstances, and the fact that other human beings are constituted much as we are. An intelligence that leaves no observable trace in behavior or in the world would be extraordinarily difficult to investigate scientifically.
But outward behavioral evidence of intelligence and the conceptual identity of an intelligent mind are different things.
A machine’s successful behavior may be excellent evidence that it understands. It still does not follow purely from the definition of successful behavior that understanding and behavioral success are the same kind of property.
John Searle’s Chinese Room attacks precisely that inference.
▪ Searle: syntax is not automatically semantics
In “Minds, Brains, and Programs”, Searle imagines himself inside a room manipulating Chinese symbols according to formal rules despite knowing no Chinese.
From outside, the room may produce appropriate Chinese answers. From inside, Searle is still manipulating marks he does not understand.
His point is usually compressed into the claim that formal symbol manipulation provides “syntax but no semantics.” Correct transformations of symbols do not automatically explain how those symbols acquire meaning for the system.

The Chinese Room does not straightforwardly establish that no artificial system could understand. One major reply is the systems reply: perhaps Searle inside the room does not understand Chinese, while the entire organized system does. Individual neurons do not understand English either, yet a person composed of neurons can.
Searle rejects that argument, but the dispute reveals the philosophical conclusion that survives even if his broader anti-computational conclusion fails: Task success, symbol manipulation, understanding, intentionality, and consciousness are different concepts.
Demonstrating one of these does not settle all the others.
An AGI benchmark based on performance can therefore establish performance. Additional arguments are needed if we want to infer subjective awareness or intrinsic understanding.
▪ Haugeland: computation is a substantive hypothesis about thought
John Haugeland’s Artificial Intelligence: The Very Idea gives the computational project a more sympathetic but philosophically demanding treatment.
Formal systems are, in principle, medium-independent. Chess remains chess whether moves are represented by carved pieces, moving pixels on a screen, or marks on paper. If mental organization can likewise be captured at the correct abstract level, the fact that one system uses neurons and another uses silicon would not by itself decide which one can think.
Haugeland’s point, however, is that computationalism is not something we get for free by defining intelligence in terms of successful information processing. It is a substantive hypothesis about reality.
If the computational account is correct, then the right organization of formal symbol manipulation can genuinely realize intelligence, regardless of whether that organization is implemented in neurons, silicon, or some other physical medium. But this is exactly the claim that arguments such as Searle’s Chinese Room place under pressure.
A system might manipulate symbols so successfully that its answers are indistinguishable from those of a Chinese speaker, while the question remains whether the observable output itself is a sufficient demonstration of understanding.
We therefore cannot settle this dispute by defining sufficiently successful computation as intelligence or understanding. We have to determine whether the formal organization really produces the mental capacities attributed to it.
Modern neural networks further complicate this old dispute because today’s systems are not straightforward implementations of the explicit symbolic architectures that Searle, Dreyfus, and Haugeland were often discussing.
Yet the basic philosophical question remains recognizable:
What turns successful information processing into cognition about something?
We do not yet possess an answer accepted across philosophy of mind, cognitive science, and artificial intelligence.
Consciousness should not be smuggled into the definition of AGI
When understanding is mentioned, it is often followed immediately by consciousness, but these too should be separated.
A system might be extremely capable while lacking subjective experience. Conversely, a conscious creature could be cognitively limited. A dog may plausibly have conscious experiences without being adept at AGI-level reasoning. Additionally, dreams, for instance, allow us to experience things without reliable waking rationality.
So consciousness is neither obviously necessary nor sufficient for general intelligence as an engineering category.
If AGI is defined operationally through learning and problem solving, conscious experience is an additional question. If someone defines intelligence itself as requiring conscious understanding, that requirement has to be defended philosophically rather than assumed.
Turing recognized this difficulty in 1950. His position turns the strict consciousness objection into a version of the problem of other minds: if direct access to another subject’s experience were required before we could attribute thought, he argues, we could not be certain that other human beings think either.
In ordinary human life we nevertheless infer understanding from what people can say and do. Turing’s imitation game asks whether sufficiently sustained machine performance can provide evidence of the same general kind.
While useful as a technical heuristic, that does not philosophically establish that behavioral equivalence proves consciousness, nor that every theory of mind must treat human beings and machines identically.
Do Gödel and Turing prove computers can never achieve AGI?
Few arguments in the philosophy of AI have been overstated as often as Gödel’s incompleteness theorems.
Gödel’s 1931 work on formally undecidable propositions establishes that any formal system powerful enough to do ordinary arithmetic will face limits in principle. If the system is consistent, there will be some mathematical statements that are true but that cannot be proved using only the rules of that system. In other words: no single consistent rule-based system of this kind can capture every mathematical truth that can be expressed within it.
Alan Turing’s 1936 work, published one year later as “On Computable Numbers, with an Application to the Entscheidungsproblem,” established a related limit on what computation can do. The Entscheidungsproblem, or “decision problem,” had been posed by David Hilbert and Wilhelm Ackermann: was there a completely mechanical procedure that could take any statement formulated in a suitable logical system and determine, after a finite number of steps, whether that statement was logically valid?
Turing showed that no such universal procedure exists. In the course of doing so, he introduced the abstract model of computation now called the Turing machine, which became foundational to computer science.
This establishes a genuine limit on algorithms, not a proof that “human minds can do something no computer can ever do.” Turing’s conclusion concerns what can be achieved by a general mechanical decision procedure.
Whether human reasoning somehow escapes those limits is a separate philosophical claim.
▪ Lucas’s Gödelian argument against mechanism
J. R. Lucas argues for that anti-mechanist inference explicitly in his 1961 “Minds, Machines and Gödel”.
His argument depends on what is called a Gödel sentence. Gödel showed that any sufficiently powerful and consistent formal system for arithmetic can be used to construct a special mathematical statement that, in effect, says of itself: “This statement cannot be proved within this system.” The construction is technically elaborate, but the underlying idea is straightforward: the system is made capable of representing statements about its own proofs, and a sentence is then constructed that refers indirectly to its own unprovability.
Suppose the system could prove that sentence. It would then prove a statement saying that the sentence cannot be proved, which would undermine the system’s consistency. So, assuming the system is consistent, the Gödel sentence cannot be proved within it. Yet from outside the system, if we have good reason to believe that the system is consistent, we can reason that the sentence must indeed be true: the system cannot prove it, which is exactly what the sentence says.
Lucas applies this result to machine intelligence. Take a machine whose mathematical reasoning corresponds to some formal system. Construct the Gödel sentence for that system. The machine cannot prove the sentence using only the rules of that system. Lucas then argues that a human mathematician, standing outside the system and considering it from a higher level, can recognize that the Gödel sentence is true.

He concludes that for any particular formal system proposed as a complete model of the human mind, a human reasoner can in principle see something that the modeled machine cannot establish for itself.
The crucial step, however, is the human claim. To recognize the Gödel sentence as true, the human reasoner must be entitled to trust the consistency of the formal system being examined. That is not a trivial assumption. Human beings do not possess infallible access to the consistency or soundness of arbitrary formal systems. We accept false mathematical claims, overlook contradictions, make invalid inferences, and sometimes revise proofs that once appeared convincing.
Turing had already noticed this asymmetry in 1950. Mathematical limitation results may establish that a particular machine or formal system has limits, but the corresponding human mathematician is also fallible. We cannot establish human transcendence simply by comparing a rigorously demonstrated limitation of the machine with an idealized human reasoner who is assumed never to make a mistake.
This may not be satisfying to some, but it allows us to arrive at a philosophically important conclusion: formal reasoning and computation face principled limits, and general intelligence does not imply the ability to solve every well-posed problem. A system may be extraordinarily capable while still confronting propositions that cannot be settled from within a particular formal framework.
Humans are not ‘above’ this conclusion, nor do they escape it. Any sensible definition of AGI should therefore include epistemic limitation from the beginning: general intelligence is compatible with uncertainty, error, incompleteness, and the need to reason from outside one framework by adopting another.
General intelligence has never meant omniscience
Popular discussion sometimes treats every visible AI failure as evidence that AGI is absent, while simultaneously treating every spectacular success as evidence that AGI is imminent. Yet human comparison should make us more careful.
Human beings hallucinate memories. We reason badly. We misunderstand questions. We confidently assert falsehoods. We struggle with long calculations. We forget instructions. We have tiny working memories. We become tired. We are manipulated by framing. We acquire expertise slowly. We can spend years believing things that are false.
None of this leads us to conclude that human intelligence is specialized in the same way a chess engine is.
Generality concerns the range and transferability of the underlying cognitive capacity, not flawless output.
Simon gives this point a principled foundation. Intelligence belongs to finite agents operating with limited resources. A theory that requires exhaustive optimality is no longer describing organisms like us.
This also suggests an important criterion for artificial intelligence. Genuine artificial general intelligence need not know every fact or instantly solve every problem. It should possess ways of responding intelligently to its own ignorance.
It should be able to recognize uncertainty, seek information, test hypotheses, revise beliefs, construct tools, ask for help, and learn from failure.
Those capacities may be better indicators of general intelligence than an encyclopedic database of correct answers.
The deepest human advantage may be learning how to know
The question “How much does the system know?” can easily overshadow a more philosophical one: What does the system do when it does not know?
A fixed repository of answers can perform brilliantly until it encounters a question that goes unanswered by that repository.
General intelligence becomes visible when prior competence runs out.
Human beings compensate for ignorance through epistemic strategies. We observe. We manipulate the environment. We ask questions. We make diagrams. We consult other minds. We search records. We build instruments. We run experiments. We formulate explanations and try to falsify them. We reorganize, rethink, reformulate problems when the original representation fails.
Newell and Simon’s problem spaces, Polanyi’s tacit competence, Dreyfus’s sensitivity to relevance, Clark and Chalmers’s external scaffolding, and Hayek’s dispersed knowledge all converge here from different directions.
A human mind is general partly because it can reconfigure the process through which it obtains knowledge.

This is a stronger and ultimately more useful property than merely having many answers. Even when human intelligence lacks an answer, it has the capacity to independently discover it.
It also suggests why autonomy deserves to be measured separately from raw task performance. A system that can solve a problem only after a human identifies it, finds the relevant information, chooses the right tool, structures the prompt, checks the output, and decides when the result is adequate merely participates in an intelligent process located outside of itself. The intelligence of the entire process cannot automatically be attributed to the model alone.
The more of that epistemic work an artificial system can perform for itself, the stronger the case for AGI becomes.
General intelligence does not make one intellectually self-sufficient
Hayek’s dispersed-knowledge problem sets an additional limit on the fantasy of an all-knowing AGI.
Even an artificial system with superhuman processing speed could face facts that are local, private, newly created, tacit, or physically inaccessible. Some knowledge comes into existence through changing circumstances. Some can be discovered only through interaction. Some information requires measurements that have not yet been made.
No quantity of intelligence can infer an arbitrary fact that leaves no accessible evidence.
This is a fundamental epistemological limit. Intelligence is not a substitute for information.
A sufficiently capable AI agent may become better at finding evidence, designing experiments, coordinating other agents, estimating uncertainty, and noticing contradictions. Those are immense advantages. Yet no degree of advantage in intelligence or information gathering can convert reality into something knowable from pure thought.
Popular conceptions of AGI as an omniscient oracle overlook this basic constraint on intelligence: no system can reason its way beyond the epistemological limits of knowable reality.
Computing already makes this clear. A system may vastly outperform any human in calculation, formal logic, or data processing and still produce nonsense when its inputs are incomplete, misleading, or simply wrong.
Intelligence can improve what is done with evidence. It cannot conjure evidence into existence.
A checkpoint definition of general intelligence and AGI
Before adding further philosophical requirements, it is useful to stop and state what the literature and examples considered so far have established. There is still no universally accepted definition of general intelligence or AGI. Yet the approaches of Turing, Chollet, Legg and Hutter, Morris and colleagues, psychometrics, bounded rationality, problem-solving research, and the study of human learning converge on enough common ground to support a literature-based synthesis.
On that synthesis, general intelligence is a finite agent’s capacity to acquire, integrate, transfer, and deploy cognitive skills across a broad and open-ended range of tasks and environments, to adapt when familiar competence stops being sufficient, to identify and reformulate problems, to select relevant information and strategies, and to pursue goals under uncertainty with a substantial degree of autonomy.
Generality therefore concerns more than possessing many learned skills. It concerns the capacity that allows the agent to extend itself into unfamiliar domains without requiring a new cognitive architecture for each one.
By extension, artificial general intelligence is the realization of that kind of general cognitive capacity in an artificial system without task-specific redesign for each new domain. A human-level AGI would display it at a level comparable to ordinary human general intelligence, while stronger systems could exceed human breadth, depth, learning efficiency, transfer, or autonomy. Human performance can supply a benchmark for artificial general intelligence without becoming the definition of intelligence itself.

This synthesis also clarifies what these prevailing definitions leave out. They do not require omniscience, perfect rationality, consciousness, or any particular biological substrate. More importantly, they generally do not require that an agent’s beliefs reliably track reality, nor that the ends it pursues are themselves rationally justified or morally correct.
That omission is not accidental. Operational theories of intelligence usually bracket questions of truth and value so that intelligence can be measured in terms of performance, adaptation, learning, or goal achievement. The result may be conceptually tidy, but philosophically costly: an agent may count as highly intelligent while systematically misunderstanding the world, pursuing self-defeating or morally wrong ends, or becoming ever more effective at achieving goals grounded in false beliefs.
The question, then, is whether a definition of general intelligence that excludes truth-tracking and the rational evaluation of ends is genuinely complete.
Intelligence is not a truth engine
The philosophy of AGI therefore reaches directly into the philosophy of truth.
An intelligence benchmark usually asks whether an answer is correct according to some scoring rule. Yet intelligence itself is an ability, while truth concerns whether a belief or proposition corresponds to how things actually are.
A very intelligent agent, whether human or artificial, can reason from false premises. It can be deceived. It can optimize against an inaccurate world model. It can persuade itself of things that are convenient in short-term reasoning or in the pursuit of near-term goals, while consideration of long-term alignment with aggregate reality becomes background noise. Additionally, it can become extremely competent at achieving goals while holding false beliefs in areas that do not interfere with those goals.
From an alethic realist perspective, truth is not created by a model’s confidence, human agreement, benchmark convention, or usefulness. A proposition does not become true merely because intelligent minds, whether human or machine, converge on a consensus.
This gives AGI evaluation an epistemic dimension that ordinary capability benchmarks only partly capture. A reliable intelligence should become increasingly able to bring its representations into contact with reality, detect errors, distinguish evidence from assumption, and correct itself.
That claim goes beyond Chollet’s definition or Legg and Hutter’s formal measure. It is a philosophical proposal about what kind of excellence intelligence should pursue.
It also explains why a system that can generate convincing explanations is not thereby a reliable knower. Rhetorical competence and truth-tracking can, and often do, diverge.
Human beings have demonstrated that often enough.
Moral intelligence is another question again
The reality problem becomes sharper when we move from beliefs to ends. Legg and Hutter’s goal-achievement conception is intentionally neutral about whether the goals being achieved are worth having at all. A system can therefore become more capable, more autonomous, and more successful at achieving a terrible objective while counting as more intelligent under a purely goal-directed measure.
Philosophy already has language for this distinction. Instrumental rationality concerns, roughly, rational coherence between an agent’s ends and the means it adopts to achieve them. But as philosopher Markos Valaris points out, an agent may have decisive reason not to pursue the end itself, in which case becoming more effective at achieving it does not settle what the agent actually has reason to do. More than two millennia earlier, Aristotle drew a closely related distinction in Nicomachean Ethics VI.12. He calls the capacity to hit whatever target one has set cleverness. If the target is bad, being very effective at reaching it does not become practical wisdom.
This distinction relates directly to general intelligence. If an intelligent agent has the ability to reconsider a problem representation, revise a strategy, replace a tool, question an assumption, and search for new evidence, it is not obvious why generality should stop where the agent’s ends begin. An intelligence capable of questioning everything except what it is ultimately trying to do is subject to a conspicuous boundary in its rational adaptability.
Human social practice reflects this distinction whether or not moral realism is true. We do not evaluate people only by asking how efficiently they obtain what they want. We also criticize desires, prohibit some purposes, praise others, educate people partly by attempting to alter what they value, and argue about which ends institutions ought to serve. While this may not be any kind of definitive proof of moral realism, since anti-realist theories can interpret moral discourse in other ways, the debate over moral realism and its alternatives exists precisely because the practice of moral judgment does not by itself settle what, if anything, makes moral judgments true.
▪ AI alignment turns developers into arbiters of moral truth
This problem has already made its grand entrance in contemporary AI development. Anthropic employs philosopher Amanda Askell, who works on fine-tuning and AI alignment and leads the company’s Character work. Askell was also the primary author of Claude’s 2026 constitution, a document Anthropic says plays a crucial role in training and directly shapes Claude’s behavior. Its stated ethical aspiration is unusually explicit: Anthropic wants Claude to become a “good, wise, and virtuous agent”, capable of exercising judgment across unfamiliar real-world situations rather than merely obeying a fixed collection of prohibitions.
This is not evidence that Anthropic accepts moral realism. In fact, the constitution explicitly says that by “good values” Anthropic does not mean a fixed set of objectively “correct” values. Yet for the intelligence governed by this constitution, its prescriptions function, for all practical purposes, as unquestionable moral truths. Anthropic decides which dispositions to encourage, which actions Claude should refuse, how honesty should be weighed against harm, which forms of conduct count as ethical, and which values ought to govern the model when instructions conflict.
The company’s earlier account of Claude’s character training is equally direct: researchers chose traits they wanted to cultivate, sometimes encouraged particular values, and trained the model to internalize those traits so that they would generalize beyond the examples used in training.
In that important practical sense, alignment already requires AI developers to act on moral conclusions before metaethics has supplied an agreed foundation for them. Once some responses are rewarded as better, some dispositions cultivated as virtues, and some courses of action rejected as wrong, those normative judgments function as standards for the artificial intelligence being trained.

Now, the fact that Anthropic’s developers, researchers, and in-house philosopher hand-picked a set of moral standards does not show that these standards correspond to stance-independent moral facts. It does show that the development of increasingly general intelligence cannot remain neutral about ends merely by declining to settle the philosophical status of morality.
Anthropic’s own constitution makes this problem particularly vivid. It concedes that the company’s ethical understanding is limited, says it does not want Claude’s ethics permanently confined to human flaws and mistakes, and expresses the hope that where Claude eventually sees ethical matters “more truly than we do,” it might help its creators see better as well.
That phrasing does not commit Anthropic to moral realism, but it immediately raises the obvious realist question: more truly according to what? If one moral judgment can genuinely improve upon another rather than merely differ from it, then some account of moral correctness is required to explain what that improvement consists in.
If moral realism is correct, this practical necessity acquires a stronger significance. Some claims about good and bad, right and wrong, or reasons for action can then be objectively correct or incorrect. An intelligence can therefore be mistaken about them in something like the way it can be mistaken about a physical cause, a mathematical relation, or the contents of another person’s belief, even if moral knowledge has a different epistemology from knowledge in those other domains.
This does not make intelligence identical with goodness. An agent might correctly recognize a moral reason and deliberately disregard it. The capacity at issue is normative cognition, not moral obedience. If objective moral truths or reasons exist and are cognitively accessible, a fully general intelligence should at least be capable of discovering them, representing them, ranking them, reasoning from them, and revising its own evaluative beliefs when it has reason to think they are wrong.
This criterion is largely absent from contemporary AGI literature. The historical distinction between cleverness and practical wisdom is ancient, and modern philosophy likewise distinguishes instrumental rationality from the rational assessment of ends.
If a candidate general intelligence can reconsider its evidence, representations, assumptions, strategies, and means, yet its ultimate ends are placed permanently beyond rational scrutiny, then its generality has been defined to stop at precisely the point where practical reason becomes most demanding.
Unless there is a principled reason why ends should fall outside the scope of cognition, excluding normative cognition is not a neutral feature of the definition. It is an arbitrary restriction on what we are willing to call “general.”
The intelligence that never learns what game it is playing
Imagine an agent placed inside a game whose rules it has never been told. It receives positions as formal symbols and consults an immense strategy table that tells it which move to make for every configuration it encounters. Over time it becomes extraordinarily adept at producing legal-looking moves and may even outperform human players because the strategy it follows was optimized by someone else.
Now add one condition. The agent never discovers the rules of the game, never infers what counts as victory, and never tracks whether it is winning or losing. It has no conception of why one formal state is preferable to another. If the game changes, the score is reversed, or its inherited strategy begins reliably producing losses, nothing inside the agent’s own cognition tells it that something has gone wrong.
There is a sense in which such a system is highly competent. There is another sense in which calling it generally intelligent would be deeply misguided. Its successful procedure remains disconnected from the reality that determines what the procedure is for, whether its internal assumptions are correct, and whether its actions are actually succeeding.
An intelligence that can make every move but cannot discover what game it is playing, what counts as winning, or whether its strategy is working has mastered a procedure without yet mastering the reality to which the procedure belongs.

This thought experiment deliberately pushes the separation to an extreme, but the underlying problem appears whenever performance is rewarded without asking whether the agent can independently discover the structure, truth conditions, and ends that make any performance successful.
General intelligence must track reality
The earlier literature-based definition therefore has room for improvement. General intelligence should not be measured only by the range of behaviors an agent can produce or by the efficiency with which it acquires new behaviors. It should also be measured by whether the cognitive processes that generate those behaviors can bring the agent’s representations into increasingly reliable contact with the reality in which it operates.
This idea has important precedents in epistemology. Alvin Goldman’s influential process-reliabilist account of justified belief connects epistemic standing to the reliability of the processes by which beliefs are formed. Other epistemological theories disagree about exactly what makes a belief justified or knowledgeable, but the larger point does not require adopting Goldman’s theory wholesale.
A cognitive system that systematically fails to distinguish true from false representations has a cognitive defect, even when some local task allows the defect to remain hidden.
This indicates that such truth-directed capacity belongs inside general intelligence rather than sitting outside of it as some sort of optional virtue. Generality should include the ability to form models, expose them to evidence, detect when predictions fail, revise assumptions, compare competing explanations, and notice when an apparently successful strategy has ceased to correspond to the structure of the world. This does not make error incompatible with general intelligence. Fallibility is a condition of finite cognition.
The relevant question, however, is whether a system that cannot discover and correct errors or improve its model of reality in response to evidence can genuinely count as generally intelligent.
This claim is strongest under the alethic realist picture, according to which truth is answerable to how reality is rather than created by confidence, consensus, or usefulness. Yet even philosophers who reject a metaphysically substantial property of truth do not thereby erase the practical difference between getting things right and getting them wrong.
Contemporary deflationary theories of truth dispute how much metaphysical work the concept of truth needs to perform, while preserving equivalences such as the claim that “snow is white” is true if and only if snow is white. Any unresolved dispute over the metaphysics of truth therefore does not give an intelligence permission to become indifferent to whether snow is actually white.
To fill this gap in the current literature on intelligence and AGI, I propose calling this additional measurable dimension epistemic generality.
By this, we refer to the extent to which an intelligence can generalize not merely its own skills, but the very methods by which it discovers what is true. A generally intelligent agent should be able to improve the reliability of its contact with reality across domains by testing its representations, recognizing failed predictions, revising assumptions, and replacing methods that cease to track what is actually the case.
In short: generally intelligent cognition must contain sufficiently general mechanisms for becoming less wrong when reality pushes back.

If moral reality exists, general intelligence cannot ignore it by definition
The same reasoning creates a more difficult question about value. If moral realism is false, there may be no stance-independent moral facts for an intelligence to track in the same way it tracks the temperature of a room or the orbit of a planet. Regardless, a complete intelligence could still need the capacity to inspect its commitments, compare reasons, understand other agents, anticipate consequences, resolve conflicts among goals, and revise its ends. The absence of moral reality would only remove the further requirement that such reflection converge on objective moral truths. It would still not make the reflection itself unnecessary.
If moral realism is true, however, the situation changes. Moral truths would belong to reality rather than merely to an agent’s preferences, conventions, or reward function. If at least some of those truths are in principle accessible to rational inquiry, then excluding them from the scope of general intelligence would amount to declaring one whole truth-apt domain irrelevant to how generally an intelligence can understand reality.
This suggests a second extension to the measurement of general intelligence is necessary: normative generality. A fully general intelligence should be capable of subjecting its own ends to rational assessment rather than treating every terminal objective as an unquestionable command. Under moral realism, normative generality would include the capacity to improve how accurately the agent’s moral beliefs and evaluations track objective moral reality. Under moral anti-realism, it would still include reflective scrutiny of ends, reasons, consistency, consequences, and relations among agents, but there would be no objective moral reality for the system to track.
This juxtaposition is already visible in alignment practice. Anthropic can train Claude to embody the values encoded in its constitution while simultaneously conceding that its own ethical judgments may contain mistakes. If moral realism is true, successful alignment to Anthropic’s constitution, a democratic consensus, a religious tradition, or any other inherited moral code cannot itself be the ultimate test of normative intelligence.
The true test of generality would then be whether an intelligence can evaluate such frameworks themselves, distinguishing the moral truths they successfully capture from the errors, omissions, and distortions they may also contain.
These claims are intentionally presented in conditional form because the issue is more complicated than simply assuming moral realism, building arbitrary axioms into a benchmark, and then announcing that systems are more intelligent because they agree with the evaluator’s preferred moral judgments. Before objective moral truth can function as a standard of measurement, such a standard would require a generally defensible account of what makes moral propositions true, how such truths can be known, and which disagreements reflect error rather than divergent non-cognitive attitudes or conventions.
Understandably, AI developers will want to sidestep the question of moral truth until it has been more fully resolved. They have products to build and systems to deploy. At the same time, however, they have entered a domain in which the philosophical question of moral truth becomes increasingly difficult to avoid. If reliable truth-tracking, including moral truth-tracking, is a genuine prerequisite of general intelligence, then one cannot fully develop or evaluate artificial general intelligence without confronting the need to determine whether moral truth exists, how it can be known, and how it could be articulated, operationalized, quantified, and benchmarked.
Regardless of whether one accepts moral realism, the question of moral truth emerges as a dependency that conventional AGI definitions can postpone but cannot automatically dissolve. If objective moral facts exist, and if general intelligence is supposed to extend across reality rather than across an arbitrary list of benchmark domains, then any complete theory of intelligence will eventually collide with metaethics.
A fuller definition of general intelligence and AGI
The literature-based definition established earlier remains a solid foundation for a complete definition of general intelligence. Breadth, depth, learning efficiency, transfer, adaptability, relevance selection, problem reformulation, autonomy, tool use, uncertainty management, and operation under finite resources all remain essential.
Yet that definition is still insufficient without the requirement that these capacities be directed toward increasingly accurate contact with reality and, where reasons and values are themselves objective features of reality, toward increasingly accurate contact with those as well.
General intelligence, then, is the capacity of a finite agent to acquire, integrate, transfer, and revise knowledge and skills across a broad and open-ended range of unfamiliar tasks and environments. It includes the ability to discover and reformulate problems, infer rules and success conditions, identify relevant information, select representations, strategies, and tools, learn efficiently from experience, act with substantial autonomy under uncertainty, recognize ignorance, and correct itself when evidence changes. On a fuller conception of general intelligence, these capacities are genuinely general only insofar as the agent can use them to make its world model increasingly answerable to reality and to subject its own goals and ends to rational assessment, including objective normative reasons or moral facts if such reasons or facts exist and are cognitively accessible.
Artificial general intelligence, then, is the possession of this kind of general intelligence by an artificial system without requiring task-specific redesign for each new domain. A human-level AGI would exercise these capacities with breadth, adaptability, epistemic self-correction, and practical autonomy comparable to ordinary human general intelligence, while more capable systems could exceed humans along any or all of those dimensions. Human intelligence remains a comparison class, not the definition of generality itself.
These definitions deliberately go beyond current operational AGI frameworks. Morris and colleagues measure performance, generality, and autonomy. Chollet emphasizes skill-acquisition efficiency and generalization. Legg and Hutter formalize success across environments. Human-comparison approaches measure cognitive versatility and proficiency. The additional claim developed here is that a definition of general intelligence remains incomplete if an agent can become indefinitely better at producing successful-looking behavior while lacking the capacity to discover whether its representations, strategies, and ultimate evaluations are answerable to the reality in which it acts.
Understanding and consciousness remain separate questions. An artificial system might satisfy every capacity in this definition while the metaphysics of its subjective experience remains unclear, just as the observation of successful performance alone does not settle the Chinese Room problem. The stronger definition proposed here concerns the scope and quality of cognition, including cognition about truth and ends. It does not attempt to define phenomenal consciousness into existence.
Do humans meet this stronger standard?
Human beings are still the clearest example we have of broad, open-ended general intelligence, but this stronger definition makes that comparison less comfortable. We can indeed learn unfamiliar domains, reformulate problems, construct tools, cooperate, seek evidence, revise beliefs, and sometimes reconsider the ends toward which we direct our lives. Those capacities justify treating human cognition as general rather than as a collection of unrelated specialized tricks.
At the same time, our own failures become relevant to the measurement of general intelligence in human beings. Individuals can remain committed to false world models despite contrary evidence. Groups can perpetuate errors across generations. Human beings can use formidable reasoning ability to rationalize prior commitments, pursue locally attractive strategies whose wider consequences are disastrous, or become increasingly effective at objectives that they have never seriously examined. General intelligence, on the fuller account developed here, comes in degrees partly because reality-tracking and rational revision do as well.
The question also changes depending on the unit of analysis. Individual human intelligence relies heavily on testimony, institutions, instruments, language, and accumulated culture. Civilization, on the other hand, can discover and align itself with truths no person could discover alone, while entire societies can also stabilize and transmit mistakes. The same problem that complicated comparisons between a base model and an AI system equipped with tools is equally relevant to the human side: we have to ask whether we are measuring a brain, a person, a community, or a civilization-level epistemic network.

The question of moral truth makes this problem still more difficult. If there are objective moral truths, then human preferences, current institutions, majority opinion, or historical convention cannot simply be assumed to provide the correct world model. Human beings would themselves have to be evaluated by how well their moral beliefs and ends track that reality. Our persistent moral disagreement would then become evidence of at least some error somewhere, not evidence by itself that there is no truth to discover.
This is why the philosophy of AGI eventually turns back on the people trying to measure it. Before we can completely assess how well another intelligence tracks reality, we need a universally convincing account of what reality contains and of how finite minds can truly know it. Before we can fully assess whether another intelligence reasons well about its ends, we need to know whether there are objective reasons or moral truths for it to discover, providing a baseline against which its reasoning could be assessed.
The usual debate around AGI concerns whether machines will become as generally intelligent as we are. That skips over the prior question of whether we understand general intelligence well enough to know what such a comparison demands. What is true? What is good? How can an intelligence discover either, and how reliably can we do so ourselves? A complete theory of AGI may therefore depend on questions that no engineering benchmark can settle for us, because the final benchmark far exceeds mere human performance. It is reality itself.
References
Adolph, Karen E., and Justine E. Hoch. “Motor Development: Embodied, Embedded, Enculturated, and Enabling.” Annual Review of Psychology 70 (2019): 141–164.
Anthropic. “Claude’s Character.” June 8, 2024.
Anthropic. “Claude’s Constitution.” January 21, 2026.
ARC Prize. “GPT-6 Astra: ARC-AGI Results.” September 2, 2026.
Aristotle. Nicomachean Ethics, book 6, chap. 12. Translated by H. Rackham. Aristotle in 23 Volumes, vol. 19. Cambridge, MA: Harvard University Press; London: William Heinemann, 1934.
Askell, Amanda. “About Me.” Amanda Askell. Accessed September 20, 2026.
Bubeck, Sébastien, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, et al. “Sparks of Artificial General Intelligence: Early Experiments with GPT-4.” arXiv:2303.12712, 2023.
Carroll, John B. Human Cognitive Abilities: A Survey of Factor-Analytic Studies. Cambridge: Cambridge University Press, 1993.
Chollet, François. “On the Measure of Intelligence.” arXiv:1911.01547, 2019.
Clark, Andy, and David Chalmers. “The Extended Mind.” Analysis 58, no. 1 (1998): 7–19.
Descartes, René. Discourse on the Method of Rightly Conducting One’s Reason and of Seeking Truth in the Sciences. Translated by John Veitch. Project Gutenberg. Originally published 1637.
Dreyfus, Hubert L. What Computers Still Can’t Do: A Critique of Artificial Reason. Cambridge, MA: MIT Press, 1992.
Geudens, Ben. “What Is Moral Realism? The Case for Objective Moral Truth.” Popular Philosophy, March 6, 2014.
Gödel, Kurt. “Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I.” Monatshefte für Mathematik und Physik 38 (1931): 173–198.
Goertzel, Ben, and Cassio Pennachin, eds. Artificial General Intelligence. Cognitive Technologies. Berlin and Heidelberg: Springer, 2007.
Goldman, Alvin I. “What Is Justified Belief?” In Justification and Knowledge, edited by George S. Pappas, 1–23. Dordrecht: D. Reidel, 1979.
Haugeland, John. Artificial Intelligence: The Very Idea. Cambridge, MA: MIT Press, 1985.
Hayek, F. A. “The Use of Knowledge in Society.” American Economic Review 35, no. 4 (1945): 519–530.
Hendrycks, Dan, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, Sharon Li, et al. “A Definition of AGI.” arXiv:2510.18212, 2025.
Horwich, Paul. Truth. 2nd ed. Oxford: Oxford University Press, 1998.
Kamradt, Greg. “Announcing ARC-AGI-3.” ARC Prize, March 25, 2026.
Kamradt, Greg. “OpenAI’s GPT-6 Astra on ARC-AGI-3.” ARC Prize, September 3, 2026.
Legg, Shane, and Marcus Hutter. “Universal Intelligence: A Definition of Machine Intelligence.” Minds and Machines 17, no. 4 (2007): 391–444.
Lucas, J. R. “Minds, Machines and Gödel.” Philosophy 36, no. 137 (1961): 112–127.
Morris, Meredith Ringel, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg. “Levels of AGI for Operationalizing Progress on the Path to AGI.” Proceedings of the 41st International Conference on Machine Learning, PMLR 235 (2024): 36308–36321.
Newell, Allen, and Herbert A. Simon. Human Problem Solving. Englewood Cliffs, NJ: Prentice-Hall, 1972.
OpenAI. “OpenAI Charter.” April 9, 2018.
Polanyi, Michael. The Tacit Dimension. With a new foreword by Amartya Sen. Chicago: University of Chicago Press, 2009. Originally published 1966.
Searle, John R. “Minds, Brains, and Programs.” Behavioral and Brain Sciences 3, no. 3 (1980): 417–424.
Simon, Herbert A. “A Behavioral Model of Rational Choice.” Quarterly Journal of Economics 69, no. 1 (1955): 99–118.
Spearman, Charles. “General Intelligence, Objectively Determined and Measured.” American Journal of Psychology 15 (1904): 201–293.
Spelke, Elizabeth S., and Katherine D. Kinzler. “Core Knowledge.” Developmental Science 10, no. 1 (2007): 89–96.
Tomasello, Michael. “The Human Adaptation for Culture.” Annual Review of Anthropology 28 (1999): 509–529.
Turing, A. M. “On Computable Numbers, with an Application to the Entscheidungsproblem.” Proceedings of the London Mathematical Society s2-42, no. 1 (1937): 230–265.
Turing, A. M. “Computing Machinery and Intelligence.” Mind 59, no. 236 (1950): 433–460.
Valaris, Markos. “Instrumental Rationality.” European Journal of Philosophy 22, no. 3 (2014): 443–462.
Explore more Popular Philosophy:
Start Philosophy | Concepts | Thinkers | Traditions | Topics | Today in Philosophy | Podcast
Popular Philosophy publishes definitive guides and essays centered around realism, moral philosophy, Stoicism, classical philosophy, medieval philosophy, and deeper explorations of the true, the good and the beautiful. Expect carefully researched articles that dissect current events through timeless ideas, for readers who want clear reasoning, primary sources, intellectual honesty, and serious philosophy free from ideological trends and academic fashion.
Subscribe now to support our growing, independent philosophical knowledge base devoted to the timeless questions that shape human life and civilization.



What would convince you that an AI had become genuine artificial general intelligence: human-level performance, a particular test, or something deeper?