AI defined: the terms with a legal test and the ones without
AI defined is a search that returns a thousand paraphrases of the same sentence, and almost none of them tell you who is doing the defining. That matters, because the answer changes depending on whether you are reading a statute, a standards document, a research paper or a pricing page. One of those definitions can be enforced against a company. The others cannot. So the useful question is not what AI means in the abstract, but which terms carry a test that somebody has to pass.
Some of the vocabulary turns out to be precise. “General-purpose AI model” has a compute threshold attached to it in EU law. “Open source AI” has a published checklist that most released models fail. And some of it is genuinely empty: AGI has no agreed definition in any binding document, only competing ladders proposed by researchers. What follows is the vocabulary sorted by how much weight each word actually carries, and for how far those same terms drift once a vendor picks them up, see our AI terminology explained.
The only definition of AI that can be enforced against you
Regulation (EU) 2024/1689, the AI Act, entered into force on 1 August 2024, and its Article 3 contains 68 numbered definitions. The first one sets the scope of the entire regulation, because the Act applies only to systems that meet it. The European Commission said as much in its guidelines on that definition, first published in February 2025 and adopted as a formal communication on 29 July 2025.
a machine-based system that is designed to operate with varying levels of autonomy and that may exhibit adaptiveness after deployment, and that, for explicit or implicit objectives, infers, from the input it receives, how to generate outputs such as predictions, content, recommendations, or decisions that can influence physical or virtual environments
Article 3(1), Regulation (EU) 2024/1689, as quoted in the Commission’s guidelines
The Commission breaks that sentence into seven elements: a machine-based system, varying levels of autonomy, possible adaptiveness after deployment, explicit or implicit objectives, inference from input, outputs of a stated kind, and influence on physical or virtual environments. The guidance adds a detail that most summaries drop. The seven elements “are not required to be present continuously throughout both phases of that lifecycle”, meaning a trait can show up during building and vanish during use, and the system still counts.
The same law says a chess engine is not AI
The more surprising half of that guidance is the exclusions, and they cut deeper than you’d expect. Systems that follow “predefined, explicit instructions or operations” without any learning, reasoning or modelling fall outside the definition. The Commission names examples: a database query that finds every customer who bought a product last month, standard spreadsheet software, and code that calculates a population average from a survey.
Descriptive analysis is out too. A sales dashboard showing totals, regional averages and trends is not an AI system, because, in the guidance’s words, it “does not recommend how to improve sales or which products to promote”. Classical heuristics are excluded on the same logic, and the Commission’s chosen illustration is a chess program using a minimax algorithm with heuristic evaluation functions. Even simple machine learning is out where a basic statistical learning rule would achieve the same performance, such as a stock forecast built from a mean-strategy baseline estimator.
That produces a clean contradiction with the research vocabulary, which is the point worth holding on to. Google DeepMind’s Levels of AGI paper places Deep Blue at Level 4, “Exceptional Narrow AI”, and Stockfish at Level 5, “Superhuman Narrow AI”. Under EU law a minimax chess engine isn’t regulated as AI at all. Both descriptions are correct within their own frame, which is why arguing about whether something “is really AI” without naming the frame goes nowhere.
Three official definitions, and only the newest mentions content
The EU wording did not appear from nowhere. It tracks the OECD definition, which member countries revised in late 2023. OECD.AI’s own account of the update, published in November 2023, lists what changed: objectives became “explicit or implicit”, the phrase “infers, from the input it receives” went in, “real” environments became “physical”, the autonomy and adaptiveness language was restructured, and, critically, “content” was added to the list of outputs so that generative systems were covered explicitly.
That single word is the fault line running through every older definition. US federal statute still carries the pre-generative wording. Under 15 U.S.C. 9401(3), enacted as part of Public Law 116-283 on 1 January 2021, artificial intelligence means a machine-based system that can “make predictions, recommendations or decisions influencing real or virtual environments”. No content. The same section defines machine learning as “an application of artificial intelligence” that lets systems “automatically learn and improve on the basis of data or experience, without being explicitly programmed”.
| Source | Date | Mentions generated content | Binding |
|---|---|---|---|
| 15 U.S.C. 9401(3), US federal statute | 1 January 2021 | No | Yes, in US federal law |
| NIST AI RMF 1.0 (AI 100-1) | January 2023 | No | No, voluntary framework |
| OECD revised definition | November 2023 | Yes | No, a recommendation |
| EU AI Act, Article 3(1) | In force 1 August 2024 | Yes | Yes, across the EU market |
NIST sits in the middle of that table and is explicit about its borrowing. The AI Risk Management Framework 1.0, published in January 2023, describes an AI system as “an engineered or machine-based system that can, for a given set of objectives, generate outputs such as predictions, recommendations, or decisions influencing real or virtual environments”, and credits the 2019 OECD recommendation and ISO/IEC 22989:2022. It is voluntary, so nothing turns on whether you meet it. Its seven characteristics of trustworthy AI are the part practitioners actually cite: valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy-enhanced, and fair with harmful bias managed.
AGI has no definition anywhere, only a ladder
Article 3 has no entry for artificial general intelligence, and the Commission’s guidance never mentions it. The US statute doesn’t define it either. The most cited attempt at structure is the Levels of AGI paper by Meredith Ringel Morris and seven co-authors at Google, submitted to arXiv on 4 November 2023 and published at ICML 2024. Its authors analysed existing definitions of AGI and distilled “six principles that a useful ontology for AGI should satisfy”, which is a careful way of saying the field had no single one.
Their framework is a grid rather than a line. Performance is the depth axis and generality is the breadth axis, so a system gets a level on each.
| Level | Performance bar in the paper | General column status |
|---|---|---|
| 0: No AI | Not applicable | Human-in-the-loop computing |
| 1: Emerging | Equal to or somewhat better than an unskilled human | ChatGPT, Bard, Llama 2, Gemini |
| 2: Competent | At least 50th percentile of skilled adults | Not yet achieved |
| 3: Expert | At least 90th percentile of skilled adults | Not yet achieved |
| 4: Exceptional | At least 99th percentile of skilled adults | Not yet achieved |
| 5: Superhuman | Outperforms 100% of humans | Not yet achieved |
Read the right-hand column and the claim gets concrete. The paper puts frontier language models at Level 1 in the general column, “Emerging AGI”, and marks every level above it as not yet achieved by any public system at the time of writing. It also notes that Level 2, Competent AGI, “best corresponds to many prior conceptions of AGI”. So when a company says it is building AGI, the most precise reading available is modest. It wants to move one rung on a taxonomy its own industry published, and nobody has a standardised benchmark for that rung.
Every level above Emerging is marked “not yet achieved”. The ladder is real; the rung nobody has reached is where the argument lives.
On Table 1 of the Levels of AGI paper
General-purpose and frontier: the one number in any definition
“Frontier model” has no legal definition, but the EU built a proxy for it out of arithmetic. Article 3(63) defines a general-purpose AI model as one that “displays significant generality and is capable of competently performing a wide range of distinct tasks”, excluding models used only for research, development or prototyping before they reach the market. Article 3(66) then defines a general-purpose AI system as one built on such a model. Neither sentence contains a number.
The number sits in Article 51. A general-purpose model is presumed to have high-impact capabilities, and therefore to pose systemic risk, “when the cumulative amount of computation used for its training measured in floating point operations is greater than 1025“. That is the closest thing to a definition of a frontier model in any statute we could find. It is a presumption rather than a fact, and the Commission can amend the threshold by delegated act as hardware and algorithms improve. Training compute at that scale is expensive enough that the threshold doubles as a cost floor, a point we worked through in what it really costs to train a frontier model.
“Systemic risk” is defined too, at Article 3(65), as risk specific to those high-impact capabilities with significant impact on the Union market and effects that “can be propagated at scale across the value chain”. Even “deep fake” gets a definition, at Article 3(60): AI-generated or manipulated image, audio or video content resembling real people, places or events that “would falsely appear to a person to be authentic or truthful”. These are drafting choices with teeth, because obligations attach to them.
Foundation model, agent and token: defined by whoever shipped them
Below the statutory layer, the vocabulary belongs to whoever published first. “Foundation model” was coined in On the Opportunities and Risks of Foundation Models, submitted by a Stanford-led group on 16 August 2021. It named models “trained on broad data at scale and are adaptable to a wide range of downstream tasks”. The paper says the name was chosen to underscore the models’ “critically central yet incomplete character”, which is a narrower claim than the marketing use of the word suggests.
“Agent” is the term with the widest spread between speakers, and the most useful line we’ve read comes from Anthropic’s Building effective agents, published on 19 December 2024. It groups everything under “agentic systems”, then splits them: workflows are “systems where LLMs and tools are orchestrated through predefined code paths”, while agents are “systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks”. Plenty of what gets sold as an agent is a workflow by that test, which is the distinction doing the work in how to build an AI agent.
Two more terms are settled enough to use without hedging, and vendor documentation is the right source for both. Anthropic’s glossary defines tokens as “the smallest individual units of a language model”, which can be words, subwords, characters or bytes, and notes that for Claude a token represents roughly 3.5 English characters. It describes the context window as the text a model “can look back on and reference when generating new text”, a working memory distinct from the training corpus. That distinction is the one readers most often collapse, and we pulled it apart in context windows explained.
Open source became a definition a model can fail
“Open source AI” was the loosest term in the set until 28 October 2024. That day the Open Source Initiative published version 1.0 of the Open Source AI Definition, announced at the All Things Open conference in Raleigh. It names three components a system has to release: data information, code and parameters. Data information means “sufficiently detailed information about the data used to train the system so that a skilled person can build a substantially equivalent system”. Code means “the complete source code used to train and run the system”. Parameters means the weights and other configuration settings.
The data requirement is the one that bites, and OSI’s own announcement quotes Ayah Bdeir of Mozilla saying it “goes further than what many proprietary or ostensibly Open Source models do today”. A model shipped with downloadable weights and a licence, but no account of its training data, is open weights rather than open source under that definition. The two are not interchangeable, and the practical gap between them is the subject of open-weight vs closed AI models.
Hallucination names the symptom and hides the cause
“Hallucination” is the term with the worst fit to the thing it describes, because it implies perception going wrong rather than statistics working as designed. The clearest primary account we’ve read is Why Language Models Hallucinate, submitted to arXiv on 4 September 2025 by Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang. Its argument is that models hallucinate “because the training and evaluation procedures reward guessing over acknowledging uncertainty”.
The mechanism they set out is deflationary on purpose. Hallucinations “originate simply as errors in binary classification”, arising from natural statistical pressure when incorrect statements cannot be distinguished from facts. They persist, the paper argues, because most evaluations are graded in a way that rewards a confident guess over an admission of uncertainty. The fix the authors propose is socio-technical: rescore the leaderboards that dominate the field rather than add more hallucination evaluations. That framing is why the word itself does damage, which we’ve argued separately in why calling it a hallucination makes AI errors harder to fix.
Saying AI when you don’t have it has a price
None of this would matter commercially if the words were free to use. They aren’t. On 18 March 2024 the SEC settled charges against two investment advisers, Delphia (USA) Inc. and Global Predictions Inc., for false and misleading statements about their use of AI. Delphia paid a civil penalty of $225,000 and Global Predictions paid $175,000, a total of $400,000. Global Predictions had claimed to be the “first regulated AI financial advisor”.
Investment advisers should not mislead the public by saying they are using an AI model when they are not. Such AI washing hurts investors.
Gary Gensler, then SEC Chair, in the Commission’s 18 March 2024 press release
Neither firm admitted or denied the findings, and the penalties are small in absolute terms. The signal is still useful, because it converts a definitional argument into a disclosure question. The regulator did not rule on what AI is. It ruled on whether these firms had what they said they had, which is the version of the question a buyer can actually check.
What would change any of this
Three specific things, all written into the documents above. The 1025 threshold is amendable by delegated act, so the line between a general-purpose model and a systemically risky one can move without new legislation. Systems placed on the market before 2 August 2026 sit under the grandfathering clause at Article 111(2), so the definition’s practical reach is narrower than its text for several years. And the Levels of AGI framework is explicit that unambiguous classification “will require a standardized benchmark of tasks” that does not yet exist.
The honest limit on all of it is that no definition here measures capability directly. The EU counts floating point operations because counting capability is harder. NIST lists characteristics because it cannot set thresholds. The AGI ladder describes percentiles of skilled adults without a test that produces them. Each of those is a proxy standing in for something nobody has learned to measure yet, and the vocabulary would tighten fast if one of them arrived.
A sceptical reader might say the whole exercise is lawyers renaming software, and on the exclusions they’d have a point: a spreadsheet was never AI. On the 68 definitions, though, the counter is simple. Obligations attach to words, and those words now have tests.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
