Holy Cow Studios · AI Research · Paper One
holycowstudios.inHow it is defined, how it is built, and how it got here — a three-part empirical study for business leaders without a technical background.
Second edition · August 2026
The short version
That is not loose talk. It is a documented disagreement between the very institutions that build, sell and regulate the technology — and it has commercial and legal consequences.
We took 28 published definitions of artificial intelligence, from scholars, companies and governments, and coded each one against the same eight themes. The three groups are not using different words for the same idea. They are emphasising systematically different things, and each in the way that suits what they need the definition to do.
Vendors lean on comparison to human intelligence — useful for marketing, imprecise in a contract. Regulators avoid it almost entirely and talk instead about autonomy, dependence on data, and how general the system is: the properties you need to draw a legal boundary. Scholars sit between the two.
Then we took 27 real, documented AI systems built between 1956 and 2025 and let statistics, rather than opinion, sort them. They fall into three groups — and the interesting one is small. Four systems cluster together not because they are new or large but because they act in the world: a board-game player, a self-driving perception stack, an autonomous coding agent. Autonomy, not size or recency, is what should trigger your closest scrutiny.
Finally we checked the familiar boom-and-bust story of AI against measured research output. It holds. Publications rose from roughly 88,000 a year in 2010 to over 240,000 by 2022, and the field's two documented collapses — the “AI winters” — are visible in the record rather than being a story told afterwards.
What follows that is practical: how to read a vendor's claim, how to write a governance policy that survives the next model, and why the pace of this field has never been steady.
Three linked studies, one message
Boardrooms now discuss artificial intelligence as routinely as they discuss cash flow, yet the term means something different depending on who is speaking. A regulator, a software vendor and a computer scientist can use the identical word to describe three substantially different things.
A purposively constructed corpus of 28 definitions, stratified across academic, industry and policy sources. Coding each against eight recurring themes shows the three sectors are not merely using different words for the same idea — they are emphasising systematically different things. Industry definitions lean heavily on human comparison, useful for marketing but legally imprecise. Policy definitions almost entirely avoid it and instead foreground autonomy, data-dependence and task-generality — the properties a regulator actually needs to draw a legal boundary. Academic definitions sit in between.
An empirical taxonomy built from 27 well-documented AI systems spanning 1956 to 2025, coded on paradigm, learning method, autonomy level, data modality and parameter scale. Statistical clustering, rather than expert opinion, sorts them into three clean groups: hand-coded symbolic systems with no learned parameters; the broad family of parametric statistical learners running from early neural networks to today's large language models; and a small, historically continuous cluster of high-autonomy agentic systems that share a reinforcement-learning lineage and a willingness to act in the world, regardless of era or task. The practical implication is that autonomy, not scale or modernity, is what should trigger the closest scrutiny.
The history of AI mapped against measured publication data. Global AI-related scholarly output grew from roughly 88,000 papers in 2010 to more than 240,000 in 2022, and papers at five leading AI conferences rose tenfold between 2014 and 2024. These measured shifts line up with three qualitative shocks — the 1956 Dartmouth workshop, the 2012 deep-learning breakthrough and the 2022 generative-AI moment — each followed by an acceleration in research output, separated by two earlier periods of contraction now known as the AI winters.
AI is not a single technology with a fixed definition, but a fast-moving family of systems whose meaning, shape and pace of change depend on who is describing it and why. The report closes with a practical framework for cutting through that ambiguity when evaluating vendors, drafting governance policy, or setting the pace of AI investment.
Background, the research gap, and what each study tests
Artificial intelligence has moved from a specialist research subject to a standing item on the board agenda. Investment committees allocate capital to it, procurement teams write contracts around it, and regulators increasingly write law about it — yet the word ‘AI’ is applied to a spreadsheet macro, a recommendation engine, a self-driving car and a large language model with equal confidence. This is not simply loose talk. It reflects a genuine, longstanding and well-documented disagreement among the very institutions that build, sell and regulate the technology.1
That disagreement has practical consequences. A vendor's marketing definition of AI is not the same instrument as a regulator's legal definition of an ‘AI system’, and confusing the two creates real commercial and compliance risk: a system a supplier happily calls ‘AI-powered’ may or may not meet the legal threshold that triggers obligations under, for example, the European Union's Artificial Intelligence Act.4 Equally, executives making build-or-buy or governance decisions are routinely presented with a bewildering variety of AI systems with no common, non-technical yardstick for comparing how much genuine autonomy or risk each one carries. And strategic planning is frequently pitched as though the technology's progress were smooth and continuous, when its own history is one of sharp accelerations punctuated by lengthy periods of disillusionment.
A substantial literature exists on each of these three problems individually. What is largely missing is a single, evidence-based account that connects all three questions — what AI is called, how it is actually built, and how it got here — in language accessible to a commercially-minded reader rather than a computer scientist or a lawyer. The most comparable prior effort, the European Commission Joint Research Centre's ‘AI Watch’ programme, conducted a rigorous definitional and taxonomic exercise, but it was built to serve European research monitoring and policy needs, not business decision-making, and it did not extend to a historical or bibliometric dimension.5
The report pursues three linked empirical objectives, one per study. Each is stated as a testable proposition rather than a theme, so that the results in Section Four can be read as confirming or disconfirming something specific.
Tests whether the well-known variation in definitions of AI across academic, industry and policy sources is essentially random noise, or whether it is systematically patterned according to each sector's institutional purpose. Working hypothesis: definitions are not arbitrary — a regulator, a vendor and a scholar each define AI in the way that best serves what they need the definition to do.
Result: supported — see §4.1Tests whether real, documented AI systems fall into a small number of empirically distinct categories when classified on measurable architectural and behavioural attributes, or whether they instead form a smooth, undifferentiated continuum. Working hypothesis: a small number of coherent clusters exist, and at least one is defined more by autonomy of action than by technical sophistication or recency.
Result: supported — see §4.2Tests whether the popular boom–bust–boom narrative of AI history is reflected in measurable research activity, or whether it is a retrospective simplification imposed on what was actually smooth, continuous growth. Working hypothesis: specific, identifiable technical events coincide with measurable step-changes in research output.
Result: supported, with the data-quality caveats in §4.3Three literatures that rarely speak to one another
The difficulty of defining AI is as old as the field itself. The term was proposed in the 1955 Dartmouth workshop proposal, which described the ambition as building ‘the science and engineering of making intelligent machines’2 — a definition broad enough to cover almost any subsequent development, and precise enough to commit to almost none. Russell and Norvig's textbook remains the standard academic reference point, organising the field's competing definitions into four families along two axes: thinking versus acting, and human-like versus rational.1 That framework is analytically elegant but was never intended to resolve the definitional question for a business or legal audience.
A parallel and largely separate literature has grown up around the policy definition of AI, driven by the practical need to set the legal scope of regulation. The OECD revised its definition of an ‘AI system’ in 2023 specifically to keep pace with generative AI,3 and the EU's Artificial Intelligence Act ultimately adopted a definition closely aligned with it.4 Notably, both the United Kingdom's 2023 AI White Paper and UNESCO's 2021 Recommendation deliberately declined to commit to a single fixed definition, preferring to characterise AI by a small number of properties — adaptivity and autonomy, in the UK's case — precisely so the framework would not become obsolete as the technology changed.6
Where definitional literature is abundant, empirically-derived taxonomies of AI systems are comparatively rare. Most classification schemes in textbooks and industry commentary — narrow versus general, symbolic versus connectionist, supervised versus unsupervised versus reinforcement — are conceptual distinctions proposed by domain experts rather than categories that emerge from statistical analysis of real systems' measured attributes. That is reasonable for a technical audience already familiar with the architectures, but it offers little to a business reader trying to compare a customer-service chatbot against an autonomous vehicle on a common, defensible scale.
The narrative history of AI is well established: the 1956 Dartmouth workshop; the optimism of the 1960s; the first AI winter following the UK's 1973 Lighthill Report; the rise and collapse of expert systems through the 1980s and the second winter; the resurgence of statistical machine learning through the 1990s and 2000s; and the deep-learning revolution ignited by the 2012 ImageNet result,9 followed by the 2017 Transformer architecture10 and the generative-AI boom after ChatGPT's launch. What this narrative literature rarely does is triangulate itself against measured research output — which is the gap Study Three addresses.
Documented so another researcher could reproduce the analysis
All three datasets constructed for this report are purposive, expert-curated samples rather than exhaustive censuses. They were assembled to be defensible and replicable, not representative in a statistical sense. No confidence interval should be attached to any proportion reported below, and the sector comparisons in Study One rest on single- and low-double-digit denominators. This is revisited in §5.4.
A corpus of 28 definitional statements was assembled, stratified across three sectors: academic (13), industry (8) and policy (7). Sources were selected purposively using three inclusion criteria: the source is a named, identifiable, citable institution or scholar rather than an anonymous or aggregator source; it is widely cited within its sector, commercially prominent, or legally authoritative as appropriate; and its own definitional text, or a directly attributable paraphrase, could be verified against a primary or reputable secondary document. The corpus spans 1950 to 2026 across six geographies.
Each definition was coded against eight binary themes, chosen because they recur across the definitional literature: human comparison, learning/adaptation, autonomy, rationality/optimality, task generality, data dependence, decision/action framing, and perception.
Twenty-seven well-documented AI systems spanning 1956 to 2025 were coded on paradigm, learning method, autonomy level, data modality and parameter scale. Categorical features were encoded and numeric features scaled before clustering. Principal component analysis was applied to the same feature matrix, retaining the first two components for visualisation.
A timeline was constructed combining documented historical events with publication counts drawn from the Stanford AI Index and Center for Security and Emerging Technology data. Data points are labelled throughout as measured, interpolated or illustrative, and that labelling is material to how the timeline should be read — see the correction at §4.3.
Presented without interpretation; discussion follows in Section Five
The final corpus comprised 28 definitional statements: 13 academic (46%), 8 industry (29%) and 7 policy (25%), spanning 1950 to 2026 and six geographies.
The first edition's executive summary described the corpus as “drawn evenly from academic, industry and policy sources.” It is not even: the split is 13 / 8 / 7, or 46% / 29% / 25%. The imbalance matters, because every sector comparison below is a proportion of these unequal denominators — a single policy definition moves that sector's figure by fourteen percentage points. The summary has been corrected to say stratified, and each proportion is now reported with its base.
Figure 1 — The largest gaps between sectors
Share of each sector's definitions exhibiting a given theme, with the count behind each proportion shown on the bar. Denominators are small — 7 policy and 8 industry definitions — so these are descriptive of the corpus and carry no confidence interval. Direction is carried by the label and the count, never by colour alone.
Academic definitions show the highest rate of decision/action framing (69%, 9 of 13) and the only meaningful presence of rationality/optimality framing (38%, 5 of 13).
Principal component analysis produced a first component explaining 39.9% of variance, loading most heavily on parameter scale and autonomy level, and a second explaining 34.1%, loading positively on autonomy level and negatively on parameter scale — together accounting for 74.0% of total variance in the feature matrix.
No learned parameters. Dominant until the late 1990s. Behaviour is written by hand, so what the system will do is inspectable in principle by reading it.
Spans 1998 to 2023 and includes everything from early neural networks to today's large language models. Trailing autonomy level, mean 1.2 on the four-point scale — these systems predict, classify and generate, but mostly do not act.
Historically and technically diverse — it groups a board-game-playing agent, self-driving perception stacks and autonomous coding agents. They share a reinforcement-learning lineage and a willingness to act in the world, regardless of era or task.
The three cluster sizes sum to the 27 systems sampled. Figure 2 in the source data plots all 27 on the two retained components, coloured by cluster assignment; the substantive finding is reported here rather than depending on that projection.
Three measured figures are reported directly by the underlying sources without adaptation. Total AI-related scholarly publications rose from approximately 88,000 in 2010 to more than 240,000 in 2022.8 Combined annual output at five leading conferences — AAAI, ICLR, ICML, IJCAI and NeurIPS — rose from 1,206 papers in 2014 to 12,026 in 2024, a tenfold increase. And AI-related computer-science publications specifically rose from approximately 102,000 to 258,000 over the decade to 2025/26.7
| Year | Event | Index | Data type |
|---|---|---|---|
| 1950 | Turing proposes the imitation game | 5 | illustrative |
| 1956 | Dartmouth workshop coins ‘artificial intelligence’ | 10 | illustrative |
| 1966 | ELIZA raises expectations of near-term general intelligence | 22 | illustrative |
| 1973 | Lighthill Report; UK funding cut sharply — first winter | 14 | illustrative |
| 1980 | Expert systems such as MYCIN and XCON spread through industry | 26 | illustrative |
| 1987 | Collapse of specialised AI hardware — second winter | 18 | illustrative |
| 2000 | Statistical learning and early probabilistic methods mature | — | not estimated |
| 2010 | Statistical machine learning becomes the dominant paradigm | 88,000 | measured |
| 2012 | AlexNet wins ImageNet by a ten-point margin | 120,000 | interpolated |
| 2014 | Five leading conferences record 1,206 papers | 150,000 | interpolated |
| 2017 | Attention Is All You Need introduces the Transformer | 190,000 | interpolated |
| 2020 | GPT-3 demonstrates large-scale self-supervised modelling | 225,000 | interpolated |
| 2022 | ChatGPT reaches ~100m monthly users within two months | 240,000 | measured |
| 2024 | Five leading conferences record 12,026 papers — tenfold since 2014 | 258,000† | different series |
The 2000 row carried 88,000, labelled “measured” — the identical value and label as 2010. The cited series begins in 2010, so no measured figure for 2000 exists, and two points a decade apart cannot share a value and both be measurements. The value has been withdrawn rather than replaced, because inventing a plausible number for 2000 would be worse than admitting the series does not reach back that far.
† The 2024 figure of 258,000 comes from a different series. It is AI-related computer-science publications over the decade to 2025/26, not the general AI-related scholarly series that gives 88,000 and 240,000. Placing it as the 2024 point of the general series silently changed what was being counted. It is retained for context and labelled, rather than deleted, because the underlying growth it describes is real.
Neither correction disturbs the study's finding. The two genuine measurements — 88,000 in 2010 and 240,000 in 2022 — and the independently reported tenfold rise in leading-conference output are what the boom–bust–boom conclusion rests on.
What the results mean, and what they cannot support
Policy definitions overwhelmingly avoid human comparison — 14%, one of seven — and instead foreground data-dependence, autonomy and task-generality: properties that can, in principle, be evidenced and adjudicated. That is exactly what a definition must do if it is to carry legal weight. Industry definitions show the mirror image: the highest rate of human comparison of any sector at 75%, six of eight, which is unsurprising given that a vendor's definition typically has to communicate capability to a non-technical buyer rather than survive cross-examination.
Academic definitions sit apart from both, with the highest rate of decision/action framing (69%) and the only meaningful presence of rationality/optimality framing (38%) — the vocabulary of a field organising itself around agents that choose well, rather than around products or statutes.
The three-cluster solution did not separate systems primarily by era or by task domain. The largest cluster, parametric statistical learners (n=18), spans 1998 to 2023 and carries a trailing mean autonomy level of 1.2 on the four-point scale. The smallest, high-autonomy agentic systems (n=4), is historically and technically diverse: it groups a board-game-playing agent with self-driving perception stacks and autonomous coding agents, systems separated by decades and by application, united by the fact that they act.
For a governance policy, this is the operative finding. A risk framework keyed to parameter count or to how recently a system was built will place a very large language model in the highest tier and a modest agentic system in a lower one — when the evidence here suggests the second is the one that can take actions with consequences.
The familiar story is not a retrospective simplification. Measured research output rose from roughly 88,000 papers in 2010 to more than 240,000 in 2022, and the tenfold rise in leading-conference output between 2014 and 2024 indicates that growth has been concentrated disproportionately in the highest-impact venues since the 2012 deep-learning shock. The two winters are documented events with documented funding consequences, not a narrative convenience.
The practical implication for investment pacing is that this field's history contains no long period of steady, predictable improvement. Planning that assumes a smooth trend line is planning against the one pattern the record does not show.
All three samples are purposive, not representative. The 28 definitions, 27 systems and the constructed timeline were curated to be defensible and replicable, not drawn randomly from a defined population. No proportion reported here should carry a confidence interval, and none generalises to “all definitions” or “all AI systems”.
The sector denominators are small. Seven policy and eight industry definitions mean a single case moves a proportion by twelve to fourteen percentage points. The direction of the sector differences is the finding; the precise percentages are not.
Coding was performed by the research team, not by independent coders, so no inter-rater reliability statistic is available. The coding scheme and the full corpus are published so the exercise can be repeated and disputed.
The timeline mixes data of three qualities — illustrative, interpolated and measured — and only two points in it are measured. It is a narrative aid anchored to real figures, not a continuous measured series, and §4.3 records where an earlier edition blurred that distinction.
Practical implications for business leaders
Together the three studies support one message: AI is not a single technology with a fixed definition, but a fast-moving family of systems whose meaning, shape and pace of change depend on who is describing it and why. Three practical consequences follow.
Ask which definition they are using, and whether the system meets the regulator's definition rather than the marketing one. The gap between “AI-powered” and an ‘AI system’ in law is where compliance risk lives.
Key the risk tiers to autonomy — can it act, and with what independence — rather than to model size or novelty. That is what the cluster analysis actually separates.
Assume discontinuity. This field's measured history is sharp accelerations separated by two documented collapses, and no long period of steady, predictable improvement. A plan that depends on a smooth trend line is depending on the one pattern the record does not contain.
A regulator, a vendor and a scientist can say “AI” and mean three different things — and each is being rational, because each needs the definition to do a different job. Knowing which job is being done is the whole of reading this field clearly.
Every entry below is cited in the text; every citation resolves here
Holy Cow Studios Pvt Ltd (2026) Making Sense of Artificial Intelligence. AI Research, Paper One. Second edition, August 2026.