Sergei Voronin
This page is the Version of Record of the article (English). An official Russian translation is available: iaexperts.com/journal/gonka-k-iskusstvennomu-intellektu-obshchego-naznacheniya.
2026 · online publication
UDC 004.8
Published online: August 30, 2026
Peer-reviewed: two independent reviewers
IAE · IAE REVIEW · Online publication · August 30, 2026
Comparative analysis
The Race to Artificial General Intelligence (AGI): A Comparative Analysis of Leading Contenders, Research Paradigms, and Safety Strategies (2024–2026)
The Race to Artificial General Intelligence (AGI): A Comparative Analysis of Leading Contenders, Research Paradigms, and Safety Strategies (2024–2026)
ARTICLE ID
IAE-2026-AGI-001

DOI
Pending registration · Crossref, prefix 10.68034

JOURNAL
IAE Review · online (rolling)

YEAR OF PUBLICATION
2026

ARTICLE TYPE
Comparative review

PUBLISHER
International Association of Experts, Inc.
LANGUAGE OF THIS VERSION
English · Version of Record

VERSION STATUS
Version of Record (this page) · official Russian translation

DATA CUTOFF
June 1, 2026

UDC
004.8

LICENSE
CC BY 4.0

PEER REVIEW
Independent · two reviewers
Article details
RECEIVED
June 10, 2026

REVISED
July 4, 2026
ACCEPTED
July 24, 2026

PUBLISHED
August 30, 2026 · online
Article history
Sergei Voronin
PhD in Law (Candidate of Legal Sciences) Rector of NNII · Founder · Director, IAE
ORCID: 0000-0002-8267-3502
About the author
The article presents a fact-checked comparative analysis of the race toward Artificial General Intelligence (AGI) over the period from 2024 through the first half of 2026. The landscape of key contenders — industry leaders, specialized laboratories, research institutes, and national programs — is systematized on the basis of primary sources, with a consistent distinction between figures reported by the developers themselves (self-reported) and independently verified results.

Six competing research paradigms are critically examined. The methodology and findings of the Stanford HAI AI Index reports are assessed. The article shows that by mid-2026 the frontier had entered a phase of multipolar parity, with capabilities growing faster than the ability to measure and govern them; staged criteria are proposed for sound conclusions about approaching AGI.
Abstract
В статье представлен фактологически выверенный сравнительный анализ состояния гонки за искусственный интеллект общего назначения (AGI) в период 2024 — первой половины 2026 года. Систематизирован ландшафт ключевых претендентов на основе первичных источников с последовательным разграничением данных, заявленных самими разработчиками (self-reported), и независимо проверенных результатов; критически рассмотрены шесть конкурирующих исследовательских подходов и дана оценка методологии отчётов Stanford HAI AI Index.
Аннотация
AGI
large language models
benchmarks
autonomous agents
AI alignment
open models
neuromorphic computing
AI Index
AI regulation
1. Introduction
Artificial General Intelligence (AGI) is usually defined as systems capable of solving a broad class of intellectual tasks at the level of a competent human and of transferring skills across domains without task-specific retraining. AGI differs from “narrow” AI in its universality, and from hypothetical superintelligence (ASI) in the absence of systematic superiority over humans. In practice, no strict, operationalizable, and generally accepted criterion for reaching AGI exists, which makes any claim of its “achievement” a matter for methodological caution.

The period from 2024 through the first half of 2026 proved pivotal: the center of gravity of development shifted definitively from academic teams to industry, while capital, access to compute, and the quality of internal evaluation systems became decisive competitive factors alongside model architecture.
Key thesis: by mid-2026 the race had entered a phase of “multipolar parity” — the gaps between flagship models on standard knowledge benchmarks have narrowed to a few percentage points, while system capabilities are growing faster than the tools for measuring and governing them.
2. Materials and Methods
The methodological basis of this review is a structured analysis of primary and authoritative secondary sources. Primary sources were prioritized for model capabilities, benchmark results, system characteristics, regulatory developments, and official corporate announcements; these included developer system cards, technical reports, peer-reviewed publications, regulatory documents, and official disclosures. Reputable secondary sources, including major financial and news organizations, were used where primary documentation was unavailable or where independent reporting was required to establish transaction values, company valuations, financing events, and other market developments.

Throughout the analysis, developer-reported performance was distinguished from independently verified results. Benchmark figures were treated as independently verified only where the evaluation methodology, model version, and source of the reported result could be identified. Where evidence was incomplete or materially inconsistent across sources, the relevant claim was qualified rather than treated as established fact.

Material discrepancies between developer-reported and independently reproduced results were examined individually, taking into account differences in model version, evaluation harness, benchmark configuration, sampling procedure, and reporting date. No developer-reported benchmark result was treated as independently verified solely on the basis of numerical proximity.
3. Landscape of Contenders
Note added after the data cutoff: after the review's data-cutoff date (June 1, 2026), OpenAI introduced the GPT-5.6 model family (Sol, Terra, Luna): a limited preview on June 26, 2026, and public release on July 9, 2026. These models are not included in the comparative analysis and data of this article, which are fixed as of June 1, 2026.
Google DeepMind. The flagship is Gemini 3.1 Pro (February 19, 2026). At release, Gemini 3 Pro topped LMArena (1501 Elo) and scored 91.9% on GPQA Diamond; Gemini 3.1 Pro more than doubled the result on ARC-AGI-2 (77.1%) [5].

Anthropic. The flagship is Claude Opus 4.8. Opus 4.5 became the first model to break 80% on SWE-bench Verified. With a $965 billion valuation after a $65 billion Series H round (May 28, 2026), Anthropic surpassed OpenAI for the first time [6–10].

Meta, xAI, Microsoft. Meta released its first proprietary frontier model, Muse Spark (April 8, 2026). xAI merged with SpaceX with a target IPO valuation of $1.75–2 trillion. Microsoft operates primarily through its partnership with OpenAI, with reported AI revenue on the order of $13 billion.
As of early 2026 the frontier is shared by three companies — OpenAI, Google DeepMind, and Anthropic — followed at varying distances by Meta, xAI, and Microsoft.

OpenAI. The flagship is GPT-5.5 (reportedly code-named “Spud” internally), released on April 23, 2026; it is the first fully retrained base model since GPT-4.5, natively multimodal, with a context window of up to 1 million tokens in the API. The company's valuation reached $852 billion after a $122 billion round (March 31, 2026) [3, 4].
3.1. Industry Leaders
Safe Superintelligence (SSI) was founded by Ilya Sutskever (June 2024). As a matter of principle it has no public models or products; its valuation reached ~$32 billion with a staff of about 20. DeepSeek V4 (MIT license, April 2026) reaches 80.6% on SWE-bench at a price orders of magnitude below proprietary counterparts; its reported lag behind the closed frontier is 3–6 months. OLMo 3.1 Think 32B with almost 90 times fewer parameters than Grok 4, achieves comparable results on a number of benchmarks — through careful data curation.
3.2. Specialized Laboratories and Research Institutes
Table 1. Flagship Models and Benchmarks, Mid-2026. Most values were reported by the developers themselves. Independent verification primarily covers ARC-AGI-2 through BenchLM and Epoch, the Artificial Analysis Intelligence Index, and LMArena.
Model
GPT-5.5

Gemini 3.1 Pro

Claude Opus 4.5

DeepSeek V4-Pro

Kimi K2.6

GLM-5

Mistral Large 3
OpenAI

Google


Anthropic


DeepSeek


Moonshot

Zhipu AI

Mistral
23.04.2026

19.02.2026


24.11.2025


24.04.2026


20.04.2026

11.02.2026

02.12.2025
~93 (5.2 Pro)

91,9 (3 Pro)








90,5

86,0

~44, assessment


76,2 (3 Pro)


80,9%


80,6%


80,2

77,8

85,0%

77,1%


37,6%









Proprietary

Proprietary


Proprietary


MIT, open


Modified MIT, open

MIT, open

Apache 2.0, open
Developer
Date
GPQA Diamond
SWE-bench Verified
ARC-AGI-2
License
Industrial scaling remains the dominant paradigm: training compute doubles roughly every five months, and dataset sizes every eight. Inference has become an estimated 280 times cheaper over about a year and a half at GPT-3.5-level quality. The weaknesses are the approach to informational and energy limits, opacity, and a colossal capital barrier to entry [1].
4.1. Academic Research Versus Industrial Scaling
On multi-step tasks (τ-bench / τ2-bench) even the best models succeed less than 50% of the time; the strict pass^8 robustness metric falls below 25%. On the new ARC-AGI-3 test (March 2026) all frontier systems score below 1% [31, 32]. Validity defects were found in 8 of 10 popular tests — a “do-nothing” agent passes up to 38% of tasks in one τ-bench subset [33, 34].
4.2. Autonomous Agents
Anthropic positions Opus 4.5 as “the most aligned frontier model in the industry,” approaching ASL-4 thresholds. The number of documented AI incidents rose to 362 (versus 233 a year earlier). The AI Index records a fundamental trade-off: improving one dimension of responsible AI can degrade another [1].
4.3. Safety and Alignment
Open Chinese models have reduced their gap behind the closed frontier to an estimated three to six months.
Open Chinese models have narrowed the gap with the closed frontier to 3–6 months. The Foundation Model Transparency Index fell to 40 points out of 100 (from 58 a year earlier). xAI's strategy of training on X-platform data carries serious legal risks: a Dutch court found a GDPR violation. Neuromorphic solutions (Intel Loihi 2, IBM NorthPole) demonstrate up to a 25-fold gain in performance per watt but remain at an early stage of commercialization [35–40].
4.4–4.6. Open Models, X Data, and Biologically Inspired Architectures
4. Comparison of Research Approaches
5. Critical Assessment of the Stanford HAI AI Index
The AI Index 2026 (published April 13, 2026) records: private AI investment in the United States reached $285.9 billion in 2025 (more than 23 times China's figure); generative AI reached 53% penetration within three years — faster than personal computers and the internet; 73% of experts expect a positive impact of AI on employment, versus 23% of the general public [1].
The AI Index authors themselves acknowledge key limitations: methodology changes make the series not fully comparable; some data are a “snapshot” as of a specific date; reviewers warn of the risk of models being “fit” to benchmarks (overfitting) [1, 41, 42].
6. Discussion
Taken together, the data support several generalizations. First, industry's lead over academia has become definitive. Second, there is a paradox: alongside rapid capability growth, a persistent failure on long-horizon agentic tasks remains, and transparency is declining. Third, the regulatory factor (the AI Act, the Grok case) is becoming an independent variable.

Conclusions about approaching AGI are best drawn in stages. A trigger for revision should be the sustained crossing of thresholds — nominally, above 50% on the pass^k metric on τ2-bench and above 30% on ARC-AGI-2 — by several independently verified systems. Marketing claims about a specific “probability of AGI” should be excluded from the evidence base.
7. Conclusion
By mid-2026 the race to AGI is a multipolar parity among a small group of leaders (OpenAI, Google DeepMind, Anthropic), with open Chinese models closing the gap rapidly and the influence of capital, infrastructure, and regulation growing.

The main methodological conclusion is the need for a strict distinction between reported and independently verified results, and for a cautious, staged interpretation of any claims of reaching AGI. Further research should focus on the reliability of agentic systems, the validity of benchmarks, and restoring transparency at the frontier.
The author is the founder and director of International Association of Experts, Inc., the publisher of IAE Review, and also serves as Editor-in-Chief of the journal. The manuscript was peer-reviewed by two independent reviewers. Because the authorial and editorial roles overlap, the journal does not claim independent editorial handling of this manuscript. This overlap is disclosed in the interest of transparency.
8.1. Competing interests and editorial roles
This work received no external funding.
8.2. Funding
No new datasets were created for this review; all data used are available in the cited sources.
8.3. Data availability
This article was peer-reviewed by two independent reviewers.
8.4. Peer Review
8. Declarations
References
  1. Stanford HAI. The AI Index 2026 Annual Report. Stanford University, Human-Centered AI Institute, April 13, 2026.
  2. Stanford HAI. The AI Index 2025 Annual Report (8th ed., 456 p.). Stanford University, April 2025.
  3. OpenAI. Introducing GPT-5.5. OpenAI official blog, April 23, 2026.
  4. Bloomberg. OpenAI valued at $852 billion after $122 billion round. March 31, 2026.
  5. Google DeepMind. Introducing Gemini 3 Pro. Google blog, November 18, 2025 (Gemini 3.1 Pro — February 19, 2026).
  6. Anthropic. Claude Opus 4.5 System Card. Anthropic, November 2025.
  7. CNBC. Anthropic raises Series H ($65B) at $965B valuation. May 28, 2026.
  8. Axios. Anthropic closes $65 billion Series H round. May 28, 2026.
  9. Babu J., Seetharaman D. Anthropic's valuation surges to $965 billion, surpassing OpenAI. Reuters, May 28, 2026.
  10. Bloomberg. Anthropic's Valuation Nears $1 Trillion After Raising $65 Billion. Bloomberg News, May 28, 2026. URL: https://www.bloomberg.com/news/articles/2026-05-28/anthropic-raises-at-965-billion-valuation-eclipsing-openai. Accessed August 30, 2026.
  11. CNBC. Tech AI spending approaches $700 billion in 2026, cash taking big hit. CNBC, February 6, 2026. URL: https://www.cnbc.com/2026/02/06/google-microsoft-meta-amazon-ai-cash.html. Accessed August 30, 2026.
  12. Babu J. Musk's SpaceX acquires xAI. Reuters, February 2, 2026.
  13. Babu J. SpaceX could seek IPO valuation of over $1.75 trillion, Bloomberg says. Reuters, February 27, 2026.
  14. Financial Times. Safe Superintelligence valuation. March–April 2025.
  15. The Wall Street Journal. Safe Superintelligence valuation talks. March–April 2025.
  16. Bloomberg. Thinking Machines Lab funding talks (~$50B). November 13, 2025.
  17. TechCrunch. Mistral closes in on Big AI rivals with new open-weight frontier and small models. TechCrunch, December 2, 2025. URL: https://techcrunch.com/2025/12/02/mistral-closes-in-on-big-ai-rivals-with-mistral-3-open-weight-frontier-and-small-models/. Accessed August 30, 2026.
  18. DeepSeek. DeepSeek V4 release and model card. April 24, 2026.
  19. Zhipu AI. GLM-5 / GLM-5.1 release notes. February–April 2026.
  20. Alibaba. Qwen3.5 release notes. February–May 2026.
  21. Moonshot AI. Kimi K2.6 release notes. April 20, 2026.
  22. BAAI et al. Emu3: native multimodal generation. Nature, January 2026.
  23. State Council of the PRC. Opinions on Implementing the “Artificial Intelligence+” Initiative (Guo Fa [2025] No. 11). August 26, 2025.
  24. The White House. Executive Order 14179. January 23, 2025.
  25. The White House. Winning the AI Race: America's AI Action Plan. July 23, 2025.
  26. European Commission. InvestAI. February 11, 2025.
  27. European Parliament and Council. Regulation (EU) 2024/1689 (AI Act). 2024.
  28. Government of the Russian Federation. National AI Development Strategy through 2030 (2019, rev. 2023).
  29. AI Journey 2025 conference: plenary materials. November 19, 2025.
  30. Artificial Analysis. Intelligence Index. 2026.
  31. ARC Prize. ARC-AGI-2 leaderboard. as of June 1, 2026.
  32. BenchLM.ai. ARC-AGI-3 leaderboard. as of June 1, 2026.
  33. Pan M. Z., Arabzadeh N., Cogo R., Zhu Y., et al. Measuring Agents in Production. arXiv:2512.04123. 2025. URL: https://arxiv.org/abs/2512.04123
  34. Zhu Y., Jin T., Pruksachatkun Y., Zhang A., Liu S., et al. Establishing Best Practices for Building Rigorous Agentic Benchmarks. NeurIPS 2025; arXiv:2507.02825. 2025. URL: https://arxiv.org/abs/2507.02825. DOI: https://doi.org/10.48550/arXiv.2507.02825
  35. Center for Countering Digital Hate. Grok floods X with sexualized images. 2026.
  36. Rechtbank (the Netherlands). Ruling on Grok data processing (GDPR violation). 2026.
  37. Court of Justice of the EU. Russmedia (platforms as joint controllers under the GDPR). December 2, 2025.
  38. Intel Labs. Loihi 2: neuromorphic research chip. 2023–2025.
  39. IBM Research. NorthPole: neural inference architecture. 2023–2025.
  40. Hawkins J. A Thousand Brains: A New Theory of Intelligence. Numenta, 2021.
  41. Nicoud A. Key findings from Stanford's 2025 AI Index Report (interview with V. Parli, Stanford HAI, and A. Minhas, IBM). IBM Think, April 25, 2025. URL: https://www.ibm.com/think/news/stanford-hai-2025-ai-index-report. Accessed August 30, 2026.
  42. Stanford HAI. Technical Performance. The 2026 AI Index Report. Stanford University, Human-Centered AI Institute, April 2026. URL: https://hai.stanford.edu/ai-index/2026-ai-index-report/technical-performance. Accessed August 30, 2026.
Cite This Article
Programs
Documents
Journal
International Association of Experts
An independent international professional association providing open standards for achievement verification, collegial certification, and a public registry of experts.
© 2026 International Association of Experts (IAE). All rights reserved.