Last Updated on by ICT BYTE
In recent years, OpenAI has been celebrated as the undisputed leader of the modern artificial intelligence revolution. From ChatGPT to GPT-4, the company’s generative AI systems have transformed industries, automated tasks, and captivated the global public. However, not everyone in the computer science community believes this trajectory is taking us toward true Artificial General Intelligence (AGI). In a bold critique, François Chollet, co-founder of the ARC Prize and creator of the popular Keras deep learning library, claimed that OpenAI has actually set back progress toward human-level AGI by five to ten years.
Chollet’s provocative statement highlights a growing divide in the AI research ecosystem. While tech giants pour billions of dollars into training ever-larger language models, critics argue that scaling up autoregressive text prediction will never produce genuine reasoning, adaptability, or true general intelligence. So, why does one of AI’s prominent researchers believe OpenAI delayed our path to AGI, and what does this mean for the future of synthetic cognition?
The Illusion of Intelligence: Scaling LLMs vs. Genuine Reasoning
To understand Chollet’s critique, one must first distinguish between memorization-based pattern recognition and true intelligence. Large Language Models (LLMs) operate by predicting the next token in a sequence based on vast amounts of training data ingested from the internet. They excel at producing fluent prose, summarizing complex texts, and generating code because they have seen millions of similar examples during training.
However, critics like Chollet argue that this mechanism represents a massive, highly sophisticated lookup table rather than an active cognitive engine. True intelligence, in contrast, involves fluid intelligence—the ability to face an unfamiliar problem, learn its rules on the fly, and invent novel solutions without prior exposure. While LLMs appear remarkably smart when answering queries covered in their training data, they often fail dramatically when presented with simple tasks that require novel reasoning or counterfactual logic. By equating impressive text output with human-level intelligence, the AI industry may have fallen for a convincing illusion.
How Massive Capital Diverted AI Research Off Course
The core of Chollet’s argument revolves around resource allocation and research focus. Before the dramatic rise of ChatGPT, the global AI community was actively exploring a wide variety of architectural paradigms. Researchers experimented with neuro-symbolic systems, program synthesis, causal inference engines, and modular world models, all aimed at tackling the core challenges of reasoning and generalization.
When OpenAI demonstrated the commercial success of giant transformer models, the financial incentives shifted overnight. Venture capital, corporate research budgets, and academic focus overwhelmingly concentrated on scaling up LLMs. Compute resources were monopolized by massive GPU clusters dedicated to training larger parameter models. Chollet contends that this massive drain of capital and talent away from fundamental architectural exploration created a decade-long detour. Instead of working on the missing fundamental breakthroughs required for AGI, the industry spent years squeezing marginal gains out of a single, structurally limited technique.
The ARC Benchmark: Measuring Real Adaptability
Chollet’s push back against the current LLM narrative is not merely theoretical; it is backed by concrete benchmarks. He created the Abstraction and Reasoning Corpus (ARC), an evaluation framework designed specifically to test an AI system’s capacity for general intelligence rather than memorization.
Unlike traditional AI benchmarks that evaluate tasks directly represented in common training sets, the ARC test presents abstract visual puzzles that require the AI to deduce underlying rules from only a few examples. It directly measures efficient zero-shot skill acquisition. While state-of-the-art LLMs can pass standardized bar exams and coding tests through memorization, they historically struggle on the ARC benchmark. This discrepancy underscores Chollet’s point: current models lack the dynamic problem-solving mechanisms inherent to human cognition.
Are Large Language Models Hitting a Performance Wall?
As AI companies continue to push the boundaries of model size, evidence is mounting that current architectures are approaching diminishing returns. The industry is rapidly running out of high-quality human text for training data, forcing labs to rely on synthetic data, which carries its own set of risks regarding model degradation and hallucinations.
Furthermore, the energy, financial, and environmental costs of training multi-trillion parameter models are rising exponentially relative to performance gains. If transformer-based architectures cannot achieve genuine reasoning through brute-force scaling alone, then the industry will inevitably have to return to the theoretical drawing board—the very research path that Chollet argues was interrupted by the LLM rush.
Conclusion: The Path Forward Beyond Next-Token Prediction
François Chollet’s critique serves as a vital reality check for an industry frequently caught up in marketing hyperbole. While OpenAI undoubtedly accelerated the commercial deployment and public awareness of artificial intelligence, it may have temporarily derailed the pursuit of true, adaptable AGI by fixating the world on large language models.
Achieving human-like intelligence will likely require an integration of multiple paradigms—combining the perception and intuitive fluency of deep neural networks with the rigorous logic, abstraction, and world-modeling capabilities of alternative AI architectures. Recognizing the limitations of current LLMs is not a rejection of progress, but a necessary step toward building machines that can truly think, adapt, and reason.







