OpenAI says the AGI era has begun. AI researchers aren’t so sure
Earlier this month, speaking at the launch event for OpenAI’s new GPT-6 Astra, Greg Brockman, OpenAI’s president, declared that the “AGI era” had begun. AGI, or artificial general intelligence, generally refers to AI systems that can match or surpass human abilities across virtually every cognitive task.
Long regarded as the holy grail of AI research, AGI has for decades remained a goal somewhere over the horizon. And now, apparently, it’s here—at least, if Brockman is to be believed. Nvidia CEO Jensen Huang, another one of the industry’s most influential voices, seems convinced. He congratulated OpenAI on its achievement, tweeting, “AGI has arrived.”
There is, however, a problem with declaring that AGI has arrived: There is no universally accepted threshold for what counts as AGI, and many researchers would dispute that today’s models have crossed it. That ambiguity also gives companies plenty of room to claim the milestone early, especially when doing so carries obvious marketing value. It feels, in a sense, like the race among cellular carriers to slap the next “G” on their networks before the underlying technology fully satisfies the technical standard.
Now, Astra is an impressive model. It almost aced a benchmark test designed to resist attempts by model makers to train specifically for the test, which has been a major problem with AI benchmarks. The ARC-AGI-3 benchmark, developed by AI researcher François Chollet, tests a model’s ability to encounter a completely unfamiliar situation (in this case a series of video games), figure out how it works, learn to play efficiently, and win.
But Astra’s score depends heavily on the software used to administer the test. Running through OpenAI’s own “harness,” the software layer that lets the model interact with the benchmark, Astra scored 99.9%. Using ARC-AGI-3’s standard testing setup, designed to give models a more uniform interface, Astra scored 62.7%.

In any case, the ARC-AGI-3 benchmark doesn’t represent the finish line in the race for AGI, Chollet says, because it tests a non-exhaustive set of attributes at very small scales.
“The real world features much longer time horizons for continual learning compared to ARC 3 games (decades vs minutes), much larger world modeling complexity, much greater goal ambiguity, more greater exploration spaces, etc.,” Chollet says in an email to Fast Company.
investment News
Artificial General Intelligence,Artificial Intelligence,nvidia,OpenAI
