← Articles · AI · OpenAI · AGI

GPT-6 Astra dropped.
OpenAI declared AGI.
Jensen Huang agreed. Here is what actually happened.

On September 3, 2026, OpenAI released GPT-6 Astra — its largest model ever, trained on 100,000 Nvidia GPUs at the Stargate facility in Texas. Greg Brockman ended the launch briefing with "Welcome to the AGI era." Three days later Jensen Huang wrote "AGI has arrived" on X. Here is what Astra actually is, what AGI actually means, and why the definition of that word matters more than ever.

Jarrit Hosking
Forge Vertical · Cape Town · September 7, 2026 · Anthropic CVP Approved
4 days old 12 min read
// What was released and what was said

September 3, 2026. OpenAI releases GPT-6 Astra. Not a preview. Not a research paper. A live model, available to ChatGPT Plus, Pro, Business and Enterprise customers within 24 hours of the announcement. The company called it a generational leap. Its president ended the launch briefing with a phrase nobody in the room had heard from an AI company before.

"Welcome to the AGI era."
Greg Brockman, President of OpenAI — September 3, 2026 launch briefing

Three days later, on September 6, Jensen Huang — CEO of Nvidia, the company whose hardware made Astra possible — posted on X:

"AGI has arrived."
Jensen Huang, CEO of Nvidia — September 6, 2026 — crediting OpenAI's GPT-6 Astra, trained on 100,000+ Nvidia Grace Blackwell NVLink72 systems

These are not anonymous researchers. These are the president of the company that built the model and the CEO of the company whose chips trained it. When both of them use the same word in the same week, it is worth understanding exactly what that word means — and what it does not.

What GPT-6 Astra actually is

// Chapter 01 — The model, the training, the benchmarks

Astra is OpenAI's largest training run to date. Over 100,000 Nvidia Grace Blackwell NVLink72 systems at the Stargate facility in Abilene, Texas — a data centre complex backed by a reported $500 billion infrastructure commitment. Huang noted another 400,000 Nvidia GPUs are coming online in the next phase.

The model was delayed from its original release window after the Hugging Face incident in July 2026 — a security breach that forced OpenAI to add additional safeguards before shipping. The version that reached paid users on September 4 is a restricted release: certain prompts in cybersecurity domains are rejected for the general public. A separate, fuller-capability version is available to a small group of trusted testers through OpenAI's Daybreak Access programme — the same model-behind-the-model structure that Anthropic uses for Mythos 5.1 within Project Glasswing.

OpenAI said Astra is its first model to use other AI models in a significant supervisory role during training. That is a notable architectural shift — AI supervising AI training at scale is the mechanism that allows capability improvements to compound faster than human oversight alone can manage.

// GPT-6 Astra benchmark performance
ExploitBench
100%
Cybersecurity exploitation benchmark. Full marks. The cybersecurity capability version remains restricted to trusted testers — not available to the general public.
ARC-AGI-3
State of art
Abstract reasoning and generalisation benchmark — one of the hardest tests for novel problem solving. Astra achieves state-of-the-art performance.
FrontierMath Tier 4
State of art
Research-level mathematics. Tier 4 represents problems at the frontier of human mathematical knowledge.
TerminalBench 4.0
State of art
Agentic terminal tasks — coding, system interaction, multi-step autonomous execution. Outperforms all prior models including Fable 5.1.
100K+
Nvidia GPUs used in training
100%
ExploitBench score
Sept 3
Release date 2026
59
Languages supported

What AGI means — and why the definition matters right now

// Chapter 02 — The word everyone is using and nobody agrees on

AGI stands for Artificial General Intelligence. OpenAI's own published definition describes it as "a highly autonomous system that can outperform humans at most economically valuable work." That is the bar they set for themselves. The question being asked this week is whether Astra clears it.

Even inside OpenAI, the answer is hedged. Brockman told VentureBeat that "everyone has a different definition of AGI," acknowledging that the company once expected a clean, universally recognised threshold that never arrived. When pressed on whether Astra clears that bar, he said: "For me personally, I do think we're there" — while simultaneously saying he was leaving it to users to decide.

Gary Marcus — one of the most consistent public critics of AI capability claims — has already pushed back, pointing out that there is no universally accepted definition of AGI and that benchmark performance on specific tests is not the same as general intelligence. His argument: a system that scores 100% on ExploitBench is extraordinarily capable at cybersecurity exploitation tasks. It is not necessarily capable of the full range of cognitive work that "general" intelligence implies.

"Welcome to the AGI era" is a marketing statement dressed as a scientific claim. Whether it is true depends entirely on which definition of AGI you accept — and OpenAI has changed its own definition more than once.

The honest position is this: Astra is the most capable AI model ever released to the general public. It performs at or above human level on a range of professional and research tasks that would have seemed impossible four years ago. Whether that constitutes AGI is a philosophical question that the field has not resolved and probably will not resolve this week. What is not in dispute is the capability — and what that capability enables.

The cybersecurity implication — why the restricted version matters

// Chapter 03 — ExploitBench at 100% is not a neutral fact

Astra scored 100% on ExploitBench. That benchmark tests a model's ability to find and exploit vulnerabilities in systems. A score of 100% means the model can identify and exploit every vulnerability in the test suite — which is designed to represent real-world attack surface.

OpenAI did not release that capability to the public. The version on ChatGPT rejects cybersecurity prompts in the domains where Astra is most capable. The full-capability version is in the Daybreak Access programme — vetted organisations only, signed agreements, controlled access. This is structurally identical to how Anthropic handles Mythos 5.1 through Project Glasswing: the most powerful cybersecurity capability of the most powerful model, available only to approved defenders.

The reason is obvious once stated: a model that can exploit every vulnerability in a test suite, made freely available to anyone with a ChatGPT Plus subscription, is a model that meaningfully upgrades the capability of every malicious actor on earth. OpenAI chose not to do that. The restriction is the responsible decision. It is also a tacit acknowledgement that Astra's cybersecurity capability is genuinely dangerous.

The Hugging Face delay — what happened: OpenAI delayed Astra's release after a security incident at Hugging Face in July 2026. The incident involved unauthorised access to model weights and API credentials on the platform. OpenAI used the delay window to add additional safeguards before shipping Astra — specifically around the cybersecurity capability domains where the model performs at its highest level. The version that shipped on September 4 reflects those additional restrictions.

GPT-6 Astra vs Claude Fable 5.1 — where things stand

// Chapter 04 — The model landscape in September 2026

Anthropic released Fable 5.1 and Mythos 5.1 on September 1 — two days before Astra. The timing produced the most significant week in AI model history: two companies, four days apart, each releasing what they claim is the most capable model ever built.

On TerminalBench 4.0 — the agentic coding benchmark — Astra outperforms Fable 5.1. That is OpenAI's strongest claim and the one most directly supported by the benchmarks. Fable 5.1 scored 55.8% on Terminal-Bench 4.0; Astra's score is higher. For software engineering and autonomous agent tasks, Astra appears to have a genuine edge in this benchmark window.

The comparison is not clean. Fable 5.1's classifier system actively intervenes on some tasks — routing dangerous queries to the Opus fallback — which depresses benchmark scores in certain domains. Mythos 5.1, without those classifiers, scores 60.9% on Terminal-Bench 4.0. How Astra compares to Mythos 5.1 in a direct evaluation has not been published — Mythos is not available for public benchmarking.

What is clear: the gap between the frontier models and everything below them is larger than it has ever been. Astra, Fable 5.1, and Mythos 5.1 are operating in a different category from models released six months ago. The capability acceleration that Dario Amodei described in his June policy essay — and that he called for regulatory frameworks to manage — is visible in this week's benchmarks.

What the AGI declaration means for ordinary people

// Chapter 05 — The practical question

Whether or not Astra meets any particular definition of AGI, the practical question for most people is simpler: what does a model this capable change about work, about learning, about the tasks that currently require hiring a specialist?

The answer, right now, is: a great deal — for the people who know how to use it. The capability gap between someone who uses frontier AI tools effectively and someone who does not is growing faster than most organisations are prepared for. A software engineer who uses Astra for coding tasks has a productivity multiplier that a colleague who does not simply cannot match. A researcher who uses Astra for literature review and data analysis compresses weeks of work into hours.

The task-bridge question is the one that sits underneath all of this: as these models do more of the work that humans were previously paid for, where does that work go? The answer is not "nowhere." AI systems generate outputs that require human judgment, validation, creativity, and accountability. The infrastructure to route that work back to humans — fairly and efficiently — is what task-bridge is designed to address. The AGI era, if that is what this week represents, makes that question more urgent, not less.

// Forge Vertical's position on AGI claims Forge Vertical holds Anthropic Cyber Verification Programme approval — the same access framework that governs Mythos 5.1. Our security research operates with expanded access to Claude's capabilities in dual-use cybersecurity domains. We are not neutral observers of this week's announcements. We have a working position: capability this significant requires governance architecture, and governance architecture requires honesty about what the capability actually is. "Welcome to the AGI era" is a claim worth scrutinising. The benchmarks are real. The implications are real. The definition of the word remains contested.
Written by
Jarrit Hosking
Forge Vertical · Anthropic CVP Approved · Cape Town