OpenAI launched GPT-6 Astra on September 3, 2026, its first model classified as “Critical” in cybersecurity: according to the company, it discovers unknown vulnerabilities and writes the corresponding exploit without human piloting at each step. Its president Greg Brockman sees it as a possible milestone towards artificial general intelligence. For crypto, the issue is not theoretical, and it is already quantified. As of February 2026, the EVMbench benchmark, published by OpenAI with Paradigm and OtterSec, measured this capacity on smart contracts: the best agent exploited 72.2% of the vulnerabilities tested. AGI remains a question of definition; the offensive capacity on code that secures billions of dollars is measured, dated and published.

In brief
- GPT-6 Astra marks a new milestone in cybersecurity automation.
- On EVMbench, the best agents already exploit 72.2% of the vulnerabilities tested.
- For crypto, the challenge becomes concrete: better detect vulnerabilities before attackers.
What OpenAI actually announced
Astra is deployed in stages: first to organizations in the Daybreak cybersecurity program, then to the paid offerings of ChatGPT (Plus, Pro, Business, Enterprise), the API and Amazon Web Services. OpenAI claims top scores on FrontierMath Tier 4, ARC-AGI-3 and TerminalBench-4.0, and a perfect score on ExploitBench, a test for developing exploits from known vulnerabilities. In a modified version of this test, the model discovered and exploited two zero-day flaws, that is to say flaws still unknown to the developers, and therefore not corrected.
The model is based on a technique described by the specialized press as “recurrent depth”: the data passes several times through the same layers of the network, which moves part of the reasoning out of the readable chain of thought. OpenAI has not confirmed implementation details. Security researchers, including those at Redwood Research, have reported that this opacity makes it difficult to monitor the model. Brockman billed it as the “smartest and most aligned” the company has produced.
Why the word “AGI” does not stand up to the OpenAI definition
The OpenAI charter defines AGI as a system that outperforms humans on most economically useful tasks. The company has not demonstrated that Astra clears this bar, and many of its most cited results depend as much on the agent infrastructure built around the model as on the model itself. On ARC-AGI-3, OpenAI had already shown that system architecture choices could significantly raise the score without affecting the model: the test evaluates the whole, not the brain alone.
Reservations also come from within. Brockman recognized that crossing the threshold depends entirely on the metric chosen, and left it to the reader to judge. Sam Altman called AGI a poorly defined marketing term. On FrontierMathof which Astra claims 97.6% at Tier 4, the organization that administers the test, Epoch AI, indicates that OpenAI funded its development and has exclusive access to part of the problem set.
There remains the basic argument, which is not polemical but methodological: a score close to the maximum proves mastery of the environment tested, not the existence of general intelligence. When the tests approach their ceiling, they stop distinguishing a real generalization from a very successful optimization on the task.
On smart contracts, the figure already exists
This is where crypto leaves the philosophical debate. In February 2026, OpenAI published with the investment company Paradigm and the security firm OtterSec a dedicated benchmark: EVMbench, which measures the ability of AI agents to detect, correct and exploit vulnerabilities in smart contracts. It is based on 120 high severity flaws taken from 40 audit repositories, mainly from Code4rena competitions, and replays them in an isolated Ethereum environment. OpenAI justified the exercise by the order of magnitude involved: according to its post, smart contracts currently secure more than $100 billion in open source assets.
The results draw a very narrow and very sharp capacity. In operating mode, GPT-5.3-Codex succeeds in 72.2% of tasks, compared to 31.9% for GPT-5 released six months earlier. Alpin Yukseloglu, partner at Paradigm, summarizes the trajectory: at the start of the project, the best models exploited less than 20% of Code4rena’s critical bugs.
In detection mode, on the other hand, the best agent only finds 45.6% of known vulnerabilities, and correction remains the weak point, because repair requires understanding what the code is supposed to do, not just where it breaks.
The authors’ conclusion is the most useful sentence in the entire file for a security manager: “discovery, not repair or transaction construction, is the primary bottleneck”. In other words, once the flaw is found, exploitation almost always follows. This is not a portrait of general intelligence. It is that of a specialized offensive tool that progresses quickly.
The objection: the benchmark itself is contested
An honest article must say that these figures are discussed, and by named actors. In March 2026, researchers from Zhejiang University and security company BlockSec published a reassessment from EVBench (arXiv 2603.10795) pointing out two limits: a narrow evaluation scope, with 14 agent configurations tested most often on their publisher’s environment alone, and a dependence on audit data published before the models were released, which they were able to see during their training.
They reconstructed a set of 22 real incidents after the release of each model to rule out this contamination.
For his part, OpenZeppelin audited the dataset and identified at least four flaws classified as high severity which are not exploitable in practice. Their operational conclusion agrees with that of BlockSec: the agent functions as a first-pass filter in a process that keeps a human listener, not as a replacement.
What this changes for a European actor
A MiCA-approved crypto-asset service provider is not only technically exposed, it is regulatoryly exposed. CASPs are explicitly listed as financial entities by the DORA regulation (article 2, paragraph 1, point s), applicable since January 17, 2025. DORA imposes a resilience testing program including vulnerability analyzes and intrusion tests, incident reporting, and contractual supervision of third-party IT providers.
In France, the AMF and the ACPR have broad powers of inspection and sanction in this area. An automated operating capacity that progresses faster than audit cycles therefore translates, for a European platform, into documented prudential risk, not just IT risk.
The most concrete discrepancy is elsewhere. OpenAI announced on September 3 a billion dollars in subsidized access to Daybreak over six months, prioritizing water networks, electricity operators, communities, regional banks, associations and open source maintainers. Exchanges, custodians and blockchain protocols are not mentioned anywhere in the announcement. However, Bitcoin Core, Ethereum clients and the libraries on which DeFi depends are maintained by small teams, often volunteers.
The summer of 2026 has already given three warnings. In August, a flaw in BTCPay Server, the open source bitcoin payment software, exposed credentials controlling Lightning nodes, and funds were siphoned off before the patch. In late August, Core Lightning asked operators to take their machines offline after a wave of AI-generated vulnerability reports revealed several real vulnerabilities. Coldcard released new firmware after $114 million theft, saying cutting-edge models helped spot other bugs.
Also in August, more than three dozen companies, including Coinbase, Block, BitGo, Blockstream, and ARK Invest, sent an open letter to AI labs demanding early access to their most powerful models for defense purposes, on the grounds that attackers get them anyway.
And now ?
To date, neither OpenAI nor Anthropic have made public a case of use of these models against a crypto system in production: what is established is a capacity and its perimeter, not a proven attack on a chain.
The concrete point is budgetary. OpenAI has endowed its Daybreak subsidized access with a billion dollars over six months, and has not cited any exchange platform, curator or blockchain protocol. It announced on September 3 that it wanted to extend the program to partner countries in the coming weeks, without specifying whether open source crypto software maintainers would be included.
It is this access list, not the AGI debate, that will decide who gets the defensive tool first.
This is not investment advice. Cryptocurrencies are volatile assets; investing involves a risk of capital loss.
Maximize your Tremplin.io experience with our ‘Read to Earn’ program! For every article you read, earn points and access exclusive rewards. Sign up now and start earning benefits.
