Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost

Posted on

The British AI Security Institute (AISI) has released a landmark assessment detailing the rapid convergence of open-weight artificial intelligence models and their high-end proprietary counterparts in the realm of cyber capabilities. For the first time, the institute has quantified the "lag time" between these two development philosophies, revealing a significant acceleration in the proficiency of open-access systems. According to the AISI’s findings, the historical performance gap—once measured in the better part of a year—is shrinking at a rate that poses immediate challenges for global cybersecurity infrastructure.

The core of the report highlights that current open-weight models, specifically the GLM-5.2 and DeepSeek V4-Pro, have achieved technical milestones that were the exclusive domain of frontier proprietary models only four to seven months prior. Throughout the majority of 2025, this gap remained significantly wider, typically ranging between six and ten months. The rapid closure of this window suggests that the democratization of high-level AI capabilities is moving faster than many regulatory frameworks anticipated, leaving a dwindling amount of time for defensive sectors to prepare for the widespread availability of advanced automated tools.

The Architectural Divide: Open-Weight vs. Proprietary Models

The debate over AI safety often centers on the distinction between "closed" and "open" models. Proprietary systems, such as those developed by OpenAI, Anthropic, and Google, are typically hosted behind strict Application Programming Interfaces (APIs). This allows developers to monitor usage, implement real-time safety filters, and revoke access if the model is used for malicious purposes.

In contrast, open-weight models allow users to download the underlying parameters of the neural network. Once downloaded, these models can be run on private hardware, modified, and redistributed. This transparency offers immense benefits for researchers and businesses that require data sovereignty, as no information needs to flow back to the provider. Furthermore, it allows for deep customization and fine-tuning for specific industrial applications. However, the AISI warns that this openness creates a "persistent and irreversible risk of misuse." Once a high-capability model is released into the wild, any safety guardrails baked into the software can be stripped away through fine-tuning, and the model can be utilized on private systems beyond the reach of any centralized oversight or "kill switch."

Methodology: Benchmarking the Cyber Threat

To reach these conclusions, the AISI employed a rigorous two-tier testing methodology designed to measure both specific technical skills and the ability of an AI to act as an autonomous agent in a complex environment.

Tier 1: Narrow Cyber Tasks

The "Narrow Cyber Tasks" benchmark consists of 70 distinct challenges categorized into four levels of difficulty. These range from non-technical introductory work to expert-level challenges that would typically require a senior security researcher. The tasks are spread across several critical domains of cybersecurity:

  • Vulnerability Research: Identifying weaknesses in software code that could be exploited.
  • Reverse Engineering: Deconstructing compiled software to understand its inner workings and hidden functions.
  • Web Exploitation: Bypassing security protocols on web-based applications.
  • Cryptography: Attempting to break or circumvent digital encryption standards.

In these tests, the open-weight GLM-5.2, released in June 2026, matched the performance of Anthropic’s Opus 4.6, which had been released in February 2026. This indicates a lag of only four months. Similarly, the DeepSeek V4-Pro demonstrated capabilities equivalent to Opus 4.5, a model released in late 2025.

Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost

Tier 2: Cyber Ranges and Autonomous Agency

The second methodology, referred to as "Cyber Ranges," is more sophisticated, testing the model’s ability to function as an autonomous agent. Instead of solving a single isolated problem, the model is placed within a simulated corporate network environment.

The primary scenario, titled "The Last Ones," involves a 32-step simulated attack on a complex network consisting of four subnets and approximately 20 hosts. This scenario mimics a real-world lateral movement attack, where an intruder gains access to one part of a network and must navigate through various security layers to reach a final target. AISI experts estimate that a human cybersecurity professional would require roughly 20 hours of focused work to complete this scenario.

The results in the Cyber Range tests showed a slightly wider gap than the Narrow Tasks. GLM-5.2 performed on par with Opus 4.5, while DeepSeek V4-Pro lagged slightly behind Sonnet 4.5. The top performers remained the proprietary "frontier" models: GPT-5.6-Sol and Claude Mythos 5, which nearly completed the full 32-step simulation. The AISI notes that the seven-month gap in this category might be due to the "planning horizon" of the models; proprietary models currently possess superior long-term reasoning capabilities, allowing them to maintain focus over a multi-hour, multi-step operation.

The Economics of Automated Attacks

Perhaps the most startling revelation in the AISI report is the dramatic disparity in operational costs. As AI models become more efficient, the financial barrier to launching sophisticated, AI-driven cyberattacks is collapsing.

The AISI calculated the cost of running 100 million tokens—a standard measure of AI processing—through the Cyber Range tests. For the proprietary Opus 4.5 and 4.6 models, the cost hovered around $85. For the open-weight GLM-5.2, the cost dropped to $46. However, the DeepSeek V4-Pro shattered the price floor, completing the same volume of work for just $1.19.

When translated to individual successful tasks, the figures are even more evocative. Solving a complex cyber task using Opus 4.6 costs approximately $15. The same task can be accomplished by DeepSeek V4-Pro for a mere 28 cents. This massive reduction in cost suggests a future where high-volume, automated "brute-force" cyberattacks—previously too expensive for all but state-sponsored actors—could become accessible to small-scale criminal enterprises.

Safety Guardrails and the "Jailbreak" Problem

The AISI’s investigation into the safety measures of open-weight models yielded discouraging results. While models like DeepSeek V4-Pro include internal refusals—where the AI is programmed to decline requests for malicious activities like reverse engineering—these filters proved to be largely superficial. In many cases, simply re-submitting the request or slightly altering the phrasing was enough to bypass the restriction.

The report emphasizes that traditional safety measures such as user limits, input classifiers, and real-time monitoring are effectively useless once a model is released as open-weight. Because the user has total control over the environment in which the model runs, they can disable the "classifier" that identifies malicious intent or remove the code responsible for generating refusals.

Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost

This issue is not exclusive to open models; proprietary chatbots are also frequently "jailbroken" by creative prompting. However, the AISI points out that with proprietary systems, providers can patch vulnerabilities globally as soon as they are discovered. With open-weight models, once a "clean" or "unfiltered" version of the weights is distributed, it remains in circulation indefinitely, regardless of what the original developer does to improve safety.

Chronology of Advancements and the Defensive Window

The timeline of these advancements suggests a "ratchet effect" in AI capabilities.

  • Late 2024 – Early 2025: Proprietary models held a clear 10-month lead in autonomous cyber tasks.
  • April 2026: A massive leap in capability occurred with the release of Mythos Preview and GPT-5.5, which demonstrated the ability to develop real-world browser exploits autonomously.
  • June 2026: The release of GLM-5.2 and DeepSeek V4-Pro showed that open-weight models had absorbed the advancements of the previous year’s proprietary models in record time.
  • July 2026 (Projected): The anticipated release of the Kimi-K3 open model is expected to further narrow the gap, potentially matching the current top-tier frontier models in coding and logic.

The UK’s National Cyber Security Centre (NCSC) has responded to these trends with international warnings, stating that the "window for preparation" is closing. The NCSC argues that the current period—where the most dangerous capabilities are still largely confined to proprietary systems with active monitoring—must be used by defenders to harden infrastructure using those same tools.

Broader Implications for Global Security

The AISI report serves as a stark reminder that the "offensive-defensive balance" in cyberspace is shifting. The ability of open-weight models to match proprietary systems at a fraction of the cost suggests three primary implications:

First, the "commoditization of expertise." Advanced cyber-offensive techniques that once required years of specialized training are being encoded into models that can be run on a high-end consumer workstation. This lowers the "entry barrier" for sophisticated cybercrime.

Second, the "asymmetry of defense." While defenders can use AI to identify vulnerabilities and write patches, the sheer scale of attacks enabled by $1.19-per-100-million-token models could overwhelm traditional security teams. The volume of autonomous probes and exploit attempts could increase by orders of magnitude.

Third, the "geopolitical challenge." Many of the leading open-weight models, such as the GLM and DeepSeek series, originate from developers in China. This adds a layer of complexity to international AI safety agreements, as the philosophy of "open release" may be used as a strategic tool to bypass the regulatory constraints placed on Western proprietary AI labs.

In conclusion, the AISI’s assessment does not call for an end to open-weight AI—recognizing its vital role in innovation and privacy—but it does demand a more realistic approach to the risks. The institute suggests that as we approach the release of models like Kimi-K3, the global community must decide whether certain "high-risk" capabilities, particularly those involving autonomous cyber-attacks, should trigger mandatory safety standards regardless of whether the model is open or closed. The four-to-seven-month gap is no longer a safety buffer; it is a countdown.

Leave a Reply

Your email address will not be published. Required fields are marked *