Autonomous AI Pentesting vs Traditional Red Teams in 2026: What the Glasswing Data Tells Us

Project Glasswing's 10,000+ findings and ExploitBench's 21/41 ACEs give security buyers the clearest comparison yet between AI and human offensive security

21/41
ExploitBench ACEs for Mythos (no human team achieves this scale)
10,000+
Glasswing findings vs. typical red team scope of hundreds
10.5x
More exploits than next best model in ExploitGym
2,100+
Patches generated by Claude Security public beta

SponsoredRetool

Retool's new app builder is where AI-generated code ships safely

Building apps with AI is easy. Getting them to production safely is another story.

Start building for free today

In April 2026, Anthropic began Project Glasswing with an initial assessment phase powered by Claude Mythos. By May 22, the Exploit Evals benchmark had been published. By June 2, Glasswing had expanded to 200+ organizations across power, water, healthcare, and critical infrastructure. On July 5, the 90-day report documented 10,000+ high- or critical-severity findings, 9 confirmed CVEs, and 1,596 coordinated disclosures. No traditional red team engagement has ever produced output at that scale. A typical red team engagement, scoped over 4-8 weeks, covers hundreds of attack paths across a defined scope. Glasswing covered tens of thousands of finding-worthy conditions across an entire critical infrastructure sector. This is not an incremental improvement in offensive security capability. It is a fundamental change in what is possible. This guide uses the Glasswing data to build a grounded comparison between autonomous AI pentesting and traditional human red teams, identify where each approach is genuinely superior, and provide a framework for how security buyers should allocate their offensive security budgets.

What Traditional Red Teams Deliver

A high-quality traditional red team engagement delivers several things that are difficult to replicate with autonomous tools. Scoped adversarial simulation: the red team operates within a defined scope, uses agreed-upon rules of engagement, and simulates a specific threat actor profile (e.g., a nation-state APT targeting IP theft, or a ransomware operator targeting operational disruption). This scope discipline produces findings that are directly relevant to the organization's specific threat model. Human creativity in attack scenario development: experienced red teamers develop attack scenarios based on their understanding of the target's business, its industry, its personnel, and its specific security controls. A red teamer who understands that a target company just completed an acquisition will test the integration seams between the two organizations' security systems. An autonomous AI tool operating without that business context will not generate the same scenario. Social engineering: phishing campaigns, vishing calls, and physical access attempts require human operators. Current AI systems can assist with crafting phishing content but cannot conduct real-time social engineering over the phone or physically tailored pretexting. Business logic attack development: vulnerabilities in business processes, in workflow approvals, in multi-person authorization procedures, are found by understanding what the business is supposed to do and identifying where the implementation deviates from the intent. This requires understanding organizational context that AI tools currently cannot gather autonomously. Attestation and reporting: a report signed by a named red team with documented credentials and methodology carries weight with auditors, regulators, and cyber insurers that an AI-generated report does not yet carry.

Subscribe to unlock Remediation & Mitigation steps

Free subscribers unlock full IOC lists, Sigma detection rules, remediation steps, and every daily briefing.

What Autonomous AI Pentesting Delivers

Autonomous AI pentesting tools operating at the level of Claude Mythos deliver a different set of capabilities. 24/7 continuous assessment: unlike a human engagement that operates during business hours over a defined window, AI tools can operate continuously, identifying vulnerabilities as they are introduced rather than only at the time of a point-in-time engagement. If a developer pushes a vulnerable dependency on a Wednesday afternoon, an AI tool monitoring the codebase can flag it by Wednesday evening. A red team engagement scheduled for October will not catch it. Breadth without fatigue: human testers cover hundreds of attack paths in a typical engagement. Glasswing covered tens of thousands of finding-worthy conditions across 200+ organizations. AI tools do not experience scope fatigue, billing limits, or attention degradation. Zero-day discovery through reasoning: the 9 confirmed Glasswing CVEs, including a 17-year-old FreeBSD NFS RCE, were discovered because Mythos reasoned about conditions in the code that no one had previously documented. Human red teamers rarely discover novel zero-days in engagement timeframes because the research investment required to find a genuine new vulnerability class is not economically viable in a time-boxed engagement. Exploit chain development: ExploitBench demonstrates that Mythos can take a vulnerability from identification to working exploit code autonomously. The 21/41 ACE result means Mythos produced functional exploit code for 21 previously-disclosed V8 vulnerabilities, while every other evaluated model produced zero. Reproducibility: AI-based assessment runs are reproducible in a way that human engagements are not. The same analysis can be re-run after remediation to verify that a finding has been resolved, without rebooking a firm and waiting for availability.

Free daily briefing

Briefings like this, every morning before 9am.

Threat intel, active CVEs, and campaign alerts, distilled for practitioners. 50,000+ subscribers. No noise.

Head-to-Head Comparison Matrix

A structured comparison across the dimensions that matter most to security buyers illustrates where each approach is genuinely superior. Breadth of coverage: AI wins decisively. Glasswing's scale versus a typical red team scope makes this comparison clear. Speed: AI wins. A human engagement requires weeks of scheduling, scoping, and execution. AI assessment begins immediately. Novel zero-day discovery: AI wins. The FreeBSD 17-year-old bug and 8 other confirmed CVEs demonstrate that AI can find what has been missed for decades. Exploit development depth: AI wins on technical exploit development (ExploitBench 21/41 vs. 0). Human red teamers may develop more contextually sophisticated exploitation scenarios, but raw exploit code generation favors AI. Social engineering: Humans win decisively. Real-time conversational deception requires human judgment. Business logic attacks: Humans win. Organizational context is required. Compliance attestation: Humans win. Regulatory frameworks require human assessors. Creative novel scenarios: Humans win. AI operates within the attack patterns it has been trained on; humans can develop genuinely novel scenarios. Cost per finding: AI wins dramatically. The finding count per dollar strongly favors continuous AI assessment. Time to find: AI wins. Point-in-time human engagements miss vulnerabilities introduced between engagements.

The question is not AI or humans. The question is which function each is better at, and how to combine them for maximum coverage.

Offensive security program design principle for the AI era

Where AI Outperforms Humans Today

Three specific domains show clear AI superiority based on current evidence. Technical exploit development: ExploitBench's 21/41 versus zero result is the clearest public evidence of a capability gap. Building working exploit code for complex browser engine vulnerabilities requires deep technical knowledge, creative hypothesis generation, and iterative validation. Mythos does this autonomously at a speed and scale that human exploit developers cannot match individually. Scale without proportional cost increase: a human red team of five people cannot cover 200 organizations simultaneously. Glasswing did. This breadth is only possible through AI automation. The cost of covering ten organizations with a human red team is ten times the cost of covering one. The cost of covering ten organizations with AI is not linearly scaled in the same way. Continuous monitoring for new vulnerabilities in existing systems: human red teams conduct point-in-time assessments. The three months between the Glasswing initial assessment in April and the 90-day July report included continuous AI monitoring that identified evolving conditions. A vulnerability introduced by a system update in June would not be caught by a human engagement scoped in April.

Where Human Red Teams Remain Essential

Human red teams retain clear superiority in four specific areas that no current AI system can replicate. Physical security testing requires a human to walk up to a reception desk, tailgate through a badge access door, or plug in a rogue device. These attack vectors cannot be assessed remotely by an AI. Social engineering at a conversational level, including real-time vishing calls, in-person pretexting, and adaptive manipulation based on the target's responses, requires human judgment and improvisational capability. Business logic attack discovery depends on understanding how an organization is supposed to operate. A red teamer who has reviewed the target's annual report, understands their approval workflows, and knows that the CFO travels frequently will develop attack scenarios that an AI operating without that context cannot generate. Advanced persistent threat simulation, where a red team maintains dwell time in the environment over weeks or months, adapting to defensive responses and simulating the patience of a nation-state actor, requires human judgment about when to advance, when to hold, and how to adapt to changing network conditions. Regulatory compliance attestation remains human-required for all major frameworks. This is likely to change over time as regulators update their frameworks to recognize AI-generated assessments, but it is the current reality.

ExploitBench as a Capability Baseline for Procurement

ExploitBench provides security buyers with the first objective public benchmark for AI offensive security capability. The 21/41 result for Mythos, against zero for every other evaluated model, establishes a capability baseline that procurement teams can reference when evaluating AI security vendors. Any vendor claiming AI-powered offensive security capabilities equivalent to Mythos should be able to demonstrate performance on ExploitBench or an equivalent independent benchmark. If they cannot, they are making marketing claims that are not supported by evidence. Procurement teams should ask AI security vendors three specific questions: What is your model's score on ExploitBench or an equivalent exploit development benchmark? How does your false positive rate compare to traditional DAST tools? How many novel zero-day vulnerabilities (not known CVEs) has your tool discovered in production deployments? Vendors with genuine AI-native capability will have specific answers. Vendors using 'AI-powered' as a marketing label for ML-assisted signature matching will not.

The Hybrid Model: AI Breadth and Human Creativity

The most effective offensive security programs in 2026 combine AI and human red team capabilities in a deliberate architecture. AI provides continuous breadth coverage and technical exploit development. Human red teams focus on high-value, context-dependent attack scenarios and compliance-required assessments. Operationally, this means the following structure. Continuous AI assessment runs against the production environment and code repositories, flagging new vulnerabilities within hours of introduction and monitoring for changes in the attack surface. This layer provides the 24/7 coverage that human engagements cannot. Quarterly human-AI collaborative red team engagements use AI-generated pre-assessment findings to eliminate low-hanging fruit from the human engagement scope. Human red teamers receive a full report of AI-discovered findings before the engagement begins and focus their time on business logic attacks, social engineering, and novel attack scenario development that the AI cannot generate. Annual compliance red team engagements, conducted by a named qualified third party, satisfy PCI DSS, SOC 2, FedRAMP, and other regulatory requirements. These engagements reference AI-discovered findings as background context but produce independently validated reports. Remediation verification uses AI re-assessment after each patching cycle to confirm that findings have been resolved, without waiting for the next human engagement. This structure provides better coverage than either approach alone at a combined cost that is typically lower than running two separate independent programs.

Cost and ROI Comparison

Comparing the cost of AI and human red teaming requires accounting for the different value each provides. A traditional red team engagement from a reputable firm costs $100,000 to $300,000 for a 4-week scoped assessment. This produces a point-in-time view of the environment at the time of the engagement. Vulnerabilities introduced the day after the engagement closes are not covered until the next engagement. A continuous AI assessment program at commercial platform pricing costs $20,000 to $80,000 annually. This produces continuous coverage, meaning vulnerabilities are identified within days of introduction rather than at the next annual engagement. The coverage differential is significant: if a critical vulnerability is introduced in month 3 of a 12-month engagement cycle, it is exposed for 9 months before the next human engagement would find it. AI continuous assessment would identify it within days. The ROI calculation depends on the cost of a breach versus the cost of detection. For organizations handling financial data, patient health information, or critical infrastructure, a 9-month detection gap for a critical vulnerability is difficult to justify when continuous AI assessment is available.

Regulatory and Compliance Considerations

The regulatory landscape for penetration testing is evolving but has not yet caught up with AI capability. PCI DSS 4.0 Requirement 11.4 requires annual penetration testing and ongoing automated technical scans, but the penetration testing requirement specifies qualified human assessors. SOC 2 Trust Services Criteria reference penetration testing as a control mechanism, but the specific requirement is for testing that satisfies the auditor, who typically expects human-conducted assessments. FedRAMP High requires annual penetration testing by an approved third-party assessment organization (3PAO), all of which are human teams. HIPAA Security Rule requires technical safeguards testing, but the specific modality is not prescribed, creating potential flexibility for AI-based assessment as a supplement. The practical guidance for compliance-driven organizations is to maintain human red team engagements as the primary compliance attestation mechanism while deploying AI assessment as a continuous improvement layer. As regulatory frameworks update to recognize AI-generated assessments (which is likely within the next 2-3 years based on the pace of AI security adoption), the compliance calculus will shift.

AI-Augmented Red Team Program Design and Tooling Recommendations

The complete hybrid red team program design, including AI tool selection criteria for offensive security, integration architecture for combining AI continuous assessment with human engagement cycles, and tooling recommendations based on Glasswing findings, are in the Mythos Brief.

Subscribe to unlock Remediation & Mitigation steps

Free subscribers unlock full IOC lists, Sigma detection rules, remediation steps, and every daily briefing.

The bottom line

Project Glasswing's 90-day report is the most comprehensive public dataset on autonomous AI offensive security capability available in 2026. The 10,000+ findings, 9 CVEs, and ExploitBench 21/41 ACE score demonstrate that AI-native pentesting operates at a scale and technical depth that no human red team can match in breadth. But breadth is not the whole picture. Human red teams retain decisive advantages in social engineering, business logic attack development, creative novel scenarios, and compliance attestation. The organizations that will build the strongest offensive security programs are those that treat AI and human capabilities as complementary rather than competitive, deploying AI for continuous breadth and technical exploit development while preserving human engagement budget for the scenarios where human judgment is irreplaceable. The complete hybrid program design template, compliance mapping matrix, and ROI calculator are in the Mythos Brief. Get it free at decryptiondigest.com/mythos-brief.

Frequently asked questions

Can AI replace a human penetration tester?

For specific, well-defined offensive security tasks, AI has already surpassed human performance. ExploitBench's result, 21 of 41 V8 ACE challenges for Mythos versus zero for all other models including human-comparable baselines, demonstrates AI superiority in narrow exploit development domains. However, human red teams provide capabilities that autonomous AI cannot yet replicate: business context understanding, social engineering, physical security testing, creative novel attack scenario development, and the judgment to prioritize findings by business impact rather than technical severity. The best answer in 2026 is not replacement but specialization: AI for breadth and technical depth, humans for creativity and business context.

Does an AI pentest satisfy compliance requirements?

Currently, no. Major compliance frameworks (PCI DSS 4.0, SOC 2, FedRAMP, HIPAA) reference penetration testing in terms that require human-conducted assessments or qualified assessors. PCI DSS 4.0 Requirement 11.4.3 specifies a 'qualified internal resource or qualified external third-party' for penetration testing, with qualifications defined around human expertise and certifications. AI-generated penetration test reports are not yet recognized by compliance auditors as satisfying these requirements. Organizations should treat AI pentesting as an augmentation to their compliance-satisfying human engagements, not a replacement for them.

How much does AI pentesting cost compared to a red team engagement?

A traditional red team engagement typically costs $50,000 to $500,000 depending on scope, duration, and firm, and produces findings from a team of 2-5 assessors working over 2-8 weeks. AI pentesting costs vary by deployment model: commercial platforms are typically subscription-based at $10,000 to $100,000 annually, providing continuous rather than point-in-time assessment. Project Glasswing operates as a coordinated vulnerability disclosure partnership with Anthropic rather than a paid service. The cost comparison is complicated by the fact that AI pentesting operates continuously while traditional engagements are point-in-time, making a direct per-engagement cost comparison misleading.

What is the best way to combine AI and human red teaming?

The most effective hybrid model uses AI for continuous breadth coverage and technical depth, with human red teamers focused on business logic attacks, social engineering, novel creative scenarios, and compliance attestation. Operationally, this means running AI-based automated assessment continuously against the environment while scheduling human red team engagements quarterly or annually, with the human team briefed on AI-discovered findings so they can focus on higher-order attack scenarios that AI cannot generate. AI pre-assessment also makes human engagements more valuable by eliminating low-hanging fruit, forcing human testers to focus on genuinely hard problems.

How does Claude Mythos compare to other AI pentest tools?

ExploitBench provides the clearest public comparison: Mythos solved 21 of 41 V8 ACE challenges, and every other evaluated model scored zero. ExploitGym shows Mythos producing 10.5 times more exploits than Opus 4.6, the next best model. XBOW, a leading AI security research team, publicly praised Mythos as a significant step up over all existing models. These are the most objective public comparisons available. Most commercial AI security vendors do not publish benchmark performance data, making independent capability comparison difficult. Organizations evaluating AI pentest vendors should ask specifically about ExploitBench-equivalent results.

What organizational prerequisites must be in place before running an autonomous AI penetration test against a production environment?

Before deploying any autonomous AI offensive security tool against production, organizations need four foundations. First, a complete and current asset inventory with ownership records: the AI assessment must be scoped explicitly to authorized systems, and every system in scope needs a documented owner who has consented to testing. Second, a written rules of engagement document signed by an executive sponsor that defines the testing scope, permitted attack categories, data handling requirements for discovered secrets or credentials, and the incident response contact chain if the testing triggers defensive alerts. Third, a staging or canary environment that the AI tool can use for initial validation of exploit chains before running them against production, reducing the risk that a proof-of-concept causes unintended service disruption. Fourth, a pre-notified SOC and SIEM suppression plan for expected testing noise so that the AI tool's reconnaissance and exploitation activity does not trigger a false incident response, while preserving alerting for activity outside the expected testing window and scope.

Sources & references

  1. Anthropic Project Glasswing 90-Day Report
  2. PTES Technical Guidelines
  3. MITRE ATT&CK Framework
  4. NIST SP 800-115 Technical Guide to Information Security Testing
  5. XBOW AI Security Research
  6. Anthropic Claude Security Public Beta

Free resources

25
Free download

Critical CVE Reference Card 2025–2026

25 actively exploited vulnerabilities with CVSS scores, exploit status, and patch availability. Print it, pin it, share it with your SOC team.

No spam. Unsubscribe anytime.

Free download

Ransomware Incident Response Playbook

Step-by-step 24-hour IR checklist covering detection, containment, eradication, and recovery. Built for SOC teams, IR leads, and CISOs.

No spam. Unsubscribe anytime.

Free newsletter

Get threat intel before your inbox does.

50,000+ security professionals read Decryption Digest for early warnings on zero-days, ransomware, and nation-state campaigns. Free, daily, no spam.

Unsubscribe anytime. We never sell your data.

Eric Bang
Author

Founder & Cybersecurity Evangelist, Decryption Digest

Cybersecurity professional with expertise in threat intelligence, vulnerability research, and enterprise security. Covers zero-days, ransomware, and nation-state operations for 50,000+ security professionals every morning.

Black Hat Giveaway

Win a $2,495 Black Hat pass.

Full-access to Black Hat USA 2026 in Las Vegas. Subscribe free to enter.

Joins Decryption Digest daily briefing. Unsubscribe anytime.

Giveaway: Black Hat USA 2026 Full-Access Pass ($2,495 value)

Details →
Daily Briefing

Subscribe to enter the giveaway

Every subscriber is automatically entered. You also get daily threat intel every morning: zero-days, ransomware, and nation-state campaigns. Free. No spam.

Already subscribed? You're already entered.

Giveaway

Win a $2,495 Black Hat USA 2026 pass.