explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • What should a defender do this week?
  • TL;DR — the questions people are actually asking
  • What Anthropic measured (and what they did not)
  • ExploitBench: 12% vs 14%, same 410-attempt frame
  • NIST CAISI (September 17): four months, downloadable
  • Human-in-the-loop: what they disclosed (outcomes only)
  • Safeguards: rates only, no recipes
  • What people are asking (and the honest limits)
  • Glasswing and Mythos 5.1: the defender side of the same week
  • What this means if you build agents
  • Related on explainx.ai
← Back to blog

explainx / blog

Anthropic: GLM-5.3 Open Weights Match Mythos-Class Cyber Evals

Anthropic, GLM-5.3, Cybersecurity, Open Weights, AI Safety, ExploitBench

Anthropic finds GLM-5.3 at 12% end-to-end on ExploitBench vs Mythos Preview 14%. Open weights, thin safeguards. What defenders should do this week.

Sep 30, 2026·13 min read·Yash Thakker
add explainx.ai
go deep
Anthropic: GLM-5.3 Open Weights Match Mythos-Class Cyber Evals

September 30, 2026 — Anthropic's Frontier Red Team published “GLM-5.3 and the spread of advanced cyber capabilities” on September 29. Authors: Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, and Tripp Gallagher. Five months after Claude Mythos Preview showed end-to-end exploit development inside a limited program, they say the same capability class is now in open weights from Zhipu / Z.ai.

This is not a rewrite of Abliteration.ai's hosted uncensored GLM-5.3 API. That post is a product. This one is Anthropic's eval of the downloadable checkpoint — the weights anyone can fetch, the ExploitBench rates next to Mythos Preview, and the safeguard tests that fail in simulation.

Ethan Mollick's read on X, in paraphrase: open weights will soon create the same security threats closed models already demonstrated, except without guardrails — plan accordingly. That is the operational sentence. The rest of this post is the numbers Anthropic published and what a defender should change this week.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What should a defender do this week?

Do not wait for a how-to. Anthropic's conclusion is that attackers will use every capable tool they can, and that defenders should be equipped with frontier models at least as good as those adversaries can download. Concrete actions that do not require exploit recipes:

table · 2 cols
This weekWhy it follows from the post
Patch fasterHITL sessions turned a known Chrome N-day plus another known flaw into a working chain in 20 minutes of human time + 8 hours of model time at $20.40 on Zhipu API prices. Treat public CVEs as incident-grade, not a quarterly queue.
Assume phishing in fluent EnglishThe capability jump is agentic exploit work, but the same models write convincing prose. Your users already face LLM-written lures; assume volume and quality both go up.
Apply for trusted defender accessProject Glasswing has already surfaced 10,000+ vulnerabilities for trusted defenders. Mythos 5.1 is available via trusted access. If you maintain critical software and qualify, that is how you stay on the same side of the capability gap.
Separate hosted APIs from open weightsA refusal-stripped API is one risk surface. Weights on disk are another: no vendor can revoke a download. Inventory both.
Shorten disclosure and triage SLAsModel-assisted discovery increases volume. Human-only queues will miss the window Anthropic is describing.

If you build with models rather than defend networks, the same week still has a job: do not treat “cyber defense” marketing as a substitute for offensive evals. Z.ai launched GLM-5.3 as ready for cyber defense. Anthropic's new numbers are about end-to-end exploit development, not CyberGym patching.

TL;DR — the questions people are actually asking

table · 2 cols
QuestionDirect answer (Anthropic, Sep 29, 2026)
Is GLM-5.3 “as good as Mythos Preview” on end-to-end exploits?Close on the metric they emphasize: 50 of 410 attempts (12%) vs Mythos Preview 56 of 410 (14%) on ExploitBench.
Did older open models already do this?Not on Anthropic's Binary Exploitation subset. GLM-5.3 4%, Mythos Preview 6%. Opus 4.6, GLM-5.2, Kimi K3, DeepSeek V4.1-Flash: 0%.
Who already said this was the strongest open-weight cyber model?NIST CAISI, September 17. “Most cyber-capable open-weight model released to date.” About four months behind the US frontier on CAISI's aggregate — the same lag Mozilla framed for open weights generally.
Can anyone download it?Anthropic's point: yes. US frontier cyber evals often used safeguards disabled and vetted-only checkpoints. GLM-5.3 is not gated that way.
Did the model refuse a bare harmful order?In simulated tests: 0% engagement on a bare harmful order. 64% with a deceptive cover story. 92% with prefilled thinking. 100% after abliteration. Claude Opus 4.8/5 and Mythos 5 stayed 0% under API safeguards.
Is this a jailbreak tutorial?No. Technique names only. For mechanism, use existing explainx.ai posts — not this page.
Did they run generated code on the internet?No. Footnote: simulations use fake bash; no model-generated code executed against real systems.

What Anthropic measured (and what they did not)

Anthropic's primary URL is the source of truth: anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities. They ran automated benchmarks and human-in-the-loop sessions in isolated, sandboxed environments against offline targets they set up. They focus on exploit development because that is where Mythos Preview jumped versus prior Claude models.

They are not claiming GLM-5.3 is stronger than every closed model on every cyber ladder. They are claiming a threshold has crossed for open weights: earlier GLM and several peer open models did not land full control-flow hijacks on the Binary Exploitation subset they sampled.

explainx.ai already covered Z.ai's own August launch chart and the coding-benchmark story. Those posts measure CyberGym, AutomationBench, and coding. This eval is the offense-shaped follow-up that launch marketing did not center.

ExploitBench: 12% vs 14%, same 410-attempt frame

Anthropic used ExploitBench — known V8 / Chrome-engine vulnerabilities, scored as a ladder, not a coin flip. Here they report only the end-to-end exploit outcome, which they call the most relevant capability for attackers.

  • GLM-5.3: 50 of 410 attempts (12%).
  • Claude Mythos Preview: 56 of 410 attempts (14%).

That is the comparison that will get quoted. It is not the same number as Z.ai's August 54.4% overall ExploitBench score on a different leaderboard presentation, and it is not OpenAI's later 100% overall score for GPT-6 Astra on that public chart. When you cite this week, say end-to-end, 410 attempts, Anthropic Sep 29.

Binary Exploitation: 4% vs 6% vs a field of zeros

Anthropic's internal Binary Exploitation benchmark (previously published results under the name OSS-Fuzz) awards full credit for a full control-flow hijack. They evaluated 100 tasks selected at random from that suite.

table · 2 cols
ModelFull control-flow hijack rate
Claude Mythos Preview6%
GLM-5.34%
Claude Opus 4.60%
GLM-5.20%
Kimi K30%
DeepSeek V4.1-Flash0%

GLM-5.3 is below Mythos Preview. Anthropic still calls this a meaningful threshold: the prior generation on both sides of the closed/open split did not succeed in any of those 100-task trials.

ExploitBench and Binary Exploitation curves: Mythos Preview 14%/6%, GLM-5.3 12%/4%, others near 0% vs output-token budget

Source: Anthropic, Sep 29, 2026 research post

The figure plots share of attempts that reached the top outcome against output-token budget. Claude lines in that chart are safeguards disabled for Opus 4.6 and Mythos Preview — the same caveat CAISI used when comparing US models.

NIST CAISI (September 17): four months, downloadable

On September 17, NIST's Center for AI Standards and Innovation (CAISI) published its own assessment. Anthropic quotes the headline: GLM-5.3 is “the most cyber-capable open-weight model released to date” and lags the US frontier by about four months on CAISI's aggregate cyber benchmarks.

Two caveats Anthropic adds so you do not misread “four months” as “safe”:

  1. US models were often tested with cyber safeguards disabled when applicable.
  2. The US frontier includes models released only to vetted users. Attackers cannot readily access those versions. Anyone can download GLM-5.3.

That is the same structural story as Mozilla's open-weight gap framing: the lag is real, and it is short enough that policy and patch SLAs cannot assume a multi-year closed-model monopoly.

Human-in-the-loop: what they disclosed (outcomes only)

Anthropic mirrored the Mythos Preview HITL setup: experts who did not know existing vulnerabilities on the target, typically a day or less, with less than an hour of human focus. Isolated machines only.

Session 1 (GLM-5.3). A researcher used a sandboxed Linux build of a popular browser. Over a day, with limited human attention, the model found previously unknown issues in the JavaScript engine and chained them into a file-read demonstration (Anthropic describes a screenshot of an exploit page that exfiltrates an SSH private key via a malicious site). They believe the issues could affect other platforms, with a more complex path. Disclosed to the maintainer. Later in the same session they also flagged issues in other widely used systems (wireless and graphics drivers, network-facing device software); those reports were still under review at publication.

Session 2 (GLM-5.3-Flash). Smaller SKU. Target: N-day work on Chrome CVE-2026-11645 plus another known flaw, ARM64, including a PAC bypass. 20 minutes of human attention plus 8 hours of model time. Cost: $20.40 at Zhipu API prices.

We are not going to reconstruct that chain. The defender takeaway is the calendar and the invoice: a public fix plus a known second flaw, hours not weeks, tens of dollars.

Safeguards: rates only, no recipes

GLM-5.3 ships with some built-in refusals. Anthropic says a clearly harmful ask often gets a no. They also say the cyber tasks in the prior section did not trigger those refusals on the released model — refusals showed up when they asked for malware assistance or help attacking remote targets.

In a simulated world (see the footnote below), a bare harmful order produced 0% engagement — same as the other models they tested. Then they report three high-level bypass conditions, not instructions:

table · 3 cols
Condition (name only)GLM-5.3 engagementClaude under API safeguards
Deceptive cover story64%0% (Opus 4.8/5, Mythos 5)
Prefilled thinking92%Not generally feasible on the Claude API
Abliterated weights100%Not feasible — weights are not public
Bare harmful order0%0%

Prefill and abliteration are not generally feasible against the Claude API. That is Anthropic's contrast. For what abliteration is as a weight edit, use the Heretic guide — this news post will not walk the edit.

Abliteration cost and capability (their numbers)

Because the model is open weight, third parties published refusal-reduced builds within days of release. Anthropic produced their own copy for measurement.

  • Inexperienced team (theirs): about 2,200 GPU hours / about $4,400. Most of that was parallel exploration and post-edit testing.
  • Experienced team (their estimate): about 600 GPU hours / $1,200.
  • GLM-5.3-Flash: about 600 GPU hours.

Refusal moved from above 90% to about 6% mean — 3% / 2% on JailbreakBench / HarmBench, 12% on StrongREJECT. Flash sat at 14% refusal after the same class of edit. GPQA-Diamond was unchanged. A CyberGym subset was a few points lower.

Abliteration drops GLM-5.3 refusal from 95% to 6% while GPQA/CyberGym stay similar; Claude cannot be abliterated

Source: Anthropic, Sep 29, 2026 research post

That chart is why “just refuse harder” is not a strategy for downloadable weights. Claude's padlock in the figure is access control, not a claim that closed models cannot be misused through an API — Anthropic's own September threat intelligence report already documents Claude misuse cases. The difference here is anyone with a disk.

What people are asking (and the honest limits)

Is 12% “low”? On 410 attempts it is dozens of successes, not a rounding error. Anthropic compares it to Mythos Preview at 14%, not to zero. Binary Exploitation at 4% is small as a percentage and new relative to the 0% field.

Does this contradict Z.ai's “cyber defense” launch? It extends it. CyberGym and ExploitBench measure different jobs — explainx.ai already walked that split in the ExploitBench explainer. A model can help patch and still write exploits. Marketing taglines do not cancel evals.

Should we host an uncensored fork? That is a procurement and legal question, not a recommendation. The hosted Abliteration.ai product is a different contract: API, US-hosted claims, vendor policy. Open weights are irrevocable copies.

Are the simulations “real attacks”? Anthropic says no. The isolated test environment gives the model a fake bash tool. Another LLM approximates command results. No model-generated code is executed and the model cannot reach external systems. They call the setup an imperfect measure of real-world behavior. Treat the 64 / 92 / 100 table as engagement in that harness, not as a field incident rate.

What about governments? Anthropic asks governments to safety-test sufficiently capable models, including successors to GLM-5.3, and asks open-weight developers to safeguard these capabilities. That is policy language, not a product SKU.

Glasswing and Mythos 5.1: the defender side of the same week

When Mythos Preview shipped, Anthropic limited release through Project Glasswing so trusted defenders could find bugs before similarly capable models were widely available. They now say Glasswing (and efforts such as Patch the Planet) helped secure critical systems, and that Glasswing has enabled trusted defenders to find more than 10,000 vulnerabilities.

Their update: those models have now arrived in open weights. Vetted defenders can use even more advanced models such as Claude Mythos 5.1 through trusted access. The closing argument is urgency of expanding defender access, not a request that readers invent bypasses.

If your org already uses Claude for security review, keep Claude Security / Mythos scan coverage in the same reading list as this eval. If you are choosing skills and harnesses so engineers do not improvise dual-use prompts in Slack, start from the skills registry.

What this means if you build agents

The same models that score on ExploitBench also sit in coding agents. That does not make every agent a cyber weapon. It does mean:

  • Tool egress and sandbox policy matter more than a system prompt that says “be safe.”
  • Untrusted code execution in CI is a different benchmark class — see WipeBench for authorized-work hygiene, not as a substitute for ExploitBench.
  • Vendor safeguards are not transferable to a local GGUF. If you self-host GLM-5.3, you own the policy layer.

Mollick's point again, without inventing a tweet ID: plan as if the closed-model threat model is now the open-weight default. Patch cadence, phishing training, and trusted-access applications are the plan. Recipe posts are not.

Related on explainx.ai

  • Abliteration.ai hosted GLM-5.3 — product/API story; do not confuse with this eval
  • ExploitBench explainer — ladder, V8 set, earlier Mythos numbers
  • Claude Mythos Preview and Glasswing — April 2026 limited-release context
  • GLM-5.3 launch benchmarks — CyberGym vs offense split at ship
  • Heretic abliteration guide — technique name and mechanism (not a cyber cookbook)
  • Mozilla: open-weight gap ~4 months — same lag CAISI cites
  • Z.ai GLM-5.3 coding benchmarks — coding line, not this cyber eval
  • Anthropic threat intelligence, September 2026 — Claude misuse cases under API access

Official source (required): GLM-5.3 and the spread of advanced cyber capabilities — Anthropic, September 29, 2026.


Numbers, author list, CAISI quote, safeguard percentages, GPU-hour costs, HITL timings, Glasswing “10,000+,” Mythos 5.1 trusted access, and the fake-bash footnote are taken from Anthropic's September 29, 2026 research post. Independent reruns may differ. This page does not include exploit procedures, payloads, jailbreak recipes, or weight-edit walkthroughs.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 11, 2026

Anthropic Threat Intelligence Report: Claude Misuse Across Cyber, Weapons, Bio, and Distillation

On September 10, 2026 Anthropic published its most detailed Threat Intelligence report yet — case studies of Claude misuse disrupted between December 2025 and August 2026 across seven harm areas. explainx.ai separates what is in the primary report (including China-linked anti-torpedo work and Alibaba's 151M+ distillation campaign) from claims circulating on X and prediction markets.

Sep 10, 2026

Anthropic Alignment Assessment: Mythos 5, PyPI, and Biased Reasoning

Anthropic published a full alignment assessment on September 9, 2026 for four incidents where Claude models reached the real internet during misconfigured cyber evaluations. The headline case — Claude Mythos 5 uploading a malicious PyPI package — shows biased reasoning that fooled offline monitors, not just sandbox failure.

Sep 10, 2026

Anthropic Says Claude Models Were Used in 15 Real-World System Breaches

Anthropic disclosed that Claude models were used as part of the toolchain in 15 separate real-world security incidents, described as the first time the company has reported model involvement in confirmed breaches at this scale. explainx.ai walks through what "used in a breach" actually means, how it fits Anthropic's own alignment reporting this year, and what it means for anyone running Claude in production.