Amazon announced on October 5, 2026 that GLM 5.3 from Z.ai is generally available on Amazon Bedrock, for eligible enterprise customers. It is a 753B-parameter mixture-of-experts coding model with a 1M-token context window, and you call it through the same Bedrock plumbing you may already use for other models.
For teams that were already evaluating GLM 5.3 through Z.ai's own API or third-party hosts, the news is less about capability and more about procurement: you can now run it inside an AWS account, with AWS billing, IAM and regional controls.
TL;DR: GLM 5.3 on Bedrock
| Question | Answer |
|---|---|
| When? | Announced October 5, 2026 (see the AWS What's New entry) |
| Who can use it? | Eligible enterprise customers |
| Model size | 753B total parameters, about 40B active per token (mixture of experts) |
| Context and output | 1M-token context, up to 128K output tokens |
| Model IDs | us.zai.glm-5.3 and global.zai.glm-5.3 (cross-Region inference profiles) |
| APIs | OpenAI-compatible Responses and Chat Completions, plus Bedrock Invoke and Converse |
| Caching | Implicit and explicit prompt caching |
| Service tiers | Flex, Priority and Standard |
| Reasoning | Always on, with selectable effort levels |
What GLM 5.3 is
GLM 5.3 is Z.ai's coding-first post-training release, built on the same base model as GLM 5.2 with all improvements coming from scaled post-training, according to AWS. We covered the original release and what the vendor measured in our GLM-5.3 coding benchmark explainer, and the training-side story in the infrastructure and dense-feedback write-up.
AWS repeats Z.ai's claims: a 50 percent improvement over GLM 5.2 on Z.ai's internal coding benchmark, and competitive performance on DeepSWE, Terminal Bench 3.0 and FrontierSWE. It also reports a CyberGym security-benchmark score of 84.5 at release, which is why the post positions the model for defensive security workflows. Those are vendor figures. The 50 percent claim in particular is on a benchmark Z.ai controls, and our earlier piece explains why that deserves care.
Why it matters that it is on Bedrock
Chinese-lab open-weight models have mostly reached enterprises through the lab's own API, aggregator marketplaces or self-hosting. Bedrock availability changes three things.
- Procurement. A security or legal team that has already approved AWS can approve a model on it more easily than a new vendor contract.
- Data path. Requests stay within your AWS account and chosen cross-Region profile, rather than leaving for another provider's endpoint. Check AWS's data-handling terms for the exact guarantees.
- Mixing models. Bedrock already hosts other labs' models. For example, we looked at when to leave the xAI API for Grok 4.7 on Bedrock, and the same routing logic now applies to a cost-efficient open-weight coder.
There is also a cyber-policy dimension. Anthropic argued that open-weight GLM-5.3 matches Mythos-class cyber evaluations, and a hosted-uncensored variant drew attention in the abliteration.ai write-up. Putting the model on a hyperscaler's catalog, with enterprise eligibility checks, is a different distribution posture from an uncensored re-host.
How to call it
AWS's post shows the OpenAI Python client pointed at the Bedrock runtime. In outline:
from openai import OpenAI
client = OpenAI(
api_key=provide_token(region=region),
base_url=f"https://bedrock-runtime.{region}.amazonaws.com/openai/v1",
)
resp = client.responses.create(
input="Your prompt",
model="global.zai.glm-5.3",
)
Here provide_token is the helper from AWS's post for generating a Bedrock token. Swap in us.zai.glm-5.3 if you need requests to stay within US Regions. If you already use Converse, the same model ID works there too, per AWS.
Practical settings to decide up front:
- Reasoning effort. Reasoning cannot be turned off, so test the low, medium and high settings for latency and token use on your real tasks.
- Service tier. Flex, Priority and Standard trade price against latency guarantees; batch-style code review suits Flex, interactive agents suit Priority or Standard.
- Caching. Put your stable system prompt and repository context first and add cache points there. For coding agents that resend large context each turn, this is where most of the savings come from.
What to check before you switch
Eligibility. AWS says the model is available to eligible enterprise customers, and its post does not define the criteria in the material we reviewed. Confirm in your console before planning a migration.
Pricing. The AWS post describes pay-per-token billing and no persistent infrastructure charges, but the prices we could verify were from other hosts. Look up the current Bedrock per-token rates in AWS pricing before comparing against Z.ai direct or aggregator prices; for a sense of the cost conversation around the smaller sibling, see our GLM-5.3 Flash pricing analysis.
Your own evals. Vendor benchmarks do not predict how a model handles your monorepo, build tools and review style. Run a fixed set of real tickets and compare pass rate, tokens per fix and time to green CI. Our GPT-6 Astra versus Claude Fable 5.1 comparison shows a simple structure for that kind of head-to-head.
Policy and geopolitics. Some organizations restrict models from specific countries or labs. Bedrock availability does not remove your own internal policy review.
The license question hyperscalers raise
The most interesting subtext of this launch is the license. As we reported in the coding benchmark explainer, GLM-5.3 ships under a bespoke GLM-5.3 License rather than MIT. Its key change is a clause saying any licensee running a Model-as-a-Service business with aggregate revenue above 10 billion dollars across any consecutive 12 months must pass a Z.ai security review before using the model or derivatives commercially. That threshold is aimed at hyperscaler-scale clouds, which describes Amazon.
Neither the AWS post nor the material we reviewed says how that review applied here. The reasonable reading is that Bedrock's availability reflects an arrangement between the two companies, but the commercial terms between the companies are not published in the sources we reviewed. If licensing terms matter to you, ask AWS and Z.ai directly rather than assuming the open-weight license governs your Bedrock usage in the same way as self-hosting.
Small discrepancies worth knowing
AWS describes 753B total parameters. Our earlier coverage of the Hugging Face release cited about 744B total with roughly 40B active. The difference may come from how parameters are counted (for example, whether certain layers are included), or from different checkpoints, but no source we read explains it. It does not change how you call the model; it is a reminder to cite the source when you quote specifications and to avoid mixing figures from different documents.
Likewise, the weights Z.ai published on Hugging Face in August were FP8, whereas Bedrock does not state its serving precision in the material we reviewed. If you compare quality between self-hosted and Bedrock, run the same evals on both rather than assuming they are identical.
A short migration checklist
If you are moving an existing GLM 5.2 or other coding-model workload to Bedrock:
- Confirm access in the Bedrock console and note which Regions the US and global profiles cover for your account.
- Replace the model ID with us.zai.glm-5.3 or global.zai.glm-5.3 and keep everything else constant for the first run.
- Re-tune your prompts for always-on reasoning. If your agent previously used a non-reasoning model, expect longer latencies and higher output-token counts.
- Add cache points to the stable prefix of every request and measure the hit rate.
- Log token usage per task for a week, then compare cost per merged change against your current model.
- Keep a rollback path. Because the model is reachable through an OpenAI-compatible API, switching back is usually a one-line change.
What this means for your agent stack
Coding agents are the heaviest consumers of long-context models: they resend repository context, tool results and test output on every turn. A 1M-token window with explicit caching changes the economics of that pattern, because the stable prefix is billed at the cache rate after the first call. If you run a harness that already supports OpenAI-compatible providers, trying GLM 5.3 is an afternoon of work rather than a project.
The harder question is quality under your conditions. Models tuned on vendor benchmarks sometimes stumble on unfamiliar build systems, private frameworks or unusual test runners. Treat Bedrock availability as permission to test, not a reason to migrate. Start with a small, well-understood slice of work, such as dependency upgrades or test generation, where success is easy to measure and mistakes are cheap to revert.
Who should try it
- Teams already on Bedrock that want a long-context coding model without adding a vendor.
- Agent builders whose workloads resend large contexts, since explicit caching and a 1M window both help.
- Security teams evaluating open-weight-class cyber capability inside a governed environment, subject to the usage policies on both sides.
Who probably should not rush: teams without an enterprise AWS relationship, and anyone whose compliance posture excludes the model's origin.
What to watch
- Published Bedrock pricing and quotas for GLM 5.3, which determine whether it beats other options on cost per fix.
- Independent coding benchmarks on the Bedrock-hosted version, since quantization or serving choices can change results.
- Whether the Flash variants follow. The smaller GLM 5.3 models are already spreading on routers like OpenRouter and Nous Portal.
Specifications reflect AWS's October 5, 2026 announcement and may change.
