GLM 5.3: Benchmarks, API Access, and Public Weights (2026)
The GLM 5.3 API is open at $1.4 in / $4.4 out, same price as 5.2. 1M context, 128K output, Terminal Bench 3.0 4.6 to 28.3, public weights checked September 9.
GLM 5.3 shipped on 2026-08-14 with no metered API, and five days later that gap has closed: the standalone API is open at $1.4 in / $4.4 out, the same price as GLM 5.2, and the model is on OpenRouter and third-party gateways. The weights are now publicly listed; the September 9 check is documented below. The model reuses the GLM 5.2 base, and the whole jump comes from post-training.
Released: 2026-08-14 (Z.ai release note)
Base model: same as GLM 5.2, all gains from post-training
Context: 1M tokens, 128K maximum output (Z.ai model page)
Price: $1.4 in / $0.26 cached in / $4.4 out, same as GLM 5.2
Open weights: public repository verified 2026-09-09
Coding: Terminal Bench 3.0 4.6 to 28.3, open-source SOTA
Breaking API: thinking.type "disabled" removed; reasoning_effort low/high/max
Standalone API: open, api.z.ai, three protocols
Also on: OpenRouter, gateway catalogs, GLM Coding Plan, ZCode
Snapshot: 2026-08-19
What Is GLM 5.3?
It is a post-training refresh of GLM 5.2, not a new model family. Z.ai says so in the first paragraph of its GLM-5.3 release note: the model “uses the same base model as GLM-5.2”, and “every gain comes from post-training.”
- Same base weights as
z-ai/glm-5.2, a month of extra RL on top - Trained on synthesized long-horizon environments, some representing days of work for an experienced engineer
- Aimed squarely at agentic coding: terminal tasks, repo-level work, multi-step tool use
- Also ships a large security research section, which is outside the scope of this page; the release note covers it
The r/LocalLLaMA release thread passed 800 upvotes at a 0.99 ratio within hours, with the top replies fixated on one thing: a 743B-class model trading blows with much larger ones. That parameter figure is community arithmetic, not a published spec, so we are not repeating it as fact.
When Was GLM 5.3 Released?
2026-08-14, and it reached GLM Coding Plan subscribers the same day.
The rollout is uneven in a way worth knowing before you go looking:
| Surface | Status on 2026-08-14 | Status on 2026-08-19 |
|---|---|---|
| Z.ai release note | Published | Unchanged |
| GLM Coding Plan / ZCode | Live for all subscribers | Live |
| Z.ai docs model page | Not up yet | Live, 1M context and 128K output |
| Standalone model API | ”Coming soon”, no date | Open, three protocols on api.z.ai |
| Z.ai docs pricing table | Stopped at GLM 5.2 | GLM-5.3 row, $1.4 / $0.26 / $4.4 |
| OpenRouter catalog | No entry | z-ai/glm-5.3, 1,048,576 context |
| Gateway catalogs including ofox | Not listed | z-ai/glm-5.3 live |
| HuggingFace weights | Not verified publicly downloadable in the original check | Original check returned HTTP 401 |
Five days took the launch from subscription-only to fully metered everywhere except the weights. That is the ordinary shape of a Z.ai release, and it is worth knowing the direction of travel: the plan gets the model first, the per-token API follows, the weights land last.
Where Can You Download GLM 5.3 Weights?
The official zai-org/GLM-5.3 repository is publicly accessible as of September 9, 2026. The Hugging Face API returned HTTP 200, private: false, and a file listing. This replaces the August 19 availability check; an earlier 401 response alone did not establish why access failed.
Read the repository’s model card, license and supported runtime instructions before downloading. Public availability does not establish that it fits your hardware. This update did not download or run the full weights. Our GLM 5.2 local guide and self-hosting cost guide remain guides to that older model, not verified GLM 5.3 hardware specifications.
How Much Does GLM 5.3 Cost?
$1.40 per million input tokens, $0.26 cached input, $4.40 output. Exactly what GLM 5.2 costs. The pricing table in Z.ai’s developer docs now carries a GLM-5.3 row, and it lines up rate for rate with 5.2 and 5.1.
| Item | Rate |
|---|---|
| Input | $1.40 / 1M tokens |
| Cached input | $0.26 / 1M tokens |
| Cached input storage | Free, marked limited-time |
| Output | $4.40 / 1M tokens |
| GLM Coding Plan | Points quota, input, cached input and output counted separately |
| Off-peak discount | 50% of standard points |
| Peak window | 14:00 to 18:00 UTC+8, Monday to Friday |
A same-price upgrade is the rare case where the pricing question answers itself: there is no cost argument for staying on 5.2 unless your runtime, evaluations or thinking-off workflow favor it. What the flat price hides is that GLM 5.3 is also the more token-efficient model on Z.ai’s own numbers, but that does not guarantee lower cost on your tasks.
The cache line is where the real money sits. At $0.26 against $1.40, a repeated prefix costs 19% of a cold read, and cache writes are free during the promotion. OpenRouter lists the same $1.4 / $4.4 pair with a 1,048,576-token context, so third-party routes are passing the first-party rate through rather than marking it up. Our GLM 5.3 API guide has the endpoint table, the migration order and measured per-call costs at each reasoning_effort level.
What Is the GLM 5.3 Context Window?
1M tokens, with a documented maximum output of 128K. Z.ai’s GLM-5.3 model page lists both, and Zhipu’s Chinese documentation carries the same pair.
Z.ai’s own evaluations did not all run at that ceiling, which is a useful signal about where the model was actually exercised:
| Benchmark | Context used | Max output |
|---|---|---|
| Terminal Bench 3.0 | 400K | 128K |
| Terminal Bench 2.1 | Not stated | 65,536 |
| DeepSWE v1.1 | 400K | Not stated |
| Agents’ Last Exam (CLI) | 1M | 64K |
| SWE-Marathon, PostTrainBench | 1M | 128K |
| NL2Repo | 1M | 64K |
| HLE with tools | 300,000, with a context management strategy | 163,840 |
The spread matters if you were planning to lean on the full million. A 1M window is what the model accepts; 400K is what Z.ai chose for its headline coding suite. Whatever endpoint eventually serves 5.3 may also publish a lower cap of its own, so read the provider’s number as the operative one.
How Much Better Is GLM 5.3 at Coding?
Large on the newest agentic benchmarks, modest on the saturated ones. All figures below are Z.ai’s own. The Terminal Bench, SWE-Marathon and Agents’ Last Exam runs used the Claude Code 2.1.207 harness at max reasoning effort; DeepSWE used mini-swe-agent.
| Benchmark | GLM 5.3 | GLM 5.2 | Kimi K3 | Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | 21.1 | 34.6 |
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 85.0 | 88.8 |
| DeepSWE v1.1 | 66.9 | 46.2 | 67.5 | 58.0 | 72.7 |
| SWE-Marathon v1.1 | 42.5 | 19.4 | 48.1 | 48.8 | 42.5 |
| Agents’ Last Exam (CLI) | 28.5 | 23.8 | 27.6 | 25.7 | 28.6 |
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1588 | 1730 |
The Terminal Bench 3.0 line does the heavy lifting: 4.6 to 28.3 on a suite where GLM 5.2 was effectively not competing. On Terminal Bench 2.1, where everyone sits between 85 and 89, the same upgrade is worth 7 points and changes no ranking. Newer benchmarks have room to show movement; old ones do not.
The number Z.ai leads with is not a score. At High effort GLM 5.3 hits 31.4% on its in-house Code Bench using about 50K output tokens per task, against Claude Opus 4.8 at 29.5% using 120K.
That is the claim worth testing yourself, because token count is what you actually pay for. Z.ai also reports 34.5% at roughly 75K tokens at Max effort, up from GLM 5.2’s 23.4% at 96K. Claude Fable 5 still leads that private benchmark at 39.5%, and GPT-5.6 Sol still leads Terminal Bench 3.0, so “open-source SOTA” is the accurate framing rather than “SOTA”.
One caution on the whole table: it is a private benchmark plus vendor-run public ones. Kimi K3 and Claude Opus 4.8 numbers here were produced by Z.ai, not by their vendors.
Do I Need to Change My Code for GLM 5.3?
Yes, if you ever disabled thinking. thinking.type: "disabled" is gone, and Z.ai says the request will simply fail.
The new contract:
thinking.typeacceptsenabledonlyreasoning_effortacceptslow,high,max, and defaults tomax- Z.ai recommends
maxfor coding - Migration order matters: set
enabledplusreasoning_effort: "low"first, then swap the model ID
{
"model": "glm-5.3",
"thinking": { "type": "enabled" },
"reasoning_effort": "max"
}
The cost of losing disabled is easy to underestimate, so we measured it on GLM 5.2 through an OpenAI-compatible endpoint on 2026-08-14. One trivial prompt, “Reply with the single word: ok”, three runs per configuration:
| Request | Output tokens (3 runs) |
|---|---|
| Default (thinking on) | 144, 138, 90 |
reasoning_effort: "low" | 69, 86, 122 |
reasoning_effort: "max" | 101, 120, 102 |
thinking.type: "disabled" | 2, 2, 2 |
Run-to-run variance between the thinking modes swamps the difference between low and max on a prompt this small, so do not read an ordering into those three rows. The row that matters is the last one. Thinking off answered in 2 output tokens every time; the cheapest thinking-on setting still spent between 69 and 122. On GLM 5.3 that last row no longer exists.
That table is GLM 5.2, and the obvious inference from it turned out to be wrong. Measured on GLM 5.3 itself on 2026-08-19, a one-word classification prompt came back at a median of 3 output tokens at reasoning_effort: "low", ten runs, against 105 at max. The dreaded floor is not 70 tokens on 5.3, it is roughly what disabled used to cost. The catch is that max is the default, so a workload ported without setting the parameter gets the 105-token version. Our GLM 5.2 versus GPT-5.5 cost comparison has the workload math for that shape of traffic.
Which Capabilities Does GLM 5.3 Support?
Text in, text out, plus thinking modes, streaming, function calling, context caching and structured output. Those five are what Z.ai’s model page lists today.
| Capability | Status |
|---|---|
| Input / output modalities | Text only |
| Thinking modes | low, high, max; no disabled |
| Streaming output | Supported |
| Function calling | Supported |
| Context caching | Supported |
| Structured output | Supported, JSON included |
| MCP | No longer listed in the capability section |
The MCP entry is worth a note because it moved. The model page carried an MCP link in its capability list on 2026-08-15 and does not on 2026-08-19. Nothing announced that removal, so treat MCP support as unconfirmed rather than dropped, and test it against your own tooling before designing around it.
Nothing here is new relative to GLM 5.2, which is consistent with a post-training-only release. Vision, image and audio stay on the separate GLM-5V and GLM-Image lines.
How Do I Access GLM 5.3?
Four routes now: the metered Z.ai API, the GLM Coding Plan, ZCode, or any aggregator carrying z-ai/glm-5.3.
| Route | What you get | Catch |
|---|---|---|
| Z.ai model API | Per-token billing at $1.4 / $4.4, three protocols | Coding Plan accounts are capped to one protocol, see below |
| GLM Coding Plan | Points quota, 50% off-peak | Subscription, not pay as you go |
| ZCode | Historical promotion ended 2026-08-31; check current terms | Z.ai’s own client |
| Claude Code / Cline / OpenCode | Existing agent, new model | Runs against the plan quota, not a metered key |
| Aggregators and gateways | One key across models | Model IDs differ, z-ai/glm-5.3 on OpenRouter and ofox |
The endpoints are not the ones Zhipu previewed on launch day. The model page now lists https://api.z.ai/api/coding/paas/v4 for the OpenAI Chat Completions protocol, https://api.z.ai/api/v1 for the OpenAI Responses protocol, and https://api.z.ai/api/anthropic for the Anthropic Messages protocol. The open.bigmodel.cn/api/paas/v4 base URL quoted in the launch-day note is not what shipped. The same page’s own Quick Start sample then posts to https://api.z.ai/api/paas/v4/chat/completions, without the /coding segment, so the two halves of one page disagree; if one 404s, try the other before assuming your key is wrong.
One restriction is easy to miss: accounts that have ever subscribed to a GLM Coding Plan, including expired subscriptions, can currently reach the model API only through the OpenAI Chat Completions protocol.
On an aggregator it is a two-line change:
from openai import OpenAI
client = OpenAI(api_key="YOUR_KEY", base_url="https://api.ofox.run/v1")
r = client.chat.completions.create(
model="z-ai/glm-5.3",
messages=[{"role": "user", "content": "Refactor this module and run the tests."}],
reasoning_effort="low", # "disabled" is not an option on 5.3
)
print(r.usage)
The model ID prefix is z-ai, not zai, which is the single most common reason that call comes back as:
{"error":{"message":"Model 'zai/glm-5.3' not found","type":"model_not_found","code":404}}
z-ai/glm-5.2 answers on the same endpoint with the same key, so a switch between the two is one string. Our GLM 5.2 API access guide covers the key setup end to end.
How Do You Point a Coding Agent at a Model That Launched Yesterday?
By changing one string, if your agent lets you set the endpoint yourself. Launch week always looks the same. The subscription product gets the model on day one, the metered API arrives days later, the docs page lands a day after the blog post, and every tool you use keeps a hardcoded model list. GLM 5.3 ran that whole sequence between 2026-08-14 and 2026-08-19, and the weights are publicly listed at the September 9 check.
The part you can control is how much re-plumbing each of those steps costs you. Agents that let you set base_url and a model string treat a launch as a one-line edit. Agents that ship a fixed dropdown make you wait for their release cycle, which is why the r/opencodeCLI thread appeared within an hour of the announcement.
Because Claude Code, OpenCode and most other agents speak OpenAI-compatible HTTP, one key behind one base URL keeps every one of them on the same billing and the same failover. ofox covers a catalog of about 130 models on that single endpoint, with z-ai/glm-5.3 and z-ai/glm-5.2 both live, so swapping between them is the code block above and nothing else.
GLM 5.3 vs GLM 5.2: What Actually Changed?
The post-training, and three things you can act on: no thinking-off switch, public weights now available, and a Terminal Bench 3.0 score that moved from 4.6 to 28.3.
| GLM 5.3 | GLM 5.2 | |
|---|---|---|
| Base model | Same | Same |
| Context / max output | 1M / 128K | 1M / 128K |
| Published per-token price | $1.40 / $0.26 / $4.40 | $1.40 / $0.26 / $4.40 |
| Metered API | Available | Available |
| On aggregators | Yes, z-ai/glm-5.3 | Yes, z-ai/glm-5.2 |
| Weights | Public repository; check its license | Public, MIT |
| Thinking off | Not supported | Supported |
| Effort levels | low / high / max, default max | Effort plus disabled |
| Terminal Bench 3.0 | 28.3 | 4.6 |
Identical token rates do not establish identical task costs. Keep GLM 5.2 if its tested runtime, output quality or thinking-off behavior fits your workload. Before moving to GLM 5.3, check the API and reasoning migration guide and evaluate representative tasks.
References
- Z.ai release note: GLM-5.3
- Z.ai developer docs: GLM-5.3 model page
- Z.ai developer docs: pricing
- Z.ai devpack docs: GLM Coding Plan overview
- Zhipu AI docs: GLM-5.3
- HuggingFace: zai-org
- OpenRouter: z-ai/glm-5.3
- r/LocalLLaMA: GLM 5.3 Released
- r/opencodeCLI: GLM 5.3 is there!
- Simon Willison: GLM-5.2 is probably the most powerful text-only open weights LLM
- GLM 5.3 API: pricing, endpoints and reasoning_effort (ofox blog)
- ofox model page: GLM 5.3
- ofox model page: GLM 5.2
Frequently Asked Questions
- Is the standalone GLM 5.3 API available now?
- Yes. Z.ai's pricing table now carries a GLM-5.3 row at $1.4 per million input tokens, $0.26 cached input and $4.4 output, identical to GLM 5.2. The model also appears on OpenRouter and on gateways including ofox as z-ai/glm-5.3. The GLM Coding Plan remains a separate points-based route.
- What is the GLM 5.3 max output length?
- 128K tokens, alongside a 1M-token context window, per Z.ai's model page. Z.ai's own benchmark runs sit at or below that: 128K output on Terminal Bench 3.0 and PostTrainBench, 64K on Agents' Last Exam. Whatever endpoint eventually serves the model may publish a lower cap of its own.
- Can I use GLM 5.3 with Claude Code?
- Yes. Z.ai lists Claude Code, Cline and OpenCode as supported coding tools for the GLM Coding Plan, and its own benchmark runs used the Claude Code 2.1.207 harness. Set reasoning_effort rather than disabling thinking, because the model no longer accepts thinking turned off.
- Is GLM 5.3 better than Kimi K3?
- On Z.ai's own table it wins Terminal Bench 3.0 by 28.3 to 17.4 and AutomationBench by 48.2 to 46.7, while Kimi K3 leads Toolathlon 76.5 to 73.0 and SWE-Marathon 48.1 to 42.5. Every one of those numbers is vendor-reported by Z.ai, so read it as a claim to verify rather than an independent result.
- Does GLM 5.3 use a new base model?
- No. Z.ai states GLM 5.3 uses the same base model as GLM 5.2 and that every gain comes from post-training. The release note opens with the line 'Scaling post-training is all we did for GLM-5.3.'
- Is GLM 5.3 free?
- There is no free tier. Metered access costs $1.4 per million input tokens and $4.4 output. The two ways to pay less are cached input at $0.26, with cache writes free for a limited time, and the GLM Coding Plan, where calls outside the peak window of 14:00 to 18:00 UTC+8 on weekdays consume half the standard points.
- Where can I download GLM 5.3 weights?
- The official zai-org/GLM-5.3 repository is publicly accessible as of September 9, 2026. Its Hugging Face API returned HTTP 200 with private=false. Check the model card, license and runtime requirements before downloading; this update did not run the weights locally.


