Ox Alpha Was GLM-5.3-Flash

9 min read · · Updated · Zentor Editorial
Ox Alpha Was GLM-5.3-Flash

Ox Alpha, the free stealth model on OpenRouter, was Z.ai's GLM-5.3-Flash. The confirmation, the 3.94T tokens it pulled, and who guessed right.

Contents

Ox Alpha was GLM-5.3-Flash. Z.ai confirmed it on 26 August 2026, six days after the model appeared on OpenRouter with no name attached, and OpenRouter now carries a de-cloaking banner on the same listing that spent a week saying the provider wished to stay anonymous.

Key Takeaways:

  • Ox Alpha ran as a free, unnamed model on OpenRouter from 20 August. On 26 August it was revealed as Z.ai's GLM-5.3-Flash
  • Two first-party confirmations: OpenRouter's banner, and Z.ai's own launch post saying it "tested GLM-5.3-Flash anonymously as ox-alpha"
  • It pulled 3.94 trillion prompt tokens during the preview, up from the 173 billion showing when we first wrote this
  • GLM-5.3-Flash is not the GLM-5.3 that launched on 14 August. Different base model, natively multimodal, roughly a twentieth of the price
  • The weights are on Hugging Face under MIT. The flagship GLM-5.3's weights, promised for the same week, still are not

The identity question ran for six days and produced two named theories in the press, one of which was right. What follows is the reveal first, then the record of the week, kept because the guessing is the more interesting half.


Confirmed: Ox Alpha was GLM-5.3-Flash

Two sources, both first-party, both dated 26 August 2026.

The OpenRouter listing now opens with a de-cloaking banner: "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash." The model has been removed from OpenRouter's live model list; the page survives as a record.

Z.ai said it themselves in the GLM-5.3-Flash launch post, and were more specific than they needed to be: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week, with all of this traffic served on Chinese AI chips."

That last clause is the part worth sitting with. The traffic everyone was measuring, the 3.94 trillion prompt tokens the activity graph finished on, was being served on domestic Chinese accelerators the whole time, and nobody watching the latency numbers guessed it.

What it turned out to be: 320B total parameters with 18B active, the first natively multimodal model in the GLM-5 series, built on a newly trained base rather than the one GLM-5.3 shares with GLM-5.2. Weights are published under MIT. Pricing on OpenRouter came out at $0.075 per million input tokens and $0.25 output, which is roughly a twentieth of flagship GLM-5.3.

The name that confused everyone: GLM-5.3-Flash sounds like a smaller GLM-5.3. It isn't. GLM-5.3 is text-only, shares GLM-5.2's base, and got its gains from post-training. Flash is a separate architecture with a separate base. Sharing a version number is a marketing decision, not a technical one.

A million tokens of context is a job, not a chat.
The window Ox Alpha was handing out free implies work measured in hours, and a paste-into-playground session dies when the tab does. Zentor is a hosted cloud AI computer that stays powered on for the length of the run, alongside the tools you already use.
Keep the long run alive after I close the tab…Try Zentor →

Who called it, and who didn't

The thread got there before the press did.

@teortaxesTex wrote that they could believe it was the next GLM with vision, guessed the working names GLM 5.3-Vision or 5.5, and said the speed and cache behaviour looked like GLM. Vision was right, GLM was right, the version number was one word off. They also floated a stolen Claude checkpoint, which was not right.

The press moved the other way. TechCrunch ran the question as a headline on 23 August. Wccftech pointed at GLM first and then revised toward an unreleased Microsoft MAI build, which was a revision away from the correct answer. AI analyst Andrew Curran's summary that week, that people "seem less sure of anything", was accurate about the mood and wrong about the trend.

We wrote at the time that "if a thread tells you Ox Alpha is GLM, that person is theorizing." That was the right call on 21 August and the theory turned out to be correct on 26 August. Both things can be true: a guess being right does not retroactively make it evidence.

What the listing actually says

The OpenRouter card is the only first-party page. It pitches Ox Alpha for coding and long agent jobs, including ones that mix text with pictures or video. We didn't run it; this is a read of the card and the thread, not a review.

Official OpenRouter listing for Ox Alpha (stealth/ox-alpha): Free, 1M context, released Aug 20, 2026. The banner states prompts and completions are retained by the anonymous provider and are not used for training.
Official OpenRouter listing for Ox Alpha (stealth/ox-alpha): Free, 1M context, released Aug 20, 2026. The banner states prompts and completions are retained by the anonymous provider and are not used for training.

OpenRouter's own post used louder language ("frontier," "real-world production use") and a card that says CONTEXT 1.05M, released Aug 20, 2026. Around 5:02 AM on August 21 that post was sitting near 522,000 views.

OpenRouter's August 20 post announcing Ox Alpha, with the model card showing CONTEXT 1.05M and Released Aug 20, 2026.
OpenRouter's August 20 post announcing Ox Alpha, with the model card showing CONTEXT 1.05M and Released Aug 20, 2026.

You can skip the rest of the spec sheet. Context is 1M, max output is 131K, inputs are text, images, and video, and the price is $0 per million tokens in and out. Latency on the page will have moved by the time you click through.

A follow-up from the same account is the sentence people will paste into Slack: "It is free," and "This time, the provider does not train on your prompts or completions."

OpenRouter follow-up notes on the Ox Alpha thread: the model is free, and this time the provider does not train on prompts or completions. A joke reply underneath says the model is three Gemini flashes wearing a trenchcoat.
OpenRouter follow-up notes on the Ox Alpha thread: the model is free, and this time the provider does not train on prompts or completions. A joke reply underneath says the model is three Gemini flashes wearing a trenchcoat.

The naming fight

@thdxr posted the card and asked people to name the model, @theo asked the same thing, and @Suhail sounded surprised that nobody knew. Digg scooped the thread out of X and into a headline.

@teortaxesTex did the longest public guess, and at the time it was only a guess. After more testing they said they can believe it's the next GLM with vision, that Zhipu sat on GLM-5.3 for a few weeks, and that it's faster than they expected in places Chinese models usually lag (they named Kimi K3). They also said they could half-believe a stolen Claude checkpoint. Speed and cache behavior look like GLM to them; DeepSeek, they wrote, it clearly isn't. Their working names are GLM 5.3-Vision or 5.5, with a rough 61 on Artificial Analysis. We haven't run those evals.

If a thread tells you "Ox Alpha is GLM," that person is theorizing.

By August 23 the guessing had left X. TechCrunch ran the question as a headline and quoted Stripe CEO Patrick Collison calling the model "very impressive". The piece also carried a second name: Wccftech first pointed at GLM, then revised toward an unreleased build of Microsoft's MAI. AI analyst Andrew Curran, who had watched the early GLM speculation, summed the week up by saying people "seem less sure of anything". Two named theories by then, neither confirmed, and OpenRouter's card still describing the provider as anonymous. That held for another three days.

One joke, because it is in the screenshot above and it will get quoted as analysis: @PeerReview wrote "The model is three Gemini flashes wearing a trenchcoat."

Who is already using it

The Apps tab was the part of the listing that wasn't marketing copy. These are the figures as they stood on 21 August, early in the preview; the run finished on 3.94 trillion prompt tokens, 46.5 billion completion tokens and 8.52 billion reasoning tokens.

App Tokens
Claude Code 36.2B
Hermes Agent 28.8B
Oh-My-Pi 26.1B
DeepSeek Harness (multimodal-bridge) 23.7B
ZCode 18.4B

Claude Code on a free mystery model is the line that travelled. ZCode sitting fifth is the tell nobody weighted properly: Z.ai's own coding app was quietly burning 18.4 billion tokens on a model Z.ai had built and would not admit to owning.

Logs

OpenRouter was not the lab, and for six days the banner said a third-party provider ran Ox Alpha and had chosen to stay anonymous. We now know that provider was Z.ai.

They kept prompts and completions, and said those logs were not used for training. "This time" in the thread note is doing work, because earlier stealth drops were more candid about using logs to improve the model. Kept still means kept. If you would not send a private repo to an unnamed lab, a free 1M window did not fix that, and the reveal does not retroactively fix it either. Anyone who ran proprietary code through Ox Alpha during the preview sent it to Z.ai without knowing that is who they were sending it to. That is the real cost of a stealth launch, and it lands on users rather than on the lab.

This is not the first Alpha

OpenRouter has done nameless Alpha drops before. Pony Alpha in February 2026 later turned out to be Zhipu / Z.ai GLM-5; OpenRouter's own Pony page now says so.

Owl Alpha was later reported as Meituan's LongCat-2.0-Preview. Meituan's LongCat account said that in June 2026, and Decrypt repeated it. That was the lab naming itself, not OpenRouter unmasking anyone.

Pony to GLM-5 we can pin. Owl to LongCat we can report. Ox Alpha to GLM-5.3-Flash now joins them, which makes three consecutive named Alphas from Chinese labs: Z.ai twice and Meituan once. That is a pattern rather than a coincidence, and it is a reasonable prior for whatever the next nameless Alpha turns out to be.

The 1M window still needs a machine that stays on

A million tokens of context is not a job that survives you closing the laptop. Paste-into-playground dies when the tab dies, which is the same hole we keep hitting when we look at long-horizon agent evals.

Keep the tools you already pay for. Put the long job on a machine that stays up.

FAQ

What was Ox Alpha?

Z.ai's GLM-5.3-Flash, running anonymously on OpenRouter and OpenCode from 20 August 2026 as a free preview with a 1M-token context window. Z.ai confirmed the identity on 26 August, the day it published the model.

Who made Ox Alpha?

Z.ai, the lab formerly trading as Zhipu. Both OpenRouter's de-cloaking banner and Z.ai's own launch post confirm it. Of the theories circulating during the preview, the next-GLM-with-vision guess was correct and the unreleased-Microsoft-MAI theory was not.

Is Ox Alpha the same as GLM-5.3?

No, and the naming makes this easy to get wrong. GLM-5.3 launched 14 August, is text-only, and shares its base model with GLM-5.2. GLM-5.3-Flash is a separately trained base, natively multimodal, 320B total parameters with 18B active, and prices at roughly a twentieth of the flagship.

Can I still use Ox Alpha for free?

Not under that name. The stealth listing has been removed from OpenRouter's live model list and the free preview is over. GLM-5.3-Flash is now available as a paid model, listed on OpenRouter at $0.075 per million input tokens and $0.25 output, and the weights are downloadable from Hugging Face under an MIT licence.

Did the Ox Alpha provider train on prompts?

Z.ai said no, through OpenRouter, both at listing time and in the follow-up thread. Prompts and completions were retained by the provider throughout. The retention was disclosed; the identity of who was doing the retaining was not.

How much traffic did Ox Alpha get?

The activity graph finished on 3.94 trillion prompt tokens, 46.5 billion completion tokens and 8.52 billion reasoning tokens across the preview. Z.ai describes it as the most popular model of the week, and says the entire load was served on Chinese AI chips.

Updated 27 August 2026. The identity confirmation, final token counts and pricing were read from OpenRouter's listing and Z.ai's GLM-5.3-Flash launch post on that date. The sections from “What the listing actually says” onward are preserved as written on 21 and 24 August, while the model was still unnamed.

Zentor Editorial
Zentor Editorial Zentor editorial team

The Zentor editorial team writes about workflow automation, AI agents, and the tools we build. Default byline for industry overviews, listicles, and collaborative pieces.

Share

Ready to put this into practice?

Zentor runs browser tasks, research, and schedules automatically. Try it free.

References https://openrouter.ai/stealth/ox-alpha · https://x.com/OpenRouter/status/2090544970923184269 · https://digg.com/tech/iuf1vp3x · https://openrouter.ai/openrouter/pony-alpha · https://decrypt.co/372579/longcat-2-0-meituan-ai-stealth-model-openrouter · https://z.ai/blog/glm-5.3-flash · https://huggingface.co/zai-org/GLM-5.3-Flash