Zhipu AI claims anonymous model dubbed ‘Niu Lai’ after a six-day run that drove tens of trillions of tokens

Picture shows an illustration of an Ox carrying a pile of semiconductors

Stealth model topped OpenRouter with 23.2 trillion tokens used by developers in less than a week, beating DeepSeek; company confirms it’s GLM-5.3-Flash.

By Xiao Jing

After nearly a week of speculation, the mysterious model Ox Alpha finally has an owner. 

On Aug. 26, Bloomberg reported that Zhipu AI (Z.AI) (2513.HK) confirmed Ox Alpha, which launched anonymously recently was developed by the company and its part of its new-generation GLM-series model . The name plays on a viral Chinese movie title Niu Lai, which translates as “the ox is coming,” and in the developer community it became a playful shorthand for Ox Alpha.

Zhipu said it would release model weights that evening, but has not yet disclosed the model’s official name, parameter count or precisely how it relates to GLM-5.3.

From its anonymous launch on Aug. 20 to the company’s confirmation on Aug. 26, Ox Alpha had already driven global developers to use tens of trillions of tokens, the unit used to price model usage or measure inference cost.

Free and anonymous, an ‘ox’ takes the top spot

Ox Alpha appeared on model-aggregation platform OpenRouter on Aug. 20 as a “stealth model,” meaning its developer was not identified.

OpenRouter initially described it only as a reasoning model developed and operated by an anonymous third party, aimed primarily at coding, long-running agent tasks and production environments.

Its publicly available specifications show a context window of 1.0486 million tokens and a maximum output length of 131,000 tokens per request. It supports text, image and video inputs, as well as tool calling and structured outputs. In practice, that means it can process large code repositories, project documentation and lengthy agent execution histories in a single context.

What really drove usage was that it was free. The open-source AI coding tool OpenCode announced a one-week free preview with “near-unlimited” calls, claiming the provider had prepared 100 trillion tokens of daily service capacity – equivalent to about 11.6 billion tokens per second. Developers rushed to test the model while asking a bigger question: which company had enough inference resources to support it?

However, 100 trillion tokens is theoretical capacity, not actual usage. Code agents repeatedly read the same codebase and context, and heavy cache reuse can make recorded token volumes far exceed actual new computation. By Aug. 25, Ox Alpha had processed about 23.2 trillion tokens during OpenRouter’s seven-day rolling window, ranking first – more than double DeepSeek V4 Flash 0731’s 11.6 trillion in the same period. While this proves broad trial adoption, OpenRouter counts token volume, not users, requests, success rates, or revenue.

Computing power questions

Before Zhipu claimed the model, computing power was the main reason many doubted it was the Chinese firm’s work. The community assumed Ox Alpha might be a trillion-parameter flagship; running it free for a week at 100 trillion tokens daily would require massive chip clusters and electricity. Some guessed Google or Microsoft, or a “Zhipu model hosted by a tech giant” hybrid.

Bloomberg reported in July that Zhipu had built a 1-gigawatt-scale AI data center using domestically produced chips and had begun partial operations. Data centers are measured by power capacity, but 1GW does not equal actual compute, and the report did not disclose chip count, IT load, or utilization. Zhipu also acquired Zhongke Jiahe, a heterogeneous-compute software firm, to bolster compilers, runtimes, and inference engines. While this could explain a large-scale stress-testing capability, Bloomberg’s confirmation of ownership did not specify where Ox Alpha was deployed, which chips were used, or whether third-party compute supported it. Thus, 1GW remains an explanation of capacity, not a verification of 100-trillion-token peak deployment.

Six days of developer sleuthing

Beyond its free access and huge usage allowance, Ox Alpha attracted attention for its performance on coding and agent tasks. Many calls occurred within Claude Code, Hermes Agent, DeepSeek Harness, and OpenCode. Stripe CEO Patrick Collison described it in an X post as “very impressive.”

A widely shared community micro-benchmark claimed Ox Alpha achieved 80% on 10 DeepSWE software-engineering tasks, outperforming Claude Fable 5 and GPT-5.6 Sol in that test. However, this was just 10 questions – not a formal leaderboard. Some developers later found it mediocre on complex backend and vision tasks .

Other developers investigated the model’s origins. Across multiple tests involving Chinese and English text, code and emojis, Ox Alpha’s token counts showed a high degree of consistency with GLM-5.3. Developers said the differences could largely be explained by hidden system prompts. In controlled video tests, its frame-sampling behavior and the relationship between video duration and token consumption also appeared similar to Zhipu’s GLM vision models. Error codes and wording returned by the application programming interface when unusual parameters were submitted were likewise seen as characteristic of Zhipu’s services.

Release timing strengthened the case: on Aug. 14, Zhipu announced GLM-5.3 and said weights would come in about two weeks; six days later, Ox Alpha appeared; on Aug. 26, Bloomberg reported Zhipu would release the weights that night .

Zhipu has used this playbook before. In February, anonymous model Pony Alpha appeared on OpenRouter and was later revealed as an early GLM-5 test version . From “pony” to “ox,” anonymous releases let Zhipu strip away brand expectations, let developers judge raw capabilities, and stress-test models, caches, task schedulers, and domestic clusters under real agent workloads. The identity mystery itself turned the community into unpaid promoters and testers .

What ultimately matters, however, is what Ox Alpha becomes next: its official name, its relationship with GLM-5.3, its eventual pricing and how many developers continue using it once the free period ends. Those factors will determine whether the anonymous launch was merely a flash in the pan or a new entry point for Zhipu into the global developer market.

Source: 
Tencent Technology

Share the story: