China’s open-source LLMs are looking for new ways to make money overseas

Photograph shows an illustration of a signpost with the names of Chinese AI companies pointing in different directions

AI companies are tightening open-source licenses as soaring agent workloads turn token consumption into a potential source of revenue.

China’s leading AI model companies are beginning to rethink one of the industry’s biggest assumptions: that giving models away is the best way to build a business.

For the past two years, Chinese developers have largely followed the same playbook overseas: release model weights, offer cheap application programming interface (API) access and distribute models through platforms such as Hugging Face and OpenRouter. The strategy has helped them gain global reach —  by mid-year they accounted for more than 60% of token traffic on OpenRouter.

But reach does not necessarily translate into revenue. Once model weights have been released, the inference — the process of running the model to generate responses — can happen on somebody else’s servers, with the resulting revenue going to the platform hosting it.

That is prompting a shift among companies including Z.AI (Zhipu AI) (2513.HK), Moonshot AI and MiniMax (0100.HK). Their approaches differ, but the common thread is that licensing is increasingly being used as a commercial tool.

The change is also being accelerated by the rise of AI agents. A single automated task can consume tens or even hundreds of times more tokens than a simple question-and-answer exchange, and agents can operate without a user manually initiating every interaction. Token consumption is therefore becoming a much more valuable commercial asset.

Zhipu bets on free distribution

The clearest example of the old model came in August, when an unnamed model called Ox Alpha appeared without explanation on OpenRouter and OpenCode, an open-source terminal coding agent. It was simply described as a frontier model designed for efficient coding, long-running agent tasks and production environments. Input and output were free, it supported million-token contexts and could accept images and video.

Within a day, it had become OpenRouter’s most-used model, setting a platform record. Developers began trying to identify its creator. Six days later, Zhipu claimed the model and simultaneously released GLM-5.3-Flash.

The stunt was not entirely new. In February, an anonymous model called Pony Alpha appeared online and was identified five days later as Zhipu’s GLM-5. The company says anonymous releases allow developers to test a model in real-world conditions without being influenced by its identity or reputation.

The strategy also gives Zhipu an obvious commercial advantage because its main business is selling API access. 

GLM-5.3-Flash uses a fully permissive MIT license, allowing commercial use, modification and redistribution without additional conditions. Zhipu’s model-as-a-service (MaaS) platform had reached an annualized recurring revenue run rate of about $1.6 billion by the end of August, up 60% from early July. In the first half, revenue from its open platform and AI related services soared 2,736% year on year to 825 million yuan, accounting for 86.5% of total revenue.

For a company whose revenue depends on developers calling its APIs, wider distribution can be an asset rather than a threat. The more environments developers use the model in, the greater the potential pool of users that can eventually be directed back to the company’s own services.

In that sense, MIT is not simply a giveaway. It is a way of minimizing distribution costs.

There is a price, however. Zhipu absorbed the computing costs of making the model freely available during its anonymous testing period, which it says attracted more than 500,000 unique users and 13 million sessions.

The company is effectively paying today’s infrastructure bill in the hope of establishing tomorrow’s user habits.

Moonshot wants a cut from the cloud

Moonshot AI is taking almost the opposite approach.

When it released Kimi K3 in July, the 2.8 trillion-parameter model was made available for download, modification and commercial use. But its license was not a standard open-source license. Hugging Face labelled it simply “other”.

The distinction matters. K3 is better described as an open-weight model than a conventionally open-source model because its license imposes commercial conditions.

The most important is aimed at companies providing MaaS. Businesses and their affiliates whose MaaS revenue exceeds $20 million over a consecutive 12-month period must sign a separate agreement before using K3 commercially.

A second condition is aimed at branding: products with more than 100 million monthly active users or more than $20 million in monthly revenue must prominently identify the model as Kimi K3.

The $20 million threshold is a trigger for negotiations, not an automatic 30% revenue share. The 30% figure reported in overseas media is one commercial condition Moonshot has proposed to large customers, with the final percentage dependent on negotiations.

The structure effectively leaves individual developers, small teams and internal corporate users in the free zone while reserving the right to negotiate with companies that turn K3 into a large commercial business.

That is particularly relevant to cloud providers such as Microsoft Azure, AWS, and Google Cloud. If they offer K3 as an external inference service, they would have to reach an agreement with Moonshot rather than simply deploy the model under a permissive open-source license. Negotiations are already underway, although no agreements have been reached. 

The company is also building a wider distribution network. K3 is available through inference providers including Together AI, Fireworks, DigitalOcean, Modal, Baseten and DeepInfra, with more than 10 third-party channels identifiable.

This arrangement is potentially attractive to both sides. Moonshot can reach enterprise customers without building and financing its own global inference infrastructure, while cloud providers can offer customers another competitive model and keep their workloads on their platforms.

But the economics become considerably harder to calculate with large cloud contracts.

An inference provider selling K3 directly by token can relatively easily establish how much revenue the model generates. Large cloud platforms, by contrast, sell enterprise customers annual contracts, discounts, reserved capacity and bundles of AI products. K3 may account for only one component of a much larger contract.

Determining what proportion of that revenue belongs to Kimi becomes a commercial and auditing problem. Moonshot needs enough visibility into usage to calculate its share, while cloud providers need to protect their customers’ data.

The negotiations therefore have to resolve three difficult questions: how revenue is allocated, what usage data can be shared and how consumption can be audited.

MiniMax takes a third route

MiniMax is pursuing a more direct model: get paid to deliver the technology.

The company released its M3 model in June, combining advanced coding capabilities, million-token context and native multimodality. It was quickly adapted for domestic chip platforms and entered OpenRouter’s top three models by usage.

In September, Saudi sovereign wealth fund PIF’s AI company HUMAIN released an Arabic-language model called humain-m3, which it said was commissioned by HUMAIN and delivered by MiniMax.

For Saudi Arabia, the project turns the concept of sovereign AI into a concrete product. For MiniMax, it provides an external demonstration that its general-purpose model can serve as the technological foundation for a country-specific AI system.

M3’s license is also a customised “Community License Agreement” rather than a standard open-source license. Like K3, it uses a $20 million threshold, although the calculation is different: K3 looks at the overall MaaS revenue of a company and its affiliates, while M3 focuses on annual revenue from the relevant product or service.

M3 also imposes obligations on some commercial users below the threshold, including displaying “Built with MiniMax M3” and notifying MiniMax once.

The commercial logic of a commissioned model is straightforward. The customer, price, and scope are agreed in a contract, rather than relying on token consumption to generate revenue over time.

The trade-off is that the relationship is deeper. A project built into another country’s sovereign AI strategy creates a stronger commercial and geopolitical dependency, and potentially higher exit costs.

Why tokens are changing the equation

The three companies are not simply choosing different licensing strategies. Their different business models are determining what they need from open distribution.

Zhipu has a strong API business and therefore benefits from maximising usage. Moonshot is trying to capture value from third-party distribution. MiniMax is monetizing the technology through direct delivery.

The deeper change, however, is the economics of AI agents.

Historically, model revenue could be approximated as the number of users multiplied by the tokens consumed in each conversation. Growth in usage was closely tied to growth in the user base. Agents break that relationship.

An automated task can consume tens or hundreds of times as many tokens as a simple exchange, while hundreds of enterprise users running agents continuously can generate more inference demand than millions of consumer users.

MiniMax’s own figures illustrate the scale of the change: platform token consumption in July 2026 was 20 times January’s level.

Under a completely open license, much of that additional consumption can take place on third-party infrastructure. The model provider gets the benefit of having a widely used model, but not necessarily the revenue generated by running it.

That changes the calculation behind “free”.

Previously, giving away tokens could be justified because higher usage was expected to produce more users, and those users could eventually be monetized. With agents, free usage may simply generate inference revenue for another platform.

The licensing changes are therefore less about suddenly deciding that AI models need to make money than about who gets to monetize the tokens they generate.

That provides a useful way to assess future licenses. Companies with their own API and MaaS platforms have an incentive to remain relatively permissive because more deployments can ultimately generate more calls to their services.

Companies dependent on cloud providers and third-party inference platforms have a stronger incentive to reserve commercial rights in their licenses.

Alibaba’s (9988.HK) (BABA.US) Qwen3.8-Max, released in August, points in the same direction. Its commercial threshold is $50 million, more than twice the level set by Kimi and MiniMax, and it specifically defines an “AI work assistant” category covering coding and office productivity products.

Licensing is therefore beginning to define not just whether a model can be used commercially, but where the model provider believes commercial competition should begin.

The license is part of the business model

For model companies, the lesson is that licensing should be designed before distribution begins, not renegotiated after a model has become popular.

Kimi is an example of the problem. By the time Moonshot began negotiating revenue-sharing arrangements, the model had already been distributed across more than 10 third-party channels. Once partners have invested in computing infrastructure and built distribution around a model, the model provider’s negotiating leverage can actually weaken.

The implications extend to users of open models.

Companies evaluating an LLM often focus on performance benchmarks and API prices. But a license can be changed, and the economics of open-weight models are evolving quickly. MongoDB, Elastic and HashiCorp have all changed licenses in response to the economics of cloud providers hosting their products. AI models face a similar conflict, except that the asset being distributed is a set of model weights rather than software code.

The practical response is to assess not just today’s license but the cost of being forced to leave it.

How much retraining would be required? How much of the inference stack would need to change? How quickly could performance benchmarks be rebuilt against another model?

The higher those switching costs, the more important it is to maintain alternatives.

For model companies, meanwhile, tighter licensing is no guarantee of lasting pricing power. Open-weight models can be replaced quickly. A 30% revenue share may be achievable while a model is clearly ahead of its competitors, but a cheaper model with similar performance and a more permissive license could weaken that bargaining position within months.

The same principle applies to users: the safest strategy is not to bet entirely on any one company’s licensing policy.

An abstraction layer between applications and models, regularly rerun evaluation benchmarks and a viable alternative model can all reduce the cost of switching.

As AI agents drive token consumption higher, the question is no longer simply which model is best or cheapest. Increasingly, it is who owns the right to make money when that model is used.

Source: 
Silicon Valley Tech News/TMTPost

Share the story:
,