
Stanford-educated founder Lin Qiao builds a $17.5 billion platform in four years as capital pivots from training to inference
By Zhang Nan
When Fudan University alumna Lin Qiao left tech giant Meta in October 2022, just a month before ChatGPT’s launch, she assembled a seven-person team in Redwood City, California, to found Fireworks AI. Four years on, the company is valued at $17.5 billion, backed by venture capital firms including Sequoia, TCV, and Lightspeed, and hailed by Nvidia co-founder Jensen Huang as “in a lot of ways, the TSMC of AI factories.”
Fireworks AI does not train frontier models or build consumer-facing AI applications. Instead, it focuses on inference — the process of running and applying trained models to answer user queries — by fine-tuning and hosting open-source models for enterprises on a pay-per-use basis.”
The company has just closed a $1.5 billion Series D funding at a $17.5bn post-money valuation, led by Atreides Management, Index Ventures and TCV, with Nvidia participating and existing backers including Lightspeed and Sequoia adding to their stakes. The new funding will primarily expand compute capacity — Fireworks sources GPUs from more than 20 suppliers, including Microsoft — and grow its engineering team from roughly 200 to 600 by year-end.
The launch of Moonshot AI’s Kimi K3, combined with the pending release of DeepSeek V4, is undermining the high valuations and competitive advantages long enjoyed by top model builders. Instead, investor capital is quietly shifting to the inference layer—moving away from the race to build the most powerful model toward the platforms that can run these models reliably and affordably.
From PyTorch to unicorn
Lin Qiao grew up in China, the daughter of a shipbuilding engineer. She earned a bachelor’s and master’s in computer science from Fudan University before moving to the U.S. for a PhD at UC Santa Barbara. After stints at IBM and LinkedIn, she joined Meta in 2015, rising to senior engineering director and co-founding PyTorch — a project that stretched into a five-year foundational rewrite of Meta’s entire AI workload, supporting over 5 trillion inferences daily by the time she left.
Her conviction was that every enterprise wanted to adopt AI but lacked the infrastructure to deploy models into production. At Meta it had taken five years to solve that problem; Fireworks would compress that into five weeks. Six of the seven founding team members came from Meta’s PyTorch group, with one from Google’s Vertex AI.
Fireworks AI operates a cloud platform that gives corporate developers access to more than 200 open-source artificial intelligence models. Rather than simply renting out raw computing hardware, the startup differentiates itself on software efficiency. Its proprietary processing stack allows AI models to process data up to five times faster than competitors at comparable prices
The platform allows enterprises to securely fine-tune models on their own proprietary data while meeting strict regulatory compliance standards. In a sign of how corporate demand is evolving, 95% of the processing volume on Fireworks AI today comes from these customized, business-specific models rather than off-the-shelf software.
Lin’s founding philosophy was compound AI: using hundreds of small expert models to solve narrow problems rather than betting on a single giant model. At comparable quality, Fireworks AI claims to be five to ten times cheaper than closed-source models. That price advantage is the engine of its growth: as bills for the latest models unsettle finance chiefs, and as open-source models narrow the capability gap with closed alternatives, enterprises are seriously considering switching.
Global token factories rising
Fireworks AI is not alone. Together AI, a cloud-based artificial intelligence platform and infrastructure provider specializing in open-source generative AI, closed an $800 million Series C in July at an $8.3 billion valuation, led by Aramco Ventures with Nvidia and Salesforce participating. Baseten, an AI infrastructure company that specializes in machine learning inference, announced a $1.5 billion Series F in June at a $13 billion valuation, having raised four rounds in 18 months with revenue up roughly 20-fold year-on-year; Nvidia is also a shareholder.
In China, the focus is on heterogenous architecture — making Nvidia and domestic chips work together. SiliconFlow, founded by Tsinghua PhD Yuan Jinhui, has filed for a Hong Kong IPO after seven funding rounds in under three years, with Alibaba as its largest institutional shareholder. The company was the first to deploy a full DeepSeek version on Ascend chips designed by HiSilicon, a fabless semiconductor subsidiary wholly owned by Huawei Technologies, when DeepSeek’s site was overwhelmed earlier this year. Infinigence AI, co-founded by Tsinghua professor Wang Yu, recently raised over $100 million, and its platform now hosts more than 160 models.
The margin question
The surge in token demand since early 2026, driven by agentic applications, has dramatically increased per-request token volumes. Open-source quality is approaching closed-source levels, giving enterprises viable alternatives. And companies are finally doing the math: training is a one-off cost, but inference is a perpetual drain — the more popular the product, the more frequent the calls, and the deeper the potential losses. Some investors note that AI hardware makers may prefer users to use less AI functionality because “tokens are simply too expensive to burn.”
That puts cost-per-inference as the critical line, and token factories — sandwiched between models, GPUs, cloud providers and applications — are the ones solving the puzzle.
But gross margins matter enormously. Compute procurement prices, utilisation rates and client pricing pressure all eat into profitability. If competition intensifies and players undercut each other, the business risks becoming high-revenue, low-margin resale. And if training intensity declines, token factories could be squeezed between cloud giants and large model developers.
Ultimately, whether Fireworks AI or any other player succeeds depends on whether inference consolidates like cloud computing — which converged around AWS, Microsoft and Google — or sustains an independent layer. The answer lies in whether it can turn running models into an indispensable, defensible efficiency that customers cannot live without.
Source:
Chinaventure.com.cn