The Art of the Missing Piece
How Z.ai used open models, export pressure, and academic roots to challenge the closed AI frontier
On March 3, 2021, researchers at Tsinghua University posted a paper with a title only a peer reviewer could love: GLM: General Language Model Pretraining with Autoregressive Blank Infilling. Buried under the jargon was a parlor trick. Take a passage of text, cut pieces out of it, shuffle them, then ask a machine to restore each one, letting it read everything that remains but forcing it to write the missing part one word at a time, blind to what comes next.
The trick mattered because of what it unified. The machines of that era came in two temperaments. One kind, in the lineage of Google’s BERT, read text the way you proofread a contract and was nearly useless for writing. The other, the GPT line, wrote fluently but only ever looked backward. Nobody had built one model that did both well. Blank infilling did, because filling a gap demands both at once: comprehension of everything around the hole, composition of everything inside it. The Tsinghua model, trained on this one exercise, outperformed BERT at reading, GPT at writing, and T5 at transforming.
Five years later, the descendants of that paper sit fourth on the global intelligence leaderboards, behind only the best of Anthropic and OpenAI, ahead of Google’s Gemini on several coding benchmarks, and available to anyone as a free download. The company behind them, Z.ai, became in January 2026, the first foundation-model company anywhere to go public. If you write software, its models are likely already in your toolchain.
Most people who’ve heard of GLM heard of it this year, as a rumor from the coding forums: a Chinese model that rivals Claude Opus at a sixth of the price. The assumption is that it appeared from nowhere, conjured overnight in DeepSeek’s wake. But GLM is not new. It’s one of the oldest continuous lineages in the field, older than ChatGPT. To understand what Z.ai is, start with what it grew out of: not a startup, but a research group with two decades of sediment beneath it.
A Company Made of Professors
The Knowledge Engineering Group at Tsinghua’s computer science department had been run by a professor named Tang Jie for twenty years before the world had reason to notice it. Tang spent much of his career on AMiner, a system that maps the world’s academic literature: who wrote what, who cites whom, how ideas move through the network of minds. Before he taught models to write, he spent years teaching software to read scholarship.
On June 11, 2019, Tang and his colleagues incorporated a company in Zhongguancun, the Beijing university district that plays the role Palo Alto plays in the American imagination. They called it Zhipu AI; the formal name, Knowledge Atlas Technology, gives away the founders’ cast of mind. The roster contains no dropouts and no wunderkind, just middle-aged academics who had spent their careers on the unglamorous machinery of knowledge itself, the citation graph and the knowledge base. When the large language model arrived, they didn’t pivot into it. They recognized it as the thing they had been circling all along.
The 2021 blank-infilling paper was the lab’s opening claim. In August 2022, months before ChatGPT existed, the group scaled the idea to 130 billion parameters and released GLM-130B, the first large bilingual English-Chinese model out of China’s research ecosystem that outsiders could touch. Then OpenAI lit the fuse in November 2022, and the quiet academic timeline compressed into a commercial sprint.
The Ladder
What followed is best read as a ladder, each rung a bet about what language models are for.
In March 2023, Zhipu open-sourced ChatGLM-6B, a conversational model small enough to run on a hobbyist’s graphics card. The bet: accessibility. January 2024 brought GLM-4 and a paid API. The bet: capability was now worth money. Through 2024 the family sprouted variants for voice, vision, and million-token contexts, tracking OpenAI’s product line the way a wolf tracks a caribou herd, close enough to learn from it, far enough back to conserve energy.
Then, in July 2025, the interesting rung. The company renamed itself Z.ai for the global market and released GLM-4.5 with two structural changes at once. The architecture went to mixture-of-experts, a design in which only a fraction of the parameters wake up for any given word, which slashes the cost of running it. And the license went to MIT, the legal equivalent of leaving the keys in the ignition with a note that says take it. A frontier-scale model any enterprise could download, modify, and deploy forever, with no API dependency and no kill switch.
The declared mission changed too. Z.ai stopped talking about chat and started talking about agents: models that complete tasks, working autonomously across hours of planning and repair. GLM-4.6, that September, reached a 48.6 percent win rate against Claude Sonnet 4 on human-judged coding tasks, and Z.ai published it alongside an admission that it still trailed Sonnet 4.5. A company that publishes its losses is telling you it expects to be around for the rematch.
The rematch came quickly. GLM-5, in February 2026, arrived under a report titled From Vibe Coding to Agentic Engineering, a jab at prompt-and-pray development: 744 billion parameters with 40 billion active, a sparse-attention scheme borrowed openly from DeepSeek, and a training system called Slime that lets the model learn from long agentic tasks without pausing the whole factory. GLM-5.1 followed in April and took the top score on SWE-Bench Pro, a software-engineering gauntlet, ahead of the best from OpenAI, Anthropic, and Google. Then in June 2026, GLM-5.2: a million-token context window, first among open-weight models on the Artificial Analysis Intelligence Index, fourth overall behind Claude Opus 4.8 and GPT-5.5, ahead of GPT-5.5 on coding, at about one-sixth the price. The gap between the closed frontier and the open one, once measured in years, is now measured in benchmark points.
The Weights Walk Free
The question worth asking is not how good the models are. The leaderboards answer that monthly. The question is why a company burning enormous sums to train frontier models keeps giving them away. The answer is that the giveaway is the strategy, and it works on four levels at once.
The first is distribution. A model under MIT license needs no sales force. GLM-4.6 appeared inside Claude Code, Cline, Kilo Code, and Roo Code within days because developers put it there, having found a near-drop-in replacement for Claude Sonnet at one-seventh the cost. Every open release recruits an unpaid, planet-wide integration team.
The second is the enterprise wedge. A bank in Frankfurt or a ministry in Jakarta that cannot ship its data to a US API can download GLM’s weights and run them in its own basement. The closed American labs cannot match this at the frontier. Z.ai fills the gap the US business model leaves open, the blank in the market; a model trained on blank infilling turning out to be a company that practices it is the kind of symmetry history rarely bothers to arrange.
The third is insurance. Weights, once published, cannot be recalled. Whatever export controls come next, GLM-5.2 is already on a hundred thousand hard drives.
The fourth is pressure. Every open release at near-frontier quality compresses what OpenAI and Anthropic can charge. On GDPval, a benchmark of economically valuable work, GLM-5.2 and GPT-5.5 are statistically tied. One costs six times as much as the other. That arithmetic does the marketing on its own.
The Blank in the Supply Chain
In January 2025, the US Commerce Department added Zhipu to the Entity List, the export-control roster that walls a company off from American technology, citing links to Chinese military modernization. It was the first time Washington had named a Chinese LLM company this way. Zhipu called the listing baseless and carried on.
Here the pattern of the company’s history repeats at the level of hardware. Cut off from Nvidia, Z.ai turned to Huawei’s Ascend processors, and by January 2026, announced that GLM-Image had been trained entirely on Huawei’s Atlas servers, the first claimed state-of-the-art multimodal model raised wholly on domestic silicon. Huawei’s chips still trail Nvidia’s, and nobody at Z.ai pretends otherwise. But the direction is unmistakable. Something was removed, and the company wrote in what was missing.
The capital markets have rendered their verdict. Z.ai’s Hong Kong IPO in January 2026 raised $558 million and was oversubscribed 1,159 times in the public tranche, with a second listing in Shanghai now planned. Investors are not buying profits; there are none. Fiscal 2025 brought roughly $105 million of revenue against a $683 million net loss. What they’re buying is the trajectory: revenue up 132 percent, API call volume up 400 percent in a single quarter, and a position Z.ai describes as the Global Token Factory: China’s energy costs and manufacturing scale, applied to machine intelligence, sold to the world at prices no one else can match.
Whether the money lasts long enough for the thesis to mature is a genuinely open question. The losses run six times revenue. The domestic competition, DeepSeek and Alibaba’s Qwen and Moonshot among the pack the Chinese press calls the Six AI Tigers, is ferocious, and Z.ai’s turns at number one have lasted weeks. GLM-5.2 also thinks out loud at greater length than its peers, which erodes some of the headline price advantage in exactly the agentic loops it was built for. None of this is hidden. All of it is priced in.
What the Machine Was Taught
It’s tempting to file Z.ai under geopolitics, and the Entity List and the Huawei chips do belong to that story. But the geopolitical frame arrives late and explains little. The company predates the rivalry’s AI chapter. Its strategy was set by its origins: a knowledge-engineering lab that spent twenty years mapping how ideas propagate, then built an instrument for propagating them, then noticed that the instrument spreads furthest when you let go of it.
The deeper continuity is in the training objective itself. In 2021, in a lab in Beijing, researchers taught a machine that the way to understand the world is to study what has been removed from it, and to supply the absence. The machine learned the lesson. So did its makers, who have spent every year since finding the removed thing, whether an open frontier model or a sanctioned chip, and writing it in. Fill in enough blanks and eventually you stop being the student of the exercise. You become its author.


