Computer geeks Tang Jie and Yang Zhilin were mere teacher and pupil more than a decade ago, with dreams of building a machine that could think like a human. AI pioneers looking over their shoulders with alarm. Today, they are two of the Chinese artificial-intelligence tycoons who have U.S.
DeepSeek was among the earliest adopters in China of a model design called “ mixture of experts ,” or MoE, in which an initial routing mechanism directs the problem to a specialized expert model—akin to a head chef directing a spaghetti order to the kitchen’s Italian cook. Liang’s team at DeepSeek came up with some of the most important workarounds, including multihead latent attention, or MLA. That technique slashes an AI model’s memory usage—such as in a chatbot conversation—by creating a shorthand version of what has gone before. Using less memory saves computing power and money. DeepSeek showed the Chinese industry a viable path to boost model performance while easing the demands on chips. Yang incorporated techniques that were invented or proven by DeepSeek into Moonshot’s subsequent models. Its K2 and K3 models adopted the MoE design and variants of MLA. DeepSeek, in turn, used a technique optimized by Moonshot to boost training efficiency and stability.

