No Bleeding-Edge Node, No HBM, No GPU Clone: Dongfang Suanxin's First AI Chip DF1000 Bets on "Moving Fast"

Dongfang Suanxin unveiled DF1000, China's first software-defined near-memory-computing 3D AI chip, targeting the inference era's "move fast" over "compute fast," validated with a fully domestic tape-out and a 128-card cluster.
The inference era rewrites the rules
For years AI chips followed one path: finer process, bigger GPU clusters, more FLOPS. But the inference era changes the rules — models care less about training speed and more about inference cost, token throughput and deployment efficiency, while energy, fabrication and supply-chain limits bite harder. At its first launch in July, Dongfang Suanxin introduced DF1000, China's first "software-defined near-memory-computing" 3D AI chip, with a set of engineering validations: a fully domestic supply-chain tape-out, a 128-card cluster running stably in real workloads, and a full-stack software ecosystem.
The truly scarce resource may be energy
Hong Kong Academy of Engineering Sciences fellow Zheng Guangting argued AI is constrained by the physical world. Per the IEA, AI-dedicated data-center power use rose 50% in 2025, far above the 17% for data centers overall; Gartner projects that by 2030 AI-optimized servers will consume nearly half of data-center power. He suggests the scarce resource ahead may be energy, not compute — releasing more effective compute under limited energy is the industry's shared problem.
From "compute fast" to "move fast"
VP Guo Wei noted: pre-training, ~70% of training compute, is compute-bound, so chip FLOPS still drive training. But inference differs — with parameters fixed, the system repeatedly reads weights and KV cache to generate tokens, so the contest shifts from "compute fast" to "move fast." The decode stage dominates inference time, and its token throughput hinges on memory bandwidth between compute and storage.
Two hurdles for domestic GPUs
Guo conceded training increasingly needs advanced nodes — leaders are at 3nm while domestically accessible processes remain around 14nm, a generation gap that caps gains; on inference the bottleneck moves to memory and interconnect — advanced HBM sets per-card bandwidth but is supply-limited, and high-speed I/O density is process-bound. More fundamentally, as Beijing Chaoxian Memory Institute EVP Zhao Chao noted, the classic von Neumann architecture may not be optimal for AI. DF1000's near-memory computing is a differentiated answer to "moving fast," the inference era's new bottleneck.