AI Compute Goes Multi-Chip: What Heterogeneous Data Centres Mean for Malaysia
Power bills, token costs and orchestration software are ending the GPU-only era, and the shift reaches Penang's packaging floors, Johor's server halls and your monthly AI bill.

Data centres are moving away from running AI on a single type of chip. According to Semiconductor Engineering, four pressures — power consumption, the cost per token, interconnect bandwidth, and software orchestration — are pushing operators toward heterogeneous clusters that mix CPUs, GPUs, NPUs, optical links, and custom accelerators, each doing the job it is best at. In plain terms, the industry is learning that asking one expensive, power-hungry chip to do everything wastes money and electricity. For Malaysia, this touches the Penang semiconductor corridor, the country's data centre build-out, and the price Malaysian SMEs ultimately pay for AI services.
AI Summary
Data centres are moving away from running AI on a single type of chip. According to Semiconductor Engineering, four pressures — power consumption, the cost per token, interconnect bandwidth, and software orchestration — are pushing operators toward heterogeneous clusters that mix CPUs, GPUs, NPUs, optical links, and custom accelerators, each doing the job it is best at. In plain terms, the industry is learning that asking one expensive, power-hungry chip to do everything wastes money and electricity. For Malaysia, this touches the Penang semiconductor corridor, the country's data centre build-out, and the price Malaysian SMEs ultimately pay for AI services.
Key Takeaways
- The GPU-only data centre is ending. Semiconductor Engineering reports that AI compute is shifting to heterogeneous clusters — CPUs, GPUs, NPUs, optics, and custom accelerators working as one system.
- Token cost has become an engineering problem, not just a pricing line item. Running cheap tasks on expensive silicon inflates every AI bill, which is why operators are splitting work across chip types.
- Interconnects, especially optical links, are becoming as important as the chips themselves. Mixed-silicon clusters only work if data moves fast between processors.
- Software orchestration — the layer that decides which chip runs which task — is the new competitive battleground, and the hardest engineering problem in the stack.
- For Malaysia: heterogeneous design typically means more advanced packaging work, which plays to Penang's assembly-and-test strengths, while local data centres get more AI output per megawatt of electricity.
What Happened
Semiconductor Engineering, a specialist publication covering chip design and manufacturing, reports that the future of AI compute will not run on just one kind of chip. The article's core finding is that data centres are being pushed toward heterogeneous clusters — server systems built from a mix of processors rather than racks of identical chips.
Four forces are driving this, per the report. First, power: AI workloads consume enormous electricity, and not every task justifies the draw of a top-end processor. Second, token cost: AI services are priced and measured in tokens (roughly, fragments of words), and processing every token on the most expensive available silicon is economically wasteful. Third, interconnects: for mixed chips to cooperate, data has to move between them at high speed, which is pushing adoption of optics — links that carry data as light instead of electrical signals through copper. Fourth, software orchestration: someone has to schedule which chip handles which piece of work, and that scheduling layer is becoming central to how these clusters perform.
To understand the shift, it helps to know the players. The CPU (central processing unit) is the generalist — good at branching logic and running systems. The GPU (graphics processing unit) became the AI workhorse because it does massive amounts of parallel maths, which suits training large AI models. The NPU (neural processing unit) is newer: a chip designed specifically for AI operations at lower power, the kind already found in modern laptops and phones, now moving into servers. Custom accelerators are chips that big operators design or commission for their own specific workloads, often to run AI inference — the act of using a trained model to produce answers — more cheaply than a general-purpose GPU can.
Think of it like a kitchen. The old model was one master chef doing everything: chopping, grilling, plating. The new model is a brigade — a butcher, a grill cook, a pastry chef, plus a head chef directing traffic. You get more meals out per hour and waste less. But the head chef's job, orchestration, just became the difference between profit and chaos.
Why It Matters
This matters because it changes where the money and the moats sit in the AI industry. For
Sources & References
AIBlog summarises and analyses published information. We do not reproduce full source text. Analysis is editorial and not financial or legal advice.

