AIAIBlog.com.my
Semiconductor & AI Infrastructure5 August 2026 · 12 min read

Scale Up, Out Challenges Amplified For Clusters

Scale Up, Out Challenges Amplified For Clusters
AIAI Summary

Compute clusters — the backbone of modern AI training and inference — deliver far more processing power than individual servers, but they bring a steep penalty: significantly higher power consumption. According to Semiconductor Engineering, the industry is now grappling with how to scale these clusters both "up" (more powerful individual nodes) and "out" (more nodes connected together), with the main bottleneck shifting from raw compute to software that can allocate resources, optimise networking, and manage memory efficiently. For Malaysia, which hosts a growing semiconductor corridor in Penang and a fast-expanding data centre footprint in Johor and the Klang Valley, these engineering challenges directly affect the cost and feasibility of running AI workloads locally. The practical takeaway for Malaysian businesses: AI infrastructure costs will stay volatile, software-defined resource management will become a critical skill, and any organisation planning serious AI deployment needs to understand cluster economics before committing budget. ---

Scale Up, Out Challenges Amplified For Clusters

AI Summary

Compute clusters — the backbone of modern AI training and inference — deliver far more processing power than individual servers, but they bring a steep penalty: significantly higher power consumption. According to Semiconductor Engineering, the industry is now grappling with how to scale these clusters both "up" (more powerful individual nodes) and "out" (more nodes connected together), with the main bottleneck shifting from raw compute to software that can allocate resources, optimise networking, and manage memory efficiently. For Malaysia, which hosts a growing semiconductor corridor in Penang and a fast-expanding data centre footprint in Johor and the Klang Valley, these engineering challenges directly affect the cost and feasibility of running AI workloads locally. The practical takeaway for Malaysian businesses: AI infrastructure costs will stay volatile, software-defined resource management will become a critical skill, and any organisation planning serious AI deployment needs to understand cluster economics before committing budget.


Key Takeaways

  • Compute clusters trade power efficiency for processing capacity — the more you scale, the more power-hungry and complex the system becomes, making electricity cost and cooling a primary business concern.
  • The main challenge is no longer just getting enough chips — it is software that can intelligently distribute workloads across nodes, manage memory, and optimise network traffic within a cluster.
  • Scale-up (building bigger, more powerful individual machines) and scale-out (connecting many smaller machines together) each carry distinct trade-offs in cost, complexity, and performance that are getting harder to manage as cluster sizes grow.
  • Malaysia's semiconductor and data centre investments position the country well on the supply side, but the skills gap in cluster software management could limit how much local value is captured.
  • Businesses should treat AI infrastructure planning as an operations problem, not a procurement problem — the total cost of running a cluster over three years can dwarf the initial hardware purchase price.

What Happened

Semiconductor Engineering reported that the semiconductor and AI infrastructure industry is confronting intensifying challenges around scaling compute clusters. These clusters — large groups of connected processors working together — are the physical foundation for training large language models and running inference at commercial scale. The article identifies a core tension: clusters provide substantially more processing capacity than standalone systems, but this comes at the direct cost of higher power consumption.

The report highlights that two scaling strategies are under strain. "Scale-up" means making individual nodes more powerful — adding more processors, more memory, more capability to each machine. "Scale-out" means adding more nodes to a cluster, connecting additional machines to share the workload. Both approaches amplify engineering problems. More powerful nodes generate more heat and draw more electricity. More nodes create networking bottlenecks and memory synchronisation issues. The larger the cluster, the more pronounced these problems become.

Semiconductor Engineering points to software as the critical missing piece. Hardware improvements alone cannot solve these challenges. What is needed is software capable of allocating resources effectively across a cluster — deciding which processor handles which task, when, and how data moves between nodes. This includes optimising the network fabric that connects nodes and managing the memory architecture so that data is available where and when it is needed without creating latency or bottlenecks.

The underlying issue is one of co-design. In the past, compute, memory, networking, and software were developed largely as separate layers. At cluster scale, these layers are deeply interdependent. A networking decision affects memory performance. A memory architecture decision affects how software must be written. A power constraint affects everything. The article signals that the industry is being forced to confront these interdependencies simultaneously rather than sequentially.

This is not a distant future problem. It is a present-day engineering challenge affecting every company building large-scale AI systems, from cloud providers to chip designers to the fabs that manufacture semiconductors.


Why It Matters

The cluster scaling problem matters because it sits at the intersection of two trends defining the AI industry: exploding demand for compute and physical limits on power delivery.

AI model training requires enormous processing capacity. Large language models — the systems behind tools like ChatGPT, Gemini, and Claude — are trained on clusters of thousands of processors running for weeks or months. As models grow larger and companies train more of them, clusters must scale up and out simultaneously. But every additional processor draws more power, generates more heat, and requires more cooling. In some regions, data centre operators cannot get enough electricity from the grid to power new clusters. This is no longer a theoretical constraint — it is a practical bottleneck shaping investment decisions.

The shift from hardware-as-bottleneck to software-as-bottleneck is significant. For years, the limiting factor in AI was chip performance — could you get enough processing power. That problem has not disappeared, but it has been joined by a harder one: can you use the processing power you already have efficiently. A cluster running at 60% utilisation because software cannot distribute workloads effectively wastes enormous amounts of money and energy. The companies that solve this — through better scheduling software, smarter memory management, and optimised networking — will have a structural cost advantage over those that do not.

This also matters for the semiconductor supply chain. Companies like Nvidia, AMD, Intel, and TSMC design chips, but the value of a chip in an AI cluster depends heavily on the software ecosystem around it. A processor that is theoretically fast but cannot be efficiently integrated into a large cluster is worth less in practice than its benchmarks suggest. This is why companies are investing not just in silicon but in the full stack — networking hardware like InfiniBand and Ethernet, memory technologies like HBM (high-bandwidth memory), and the software layers that tie everything together.

For investors and business strategists, the signal is clear: the AI infrastructure market is not just about who makes the fastest chip. It is about who delivers the most efficient, manageable, and cost-effective cluster. That competition will play out over the next several years and will determine which technology providers capture the most value from the AI buildout.


What This Means for Malaysia

Malaysia occupies a strategic position in this story. Penang and Kulim host major semiconductor manufacturing and packaging operations for companies including Intel, AMD, Infineon, and Bosch. Malaysia accounts for an estimated 13% of global semiconductor packaging and testing. The cluster scaling challenges described by Semiconductor Engineering create both opportunity and pressure for these Malaysian facilities — demand for advanced packaging, interconnect technologies, and memory-related components will grow as cluster architectures evolve.

On the data centre side, Johor has emerged as one of Southeast Asia's fastest-growing data centre markets. Microsoft, Google, Amazon Web Services, and Equinix have all announced or operational facilities in Malaysia. These data centres will house the kind of compute clusters described in the source material. The power consumption challenge is directly relevant here — Malaysia's grid capacity, energy pricing, and cooling infrastructure will determine how competitive these facilities remain over time. States like Sarawak, with surplus hydropower, are positioning themselves as alternative locations for power-intensive AI infrastructure.

The skills dimension is critical. Malaysia has strong semiconductor manufacturing talent, but the source article points to software-defined resource management as the key differentiator. Cluster orchestration, workload scheduling, network optimisation, and memory management require a different skill set from traditional semiconductor engineering. Malaysian universities and training institutions — under initiatives like MDEC's digital economy programmes and the National Semiconductor Strategy announced in 2024 — need to produce graduates who understand both the hardware and the software sides of cluster computing. Without this talent, Malaysia risks remaining a manufacturing and hosting location while higher-value software and design work happens elsewhere.

For Malaysian SMEs, the practical impact is indirect but real. If you are building AI applications, your cloud costs are influenced by how efficiently providers manage their clusters. If cluster software improves and utilisation rises, inference costs could fall over time. If power constraints limit cluster expansion, compute costs could rise. Either way, understanding these dynamics helps Malaysian businesses make better decisions about when to use cloud-based AI services versus investing in on-premises infrastructure.


How Your Business Can Use This

Start by auditing your current and planned AI workloads. If your business runs machine learning models — for customer service automation, demand forecasting, fraud detection, or content generation — understand where the compute happens and what it costs per month. Most Malaysian SMEs access AI through cloud APIs, which means cluster efficiency is already baked into your pricing. But knowing whether you are paying for inefficient cluster utilisation helps you negotiate better terms with providers or evaluate alternative suppliers.

For companies running their own GPU servers or considering on-premises AI infrastructure, the cluster scaling challenge should reframe your procurement thinking. Do not evaluate hardware in isolation. A high-performance GPU that sits idle 40% of the time because your software cannot distribute workloads efficiently is a poor investment. Before buying hardware, assess your software stack — your orchestration layer, your scheduling logic, your network configuration. If you lack in-house expertise, work with a managed service provider who understands cluster optimisation, not just hardware sales.

Malaysian businesses in manufacturing, logistics, and financial services should also consider phased AI infrastructure strategies. Start with cloud-based inference for pilot projects. Move to reserved cloud capacity once workloads are predictable. Only invest in dedicated cluster infrastructure when your utilisation patterns justify the capital expenditure and when you have the software talent to manage it. This staged approach protects cash flow and avoids the trap of buying hardware that underperforms because the software layer is neglected.


The Agentic AI Angle

Agentic AI — autonomous systems that plan, reason, and execute multi-step tasks without constant human supervision — has a direct relationship with cluster economics. AI agents that call multiple tools, retrieve information from databases, run calculations, and coordinate with other agents generate significantly more compute requests than a simple chatbot interaction. Each agentic step may involve a separate inference call. A single user request handled by an agent with ten steps could consume ten times the compute of a standard query.

This means businesses deploying agentic AI must think about cluster utilisation from the start. If your agent architecture is poorly designed — making redundant calls, failing to cache intermediate results, or not batching requests — you will waste cluster capacity and pay for it. Well-designed agent systems, by contrast, can be optimised to minimise compute usage by caching tool outputs, parallelising independent steps, and routing requests to the most cost-effective model for each subtask.

On the infrastructure side, agentic AI creates a compelling use case for the cluster management software described in the source article. As agent deployments scale from dozens to thousands of concurrent users, the cluster must dynamically allocate resources based on real-time demand. This is exactly the "software to allocate resources effectively" that Semiconductor Engineering identifies as the critical need. Companies building agent platforms — whether for customer service, internal operations, or supply chain management — should evaluate their infrastructure provider's cluster orchestration capabilities as a selection criterion, not an afterthought.

For Malaysian companies building custom agents, the practical step is to instrument your agent workflows with monitoring that tracks token usage, latency, and failure rates per step. This data tells you where compute is being wasted and where cluster-level optimisation would have the biggest impact.


Risks and Limitations

The cluster scaling challenges described here are real, but the timeline for solutions is uncertain. Software optimisation is a gradual process — it improves incrementally, not overnight. Businesses that delay AI deployments waiting for "better cluster efficiency" may fall behind competitors who accept current inefficiencies and build institutional learning now.

Power infrastructure is a political and regulatory risk in Malaysia. Data centre development in Johor and Selangor has already raised questions about grid capacity, water usage for cooling, and community impact. Government policy could shift — through stricter environmental permitting, revised power tariffs, or moratoriums on new data centre construction in certain states. Companies making long-term infrastructure bets should scenario-plan for regulatory change.

Finally, the source material focuses on engineering challenges and does not address cybersecurity, data sovereignty, or PDPA compliance implications of running workloads on shared cluster infrastructure. Malaysian businesses must evaluate these independently — cluster efficiency means nothing if your data governance framework cannot legally support the deployment architecture you have chosen.


The Bottom Line

The semiconductor industry's cluster scaling problem is, at its core, a software problem wearing a hardware costume. Chips are fast enough — the question is whether the software layer can orchestrate them efficiently at scale. For Malaysian businesses, this means the real cost of AI over the next three years will be determined less by chip prices and more by how well infrastructure providers — and your own engineering teams — manage power, networking, and memory across compute clusters.

The action to take this quarter: if you are spending more than RM10,000 monthly on AI compute, commission a utilisation audit. Understand what percentage of your paid compute capacity is actually being used productively. That single number will tell you more about your AI cost trajectory than any benchmark or product brochure.


FAQ

What is the difference between scale-up and scale-out in compute clusters? Scale-up means making individual machines more powerful (bigger processors, more memory). Scale-out means connecting more machines together. Both increase total capacity but create different engineering problems around power, networking, and software management.

Why should a Malaysian SME care about cluster engineering? Because it directly affects your cloud AI costs. If providers struggle with cluster efficiency, they pass those costs on to customers. Understanding this dynamic helps you choose providers and plan budgets more accurately.

Is Malaysia's semiconductor industry positioned to benefit from these challenges? Yes, particularly in advanced packaging and interconnect manufacturing in Penang and Kulim. But capturing higher-value software and design work will require targeted investment in skills that Malaysia currently lacks at scale.


Sources / References

  • Semiconductor Engineering, "Scale Up, Out Challenges Amplified For Clusters" — Primary source reporting on compute cluster scaling challenges, power consumption trade-offs, and the need for software-driven resource allocation, networking optimisation, and memory management.

Sources & References

AIBlog summarises and analyses published information. We do not reproduce full source text. Analysis is editorial and not financial or legal advice.

Related articles

Get Malaysia's AI intelligence every morning

Daily digest on Telegram and WhatsApp. Written for Malaysian business readers.

Daily AI intelligence
From RM5/month
Subscribe