AIAIBlog.com.my
International AI News10 August 2026 · 13 min read

These startups are chasing the next big thing in LLMs

These startups are chasing the next big thing in LLMs
AIAI Summary

MIT Technology Review reports that a crop of startups is actively pursuing the next major architecture breakthrough in large language models (LLMs), looking beyond the Transformer architecture that has dominated AI since Google's landmark 2017 paper "Attention Is All You Need." The piece, part of MIT Technology Review's "What's Next" series, signals that the foundational technology behind ChatGPT, Claude, Gemini, and nearly every major LLM in production today may be approaching its practical limits — and venture-backed companies are racing to build something better. For Malaysian businesses investing in AI tools and infrastructure, this matters because the architecture shift will eventually change what models cost, how fast they run, and what tasks they can handle reliably. The transition will not happen overnight, but decision-makers who track this now will avoid over-committing to today's paradigm. ---

These Startups Are Chasing the Next Big Thing in LLMs

A new wave of startups is building what comes after the Transformer — and Malaysian businesses that understand the shift early will have a cost and capability advantage.


AI Summary

MIT Technology Review reports that a crop of startups is actively pursuing the next major architecture breakthrough in large language models (LLMs), looking beyond the Transformer architecture that has dominated AI since Google's landmark 2017 paper "Attention Is All You Need." The piece, part of MIT Technology Review's "What's Next" series, signals that the foundational technology behind ChatGPT, Claude, Gemini, and nearly every major LLM in production today may be approaching its practical limits — and venture-backed companies are racing to build something better. For Malaysian businesses investing in AI tools and infrastructure, this matters because the architecture shift will eventually change what models cost, how fast they run, and what tasks they can handle reliably. The transition will not happen overnight, but decision-makers who track this now will avoid over-committing to today's paradigm.


Key Takeaways

  • The Transformer architecture has been the backbone of every major LLM since 2017, but startups are now building alternatives — suggesting the industry sees diminishing returns from simply scaling up current models.
  • "Attention Is All You Need" (the 2017 Google paper) introduced the Transformer, and it shaped the entire current AI boom — but what replaced prior architectures can itself be replaced, and that replacement is now being funded.
  • The startups pursuing next-generation architectures are venture-backed, meaning investors with deep technical conviction are placing real capital on the belief that post-Transformer models are commercially viable within years, not decades.
  • For Malaysian businesses, the practical signal is this: do not lock your AI strategy permanently to one model provider or one architecture — design workflows that are model-agnostic and can swap underlying engines.
  • This shift will eventually lower inference costs and enable on-device or locally deployable AI, which matters for Malaysian SMEs concerned about data sovereignty, PDPA compliance, and API dependency on foreign providers.

What Happened

MIT Technology Review, through its "What's Next" series — which surveys industries, trends, and technologies to identify developments on the horizon — reported that a group of startups is actively building what they believe will be the successor to the Transformer architecture in large language models.

To understand why this is significant, a bit of context. In the summer of 2017, researchers at Google published a paper titled "Attention Is All You Need." That paper described a new neural network architecture called the Transformer. Before the Transformer, AI language models relied on older architectures — recurrent neural networks (RNNs) and long short-term memory networks (LSTMs) — which processed text sequentially, one word at a time. This made them slow and limited in how much context they could remember. The Transformer solved this by processing all words in a sentence simultaneously and using an "attention mechanism" to weigh the importance of different words relative to each other, regardless of their position.

That single architectural change unlocked everything that followed. GPT models, BERT, ChatGPT, Claude, Gemini, Llama — every major LLM in use today traces its lineage back to that 2017 paper. The Transformer became the default, the way the internal combustion engine became the default for cars. It worked so well that for nearly eight years, the entire field refined it rather than replaced it. Companies got better results mainly by making Transformer models bigger — more parameters, more training data, more computing power.

The MIT Technology Review piece signals that this scaling approach is running into practical walls. Training the largest Transformer models now costs hundreds of millions of dollars in compute. Inference — the cost of running a trained model to generate responses — remains expensive at scale. Context windows (how much text a model can process at once) have grown but still hit ceilings. And some researchers believe the architecture itself has inherent limitations that more data and more compute cannot fully overcome.

Enter the startups. These companies, backed by venture funding, are exploring alternative architectures — new ways of structuring neural networks that could process language more efficiently, handle longer context, require less training data, or run on smaller hardware. The MIT Technology Review article frames this as the next frontier: not a better version of what exists, but a fundamentally different approach.


Why It Matters

This development matters because architecture shifts in computing are rare and consequential. When they happen, they redistribute power across the entire technology stack.

Consider a historical parallel. In the early 2000s, Nokia and BlackBerry dominated mobile phones because they were excellent at the prevailing architecture: hardware-keyboard devices with phone-first design. When the touchscreen + app ecosystem architecture arrived (led by Apple's iPhone in 2007), it did not merely improve the phone — it redefined what a phone was. Companies that bet too hard on the old architecture lost everything within five years.

A similar dynamic could play out in AI. If a startup successfully develops a post-Transformer architecture that is, say, ten times more efficient at inference, the economics of every AI application change. Models that currently require cloud API calls to OpenAI or Anthropic could run locally on a laptop or a phone. SMEs that cannot afford USD $0.01 per 1,000 tokens for API usage might get equivalent capability for a fraction of the cost — or free, via open-source releases.

This also matters for the competitive dynamics among AI providers. Today, OpenAI, Google, Anthropic, and Meta hold enormous influence because they have the capital and engineering talent to train massive Transformer models. A new architecture that requires less compute to train could lower the barrier to entry, allowing smaller players — including those in Southeast Asia — to build competitive models. The moat of "we have the most GPUs" shrinks if the next architecture does not need as many.

The MIT Technology Review framing as a "What's Next" piece is itself a signal. This series looks for developments that are past the research-paper stage and entering commercial reality. The fact that startups, not just university labs, are pursuing these architectures means the technology is far enough along that investors are willing to fund product development. That narrows the timeline from "interesting theory" to "something you might actually deploy" to roughly the 2026–2028 window, in my assessment — though readers should treat that as analytical inference, not a date from the source.


What This Means for Malaysia

For Malaysian businesses, this development has several layers of relevance.

First, cost. Malaysian SMEs adopting AI today typically do so through API subscriptions to foreign providers — OpenAI, Anthropic, Google. These costs are denominated in US dollars and are subject to exchange rate exposure. A Malaysian SME paying USD $200 per month for API access is paying close to RM950 at current rates. If post-Transformer architectures dramatically reduce inference costs, or enable on-device models, that cost burden drops substantially. This would make AI adoption viable for a much larger segment of Malaysian small businesses — the 97% of the economy classified as SMEs.

Second, data sovereignty. Under Malaysia's Personal Data Protection Act (PDPA) and the evolving regulatory framework around AI governance, businesses handling customer data face restrictions on cross-border data transfer. When you send a prompt to an overseas API, you are sending data outside Malaysian jurisdiction. If next-generation architectures enable capable models to run locally — on company servers, on devices — Malaysian businesses in regulated sectors like banking, healthcare, and government can deploy AI without that compliance headache. This aligns with national objectives under MyDIGITAL and the National AI Roadmap, which emphasise building domestic AI capability.

Third, Malaysia's semiconductor industry. Penang is a global semiconductor manufacturing hub, home to assembly and testing operations for Intel, AMD, Infineon, and others. If the next generation of AI requires different types of chips — optimised for new architectures rather than Transformer-specific matrix multiplication — Malaysia's position in the semiconductor supply chain becomes even more strategically important. The government's push to attract advanced packaging and chip design investment, supported by MDEC and the Ministry of Investment, Trade and Industry (MITI), could benefit if Malaysia becomes a production base for the hardware that powers post-Transformer AI.

Fourth, talent. Malaysian AI builders — developers, data scientists, ML engineers at companies like Carsome, Fave, or within GLCs like Petronas and Maybank — should be aware that the skills they are building today (fine-tuning Transformers, prompt engineering for GPT-class models) may evolve. Teams that understand the broader landscape of neural network architectures, not just one dominant paradigm, will be better positioned as the field shifts.


How Your Business Can Use This

For most Malaysian businesses, the practical response to this news is not to wait for post-Transformer models, but to build AI workflows that are architecturally flexible. Here is what that looks like.

Abstract your model layer. If your company has built customer service bots, document processing pipelines, or analytics tools directly on top of one provider's API — say, OpenAI's GPT-4 — you have created a dependency. Instead, use an abstraction layer or orchestration framework (such as LangChain, LiteLLM, or similar middleware) that lets you swap the underlying model without rewriting your application logic. This means when a more efficient architecture becomes available, you switch to it by changing one configuration, not rebuilding your system.

Pilot open-source models alongside proprietary ones. Models like Meta's Llama series and Mistral's models can be self-hosted on cloud infrastructure (AWS Singapore, Azure Southeast Asia, or local providers like TM One and Maxwell). Running a pilot with an open-source model — even if it is less capable than GPT-4 today — builds the internal muscle for model evaluation, deployment, and monitoring. When post-Transformer open-source models arrive, your team will already have the infrastructure and process to adopt them quickly.

Budget for declining AI costs over 2–3 years. If you are negotiating vendor contracts or building business cases for AI investment, model a scenario where inference costs decline 50–80% over the next 24 months due to architectural improvements. This is speculative but directionally consistent with the trend. Do not over-commit to long-term, high-cost API contracts at today's prices.


The Agentic AI Angle

The shift toward post-Transformer architectures has direct implications for agentic AI — autonomous systems that plan, reason, and execute multi-step tasks without human intervention at each step.

Today's agentic AI systems are built on Transformer-based LLMs, which creates a specific bottleneck: agents that need to process large amounts of context (reading dozens of documents, maintaining conversation memory across a long workflow, reasoning over complex data) consume enormous numbers of tokens. Each token costs money and adds latency. A Malaysian logistics company running an AI agent that reads shipping manifests, checks customs regulations, routes cargo, and communicates with port authorities might burn thousands of tokens per task. At current API prices, this adds up fast — and it limits how many agents you can deploy simultaneously.

A more efficient architecture changes the unit economics. If a post-Transformer model processes context at one-tenth the cost, you can run more agents, give each agent more context, and deploy them for tasks that are not economically viable today. For example, a Malaysian SME in e-commerce could afford to run a dedicated AI agent for every single customer — one that tracks that customer's purchase history, preferences, and complaints in full context, and acts autonomously to resolve issues, recommend products, or flag churn risk. Today, the compute cost makes this impractical for all but the largest companies.

Additionally, some of the architectures being explored by startups focus specifically on reasoning capability — the ability to plan, backtrack, and correct course. This is the core limitation of today's agentic systems. Current LLMs often fail at multi-step reasoning because they generate responses token by token without genuine planning. A new architecture designed for reasoning from the ground up could make agents dramatically more reliable, which is the single biggest barrier to deploying them in production today.


Risks and Limitations

Several cautions are warranted. First, the source material — MIT Technology Review's overview — reports that startups are pursuing these architectures. It does not claim any have succeeded commercially. The history of AI is full of promising architectures that produced interesting research papers but never achieved production-grade performance. The Transformer won in 2017 because it was genuinely better, not because it was well-marketed. Any successor will face the same bar.

Second, there is an enormous incumbent advantage. OpenAI, Google, Anthropic, and Meta have invested billions in Transformer-specific infrastructure — training pipelines, fine-tuning tooling, optimised hardware, developer ecosystems. Even if a better architecture exists, the switching cost for the industry is high. Adoption could be slow.

Third, for Malaysian businesses, the risk is not in the technology itself but in strategic missteps: over-investing in Transformer-specific tooling that becomes obsolete, or conversely, under-investing in AI today because you are waiting for something better. The right approach is to adopt now with flexibility built in.


The Bottom Line

The Transformer has had an extraordinary eight-year run, but the smartest investors and researchers in AI are now funding what comes next. Malaysian businesses should treat this as a planning signal, not an immediate action item. Continue adopting current-generation AI tools — they are powerful and commercially proven. But design your systems so that the model underneath is a replaceable component, not a permanent foundation. The companies that build flexibility into their AI stack today will be the ones who benefit first when the next architecture arrives — whenever that is.


FAQ

Should Malaysian SMEs stop adopting current AI tools like ChatGPT and wait for the next generation? No. Current Transformer-based models are production-ready and deliver real value today. Build with abstraction layers so you can adopt better models as they arrive, but do not delay deployment.

How will post-Transformer models affect AI costs for Malaysian businesses? If new architectures are more efficient at inference — which is the goal of many startups in this space — API costs and compute costs should decline meaningfully, making advanced AI accessible to more SMEs. This is directional expectation, not guaranteed.

Does this affect Malaysia's semiconductor industry in Penang? Potentially yes. New AI architectures may require different chip designs, and Malaysia's position in semiconductor assembly, testing, and increasingly design makes it relevant to any shift in what chips the AI industry needs.


Sources / References

  • MIT Technology Review — "These startups are chasing the next big thing in LLMs" (technologyreview.com, dated August 2026): Primary source for this article. Part of the "What's Next" series examining future trends across industries. Provided the core facts about startups pursuing post-Transformer LLM architectures and the historical reference to Google's 2017 "Attention Is All You Need" paper. Note: the source summary available was truncated, so this analysis draws analytical inferences from the framing and context provided. Readers should consult the full MIT Technology Review article for complete details on specific startups and architectures discussed.

Sources & References

AIBlog summarises and analyses published information. We do not reproduce full source text. Analysis is editorial and not financial or legal advice.

Related articles

Get Malaysia's AI intelligence every morning

Daily digest on Telegram and WhatsApp. Written for Malaysian business readers.

Daily AI intelligence
From RM5/month
Subscribe