Scaling AI agents with trustworthy data

MIT Technology Review reports that while business leaders are convinced agentic AI — autonomous AI systems that plan and act across multiple steps — is ready for real deployment, most organisations are discovering that their data infrastructure cannot support it. The article identifies inadequate data foundations and weak infrastructure as the primary barriers to achieving return on investment from AI agents. For Malaysian companies and government agencies investing in AI under the MyDIGITAL framework, this signals that data quality, governance, and accessibility — not model selection — will determine whether AI agent projects succeed or become expensive experiments. ---
Scaling AI Agents With Trustworthy Data: Why Malaysia's AI Push Lives or Dies on Data Foundations
The gap between AI agent ambition and actual ROI isn't the AI model — it's your data. MIT Technology Review reports that organisations adopting agentic AI are hitting a wall: inadequate infrastructure and untrustworthy data. For Malaysian businesses rushing into AI, the message is clear. Fix your data first, or your agent deployments will stall.
AI Summary
MIT Technology Review reports that while business leaders are convinced agentic AI — autonomous AI systems that plan and act across multiple steps — is ready for real deployment, most organisations are discovering that their data infrastructure cannot support it. The article identifies inadequate data foundations and weak infrastructure as the primary barriers to achieving return on investment from AI agents. For Malaysian companies and government agencies investing in AI under the MyDIGITAL framework, this signals that data quality, governance, and accessibility — not model selection — will determine whether AI agent projects succeed or become expensive experiments.
Key Takeaways
- Executive confidence in agentic AI is high, but execution confidence is not. Leaders believe in the technology. Their data foundations do not support that belief.
- The ROI bottleneck is infrastructure and data, not AI model capability. Organisations are discovering that even the best LLMs produce poor results when fed poor data.
- "Trustworthy data" means more than accuracy. It encompasses completeness, timeliness, consistency, accessibility, governance, and provenance — knowing where data came from and whether it can be relied upon.
- Scaling agents magnifies data problems exponentially. A chatbot answering one question has a narrow data footprint. An agent executing a ten-step procurement workflow touches dozens of data sources, and a single weak link breaks the chain.
- Malaysian businesses investing in AI without auditing their data readiness are building on sand. This is particularly relevant for SMEs with fragmented, spreadsheet-based systems and for enterprises with siloed legacy databases.
What Happened
MIT Technology Review published an analysis on August 12, 2026, focused on the central challenge organisations face when scaling AI agents: the quality and trustworthiness of the data those agents depend on.
According to the report, business and technology leaders no longer need convincing that agentic AI has arrived. organisations across sectors are rapidly adopting AI agents — systems that go beyond answering questions to actually taking actions, making decisions, and executing multi-step workflows. Executives broadly accept that this technology can transform how work gets done.
However, the report identifies a critical gap between expectation and outcome. Many organisations are finding that the return on investment they anticipated from AI is not materialising. The reason is not that the AI models are insufficiently capable. It is that the foundations beneath them — data infrastructure, data quality, data governance, and data accessibility — are inadequate for the demands of autonomous agents.
The article's core argument is that trustworthy data is the prerequisite for scaling AI agents successfully. Without it, organisations deploy agents that hallucinate, make incorrect decisions, access outdated information, or fail to execute workflows correctly. The technology itself is ready. The data plumbing is not.
This finding aligns with a pattern seen across enterprise AI adoption over the past several years. Each wave of AI capability — from predictive analytics to large language models to autonomous agents — has revealed that the limiting factor is rarely the algorithm. It is almost always the state of the organisation's data.
Why It Matters
The MIT Technology Review report matters because it reframes the AI conversation from a technology problem to a data problem. This distinction has significant strategic implications.
For the past two years, many organisations have focused their AI investments on model selection — choosing between GPT-4, Claude, Gemini, Llama, or other LLMs. The assumption has been that the right model, paired with a good prompt, would deliver results. The report suggests this assumption is flawed. When an AI agent is tasked with, say, reviewing supplier contracts, cross-referencing pricing against a procurement database, checking delivery timelines, and drafting a recommendation, the model's reasoning ability is only one component. If the procurement database is incomplete, if the contract repository is not searchable, if pricing data is six months out of date, the agent will produce a confident, well-reasoned answer based on bad information. That is worse than no answer at all, because it looks authoritative.
This is the "trustworthy data" problem. Trustworthy data is data that is accurate, yes, but also complete, current, consistent across systems, properly governed with clear access controls, and traceable to its source. For an AI agent operating autonomously — making decisions without a human reviewing every step — the trustworthiness of its data inputs is the difference between a useful tool and a liability.
The scaling dimension compounds this. A single AI agent performing a narrow task may function adequately with imperfect data. But organisations are moving toward multi-agent systems where several agents collaborate on complex workflows. One agent gathers data, another analyses it, a third makes recommendations, a fourth executes transactions. Each agent depends on the output of the others. If the foundational data is unreliable, errors propagate through the chain. A two percent error rate in source data can become a twenty percent error rate by the time it reaches the final decision point.
For organisations pouring budget into AI initiatives, this report is a warning. The largest investments may need to be in data infrastructure and governance, not in AI models or agent platforms. Companies that skip this step will burn money on pilot projects that never scale.
What This Means for Malaysia
Malaysia's AI push makes this report directly relevant. Under the MyDIGITAL initiative and related Budget allocations, the government has encouraged AI adoption across both public and private sectors. MDEC has promoted AI literacy programmes. Agencies are exploring AI for public service delivery. Malaysian enterprises — from banks in Kuala Lumpur to manufacturers in Penang's semiconductor corridor — are piloting AI projects.
The MIT Technology Review finding suggests that many of these initiatives will hit the same wall: data.
Malaysian SMEs face a particularly acute version of this problem. Many small and medium enterprises operate on a patchwork of spreadsheets, legacy accounting software, paper-based records, and informal knowledge stored in employees' heads. Data is siloed by department, inconsistently formatted, rarely updated, and almost never governed by formal data quality standards. When an SME owner decides to "implement AI," the model they choose is the easy part. Preparing their data so that an AI agent can reliably use it is where the real work — and cost — lies.
Larger Malaysian enterprises face a different flavour of the same problem. Banks, telcos, and government-linked companies often have decades of accumulated data spread across legacy systems built in different eras. Customer data may live in three different databases with conflicting formats. Product catalogues may be maintained manually. Internal documents may exist as scanned PDFs that are not machine-readable. These are not AI problems. They are data infrastructure problems that predate AI but become critical when AI agents need to access and reason over that data.
The regulatory dimension adds another layer. Malaysia's Personal Data Protection Act (PDPA) sets rules for how personal data is collected, used, and disclosed. AI agents that access customer data must do so within PDPA compliance boundaries. This means data governance is not just an operational concern — it is a legal one. organisations deploying agents need to know exactly what data the agent can access, how it uses that data, and whether that usage is compliant. Trustworthy data, in this context, includes the dimension of lawful, compliant access.
For Malaysia's semiconductor industry in Penang and Klang Valley, the stakes are specific. Semiconductor manufacturing generates enormous volumes of process data — yield rates, defect patterns, equipment telemetry, supply chain logistics. AI agents could optimise production lines, predict equipment failures, or automate supplier coordination. But only if that data is captured cleanly, stored accessibly, and governed properly.
How Your Business Can Use This
The practical implication of this report is straightforward: before you invest in an AI agent platform, invest in understanding your data. Here is a concrete approach.
Step 1: Conduct a data readiness audit. List the key business processes where you want AI agents to operate. For each process, identify every data source the agent would need to access — databases, documents, spreadsheets, APIs, external systems. For each source, assess: Is the data accurate? Is it complete? Is it current? Is it in a format a machine can read? Is access controlled and PDPA-compliant? Score each source. You will likely find that thirty to fifty percent of your data sources are not agent-ready.
Step 2: Prioritise one workflow, not ten. Pick a single business process where the data is in the best condition and the potential ROI is clear. Examples for Malaysian SMEs: automated invoice processing, customer service response drafting, inventory reorder alerts. Build your first agent deployment on that one workflow. Prove the concept. Learn what data gaps remain.
Step 3: Fix data at the source. Rather than building complex data cleaning pipelines, address root causes. If sales staff are entering customer data inconsistently, fix the input forms. If product information lives in someone's Excel file, migrate it to a structured database. If documents are scanned images, implement OCR (optical character recognition) to make them machine-readable.
Step 4: Establish basic data governance. Assign ownership for each major data domain — customer data, financial data, operational data. Document what data exists, where it lives, who can access it, and how often it is updated. This does not require enterprise-grade software. A well-maintained spreadsheet cataloguing your data assets is a starting point.
The Agentic AI Angle
The MIT Technology Review report underscores a specific challenge for agentic AI that does not apply to simpler AI tools. A chatbot that answers a question has a single interaction with data: it retrieves information and formulates a response. If the data is wrong, the user can often spot it and ask again.
An AI agent operates differently. It plans a sequence of actions, executes them, observes results, and adjusts. Consider an agent tasked with processing a supplier payment. It must read the invoice, check it against the purchase order, verify the goods were received, confirm the payment terms, initiate the transaction, and log the result. Each step depends on data from the previous one. If the invoice data is misread because the document quality is poor, the agent may approve an incorrect amount. If the goods receipt record is missing from the system, the agent may hold a legitimate payment.
This means trustworthy data is not a one-time setup for agents. It requires continuous data quality monitoring. Malaysian businesses deploying agents should build feedback mechanisms — logs of agent decisions, flagging of anomalies, human review of edge cases — that surface data quality issues as they arise. The agent itself becomes a diagnostic tool, revealing where your data is weakest by failing at specific steps in its workflow.
Multi-agent systems amplify this further. If one agent is responsible for data gathering and another for decision-making, the decision agent inherits all the data quality problems of the gathering agent, with no ability to assess them. Provenance — knowing where each piece of data came from and how reliable that source is — becomes essential. Agents need confidence scores attached to data, not just the data itself.
Risks and Limitations
The primary risk is that organisations treat data quality as a one-time project rather than an ongoing discipline. Data degrades. Customer information changes. Product specifications update. Pricing shifts. An agent deployed against clean data today will degrade in performance if that data is not maintained.
A second risk is over-indexing on data perfection. Some organisations will use "our data isn't ready" as a reason to delay all AI investment indefinitely. The report's message is not that you need perfect data before starting. It is that you need to understand your data's limitations, choose agent deployments that match your data maturity, and invest progressively in improvement.
PDPA compliance remains a live concern. AI agents that access personal data must do so within lawful purposes. Malaysian organisations should involve their data protection officer or legal counsel in agent deployment decisions, particularly for agents that handle customer or employee data.
The Bottom Line
The single most important takeaway from the MIT Technology Review analysis is this: your AI agent is only as trustworthy as the data it operates on. The model is not the bottleneck. Your data is.
For Malaysian businesses, the action this quarter is simple. Before you sign a contract with an AI vendor or deploy an agent platform, conduct an honest assessment of your data foundations. Identify your strongest data domain. Build your first agent deployment there. Use the experience to surface data gaps and fix them incrementally. The companies that win with AI agents over the next two years will not be the ones with the best models. They will be the ones with the cleanest, most trustworthy data.
FAQ
What does "trustworthy data" actually mean in practice? It means data that is accurate, complete, current, consistent across systems, properly governed, and traceable to its source — so an AI agent can rely on it for autonomous decisions.
We are a Malaysian SME with messy data. Should we wait before adopting AI agents? No, but start narrow. Pick one workflow where your data is in the best condition, deploy a small agent there, and use the experience to drive broader data improvements.
How does PDPA affect AI agent deployments in Malaysia? AI agents that access personal data must comply with PDPA's principles on collection, use, and disclosure. Involve your data protection officer in deployment planning, especially for agents handling customer or employee information.
Sources / References
- MIT Technology Review — "Scaling AI agents with trustworthy data" (August 12, 2026). Provided the core finding that inadequate infrastructure and data foundations are the primary barriers to AI agent ROI, and that trustworthy data is the prerequisite for scaling agent deployments. Link
Sources & References
AIBlog summarises and analyses published information. We do not reproduce full source text. Analysis is editorial and not financial or legal advice.


