AIAIBlog.com.my
International AI News1 August 2026 · 11 min read

Reddit keeps its strange DMCA fight over Google search results alive

Reddit keeps its strange DMCA fight over Google search results alive
AIAI Summary

Reddit is pressing forward with an unusual Digital Millennium Copyright Act (DMCA) lawsuit against Perplexity AI, accusing the AI answer engine of conspiring with a web scraper to access and use Reddit content without authorisation. The case persists despite Google having lost a similar legal battle, raising questions about whether DMCA — originally designed for traditional copyright takedowns — can be effectively weaponised against AI-driven content scraping. For Malaysian businesses, this case is an early signal of how content ownership, data licensing, and AI training practices are being renegotiated globally, with direct implications for any company that publishes, scrapes, or relies on web-based content for AI products. ---

Reddit Keeps Its Strange DMCA Fight Over Google Search Results Alive

How a content platform's legal war on AI scrapers signals a new era of data ownership that every Malaysian digital business needs to understand — and prepare for.


AI Summary

Reddit is pressing forward with an unusual Digital Millennium Copyright Act (DMCA) lawsuit against Perplexity AI, accusing the AI answer engine of conspiring with a web scraper to access and use Reddit content without authorisation. The case persists despite Google having lost a similar legal battle, raising questions about whether DMCA — originally designed for traditional copyright takedowns — can be effectively weaponised against AI-driven content scraping. For Malaysian businesses, this case is an early signal of how content ownership, data licensing, and AI training practices are being renegotiated globally, with direct implications for any company that publishes, scrapes, or relies on web-based content for AI products.


Key Takeaways

  • Reddit is using DMCA copyright law as a tool against AI scraping, not traditional piracy — a novel legal strategy that could reshape how content platforms protect their data from AI companies.
  • The case survives despite Google's prior loss in a similar matter, suggesting Reddit believes it has a distinct legal theory or factual basis, particularly around the alleged conspiracy with a third-party web scraper.
  • Perplexity AI is accused of working with a web scraper to bypass restrictions, which, if proven, would represent a deliberate circumvention of content access controls rather than passive data collection.
  • This is part of a much larger pattern: content platforms worldwide are increasingly treating their user-generated data as proprietary, licensable assets — not freely available web content.
  • Malaysian businesses that rely on web-scraping, content aggregation, or AI training need to audit their data sourcing practices now, because the legal landscape is shifting rapidly and exposure cuts both ways.

What Happened

Reddit is continuing to pursue a lawsuit against Perplexity AI, the AI-powered answer engine that synthesises search results into conversational responses. According to Ars Technica's reporting, Reddit accuses Perplexity of conspiring with a web scraper — a third-party tool or service that systematically extracts content from websites — to access Reddit's user-generated content without permission or compensation.

The case is described as "strange" and "weird" by the publication, reflecting the unusual legal theory at its core. Reddit is invoking the DMCA, the United States' primary copyright enforcement mechanism, which was originally designed to handle straightforward takedown requests for pirated movies, music, and software. Applying it to AI scraping — where content is ingested, processed, and re-expressed in transformed ways — represents a creative expansion of the law's intended scope.

What makes the case particularly notable is that it continues despite Google having already lost a similar legal challenge. This means a court has previously ruled that DMCA does not cleanly apply to the kind of search-related content use that Google and similar services engage in. Reddit's decision to press forward anyway suggests the company believes its situation is factually distinct — specifically, the allegation that Perplexity actively conspired with a scraper to circumvent Reddit's content access barriers may provide a different legal footing than a straightforward indexing or caching argument.

The lawsuit is ongoing as of July 2026, and its outcome could establish important precedent for how content platforms protect their data assets in the AI era.


Why It Matters

This case matters because it sits at the intersection of three massive shifts happening simultaneously in the digital economy — and each one directly affects how businesses operate online.

First, the shift from open web to walled data gardens. For over two decades, the internet operated on a largely implicit social contract: content was published openly, search engines indexed it, and traffic flowed back to publishers. That system is breaking down. AI answer engines like Perplexity don't send traffic back — they consume the content and present it directly to the user. This eliminates the reciprocal value exchange that sustained the open web. Reddit's lawsuit is one battle in a war where every major content platform — from news publishers to social networks to specialised databases — is now asking: if our data is the raw material for AI products worth billions, why are we giving it away?

Second, the legal framework is lagging badly behind the technology. The DMCA was passed in 1998 — a world of Napster and peer-to-peer file sharing. Applying it to large language model training and AI answer synthesis is like using telecommunications law from the 1930s to regulate social media. The fact that Google already lost a similar case suggests courts are sceptical of stretching DMCA this far. But Reddit's persistence, particularly its conspiracy allegation around deliberate scraper coordination, signals that plaintiffs are still searching for the right legal theory that will stick. When one does — and it eventually will, whether through DMCA, breach of contract, trespass to chattels, or new legislation — the entire data supply chain for AI will be affected.

Third, this signals that data licensing is becoming a major business category. Reddit has already signed a lucrative licensing deal with Google reported to be worth approximately USD 60 million annually. The lawsuit against Perplexity is, in part, about enforcing that licensing model: if you want Reddit data for your AI, you pay for it. This is the emerging norm, and it extends far beyond Reddit. Every platform with proprietary data — customer reviews, forum discussions, product catalogues, professional content — is now thinking about monetising access. The era of freely scraping the web for AI training data is closing.


What This Means for Malaysia

For Malaysian businesses, this development has both direct and indirect implications that warrant attention.

For Malaysian content publishers and platform operators, the Reddit case validates a strategy of treating proprietary data as a monetisable asset. If you operate a platform with significant user-generated content — a property listing site, a review platform, a specialised forum, or an e-commerce marketplace with rich product data — you should be thinking about how to protect and potentially license that data. Malaysia's PDPA (Personal Data Protection Act) governs personal data but does not comprehensively address the scraping of non-personal but commercially valuable content. This means Malaysian platforms need to rely on terms of service, technical barriers (like robots.txt and rate limiting), and contractual protections to safeguard their data assets. The Reddit case demonstrates that even these measures may need legal enforcement backing.

For Malaysian AI builders and startups, the warning is clear. If your product relies on scraping web content — whether for training models, powering search features, or building answer engines — you face growing legal exposure. Malaysia's AI ecosystem, supported by initiatives like MyDIGITAL and MDEC's AI roadmap, is producing more local AI companies. These companies need to ensure their data sourcing is legitimate, properly licensed, or appropriately synthetic. A Malaysian startup that scrapes content from platforms without permission could face DMCA-style actions (many platforms are US-based and subject to US law), platform bans, or reputational damage that affects funding prospects.

For Malaysian enterprises and government agencies, the broader trend suggests that data governance frameworks need updating. Government open data initiatives, smart city projects, and public sector AI deployments all need clear policies on what data can be used, how it can be licensed, and what rights attach to AI outputs derived from public or proprietary sources. Malaysia's National AI Roadmap and related policy documents should account for the data ownership tensions that cases like Reddit v. Perplexity are bringing to the surface.


How Your Business Can Use This

Whether you are a content creator, platform operator, or AI consumer, there are concrete steps to take this quarter.

Audit your data supply chain. If your business uses AI tools that ingest or reference external content — whether that is a chatbot trained on web data, an analytics tool that scrapes competitor pricing, or a content aggregation service — document where every data source comes from. Identify which sources are openly licensed, which are used under implied permission (the traditional web model), and which are used without clear authorisation. This audit should cover training data, real-time retrieval sources, and any third-party data feeds.

Review and strengthen your terms of service. If you operate a platform or publish significant original content, ensure your terms explicitly address AI scraping, training use, and content licensing. Standard terms from five years ago likely do not address these use cases. Update your robots.txt file to clearly signal which automated access is permitted. While robots.txt is not legally binding in all jurisdictions, it strengthens your position in any dispute.

Explore data licensing as a revenue stream. If your business has accumulated proprietary data — customer behaviour patterns, industry-specific content, specialist databases — consider whether licensing that data to AI companies represents a viable revenue opportunity. The market for training data is growing, and niche, high-quality datasets command premium prices.


The Agentic AI Angle

The Reddit v. Perplexity case has significant implications for how agentic AI systems — autonomous agents that plan, reason, and take multi-step actions — access and use web content in real-time.

Many agentic AI workflows involve web browsing capabilities: an agent might research competitor pricing, compile market intelligence, or gather customer sentiment from online forums. If the legal environment tightens around web scraping, these agentic workflows face direct disruption. An agent that autonomously browses Reddit to extract product feedback for a Malaysian SME's competitive analysis could be engaging in the exact behaviour that Reddit is suing Perplexity over.

This means businesses deploying agentic AI need to build data compliance into their agent architecture. Rather than allowing agents to freely browse and scrape, responsible implementations should use licensed APIs (like Reddit's official API), purchased data feeds, or content sources with clear usage rights. Agent orchestration platforms should include data provenance tracking — a record of where every piece of information came from and under what terms it was accessed.

For Malaysian businesses building agentic workflows, this is not a distant concern. If your AI agent retrieves information from the web to complete tasks — whether that is market research, lead generation, or customer support — you need to ensure the data access method is defensible. The alternative is building agents that operate only on internal, owned data, which is safer but more limited.


Risks and Limitations

The outcome of Reddit's case is uncertain. Google's loss in a similar matter suggests courts may be reluctant to extend DMCA to AI-related content use, and Reddit's conspiracy theory around scraper coordination remains unproven. If Reddit loses, it could actually strengthen the position of AI companies and weaken content platforms' ability to control their data through copyright law.

Additionally, the case is US-specific. Malaysian businesses operating under Malaysian law face a different regulatory environment where DMCA does not directly apply. However, many platforms and AI services are US-based, so American legal precedents still affect Malaysian users indirectly through platform policies, API terms, and service availability.

Finally, over-indexing on data restriction carries its own risks. Businesses that lock down their content too aggressively may lose visibility in AI-mediated search results, reducing discoverability at a time when AI answer engines are increasingly shaping how users find information.


The Bottom Line

Reddit's persistent legal battle against Perplexity AI is not just a Silicon Valley dispute — it is a defining moment in the renegotiation of who owns, controls, and profits from data in the AI era. The old model of the open web is being replaced by a licensed data economy, and the legal rules are being written in real-time through cases exactly like this one.

Malaysian businesses should take two actions this quarter: audit your data inputs to ensure they are defensibly sourced, and review your data outputs to determine whether your proprietary content has untapped licensing value. The businesses that understand this shift early will be better positioned whether they are building AI products, protecting their content, or both.


FAQ

Does this affect Malaysian businesses given that DMCA is a US law? Yes, indirectly. Most major AI platforms and content networks are US-based, so their compliance with US legal precedents affects what services, APIs, and data access Malaysian businesses can use.

If my Malaysian company scrapes public web data for AI training, are we at risk? Potentially, yes. Even if Malaysian law does not prohibit it, you may violate the terms of service of US-based platforms, face API access revocation, or encounter legal action if you operate in or serve US markets.

Should Malaysian content publishers block AI scrapers from their sites? It depends on your business model. Blocking protects your data but may reduce visibility in AI-powered search results. A balanced approach is to allow indexing for discoverability while restricting bulk scraping through technical measures and clear terms of service.


Sources / References

Sources & References

AIBlog summarises and analyses published information. We do not reproduce full source text. Analysis is editorial and not financial or legal advice.

Related articles

Get Malaysia's AI intelligence every morning

Daily digest on Telegram and WhatsApp. Written for Malaysian business readers.

Daily AI intelligence
From RM5/month
Subscribe