Reddit vs Perplexity: The Copyright Fight That Could Reshape AI's Data Supply
Reddit's persistent lawsuit against AI search engine Perplexity signals a tightening squeeze on how AI companies access content — and Malaysian businesses with a web presence should be paying attention.

Reddit is continuing an aggressive copyright lawsuit against Perplexity AI, accusing the AI-powered search company of conspiring with a web scraper to access Reddit's content without permission. The case persists even though Google reportedly lost a similar legal challenge, suggesting Reddit believes its legal footing is stronger or that it is willing to fight on principle. The dispute is part of a escalating global battle over who controls the data that fuels AI models — the platforms that host content, the AI companies that consume it, or nobody at all. For Malaysian businesses, this matters because it signals a future where scraping public web content for AI training or AI-powered search may increasingly require licensing agreements, creating both risks and opportunities for content-rich companies. ---
AI Summary
Reddit is continuing an aggressive copyright lawsuit against Perplexity AI, accusing the AI-powered search company of conspiring with a web scraper to access Reddit's content without permission. The case persists even though Google reportedly lost a similar legal challenge, suggesting Reddit believes its legal footing is stronger or that it is willing to fight on principle. The dispute is part of a escalating global battle over who controls the data that fuels AI models — the platforms that host content, the AI companies that consume it, or nobody at all. For Malaysian businesses, this matters because it signals a future where scraping public web content for AI training or AI-powered search may increasingly require licensing agreements, creating both risks and opportunities for content-rich companies.
Key Takeaways
- Reddit is escalating, not retreating. Despite Google losing a similar legal fight, Reddit is pressing forward with its DMCA lawsuit against Perplexity AI, signalling that content platforms see copyright enforcement as a viable strategy against AI scrapers.
- The weapon is the DMCA. The Digital Millennium Copyright Act — a US copyright law — is being used not just to remove infringing content from search results, but to challenge the fundamental way AI companies build their products by scraping the web.
- "Conspiring with a web scraper" is the core accusation. Reddit alleges that Perplexity worked with third-party scraping infrastructure to access Reddit data, rather than negotiating access through Reddit's official API or licensing terms.
- Google's loss doesn't deter Reddit. The fact that a similar challenge involving Google apparently failed suggests Reddit either has different evidence, different legal arguments, or different strategic goals — possibly including deterrence through legal cost.
- Malaysian content owners should take note. If you operate a content-rich website — e-commerce listings, forums, media, educational content — the legal landscape around who can scrape and reuse your content is shifting, and your terms of service may matter more than you think.
What Happened
Reddit has chosen to keep alive a lawsuit that accuses Perplexity AI of conspiring with a web scraper to obtain Reddit content without authorisation. The lawsuit, which is still working its way through the US legal system, represents one of the most aggressive uses of copyright law against an AI company to date. Reddit's argument centres on the claim that Perplexity — which operates an AI-powered "answer engine" that summarises web content in response to user questions — did not obtain Reddit's content through legitimate channels but instead relied on scraping infrastructure that circumvented Reddit's technical and legal barriers.
The case is notable because it comes after Google reportedly lost a similar legal battle, which many observers might have expected to discourage further copyright-based challenges against AI companies. Reddit's decision to press forward suggests the company believes its circumstances are materially different — whether because of the specific nature of Perplexity's scraping, the evidence of conspiracy with a third-party scraper, or simply Reddit's determination to set a precedent that its content cannot be freely harvested.
The legal mechanism at the heart of the case is the DMCA, or Digital Millennium Copyright Act. This is a United States copyright law passed in 1998 that, among other things, provides a framework for copyright holders to demand the removal of infringing content from websites and search results through what are called "takedown notices." Reddit has been using DMCA takedown notices aggressively to remove content from Google search results that it believes originates from unauthorised scraping. The lawsuit against Perplexity takes this a step further — moving from takedown notices to full litigation, targeting not just the content that appears online but the process by which it was obtained.
At its core, this is a dispute about control. Reddit hosts enormous volumes of user-generated content — discussions, reviews, answers, opinions — that are highly valuable for training AI models and for powering AI search experiences. AI companies like Perplexity want to access that content to improve their products. Reddit wants to control who accesses that content and under what terms, ideally generating revenue through licensing agreements. When Perplexity allegedly bypassed Reddit's API and licensing framework by working with a web scraper, Reddit responded with litigation.
Why It Matters
This case matters because it represents a critical test of whether copyright law can effectively regulate the AI industry's hunger for training data. For the past several years, AI companies have operated largely on the assumption that publicly available web content is fair game for scraping, training, and repurposing. That assumption is now under sustained legal attack from multiple directions — from news organisations suing OpenAI, from authors suing over book training data, and from platforms like Reddit suing over user-generated content.
The Reddit-Perplexity dispute is particularly significant because it targets not just the output of an AI system (what it generates in response to a query) but the input pipeline (how it gets the data in the first place). If Reddit succeeds, it could establish a legal precedent that AI companies must negotiate licences with content platforms before scraping their data, even if that data is technically publicly accessible. This would fundamentally change the economics of AI development — making data acquisition a major cost centre rather than a free resource.
Conversely, if Reddit loses — as Google apparently did in a similar challenge — it would reinforce the current norm where AI companies can scrape public web content with relatively little legal risk. This would embolden more scraping, more AI products built on third-party content, and more tension between content creators and AI platforms.
The broader signal here is that the "data commons" era of the internet — where content was freely accessible and freely scrapable — is ending. What replaces it is still being negotiated through lawsuits like this one. The outcome will determine who profits from the content that feeds AI systems: the original creators and platforms, the AI companies that build products on top of it, or some combination through licensing deals.
For any business that publishes content online, this case is a leading indicator of how the rules of the road are being rewritten. The days of assuming your public website content can be freely consumed by any AI system without consequence — either to you or to the AI company — are numbered.
What This Means for Malaysia
For Malaysian businesses, the Reddit-Perplexity case may seem like a distant American legal drama, but its implications reach Malaysian shores through several channels.
First, Malaysian companies that use AI tools powered by scraped web data — including AI search engines, content generation tools, and customer service chatbots — should understand that the legal foundation of these tools is unsettled. If you are a Malaysian SME using an AI-powered search or content tool that summarises information from across the web, the content it surfaces may come from sources that did not consent to being used this way. While Malaysian businesses are unlikely to be the primary targets of international copyright lawsuits, using tools built on legally disputed data pipelines introduces reputational and compliance risk, particularly if your business operates internationally or serves multinational clients.
Second, Malaysian content creators, media companies, e-commerce platforms, and educational institutions should be thinking proactively about protecting their content. Malaysia's Personal Data Protection Act (PDPA) governs personal data, but copyright protection for content falls under the Copyright Act 1987, which provides legal recourse against unauthorised reproduction. If international precedent shifts toward requiring licences for AI scraping, Malaysian content owners who have clear terms of service, copyright notices, and technical barriers against scraping will be better positioned to enforce their rights — or to negotiate licensing deals with AI companies.
Third, Malaysia's ambitions under the MyDIGITAL framework and MDEC's AI initiatives are closely tied to building a local AI ecosystem. If global data access becomes more restricted and more expensive, Malaysian AI startups and researchers will face the same data acquisition challenges as their international counterparts. This could actually be an opportunity: Malaysian AI companies that partner directly with local content owners — media companies, universities, government agencies — to licence data could build a competitive advantage over international players who are locked out.
The Klang Valley and Penang tech corridors, which host significant numbers of digital economy companies, should be watching this case as part of a broader trend toward data localisation and data licensing. Companies in these hubs that treat data as a strategic asset — both protecting their own and securing proper access to others' — will be better positioned as the regulatory environment tightens globally.
How Your Business Can Use This
If you operate a content-rich website — and most Malaysian businesses do, whether through product listings, blog posts, customer reviews, or support documentation — here is what you should do this quarter.
Audit your content exposure. Identify what publicly accessible content you have on your website that could be valuable to AI systems. Product catalogues, expert articles, customer reviews, community forums, and proprietary databases are all prime targets for scraping. Understand that this content has value not just to your customers but to AI companies building search and generation tools.
Review your terms of service. Ensure your website's terms explicitly address AI scraping and content reuse. Many Malaysian business websites have generic terms of service that do not clearly prohibit automated scraping or AI training use. While a terms-of-service clause alone will not stop a determined scraper, it provides a legal foundation for enforcement and for future licensing negotiations.
Implement technical barriers. Use robots.txt files — a standard web protocol that tells automated crawlers which parts of your site they may access — to explicitly disallow known AI scrapers if you do not want your content used for AI training. Consider rate limiting and bot detection tools if your content is particularly valuable. These are not perfect solutions, but they signal your intent and provide evidence of your position if a dispute arises.
Evaluate your AI tool usage. If your business uses AI-powered tools — search engines, content generators, research assistants — understand where those tools get their data. Ask your vendors about their data sourcing practices. This is particularly important for Malaysian companies serving international clients who may have stricter expectations about data provenance and copyright compliance.
Explore licensing opportunities. If you have proprietary content that AI companies might want — industry data, expert knowledge bases, local market insights — consider whether licensing that content could become a revenue stream. The trend is moving toward paid data access, and early movers who establish licensing frameworks will have an advantage.
The Agentic AI Angle
This legal battle directly affects the future of agentic AI — autonomous AI agents that plan, reason, and take actions across multiple steps to complete business tasks. Here is why.
AI agents rely on real-time access to web content to function. An AI agent tasked with researching competitors, monitoring market trends, summarising industry news, or answering customer questions needs to fetch and process live web content. If the legal landscape makes it riskier for AI companies to access that content, the agents built on those platforms become less capable, less reliable, or more expensive to operate.
For Malaysian businesses building or deploying AI agents, this means you should think carefully about data sourcing. Rather than relying solely on agents that scrape the open web, consider building agents that operate within a controlled data environment — your own databases, licensed data feeds, or partner APIs. For example, a customer service agent for a Malaysian e-commerce company should be trained on and connected to your own product database and knowledge base, not reliant on scraping external sources that could become legally inaccessible.
The companies that will succeed with agentic AI are those that control their data pipelines. If your AI agent's effectiveness depends on content that someone else owns and could legally restrict access to, your agent is built on uncertain ground. The Reddit-Perplexity case is a reminder that data access is not guaranteed — it must be secured through ownership, partnership, or licence.
Risks and Limitations
The outcome of the Reddit-Perplexity case is uncertain. Google reportedly lost a similar challenge, which suggests that courts may be reluctant to extend copyright law to cover AI scraping of publicly available content. Reddit may lose, in which case the current norms around web scraping for AI purposes would be reinforced.
Even if Reddit wins, enforcement will be complex. The global nature of the internet means that AI companies can operate from jurisdictions where US copyright law has limited reach. Malaysian businesses should not assume that any single lawsuit will dramatically change the practical landscape overnight.
Additionally, technical countermeasures against scraping — robots.txt, bot detection, rate limiting — are imperfect. Determined scrapers can bypass these measures, and enforcement requires resources that most SMEs do not have. The legal and technical tools available to protect content are improving but remain incomplete.
The Bottom Line
Reddit's decision to press forward with its copyright lawsuit against Perplexity AI, despite Google's loss in a similar case, is a clear signal that the battle over AI data access is intensifying, not fading. The era of freely scraping public web content to build AI products is being challenged, and the outcome will reshape how AI companies operate globally.
For Malaysian businesses, the action items are concrete: audit your content assets, update your terms of service, implement technical barriers where appropriate, evaluate your AI vendors' data practices, and explore whether your proprietary data could become a licensing opportunity. The businesses that treat data as a strategic asset — both protecting their own and securing proper access to others' — will be better positioned regardless of how this lawsuit ends.
FAQ
Does this lawsuit affect Malaysian businesses directly? Not immediately through legal obligation, but Malaysian companies using AI tools built on scraped data face growing reputational and compliance risks, particularly if they serve international clients.
Can I stop AI companies from scraping my Malaysian website? You can implement technical barriers like robots.txt files and bot detection, and your Copyright Act 1987 protections apply, but enforcement against international scrapers is difficult without significant resources.
Should my Malaysian business licence its content to AI companies? If you have proprietary, high-value content — industry data, expert knowledge, local market insights — licensing is worth exploring as the trend moves toward paid data access and AI companies seek legitimate content sources.
Sources / References
- Ars Technica — "Reddit keeps its strange DMCA fight over Google search results alive" (arstechnica.com): Primary source for all factual details about the ongoing Reddit vs Perplexity AI lawsuit, the DMCA legal mechanism, the accusation of conspiracy with a web scraper, and the context of Google's similar legal loss.
Sources & References
AIBlog summarises and analyses published information. We do not reproduce full source text. Analysis is editorial and not financial or legal advice.


