TLDR: Encyclopedia Britannica, Merriam-Webster, and Japanese news giants have sued Perplexity AI for copyright infringement and content scraping. These lawsuits signal the end of unchecked data scraping for AI training, forcing investors to re-evaluate their due diligence and valuation methods for AI companies. The legal actions challenge the ‘fair use’ defense and highlight the rising costs of data acquisition, shifting the focus to licensing and ethical data strategies.
The artificial intelligence landscape is witnessing a seismic shift, underscored by the recent legal actions taken by Encyclopedia Britannica and Merriam-Webster against Perplexity AI. These prominent reference publishers, alongside Japanese news giants Nikkei and Asahi Shimbun, have initiated lawsuits alleging widespread copyright infringement and unauthorized content scraping. This seemingly tactical legal maneuver is, in fact, the clearest signal yet that the era of unchecked data scraping for AI training is drawing to a close, compelling Investment and Venture Capital Professionals to fundamentally re-evaluate their due diligence, valuation methodologies, and long-term investment strategies for AI companies. For a deeper dive into the immediate news, you can find further details here.
The Unraveling of "Fair Use": A New Risk Frontier
For years, many AI developers operated under the implicit assumption that web scraping for model training fell within the bounds of "fair use." The Perplexity lawsuits, however, forcefully challenge this notion. Allegations extend beyond mere data ingestion, citing direct, often verbatim, reproduction of copyrighted articles and the systematic siphoning of web traffic and revenue from original publishers. Perplexity is accused of deploying crawlers that intentionally bypass `robots.txt` directives and of attributing AI-generated "hallucinations" to reputable sources, creating significant reputational risk for content owners.
This legal offensive follows a growing wave of litigation, including high-profile cases against OpenAI by The New York Times and others, and Dow Jones and the New York Post against Perplexity AI, the latter of which has already cleared preliminary hurdles and is proceeding to trial. The courts are increasingly scrutinizing the transformative nature of AI’s use of copyrighted material, and the defense of "fair use" is facing an uphill battle, particularly when AI outputs directly compete with, or act as substitutes for, original content. For investors, this translates into a heightened legal risk profile for any AI startup relying heavily on unlicensed, scraped data.
From Scraping to Licensing: The Rise of the Data Economy
The writing is now on the wall: the cost of data for AI training is set to rise dramatically. The casual approach to data acquisition is being replaced by a burgeoning, yet complex, data licensing ecosystem. Content creators, from major news conglomerates like News Corp and Axel Springer to individual authors and artists, are recognizing the inherent value of their intellectual property and are actively pursuing licensing agreements with AI developers. We’re already seeing significant deals, with companies like HarperCollins reportedly charging thousands per title for non-fiction works, and Shutterstock generating substantial revenue from licensing its digital assets to AI firms.
This shift represents both a challenge and an opportunity. While it introduces new operational costs for AI companies, it also creates legitimate, sustainable pathways for data acquisition and unlocks new revenue streams for publishers. Investors must now analyze an AI company’s data strategy with the same rigor they apply to its technological stack. Those with proprietary datasets or a proactive approach to securing robust licensing agreements will possess a significant competitive advantage and a more defensible business model.
Redefining Due Diligence: What Investors Must Now Scrutinize
The Perplexity cases, alongside Anthropic’s recent $1.5 billion copyright settlement, signal a clear mandate for Investment and Venture Capital Professionals: elevate your due diligence on AI ventures. The traditional checklist is no longer sufficient. Key areas of intensified scrutiny now include:
- Data Provenance and Acquisition: Demand transparent documentation of all training data sources. Understand the methods used for data collection and verify legal rights to use that data.
- Licensing Agreements: Assess the breadth, terms, and cost of existing data licensing deals. Are they scalable? Are they future-proof against evolving copyright interpretations?
- IP Risk Assessment: Evaluate potential exposure to copyright and trademark infringement lawsuits. Look for robust legal teams and proactive measures to mitigate these risks.
- "Hallucination" and Attribution Policies: Understand how AI models are designed to minimize inaccuracies and misattribution, particularly if user-facing. Reputational damage to partners can quickly translate to financial liabilities.
- Ethical AI Frameworks: Beyond legal compliance, an ethical approach to data use and content generation is becoming a competitive differentiator and a safeguard against future regulatory headaches.
Valuation in Flux: Accounting for Content Acquisition Costs
The "free data" paradigm that underpinned many early AI valuations is effectively over. The increased cost of acquiring legitimate training data will directly impact profitability and, consequently, valuation multiples. Private equity analysts and VCs must recalibrate their financial models to account for these rising expenses, which will shift from a perceived negligible cost to a significant and ongoing operational expenditure.
Companies that strategically invest in legally sound, high-quality datasets will build more resilient and valuable businesses. Those that continue to rely on ambiguous data sourcing methods face not only the risk of crippling legal settlements but also the potential for their core models to be deemed unusable, rendering their entire enterprise uninvestable. The ability to demonstrate a clear, ethical, and legally compliant data strategy will become a cornerstone of investor confidence and a critical determinant of long-term success.
A New Era of Sustainable AI Investment
The Perplexity AI lawsuits are more than just a legal skirmish; they represent a pivotal moment in the maturity of the AI industry. The era of the "Wild West" for data scraping is concluding, giving way to a more structured and legally compliant landscape. For Investment and Venture Capital Professionals, this demands an immediate re-evaluation of current holdings and future investment theses. Success in this new environment will hinge on backing AI companies that prioritize ethical data acquisition, cultivate transparent licensing relationships, and build defensible models on legally sound foundations. The next wave of AI unicorns will be built not just on groundbreaking algorithms, but on scrupulous adherence to intellectual property rights, signaling a more sustainable and responsible trajectory for the entire sector.
Also Read:


