spot_img
HomeNews & Current EventsAI Models ChatGPT, Cursor, and Perplexity Utilize Google-Scraping Startup...

AI Models ChatGPT, Cursor, and Perplexity Utilize Google-Scraping Startup for Data, Raising Ethical Concerns

TLDR: A recent report by The Information reveals that prominent AI models like ChatGPT, Cursor, and Perplexity are leveraging a Google-scraping startup, identified as SerpApi, to gather data from Google’s search results. This practice has ignited a debate over data sourcing, copyright infringement, and the ethical implications of AI models republishing content, including paywalled articles, without proper attribution. Publishers are expressing concerns over the increased scraping activity and the challenges in protecting their content.

A new report from The Information, published on August 28, 2025, highlights the growing reliance of leading artificial intelligence models, including OpenAI’s ChatGPT, Cursor, and Perplexity, on third-party services to scrape data directly from Google’s search results. The startup at the center of this revelation is identified as SerpApi, an eight-year-old web scraping firm that has listed OpenAI as a customer. This practice has brought to the forefront significant ethical and legal questions regarding data acquisition for AI training and content generation.

According to the report, AI developers’ web scraping activities have more than doubled in recent months, with major AI companies like OpenAI, Perplexity, and Meta collectively scraping websites millions of times. This aggressive data collection has led to instances where AI models, particularly Perplexity, have been accused of republishing paywalled articles and other content with nearly identical wording from news outlets such as Forbes, CNBC, and Bloomberg, often without adequate attribution. Perplexity CEO Aravind Srinivas acknowledged that their republishing feature, ‘Perplexity Pages,’ had ‘rough edges’ following a cease-and-desist letter from Forbes for alleged copyright infringement. Furthermore, Perplexity has been found to cite low-quality, AI-generated blogs and social media posts containing inaccurate information.

Also Read:

Publishers face a dilemma, as blocking Google’s bots, which are reportedly used for both web indexing and AI data scraping, could negatively impact their search engine optimization (SEO). This dual-purpose use makes it challenging for content creators to discern the intent behind the scraping and protect their intellectual property. Despite efforts to block services like Perplexity, reports indicate that the AI startup continues to send referral traffic, suggesting covert scraping operations. While some studies suggest a limited overlap between ChatGPT’s generated results and Google’s search outcomes, the underlying method of data acquisition remains a contentious issue in the AI community.

Dev Sundaram
Dev Sundaramhttps://blogs.edgentiq.com
Dev Sundaram is an investigative tech journalist with a nose for exclusives and leaks. With stints in cybersecurity and enterprise AI reporting, Dev thrives on breaking big stories—product launches, funding rounds, regulatory shifts—and giving them context. He believes journalism should push the AI industry toward transparency and accountability, especially as Generative AI becomes mainstream. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -