Perplexity Accused of circumventing website Restrictions with Stealth Bots
AI-powered search engine Perplexity is facing allegations of employing deceptive tactics to bypass website restrictions, potentially violating long-standing internet protocols. network security firm Cloudflare reported Monday that Perplexity appears to be utilizing “stealth” bots to access content even after being explicitly blocked.
Cloudflare researchers began investigating after receiving complaints from customers. These customers had already taken steps to prevent Perplexity’s known crawlers from accessing their sites using robots.txt files and web submission firewalls. Despite these measures,Perplexity continued to scrape content.
Further examination revealed a concerning pattern. When confronted with blocks, Perplexity allegedly deployed an undeclared crawler designed to mask its activity. this crawler employed several techniques to evade detection.
Key Findings of the Investigation
Multiple IP addresses, outside of Perplexity’s officially listed range, were used. IP addresses were rotated frequently in response to restrictive policies.
Requests originated from diverse Autonomous System Numbers (ASNs) to further obscure the source.
This activity spanned over 10,000 domains and millions of requests daily.
This alleged behavior directly challenges established internet norms. In 1994, the Robots Exclusion Protocol was proposed by engineer Martijn Koster.This protocol provided a standardized way for websites to communicate their crawling preferences to bots.
The protocol, implemented through a simple robots.txt file placed on a website’s homepage, has been widely respected for over three decades. It formally became an Internet Engineering Task Force standard in 2022, solidifying its importance.Essentially, the robots.txt file acts as a set of instructions for well-behaved bots. It tells them which parts of your site they shouldn’t crawl. By allegedly ignoring these directives, Perplexity’s actions raise important questions about responsible AI development and respect for website owners’ wishes.
You rely on these protocols to control how your content is accessed and indexed. Circumventing these rules could have serious implications for website security, bandwidth usage, and overall control over your online presence. This situation warrants close attention as the debate around AI scraping and ethical web crawling continues to evolve.
Worth a look