Automated software crawlers now generate more than half of all internet traffic, outnumbering human users and creating severe economic and operational strains for digital publishers, according to recent research.
The Stealth Bot Threat to Digital Publishing
Bad bots—often referred to as stealth bots—are automated crawlers that intentionally conceal or falsify their identity and purpose. These malicious actors routinely circumvent access controls and rights reservations, frequently impersonating human traffic to scrape original journalism from news websites, including content locked behind paywalls.
According to media analysts, stolen content feeds a billion-dollar reseller market where artificial intelligence developers purchase scraped journalism directly from bot operators rather than compensating publishers and journalists. This unauthorized material is then utilized to build competing products that directly challenge original news sources.
News Corp Chief Executive Robert Thomson has characterized stealth bots as “the silent scavengers of the internet,” while Condé Nast CEO Roger Lynch warned that disguised crawlers scrape original journalism “with zero accountability.”
The financial and infrastructure toll on newsrooms is substantial. Uninvited bad bots can bombard a website millions of times in a single day, slowing performance for human readers, driving up bandwidth costs, and risking complete site outages. In August 2025, a UK technology website was forced offline after facing 1.6 million scrapes in a single day. The Wikimedia Foundation, which hosts Wikipedia, now blocks or throttles roughly a quarter of all automated requests hitting its infrastructure, amounting to billions of daily requests from crawlers that ignore site access policies.
While publishers have deployed sophisticated blocking technology, malicious bots continue to adapt by disguising themselves as humans. Local news organizations, operating on razor-thin profit margins, struggle to absorb these escalating technical expenses.
A.G. Sulzberger noted that “This theft isn’t just happening because publishers are leaving their toys out on the lawn; it’s happening when they are locked up safely in the house.”
Legislative Action on Both Sides of the Atlantic
Lawmakers in the United States and the United Kingdom are advancing parallel legislative measures to mandate transparency for automated crawlers.
In the U.S., a bipartisan proposal titled the Stealth Bot Prohibition Act would require automated crawlers to explicitly state their identity and purpose when visiting websites, providing publishers with the necessary tools to protect their content from disguised actors.
Meanwhile, the UK is pursuing a similar path through the Automated Online Software (Access and Transparency) Bill. Introduced as a Private Member’s Bill by MP Damian Hinds and backed by the News Media Association, the legislation would require any operator running a bot that systematically copies content from a UK website to disclose their identity, ownership, and intended use of the material.
This independent convergence across two major legislatures highlights a growing consensus that crawler transparency is foundational to a functioning digital market.
Transparency serves as an essential prerequisite for content licensing. Existing legal frameworks, such as the European Union’s Copyright in the Digital Single Market Directive, grant rightsholders the legal right to reserve their content from text and data mining in a machine-readable format. Since August 2, 2026, AI developers ignoring these reservations face penalties of up to €15 million or 3% of global turnover.
However, industry advocates note that opt-out mechanisms remain ineffective if incoming machines are permitted to falsify their identities. Effective enforcement relies entirely on the ability to distinguish truthful crawlers from dishonest ones.