
Robots.txt no longer stops AI scrapers stealing your content
Patreon proved that AI crawlers routinely ignore robots.txt rules, scraping content thousands of times weekly until active blocking was deployed. UK IT managers and digital security leads need to know what this means for protecting business IP.


Most UK businesses assume a robots.txt file is enough to keep AI scrapers away from their content. Patreon's experience this month shows it is not. The creative platform discovered that AI crawlers were ignoring its passive rules and making thousands of attempts every week to extract content. That finding matters far beyond the creator economy.
If your organisation publishes product documentation, marketing materials, or proprietary databases online, the same thing is very likely happening to you right now.
What Patreon found and what it did about it
Patreon announced in July 2026 that it had moved away from passive robots.txt instructions to actively blocking AI scraping bots using Cloudflare's AI Crawl Control tool. The decision followed a clear pattern: AI crawlers were routinely disregarding the standard protocol that tells bots which pages to avoid.
Before the switch, Patreon recorded thousands of weekly unauthorised scraping attempts. After deploying active blocking, those attempts dropped to zero.
The move signals a broader shift in how organisations need to think about content protection. Robots.txt was designed for a cooperative web, where search engines agreed to follow the rules. Many AI training crawlers do not operate on that assumption.
Why this is a direct concern for UK IT managers
Robots.txt has been the default protection layer for most UK SME websites for decades. IT managers and digital security leads often set it once and move on, assuming the problem is solved. This case study makes clear that assumption carries real risk.
AI companies building large language models need vast amounts of text data. Your product pages, case studies, knowledge bases, and service descriptions are exactly the kind of structured, high-quality content those crawlers target. If scraped without consent, that material can be used to train competitor tools or replicate proprietary knowledge at scale.
The challenge is not purely technical. Most SME teams have not updated their threat models to account for AI-era scraping. The operational discipline and security awareness required to respond effectively are often absent, not because the team lacks skill, but because no one has mapped this as a live risk.
At gecco, our Training and consultancy programme addresses exactly this gap, helping teams build the governance awareness and practical habits needed to manage AI-era risks before they become incidents.
Three actions your team can take this week
1. Audit your current robots.txt file. Log into your website's root directory and review what instructions are currently in place. Check whether AI-specific crawlers such as GPTBot, ClaudeBot, or CCBot are named. This takes under 30 minutes and gives you an immediate picture of your exposure.
2. Evaluate Cloudflare's AI Crawl Control. If your site is already on Cloudflare, this feature is available within your existing dashboard. If not, assess whether migrating DNS to Cloudflare is feasible. Your IT manager or a web developer can run a scoping review within two to three working days.
3. Map your highest-value content assets. Ask your digital security lead to identify which pages contain proprietary or commercially sensitive material. Prioritise active blocking rules for those areas first. This focused approach means you protect what matters most without waiting for a full site overhaul.
The honest limits of active blocking
Active blocking through tools like Cloudflare's AI Crawl Control is a meaningful improvement on robots.txt. It is not a complete defence. Crawlers adapt, and new bots will emerge that are not yet in any blocklist.
For IT managers considering this approach, it is worth noting that implementation requires some technical confidence. Misconfigured blocking rules can affect legitimate search engine indexing if applied too broadly. Testing in a staging environment before rolling changes to production is essential.
Active blocking also does not address content that has already been scraped. If AI crawlers have been active on your site for months, some of your material may already be in training datasets. Blocking prevents future exposure, not historical loss.
Where to go from here
The shift from robots.txt to active blocking is a technical change. But it sits on top of a deeper organisational one. IT managers and digital security leads cannot protect what they have not mapped, and they cannot sustain protection without a team that understands why it matters.
If your organisation is working through where AI governance and IP protection fit into your existing security practice, that is exactly what the AI Readiness survey is designed to surface. Taking the survey gives you access to 65+ free resources and a custom AI Readiness report, followed by a free 45-minute AI Readiness call to walk through the results with a member of our team.

Claude Code now builds live dashboards for your whole team
Anthropic updated Claude Code on 18 July 2026, adding live data integration and collaborative editing for UK software teams. This article explains what changed and how developers at UK SMEs can act on it this week.

How AI agents fix communications bottlenecks
UK SMEs are using agentic AI to automate call routing, meeting scheduling, and network monitoring. Find out how this approach reduces admin and cuts downtime risk.

