conv.

All stories
TechFading · day 4

Website owner details year-long fight against AI scraper bots on 1.5M-page site

PatronView's operator found 99% of traffic was bots, with Anthropic's Claude crawling 35,000 times for every referred visitor, sparking wide discussion of scraper economics and defenses.

Conversation activity · last 4 days peak 5/hr

Peak 5 items in one hour at Aug 7, 10 AM; 14 items over 4 days Aug 7, 10 AM — 5 items · Hacker News 5Aug 7, 11 AM — 6 items · Hacker News 6Aug 7, 12 PM — no itemsAug 7, 1 PM — no itemsAug 7, 2 PM — 1 item · Hacker News 1Aug 7, 3 PM — no itemsAug 7, 4 PM — no itemsAug 7, 5 PM — no itemsAug 7, 6 PM — no itemsAug 7, 7 PM — no itemsAug 7, 8 PM — no itemsAug 7, 9 PM — no itemsAug 7, 10 PM — no itemsAug 7, 11 PM — no itemsAug 8, 12 AM — no itemsAug 8, 1 AM — no itemsAug 8, 2 AM — 1 item · Hacker News 1Aug 8, 3 AM — no itemsAug 8, 4 AM — no itemsAug 8, 5 AM — no itemsAug 8, 6 AM — no itemsAug 8, 7 AM — no itemsAug 8, 8 AM — no itemsAug 8, 9 AM — no itemsAug 8, 10 AM — no itemsAug 8, 11 AM — no itemsAug 8, 12 PM — no itemsAug 8, 1 PM — no itemsAug 8, 2 PM — no itemsAug 8, 3 PM — no itemsAug 8, 4 PM — no itemsAug 8, 5 PM — no itemsAug 8, 6 PM — no itemsAug 8, 7 PM — no itemsAug 8, 8 PM — no itemsAug 8, 9 PM — no itemsAug 8, 10 PM — no itemsAug 8, 11 PM — no itemsAug 9, 12 AM — no itemsAug 9, 1 AM — no itemsAug 9, 2 AM — no itemsAug 9, 3 AM — no itemsAug 9, 4 AM — no itemsAug 9, 5 AM — no itemsAug 9, 6 AM — no itemsAug 9, 7 AM — no itemsAug 9, 8 AM — no itemsAug 9, 9 AM — no itemsAug 9, 10 AM — no itemsAug 9, 11 AM — no itemsAug 9, 12 PM — no itemsAug 9, 1 PM — no itemsAug 9, 2 PM — no itemsAug 9, 3 PM — no itemsAug 9, 4 PM — no itemsAug 9, 5 PM — no itemsAug 9, 6 PM — no itemsAug 9, 7 PM — no itemsAug 9, 8 PM — no itemsAug 9, 9 PM — no itemsAug 9, 10 PM — no itemsAug 9, 11 PM — no itemsAug 10, 12 AM — no itemsAug 10, 1 AM — no itemsAug 10, 2 AM — no itemsAug 10, 3 AM — no itemsAug 10, 4 AM — no itemsAug 10, 5 AM — no itemsAug 10, 6 AM — no itemsAug 10, 7 AM — no itemsAug 10, 8 AM — no itemsAug 10, 9 AM — no itemsAug 10, 10 AM — no itemsAug 10, 11 AM — no itemsAug 10, 12 PM — no itemsAug 10, 1 PM — no itemsAug 10, 2 PM — no itemsAug 10, 3 PM — no itemsAug 10, 4 PM — no itemsAug 10, 5 PM — no itemsAug 10, 6 PM — no itemsAug 10, 7 PM — no itemsAug 10, 8 PM — no itemsAug 10, 9 PM — no itemsAug 10, 10 PM — no itemsAug 10, 11 PM — no itemsAug 11, 12 AM — no itemsAug 11, 1 AM — 1 item · Techmeme 1Aug 11, 2 AM — no itemsAug 11, 3 AM — no itemsAug 11, 4 AM — no itemsAug 11, 5 AM — no itemsAug 11, 6 AM — no itemsAug 11, 7 AM — no itemsAug 11, 8 AM — no itemsAug 11, 9 AM — no itemsAug 11, 10 AM — no itemsAug 11, 11 AM — no itemsAug 11, 12 PM — no itemsAug 11, 1 PM — no itemsAug 11, 2 PM — no itemsAug 11, 3 PM — no itemsAug 11, 4 PM — no itemsAug 11, 5 PM — no itemsAug 11, 6 PM — no itemsAug 11, 7 PM — no itemsAug 11, 8 PM — no items 5 items · 10 AM
Aug 8Aug 9yesterdaytodaynow · 9:45 PM

Summary, timeline and people extracted by Claude from 14 items across 2 sources · 16h ago. Quotes are verbatim.

The operator of PatronView, a 1.5-million-page website, published a detailed account of a year spent fighting automated scraper traffic, finding a 214:1 bot-to-human page-load ratio and that Anthropic's Claude-SearchBot fetched pages roughly 35,000 times for every human visitor it referred, while an Amazon bot referred none at all. The post, amplified on Hacker News and X, drew a large discussion from other site operators reporting similar bot ratios and debating tools like Cloudflare and Anubis, the ethics of AI crawlers taking content without compensation, and the tradeoffs of aggressively blocking traffic on the open web.

  • PatronView, a 1.5-million-page site, reported a 214:1 bot-to-human page-load ratio over a year of tracking traffic.
  • Anthropic's Claude-SearchBot crawled the site roughly 35,000 times for every human visitor it referred; an Amazon bot referred none.
  • Other operators (e.g., SignalBloom) reported nearly identical patterns, fueling broader complaints that AI crawlers extract value without compensating or crediting source sites.
  • Discussion split between technical fixes (Cloudflare, Anubis, proof-of-work challenges, nginx/firewall rules) and concerns that heavy anti-bot measures centralize web access control and can also block legitimate automated users.

How it unfolded

  1. Reaction Story recirculates via Techmeme and X

    Techmeme aggregated the story and it was reshared on X by @nickgraynews, restating the 35,000-to-1 Claude crawl ratio.

    “I was looking at my website server logs and was shocked to see that 99% of my traffic is from bots Anthropic scrapes my site 35,000 times for every 1 visitor they send”

    @nickgraynews
  2. A commenter described years of tweaking nginx and firewall rules to fend off scrapers and reluctance to rely on Cloudflare.

  3. Reaction Similar case from SignalBloom owner

    Another site operator reported that Claude-searchbot fetched about 205,000 pages from their finance site in 72 hours while sending only one referral, echoing PatronView's findings.

    “Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral.”

    GodelNumbering · Hacker News ↗
  4. Reaction Debate over bot motivations

    A commenter questioned why the same crawler would refetch identical pages within an hour, prompting discussion of whether this reflects redundant independent scrapers or inefficient crawling.

  5. Reaction Hosting cost spike discussed

    A commenter questioned the site's hosting bill, noting the article's mention of a 500% cost spike during a bad month and suggesting a database/hosting change.

    “My normal bill for running this whole site is around $90 a month. During one bad spike month, it jumped about 500%.”

    PatronView (site owner) · Hacker News ↗
  6. A commenter highlighted the article's admission that PatronView itself scrapes public documents to build its data, calling out the irony.

    “And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.”

    PatronView (site owner) · Hacker News ↗
  7. Event Story hits Hacker News

    petercooper submitted the post to Hacker News, where it reached 455 points and 426 comments.

  8. The site owner detailed a year of anti-scraper efforts, reporting 99% bot traffic, a 214:1 bot-to-human page-load ratio, and that Claude crawled the site 35,000 times per referred user while an Amazon bot sent zero referrals.

What people are saying verbatim

“I was looking at my website server logs and was shocked to see that 99% of my traffic is from bots Anthropic scrapes my site 35,000 times for every 1 visitor they send”

@nickgraynews, X user sharing the story · X · Aug 10

“And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.”

PatronView (site owner), Site operator, quoted in article · patronview.com (quoted via HN comment) ↗ · Aug 6

“My normal bill for running this whole site is around $90 a month. During one bad spike month, it jumped about 500%.”

PatronView (site owner), Site operator, quoted in article · patronview.com (quoted via HN comment) ↗ · Aug 6

“Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral.”

GodelNumbering, Owner of SignalBloom.ai · Hacker News comment ↗ · Aug 6

“I'm in a similar boat. Probably 99.999% is bots. I have nearly 1m unique "visitors" according to cloudflare and my real users are in the dozens a day.”

ddxv, Hacker News commenter, site operator · Hacker News ↗ · Aug 6

“The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare).”

jwr, Hacker News commenter · Hacker News ↗ · Aug 6

“This will resonate with anyone who operates a public-facing website and doesn't work at a big tech company.”

aorth, Hacker News commenter, site operator · Hacker News ↗ · Aug 7

Voices from the web unedited

  • This will resonate with anyone who operates a public-facing website and doesn't work at a big tech company. In the past few years I have also been down this exact same road. I'm desperately trying to avoid resorting to Cloudflare, but I'm running out of time and patience to keep tweaking nginx and firewall rules every few weeks.I had initial…

    aorthHacker News3d agoview on Hacker News ↗
  • The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I…

    jwrHacker News4d agoview on Hacker News ↗
  • I just checked Cloudflare for SignalBloom (https://www.signalbloom.ai, which I own and operate).Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral. There is a lot of free financial data on the site, hoping for real users to benefit from it. It is hard to not feel a little cheated out that Claude gets…

    GodelNumberingHacker News4d agoview on Hacker News ↗
  • Can someone help me understand the underlying motivation behind this?It makes sense that some crawlers, in the style of Google, would want to index the entire internet. But what is the point of the same crawler re-fetching a page they already fetched an hour ago? Or possibly all this traffic is just independent entities, each trying to cache the…

    varencHacker News4d agoview on Hacker News ↗
  • Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software.[1]

    johnorourkeHacker News4d agoview on Hacker News ↗
  • > If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user.People trying to block bots end up keeping out all kinds of users. I get blocked frequently for using a regular browser, just with JS disabled. 99% of the time, I just close the browser tab and…

    autoexecHacker News4d agoview on Hacker News ↗
  • I am denied by cloudflare CONSTANTLY on one system.I have an old os (macos 10.11), running the highest firefox esr I can run, and I get denied by cloudflare.But not always immediately - I get to enable javascript/cookies sometimes just to be denied.they are not the folks we want gatekeeping the internet, they are opportunists increasing their OKRs

    m463Hacker News4d agoview on Hacker News ↗
  • > My normal bill for running this whole site is around $90 a month. During one bad spike month, it jumped about 500%.This is D1 - which has very surprising costs. you may just want to drop D1 and move to a static site. There’s no reason your site should cost this much.

    tarr11Hacker News4d agoview on Hacker News ↗
  • I'm in a similar boat. Probably 99.999% is bots. I have nearly 1m unique "visitors" according to cloudflare and my real users are in the dozens a day. That being said, I love the open internet and am holding on to keeping as much open as I can.

    ddxvHacker News4d agoview on Hacker News ↗
  • > I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website.I think you've hit the nail on the head here. I think that's a big reason why people want to ban bots.

    eigencoderHacker News4d agoview on Hacker News ↗