For a small content site, my answer is to let the search and answer crawlers in, to think harder about the training crawlers, and to treat charging as an experiment rather than income. That is our current reasoning, not a law, and the rest of this post shows how I got there and what would change it.

We run about a dozen content and product sites at Wbcom Designs, most of them behind Cloudflare. Several are EmDash sites on Cloudflare Workers and the rest are WordPress. Once a month or so someone on the team asks the same question: should we block the AI bots, or charge them? This post is the decision, written down so we stop re-arguing it. It is not legal or financial advice, and the numbers I quote come from the sources I name, on the dates I read them.

Two sister posts cover neighbouring ground, so I will not repeat them. For measuring how much traffic AI actually sends you, read attowp’s How Much Traffic Do You Actually Get From AI? How to Measure It, Then Grow It Back. For the mechanics of blocking on WordPress, read How to Block AI Crawlers from Scraping Your WordPress Site. This one is only about the choice.

In this guide

  • The three kinds of AI visitor, and why treating them as one leads to bad decisions
  • What we allow today, quoted from our own robots.txt
  • What blocking buys and what it costs
  • What charging means today, who can use it, and what the reported numbers say
  • Letting bots in on terms: content signals, rate limits and page-by-page rules
  • A decision table by page type
  • What I will watch, and what would change my mind

What are the three kinds of AI visitor?

An “AI crawler” is not one thing. The companies themselves split their bots into at least three jobs, and the right answer differs for each. Lumping them together is how people end up blocking the one bot that would have sent them a reader.

1. Training crawlers

These collect pages to build or improve a model. OpenAI’s documentation says GPTBot is “used to make our generative AI foundation models more useful and safe”, and that disallowing it “indicates a site’s content should not be used in training generative AI foundation models”. Anthropic describes ClaudeBot as collecting web content that helps “enhance the utility and safety of our generative AI models”. Google’s training control is a token called Google-Extended, which Google describes as a way to manage whether content it crawls “may be used for training future generations of Gemini models”. The same Google page says Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”, and that it has no user agent string of its own. It is a robots.txt token only.

2. Search and answer crawlers

These build an index so the product can show your page, with a link, when someone asks a question. OpenAI says OAI-SearchBot is “used to surface websites in search results in ChatGPT’s search features” and that sites which disallow it “will not be shown in ChatGPT search answers”. Perplexity says PerplexityBot is “designed to surface and link websites in search results on Perplexity” and “is not used to crawl content for AI foundation models”. Anthropic says Claude-SearchBot “navigates the web to improve search result quality for users”.

This is the group that can send you readers. It is also the group whose blocking has the most direct cost to visibility.

3. User-triggered fetchers

These do not crawl on a schedule. A person asks an assistant a question, and the assistant fetches a page to answer it. OpenAI’s ChatGPT-User is used “for certain user actions in ChatGPT and Custom GPTs”, and the documentation says “because these actions are initiated by a user, robots.txt rules may not apply”. Perplexity says of Perplexity-User that “since a user requested the fetch, this fetcher generally ignores robots.txt rules”. Anthropic lists Claude-User as the bot that “supports Claude AI users”, and says its bots honour robots.txt. Cloudflare’s category for these is “AI Assistant”, described as an “automated AI bot driven by user action”.

Notice what that means. A robots.txt rule you wrote for the first two groups may simply not be consulted by the third, by the vendor’s own account. If you want to stop a user-triggered fetcher you need a server-side rule, not a text file.

Why lumping them together fails

Symptom: someone reads a scary headline about AI scraping, adds a blanket block for every bot with “AI” in its name, and a month later asks why their articles have vanished from AI answers. The check is to look at which user agents you actually blocked and which job each one does. The fix is to write rules per job. To prevent it, keep a short list in one place that says, for each bot you know of, whether it is training, search or user-triggered, and revisit it when a vendor changes its documentation.

What do we allow today, and why?

Our EmDash site template serves a robots.txt that lets every crawler in. I read the file in the template and then fetched the live one from tweakswp.com to confirm they match. This is what is served:

# Every crawler, including AI crawlers (GPTBot, ClaudeBot, PerplexityBot,
# Google-Extended...), may crawl all public pages and media.
User-agent: *
Allow: /
Allow: /_emdash/api/media/file/
Disallow: /_emdash/

Sitemap: https://tweakswp.com/sitemap.xml

In plain words: everything public is open, the media files are explicitly crawlable, and the only thing off limits is the path the CMS uses for its admin and private API. The longest matching rule wins, so the media allow line beats the broader disallow. There are no per-bot rules and no content signals in it.

Why that default? Our reasoning is simple, and I want to be honest that it is a judgement and not a measured result. These sites exist to be found. People who need help describe their problem to an assistant in their own words, and an assistant can only recommend a page it has been allowed to read. For a small site, being absent from those answers looks like a bigger risk than being read without payment. Media stays crawlable for the same reason: images and files are part of how a page gets understood and quoted.

I have no traffic, revenue or experiment numbers from our own sites to put behind this, and I will not invent any. An open default is easy to defend while you do not yet know what AI referral is worth, because it is also the easy one to tighten later.

We also published about moving our WordPress sites onto EmDash in We Moved Five WordPress Sites to EmDash the Week 1.0 Shipped. This robots.txt is part of that template, so every site built from it inherits the same open policy until someone changes it on purpose.

What does blocking buy, and what does it cost?

Blocking buys you a statement of preference and, if you enforce it, a real barrier. It costs you whatever those bots would have given you.

Robots.txt is a request, not a lock

The standard that defines robots.txt, RFC 9309, describes it as a way for service owners to control how their content may be accessed “by automatic clients known as crawlers”. Its security section says plainly: “The Robots Exclusion Protocol is not a substitute for valid content security measures.” Cloudflare’s documentation says the same about its managed file: “robots.txt compliance is voluntary. The file expresses your preferences, but it does not prevent crawlers from accessing your content at a technical level.” And its advice for enforcement is to “use AI Crawl Control”.

So there are two different things people mean by “block”. One is asking politely in robots.txt, which well-behaved crawlers honour (the vendors above say theirs do, except the user-triggered ones). The other is refusing the request at the edge, which works whether or not the bot cooperates.

What blocking can cost

  • Visibility in AI answers. OpenAI states that a site disallowing OAI-SearchBot will not be shown in ChatGPT search answers. Block the search bots and you opt out of those answers, by the vendor’s own description.
  • Agents acting for your customers. Sarah Perez reported at TechCrunch on 6 October 2026 that many sites are blocking personal AI agents, some on purpose and some through ordinary anti-bot checks, and that this is frustrating the people using those agents. If you sell something, a blocked agent is a customer turned away.

What blocking can buy

  • A say over training. If you do not want your text used to train models, the training bots are the ones to refuse. Blocking GPTBot or Google-Extended does not, per the vendors, remove you from their search products (Google says so explicitly for Google-Extended).
  • Lower server load. Crawlers cost bandwidth and compute. On a cached static or edge-served site this is small. On a WordPress site running uncached queries, it can matter more. Check your own logs before assuming it is your problem. If it is, rate limiting (below) is usually a better first tool than a block.

How to check what you have blocked by accident

Symptom: your robots.txt, a security plugin or a Cloudflare rule is stricter than you meant. To check, fetch your own robots.txt in a browser, read it top to bottom, and then look at the bot analytics in your CDN dashboard to see which crawlers are being refused. On Cloudflare, Cloudflare’s documentation says AI Crawl Control is “available on all plans”, and that on the Free plan detection is limited to user agent strings. To prevent surprises, note which layer owns the rules, because the CMS, a security plugin and the CDN can each add a block unseen by the others.

What does “charge” actually mean today?

Charging AI visitors means answering a request with HTTP 402 Payment Required plus a price, and serving the page only when the visitor pays. Several schemes exist, and they are not the same thing.

Pay per crawl (Cloudflare)

Cloudflare’s documentation describes pay per crawl as a way to “control and monetize AI crawler access to content by setting a price per zone”. A crawler that arrives without payment intent receives “an HTTP 402 Payment Required response with pricing”, and one that presents valid payment headers gets the page. Cloudflare’s original announcement post (dated 1 July 2025, last modified 15 July 2026) says publishers set “a flat, per-request price across their entire site”, with the ability to exempt specific crawlers, and that crawlers identify themselves with cryptographic signatures.

Two facts matter for a small site. First, the status: Cloudflare’s pages call it “private beta” or “closed beta”, with a sign-up form or an Enterprise account contact. Second, order of operations: Cloudflare says it operates “only after existing WAF policies and bot management or bot blocking features have been applied”, so a bot you have blocked cannot pay its way in. I could not find a plan tier or a price stated on those pages, so I will not state one. In the AI Crawl Control dashboard the three actions are allow, block and “charge for crawl”, and the charge option is the one in private beta.

Monetization Gateway and Pay Per Use (Cloudflare)

Search Engine Journal (Matt G. Southern, 1 October 2026) reported that Cloudflare moved two products into closed beta on 30 September 2026, limited to U.S.-based participants. The Monetization Gateway “charges an AI agent per request, with prices set by the seller”, using the x402 protocol and HTTP 402, for APIs, MCP tools, datasets or websites. Pay Per Use is different: SEJ quotes Cloudflare saying that pay per crawl “charges for access”, whereas Pay Per Use “pays for what happens next”, such as a cited answer in AI search. The AI company defines the use and the price, and publishers decide which offers to accept.

The article lists seller requirements: a credit card on file, a verified email, an account older than 60 days, a zone proxied through Cloudflare that is older than 30 days, and passed security checks. It reports settlements from $0.001 to $100 per request, paid in USDC (a dollar-linked stablecoin) on the Base blockchain, and names four live services, none of them a content blog: Ceramic.ai, Cloudflare’s AI Gateway, Stocktwits and API2PDF. I read this on SEJ and not on a Cloudflare page, so treat it as SEJ’s account of the announcement.

A penny per page, paid by an agent

Suganthan Mohanadasan wrote on Search Engine Journal on 17 September 2026 about making his own demo site charge AI agents one cent per request for one article, using the open x402 protocol on Cloudflare Workers with roughly 300 lines of code. Five payments went through on 15 September 2026, one of them from Claude Code using a wallet set up for it. The detail that matters most is in the article: the payments used testnet USDC, which has no real monetary value, and his buyer agent enforced limits of $0.05 per request and $0.25 a day.

He is also candid about the limits. He wrote that he would not do it “expecting search crawlers to pay you”, and the article notes that the major crawlers (GPTBot, ClaudeBot, PerplexityBot, Googlebot) do not yet support payment. His suggestion is to price original research, datasets and tools rather than general explainers.

I read that as proof the plumbing works, not proof there is money in it.

Google’s pilot

Matt G. Southern, again at SEJ on 1 October 2026, summarised reporting from The Information on Google’s AI contribution pilot. Google is paying roughly 100 digital publishers. For “several small and midsize sites in the pilot, the payments amount to less than 0.1% of their advertising revenue”. Some smaller sites earned less than $1,000 over several months, while one early participant earns more than $1 million a year. Participants told The Information they do not know how payments are calculated and that amounts can change month to month without explanation. Content has to contribute to Gemini app, AI Overviews or AI Mode answers; merely confirming facts or being linked afterwards does not qualify. This is a closed pilot, not something a site can sign up for, and I am relaying a second-hand report.

Why the money at small-site scale is likely small

This is my reasoning and not a measurement. I would not call it a forecast.

  • The buyers are mostly not there yet. By the one article that tested it, the major crawlers do not pay. A price nobody can pay brings in nothing.
  • Prices in the examples are fractions of a cent to a cent. A cent per page needs a great many paid fetches to add up to anything. A small site’s pages are fetched by a small number of agents.
  • The reported payouts for small sites are tiny. The large figures in the reporting belong to outliers. The number reported for several small and midsize sites in the Google story is under 0.1% of ad revenue.
  • Access is gated. The Cloudflare products are closed betas with requirements, some limited to the U.S., so most sites cannot try them today even if they wanted to.
  • Charging can cost you answers. If a crawler cannot pay, a price wall turns into a block, and you are back to the visibility cost from the last section.

That does not make charging pointless. It makes it an experiment worth running on the one kind of page where the content has real value to a machine, and worth watching, but not a line in the budget.

Can I let bots in on terms?

Yes, and for most small sites this middle path is where the real decision sits. There are three levers.

Content signals

Cloudflare’s managed robots.txt adds machine-readable preferences called content signals. Its documentation lists three: search (“building a search index and providing search results”), ai-input (“inputting content into one or more AI models”) and ai-train (“training or fine-tuning AI models”). When you enable it, Cloudflare prepends directives with the defaults search=yes, ai-train=no, use=reference, plus a policy text. It is available on all plans. The caveat is the same as for any robots.txt: it states a preference and does not enforce one.

This maps neatly onto the three kinds of visitor. Search yes, training no, is a defensible stance for a content site that wants to be found but not folded into a model. Our template does not publish content signals yet. Whether we should add them is on my list below.

Rate limits

If the worry is load and not principle, slow the crawlers down. Anthropic documents a Crawl-delay extension for ClaudeBot in robots.txt (Crawl-delay: 1), and a rate limit at the CDN works on any bot regardless of what it promises. Symptom: a spike of requests from a single user agent slows a WordPress site. Check: your server or CDN logs, sorted by user agent. Fix: a rate rule for that agent, not a permanent block. Prevent: page caching, so the next spike costs almost nothing.

Different rules for different pages

The TechCrunch article is mostly about this. Personal agents that book flights and order groceries get turned away, some by deliberate policy (it reports Amazon blocking Meta’s Muse agent, and Yelp requiring paid data licensing) and some by accident, because a CAPTCHA built for spam also stops a legitimate agent (it names Walmart, which has a partnership with Meta). It reports that Meta and several companies including Stripe, Sierra, Genesys, Rocket, NiCE and Decagon are working on an open standard for agent communication in online commerce, aimed at telling a personal agent apart from a malicious bot. It lists options for site owners: partnerships, adopting the standard, adjusting security settings to tell good agents from bad, and using Cloudflare’s controls to block training while allowing agent activity.

For a content portfolio the practical reading is this. Reading an article and using a form are different risks. Letting an agent read an article costs you little. Letting an unknown agent submit a sign-up form, a checkout or a contact form costs you spam, fraud or support load, and it needs a rule that identifies the visitor, not a robots.txt line. At the same time, if you sell something online, treating every non-human visitor as hostile can turn away a real customer who happens to be using an assistant. Decide per page, and keep the rules at the edge where they can tell a request for an article from a request to submit a form.

What should I do for each type of page?

This is the table we are working from for a portfolio. It is a starting position that we review, not a recommendation for every site. “Limit” means rate limits and content signals. “Charge” means worth trying if you are eligible, not a plan to rely on.

Page typeSearch and answer botsTraining botsUser-triggered agentsWhy
Articles and guidesAllowAllow for now, reviewAllowThese exist to be found and quoted. Add content signals to state preferences.
Product and service pagesAllowAllow for nowAllow reading, limit actionsBeing described correctly in answers helps buyers. Reading is harmless, checkout is not.
DocumentationAllowAllowAllowDocs answer specific questions, so being quoted accurately is the point.
Original research, datasets, paid toolsAllow summaries, limit full textBlock or chargeLimit or chargeThis is the only content where a machine-readable price is plausible, per the SEJ author.
Forms: sign-up, checkout, contact, loginNot relevantNot relevantLimit or block unidentified agentsAbuse and fraud risk. Use edge rules and verification, not robots.txt.
Media: images, filesAllowReview per siteAllowHelps pages be understood and quoted. Rate limit if bandwidth bites.
Admin and private APIsBlockBlockBlockNever public. Our template disallows the CMS path for this reason, and real protection needs authentication.

What will I watch, and what would change my mind?

I am not going to promise a date for any of this. These are the signals, in the order I will check them.

  1. Whether the big crawlers start paying. The SEJ author says they do not support payment today. If that changes and a price can actually be collected, charging moves from a test to a real option. Until then, a price is a block in disguise.
  2. Whether the programmes open to us. The Cloudflare features are closed betas, one with a U.S. limit and account age rules. I do not yet know whether our zones meet the conditions, and I will not claim they do. I will check eligibility zone by zone when each opens further.
  3. What our own logs say. If a particular bot is costing real load, a rate rule comes first. If we find a training bot taking everything and answer bots sending nothing, the training default is the thing I would change first.
  4. What the content signals do in practice. Honouring them is voluntary, so I want vendor documentation saying they do before adding them to the template.
  5. Whether the agent standard lands. If a clear way to identify a legitimate personal agent appears, the forms and checkout rows in the table can move from “limit” to “allow identified agents”.

What would change my mind toward blocking training bots: a sign that they take content without sending anything back, or that our research-style pages are being reproduced in answers without a link. What would change my mind toward charging: a buyer that actually pays, a programme we qualify for, and a page type where a price makes sense. What would change my mind toward leaving everything open forever: evidence that AI answers send us readers who convert.

Questions people ask

Should a small content site block AI crawlers?

Not all of them. The search and answer bots are the ones that can show and link your pages. Block by job, not by name. Training bots are the group worth a genuine decision, and it is reasonable either way.

Does blocking Google-Extended hurt my Google rankings?

Google says no. Its documentation states that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”. It controls whether content may be used for training future Gemini models.

Will robots.txt stop an AI bot from reading my page?

Only if the bot chooses to obey it. RFC 9309 says the protocol is not a substitute for real content security. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person triggers it, and Perplexity says Perplexity-User generally ignores them. To enforce a block you need a rule at your server or CDN.

Can I get paid when an AI crawler reads my site?

Possibly, but not easily today. Cloudflare’s pay per crawl is in private or closed beta with sign-up, and the newer Monetization Gateway and Pay Per Use are closed betas limited to the U.S. The one public experiment I read used testnet money with no real value.

How much could a small site earn from charging bots?

I do not have a figure, and nobody has given one I trust for small content sites. The reported numbers point to small amounts: a cent per page in the SEJ demo, and under 0.1% of ad revenue for several small and midsize sites in Google’s pilot. Treat any income as a bonus.

If you run a handful of content sites and want a second pair of eyes on your crawler policy, or on how your sites cope with AI traffic, you can reach me through the contact page.