Cloudflare's new default could hide your site from AI | Respondyr Blog Skip to main content
Respondyr
Back to Blog
ai search local seo small business visibility

Cloudflare's new default could hide your site from AI

On September 15, 2026, Cloudflare starts blocking AI training and agent crawlers by default on new domains. How to check your site before the change.

Respondyr

On September 15, 2026, Cloudflare changes a default that decides whether AI assistants can read your website. If your site runs through Cloudflare, or your web company moves it there, this affects you. Nobody is going to call you about it.

Cloudflare sits in front of an enormous share of the web. Plenty of small business sites run through it without the owner ever knowing, because a web designer or hosting company set it up that way years ago.

What actually changes on September 15

Cloudflare announced the change as part of a larger push around how AI systems access web content. The specifics matter, so here they are plainly.

Starting September 15, domains that newly onboard to Cloudflare get AI crawlers in the “Training” and “Agent” categories blocked by default on pages that display ads. Crawlers in the “Search” category stay allowed. Existing customers keep whatever settings they have now and can adjust them in their Security settings.

Cloudflare’s reasoning is fair on its face. AI companies crawl web content, use it to answer questions, and often send nothing back to the site that supplied the answer. Cloudflare wants site owners compensated for that, and a default block creates the pressure to make that happen. For a publisher whose revenue is ad impressions, that logic holds up.

A local business runs on different math. Your website doesn’t earn money from pageviews. It earns money when someone reads it, trusts it, and calls. A setting built to protect ad revenue can quietly cost you the phone call instead.

Why a local business should care

People have started asking ChatGPT, Claude, and Gemini the questions they used to type into Google. “Who’s a good plumber near me.” “Best med spa in town that does lip filler.” The assistant answers with a short list and a reason for each pick.

Here’s the constraint: an assistant can only cite what its crawlers can read. If your site returns a 403 to those crawlers, your pages can’t be part of the answer. Your competitor’s pages can.

The part that bothers us is how silent it is. You never chose this. Your web vendor onboards your site to Cloudflare in October, the new default applies, and your site goes dark to AI assistants without a single notification reaching you.

Blocked doesn’t mean invisible

Honesty requires a caveat here. Getting crawl-blocked is a gate, not the whole game.

AI assistants also learn about businesses from third-party pages: Reddit threads, local directories, review platforms, news mentions. A blocked site doesn’t erase you from those answers. Businesses get recommended all the time on the strength of what other people wrote about them.

What you lose is your own voice in the answer. Your services page, your service area, your own description of what you do: none of it can be read or cited firsthand. You go from primary source to secondhand mention. For a business the assistant barely knows, that’s often the difference between making the list and not.

How to check your site this week

None of this requires panic. It requires fifteen minutes and a deliberate decision.

  1. Read your robots.txt. Go to yoursite.com/robots.txt and look for these names: GPTBot, OAI-SearchBot, ClaudeBot, Claude-User, PerplexityBot, Google-Extended. A “Disallow” line under any of them means you’re telling that crawler to stay out. Maybe that’s what you want. Know that it’s there.

  2. If you’re on Cloudflare, open the AI crawler settings. They live under Security in the dashboard. Look at what’s blocked and what’s allowed, and set it on purpose before September 15. If a vendor manages your site, ask them to show you the current setting.

  3. Confirm your key pages return a 200, not a 403. Your homepage and your main service pages should load normally for AI crawler user agents. Whoever built your site can test this in a minute with a curl command and the user agent strings above.

  4. Make sure your content exists without JavaScript. Many AI crawlers don’t run scripts. If your services and hours only appear after JavaScript loads, those crawlers see an empty page. View your site with JavaScript off and check what survives.

  5. Decide your posture on purpose. Blocking AI crawlers is a legitimate choice, and some businesses have reasons to make it. Inheriting a block you never heard about is not a choice. The point of this list is that whatever your setting is on September 16, you picked it.

The visibility you control either way

Whatever you decide about your website, one surface stays readable no matter what your CDN does: your Google Business Profile and the reviews on it. That content lives on Google, not on your domain, and it’s public to everyone, people and machines alike.

It’s also where customers already look. 63.6% of consumers check Google reviews before visiting a business (BrightLocal). A profile full of recent reviews, each one answered in your voice, is the most durable public record of what your business is like to work with.

That’s the work Respondyr does. Every review that lands gets a response drafted in your voice for your approval, and Get Reviews gives you a branded link and QR code so the fresh reviews keep coming. Your website’s crawl settings deserve fifteen minutes this week. Your reviews deserve an answer every week, and that part we can take off your plate.

See the plans or get started.