Check your own site for AI readiness, by hand, in about twenty minutes.
Nine checks, a browser, no signup and no tools. Most sites that are missing from AI answers fail one of the first three, and those are the fastest things to fix in all of SEO.
Work through these in order. The early ones are the cheap ones that rule out expensive explanations, so do not skip ahead to the content questions before you know a crawler can reach the page at all.
The checklist
Don't know what these are? Don't worry, we've got you!
- 1
Can AI search crawlers reach you?
Do this: Visit yourdomain.com/robots.txt and read it. Look for any Disallow lines under OAI-SearchBot, PerplexityBot, Claude-SearchBot or Google-Extended.
Good: Either no robots.txt at all (a genuine 404), or one that allows those crawlers.
Problem: A Disallow: / under any search crawler. This removes you from that platform’s answers as surely as blocking Googlebot would remove you from Google.
Blocking GPTBot or CCBot is different. Those are training crawlers, and refusing them is a legitimate editorial choice that costs you nothing in visibility.
- 2
Can user-fetch crawlers reach you?
Do this: In the same file, look for ChatGPT-User and similar. These fetch a page at the moment a person asks about it.
Good: Allowed.
Problem: Blocked, which means an assistant cannot read your site even when a customer explicitly pastes your link and asks about you.
This is the one people most often block by accident while trying to opt out of AI training.
- 3
Is your robots.txt actually readable?
Do this: Load the URL and check it returns quickly and as plain text.
Good: A 200 with text, or a clean 404.
Problem: A timeout, a server error, or an HTML error page. Careful crawlers treat an unreadable robots.txt conservatively, which can cost you more than a missing one.
A 404 means nothing is restricted. A 500 means unknown. They are not the same thing.
- 4
Do you have a sitemap, and does robots.txt point at it?
Do this: Try yourdomain.com/sitemap.xml. Then check robots.txt for a Sitemap: line.
Good: A sitemap listing your real pages, and a Sitemap: line naming it.
Problem: A 404, or a sitemap nothing references. On a newer site with few inbound links this is your main discovery path.
Generate it from your real page list rather than maintaining it by hand, or it will drift within a month.
- 5
Does your homepage answer a plain request?
Do this: Open your site in a private browsing window with no extensions.
Good: The page loads and the important text is visible immediately.
Problem: A challenge page, a cookie wall that blocks content, or a page where the text only appears after scripts run.
Aggressive bot protection is the quiet killer here: it can turn away legitimate AI crawlers along with the bad ones.
- 6
Are your key facts real text?
Do this: Use your browser’s find function to search your own page for your phone number, your prices and the areas you cover.
Good: Find locates them.
Problem: Find cannot see them, because they live in an image, a PDF or an embedded widget. Nothing can read those reliably.
This one catches a surprising number of otherwise well-built sites, especially price lists saved as images.
- 7
Does each page have one clear title and heading?
Do this: View source and look for the title tag and the first h1.
Good: A title that describes the page and states the category, and one h1 that matches what the page is about.
Problem: The same title on every page, a title that is only your business name, or no h1.
A brand-only title on your most important page is the most common wasted opportunity on the web. We had this exact problem ourselves until September 2026.
- 8
Is there any structured data?
Do this: Search your page source for application/ld+json.
Good: At least an Organization entry identifying the business, plus something describing what the page is.
Problem: Nothing. Systems then have to infer that a name refers to a company, what it costs and who publishes it, and they sometimes infer wrong.
Mark up only what a visitor can actually see on the page. Marking up prices or reviews that are not there risks a penalty.
- 9
Does an assistant actually name you?
Do this: Ask ChatGPT, Gemini and Perplexity a question a customer would genuinely ask. Not your business name: a question where you are competing to exist, like who does this in this town.
Good: You get named, ideally with a link.
Problem: Competitors get named and you do not, or the answer is about another country entirely.
Ask more than once. These systems are not deterministic and one run tells you very little.
What to fix first
If checks one, two or three failed, stop and fix those before anything else. They are usually one line in one file, they take effect as soon as the crawler next visits, and no amount of writing compensates for a page nothing can fetch.
If they passed and you are still absent from AI answers, the cause is more likely to be structural: your pages may not answer the specific question in a form that can be lifted, or the assistants may be drawing on sources that never mention you. Those take longer and are worth doing properly.
The scan is the commodity. The timeline is not.
Be sceptical of anyone selling you a one-off AI readiness score, ourselves included. Free scanners for this are everywhere now, they mostly check the same handful of things, and today's number is worth about as much as today's weather. You can get it from the checklist above in twenty minutes without paying anybody.
What a snapshot cannot tell you is when it changed. And changing is what these things do. A plugin update rewrites robots.txt. A deployment adds bot protection that turns away legitimate crawlers. A migration drops the sitemap. None of it announces itself, none of it appears in a scan you ran in March, and the next time anybody looks is usually after a bad quarter.
So we keep the score. One scan per site per day, stored: the overall number out of 100, the four category scores behind it, and every individual check with what was found. Because it is ordinary web requests rather than paid data, running it daily costs nothing, which is precisely why we can keep a history rather than selling you a button to press.
What having the history actually gets you
- A date to blame. The score is a step function: flat for weeks, then a jump when somebody edits a file. The useful output is not the line, it is the list of which checks changed status and on what day, which is what turns “we lost AI visibility somewhere” into “this deployment, this check, this Tuesday”.
- Regression detection. A fix that quietly gets undone six weeks later looks identical to a fix that held, unless something was watching in between.
- Progress you can show someone. Not ready, Minimal, Partially ready, Agent-ready, Agent-native. A line moving between those over a quarter is a report to a client or a board. A single score is a screenshot.
- Your AI assistant can read all of it. The history is available over our MCP connector, so you can ask your own assistant what changed since March, whether the readiness dip lines up with the traffic dip, and what to fix first. That is the part a scanner page cannot do at all, because it has nothing to remember.
It is included on every plan from £29 a month, alongside rank tracking, AI visibility and the security checks, which have exactly the same shape of problem: the fix is an afternoon, and knowing it is still fixed next month is the thing worth paying for.