The AI Search Optimisation Checklist: 18 Checks You Can Run Without Paid Tools
Image Source: depositphotos.com
You can find out whether ChatGPT, Gemini, Perplexity and Google's AI Overviews can see your business in about an hour, using nothing but a web browser. Most businesses have never checked.
Ofcom's Online Nation 2025 puts an AI Overview on roughly 30% of UK Google searches, and 53% of UK adults say they see them often. A GrowthSRC study of more than 200,000 keywords found click-through rate at position one fell from 28% to 19% inside a year. Ranking still matters, it just returns fewer clicks than it used to. If you want the full picture, Slingshot Marketing has pulled together SEO statistics for trades businesses with every source named and every figure labelled UK or US.
The eighteen checks below run in six stages, from establishing a baseline to measuring whether anything you changed made a difference. Every one of them is free. None needs software you do not already have.
The AI search optimisation checklist at a glance
All eighteen checks in order. Each one is explained in full below.
|
# |
Check |
How to run it |
|
1 |
Ask three assistants a question you should be the answer to |
Put the same buyer question into ChatGPT, Gemini and Perplexity. Record mentioned, cited or absent for each. |
|
2 |
Run the same prompt three times |
Three separate chats, identical question. One run is a sample, not a ranking. |
|
3 |
Read your brand-name AI Overview |
Search your company name in Google and note every factual error in the summary. |
|
4 |
Audit robots.txt for retrieval crawlers |
Read yourdomain.com/robots.txt. Check OAI-SearchBot, Claude-SearchBot and PerplexityBot, not just GPTBot. |
|
5 |
Turn JavaScript off |
DevTools, Command Menu, Disable JavaScript, reload. Note what disappears. |
|
6 |
Confirm key pages are indexed |
Run a site: search, then check the Pages report in Search Console for crawled-not-indexed. |
|
7 |
Find content trapped in PDFs and images |
List what a buyer needs to know, then check each fact exists as HTML text. |
|
8 |
First sentence answers the heading |
Read the opening sentence under every H2. Rewrite any that delays the answer. |
|
9 |
Headings shaped like questions |
Compare each heading against how someone would type the query. |
|
10 |
One paragraph survives on its own |
Copy a mid-page paragraph and read it cold. Replace backward-pointing pronouns. |
|
11 |
Prices, coverage and purpose in plain text |
Confirm each appears in a sentence, not implied by layout or hidden behind a form. |
|
12 |
Name, description and category match everywhere |
Compare site, LinkedIn, Google Business Profile, Companies House and directories side by side. |
|
13 |
Schema matches the visible page |
Search the source for application/ld+json, or use the Rich Results Test. Check every field. |
|
14 |
Named author on trust-dependent pages |
A real person with a bio, not Admin and not the company name. |
|
15 |
Note which sources get cited |
Re-read your stage one prompts and list every domain the assistant used. |
|
16 |
Review platforms and directories current |
Listing exists, details match check 12, updated within the year. |
|
17 |
AI referral traffic in GA4 |
Filter session source for chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com. |
|
18 |
Branded search lift in Search Console |
Filter queries containing your brand and track the total month on month. |
Stage one: find out where you actually stand
You cannot fix visibility you have not measured. The first three checks establish a baseline: what the assistants currently say about you, how much that answer moves between sessions, and what Google's own AI thinks you do.
1. Ask three assistants a question your business should be the answer to
Pick the question a buyer would ask, not your brand name. Something like "who are the best managed service providers in Leeds" or "what tool handles log retention at scale". Put it into ChatGPT, Gemini and Perplexity, then record one of three outcomes for each: mentioned in the prose, cited with a link, or absent.
Mentioned and cited are different results. Being named without a link means the model knows you from training data. Being cited with a link means a retrieval crawler found your page during that session, and that is the outcome you can influence. Put the three results in a spreadsheet with the date, because you will need the baseline in stage six.
2. Run the same prompt three times
Answers vary between sessions, so a single run is a sample rather than a ranking. Open three separate chats, ask the identical question in each, and compare.
Appearing in one of three means the model can find you but does not reliably choose you. Three of three means you are the default answer for that phrasing. None of three means the rest of this list is diagnostic work rather than optimisation.
Vary the wording on a second pass. "Best X in Y" and "who should I hire for X in Y" often return different companies, and the second is closer to how people type.
3. Read what Google's AI Overview says about your brand name
Search your own company name in Google and read the AI Overview in full. In our experience it is wrong or out of date a good share of the time: an old address, a service dropped two years ago, a founder who has left, or a description scraped from a directory rather than from your own site.
The Overview is assembled from whatever Google can find, weighted towards sources it trusts. If the wrong facts appear, the correction is not made in Google. It is made on the pages the Overview is drawing from, which is what stage four is for.
Write the specific errors down now. They are the fastest thing on this list to fix and the most visible when you get them wrong.
Stage two: can the crawlers actually reach you
An assistant cannot cite a page it cannot fetch. These four checks confirm your content is reachable by the bots that matter, which are not the bots most robots.txt files were written for.
4. Check robots.txt for the retrieval crawlers, not just the training ones
Type your domain followed by /robots.txt into a browser and read the file. Two families of AI crawler exist and they do different jobs.
Training crawlers build model datasets. Retrieval crawlers fetch pages live to answer a question, and those are the ones that decide whether you get cited. Blocking a training crawler keeps you out of the dataset. Blocking a retrieval crawler removes you from the answers.
The user agents worth searching your file for:
- GPTBot and CCBot collect content for model training
- OAI-SearchBot indexes pages for ChatGPT's search feature
- ChatGPT-User fetches a page when someone clicks a citation or asks ChatGPT to read a URL
- ClaudeBot trains, Claude-SearchBot indexes for search, and Claude-User fetches on request
- PerplexityBot crawls for Perplexity's index
- Google-Extended controls Gemini training and grounding only. Ordinary Google Search is handled by Googlebot and is unaffected by it
A robots.txt written in 2023 will not name half of these, because half of them did not exist. If yours carries a blanket disallow left over from the block-the-scrapers period, that decision is now costing you citations.
Then check the layer above it. Cloudflare and most CDNs offer a one-click block-AI-bots toggle that overrides robots.txt entirely, and it is easy to have switched on without anyone recording the decision.
5. Turn JavaScript off and see what survives
Open your key pages with JavaScript disabled and read what is left. In Chrome: open DevTools, hit the Command Menu with Ctrl+Shift+P or Cmd+Shift+P, type javascript, choose Disable JavaScript, then reload.
Retrieval crawlers work under time pressure. Some render JavaScript, some do not, and the ones that do give up sooner than Googlebot. If your service descriptions, prices or location details only appear after a client-side render, assume a share of the crawlers never see them.
Anything that vanishes with JavaScript off is content you are hoping a crawler will wait for.
6. Confirm your key pages are indexed
Run a site: search on your own domain in Google and count what comes back, then open the Pages report in Google Search Console and look for anything marked crawled but not indexed.
An unindexed page can still be fetched by a retrieval crawler, but it will rarely be found by one, because the assistants search the web much as a person does. If your best explainer is not in the index, it is not in the running.
Pay attention to pages that were indexed and no longer are. That usually points at a template change or a stray canonical tag, not a content problem.
7. Find anything important trapped in a PDF or an image
List what a buyer needs before they can decide: pricing, specification, service coverage, compliance detail. Then check where each one lives on your site.
Text inside an image is invisible to every crawler on this list. Text inside a PDF is readable in principle but often skipped, and it is rarely the version that ends up quoted. Price lists, spec sheets and coverage maps are the usual offenders.
The fix is not to delete the PDF. It is to publish the same information as HTML on a normal page and keep the PDF as the download.
Stage three: is your content extractable
Extractable content is content an assistant can quote as a single passage without needing the rest of the page for context. These four checks test whether yours qualifies.
8. Check that the first sentence under each heading answers the heading
Read the first sentence under each of your H2s. If the heading poses a question and the sentence answers it, that block can be lifted whole. If the sentence is a windup, it cannot.
Assistants extract passages, not pages, and the passage that gets used is the one that stands on its own. In practice that means the direct answer sits first and the detail follows it.
Work through your five most important pages and rewrite any opening sentence that delays the answer. It is the highest-return edit on this list and it takes minutes per page.
9. Check that your headings are shaped like the questions people ask
Compare each heading against the wording of a real query. "Pricing" is a label. "How much does X cost in the UK" is a question, and it matches the way both people and assistants phrase things.
Heading text is one of the strongest signals of what a section contains. A section headed with the question it answers is far easier to match to a query than one headed with a noun.
Do not do this to every heading. Navigation headings can stay as labels. It is the sections carrying answers that benefit.
10. Test whether one paragraph survives on its own
Copy a single paragraph from the middle of a page and read it cold. Does it still make sense? Does it name its subject, or does it lean on "this" and "it" and the four paragraphs around it?
Pronouns pointing backwards are the main cause of unquotable copy. A paragraph opening "This means you will need to..." is useless in isolation. The same paragraph opening "An incident response plan needs..." can be quoted anywhere.
Repeat the subject more often than feels natural. Written for extraction, that repetition reads as clarity rather than clumsiness.
11. Check that prices, coverage and what you do are stated in plain text
Search your own site for the facts a buyer needs and confirm each appears as plain text in a sentence, not implied by a design element or sitting behind a form.
Three things go missing most often: what you charge or how you charge, where you operate, and what you do in one sentence that does not rely on your own product names. A human visitor can work these out from context. An assistant lifting a passage cannot.
Write the one-sentence version of what you do and put it in the first hundred words of your homepage. If you cannot get it into one sentence, you have found a different problem.
Stage four: do you exist as an entity
An entity is something a system can recognise as the same thing across different sources. These three checks test whether your business reads as one entity or several.
12. Compare your name, description and category across every source
Open your website, LinkedIn, Google Business Profile, Companies House and the directories you appear in, then put the business name, one-line description and category side by side.
Differences that look trivial to you are ambiguity to a machine. Ltd in one place and Limited in another. A trading name on the site and a registered name at Companies House with nothing connecting the two. Three different one-line descriptions written by three different people over five years.
Pick one version of each and change the others to match. Where the registered and trading names differ, say so on your about page, so the connection is stated rather than inferred.
13. Confirm your Organization schema matches what is on the page
View the page source and search for application/ld+json, or paste the URL into Google's Rich Results Test. Read the JSON and compare each field against what a visitor sees.
Schema that contradicts the visible page is worse than no schema at all. An old phone number in the markup, an address the business moved out of, a sameAs array pointing at a dead social account: each one is a claim that fails the moment it is cross-checked.
Schema will not make you rank. It removes ambiguity about who you are, and that matters more when the reader is a machine assembling an answer from several sources at once.
14. Check that anything trust-dependent carries a named author
Open your most important guide or opinion piece and look for a name. Not Admin, not the company name, but a person with a page saying who they are and why they know this.
Anonymous content is harder to corroborate. When an assistant weighs two pages making different claims, provenance is one of the few tie-breakers available to it.
You do not need a byline on every page. Put one on the pages where a reader would reasonably ask who says so.
Stage five: third-party corroboration
What other sites say about you carries more weight than what you say about yourself. These two checks look at the sources the assistants are already reading.
15. Note which sources get cited, and whether you are in them
Go back to the prompts from stage one and, rather than looking for your own name, read the citation list. Write down every domain the assistant used.
You will usually find the same handful: a trade directory, a review platform, one or two industry publications, sometimes a forum thread. Those are the pages the answer is being assembled from, and being absent from all of them is why you are absent from the answer.
This gives you a target list built from evidence rather than guesswork. Getting into three sources that already get cited beats getting into thirty that do not.
16. Check the review platforms and directories you already appear on
Look yourself up on the platforms that surfaced in check 15, plus the obvious ones for your sector. Check three things on each: the listing exists, the details match check 12, and it has been updated within the last year.
Dormant listings are common and they do real damage, because they tend to be the ones carrying the old address or the discontinued service that turned up in the AI Overview you read in check 3.
Claim what you can claim, correct what you can correct, and ask for duplicates to be removed. Duplicate listings are a frequent cause of conflicting information about the same business.
Stage six: measure it without kidding yourself
Most AI visibility produces no click at all, so the usual traffic reports understate it. These two checks give you the closest honest proxy.
17. Look for referral traffic from the assistants in GA4
In GA4, open Reports, then Acquisition, then Traffic acquisition, and filter session source for chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com.
The numbers will be small, and that is not the point. What matters is the trend and the behaviour, because AI referrals arrive further down the funnel than organic search: the visitor has already had the question answered and is coming to verify or to buy.
Save it as a comparison so you are reading the same view each month rather than rebuilding the filter from memory.
18. Track branded search lift in Search Console as your real proxy
In Search Console, filter queries containing your brand name and watch the total over time. Because most AI visibility ends without a click, the honest signal is people who heard of you somewhere and then searched for you by name.
Compare branded impressions month on month against the baseline you took in stage one. A rise in branded search with no matching rise in non-branded is the pattern you would expect if assistants are mentioning you without linking to you.
Treat any single AI visibility score from a third-party tool with suspicion. Nobody outside the AI companies can see real click data from AI answers, which makes those figures share-of-voice estimates, and the tools disagree with each other.
How often to run this checklist
Quarterly is enough for most businesses, plus a full pass after anything that changes what you are or where you live: a rebrand, a domain move, a site rebuild, a CMS migration, a merger, or a change of address.
Stage two is the exception. Crawler user agents change several times a year, and a CDN update can reinstate a block nobody set deliberately. Re-read robots.txt and your CDN bot settings monthly, which takes about two minutes.
Stage one is worth repeating more often than the rest, because it is the only part of this list that tells you whether any of the others worked.