3 levels · read in order
The QueryWin handbook
Everything we know about ranking on Google and getting recommended by AI, in the order you should learn it.
- Level 1
Foundations
For readers new to SEO and GEO.
- 1How to check if AI can read your site (in one command)To check if AI can read your site, send one request per AI crawler and read the status code. Here is the command, the twelve crawlers worth testing, and the two you cannot test at all.
- 2robots.txt for AI crawlers: the rule that breaks most filesA robots.txt for AI crawlers names each crawler and states what it may fetch. Most files break on one rule: a named group and User-agent: * are never combined.
- 3When your CDN is blocking AI crawlers and your robots.txt is notIf your CDN is blocking AI crawlers, one response header tells you which vendor to go to. Here is how to read it, what eleven live sites returned, and where each platform hides the control.
- 4Striking distance keywords: the Search Console filter and what to fix firstStriking distance keywords are search terms where you already rank 5 to 15. Here is the exact Search Console filter that lists yours, and the rule for which one to fix first.
- 5What makes content citable by AI: the extract test and seven fixesWhat makes content citable by AI is whether one passage can be lifted out and still be true on its own. Here is the test, seven fixes ordered by effect, and the same claim written twice.
- 6Record your baseline before you change anythingRecord your baseline before you change anything, or the 14 day check has nothing to compare against. Five things to save, one command, and where to put them.
- 7A rewrite worksheet for one page: eight edits, three things not to touchA rewrite worksheet for one page: forty minutes, eight ordered edits, and three things that must not change. Includes what Google says about editing dates.
- 8Get your page indexed faster: sitemap, Search Console, IndexNowTo get your page indexed faster, do three things in order after publishing, then verify at 48 hours. None of them force indexing, and Google says so in writing.
- 9Did it work? The 14 day check, and what each result meansFourteen days after publishing, compare the same four numbers against your baseline. Three outcomes, three different next moves, and what this check cannot tell you.
- Level 2
Practice
Step-by-step playbooks you can follow immediately.
- 1The GPTBot user agent, and the seven other names in your robots.txtThe GPTBot user agent is one of eight names worth addressing in robots.txt, and three of the eight never send a request at all. Here is what each one costs you, with the full agent strings and a block you can paste.
- 2How to verify Googlebot (and every other crawler) from your own logsTo verify Googlebot you need two DNS lookups, not a user agent string. This chapter gives the reverse-then-forward procedure Google documents, plus a table of the nine published crawler IP files and one script that checks an address against all of them.
- 3Log file analysis for SEO: how to see which AI crawlers actually visitedLog file analysis for AI crawlers reads two fields your server already keeps: the User-Agent and the IP. This chapter gives the one command that counts every agent, the field table, the token list, and how to verify which lines you can believe.
- 4Unblock AI crawlers: which layer your platform actually lets you editTo unblock AI crawlers you have to fix the layer that is refusing them, and a hosted platform decides which layer you may touch. Four platforms, their own documentation, and the two fixes that silently do nothing.
- 5Can AI crawl JavaScript? How to check what the crawler actually gotCan AI crawl JavaScript? Mostly not — which is why a page can return 200 and still say nothing. Here is the one command that shows what the crawler received, and three fixes with their costs.
- 6Content security policy and SEO: is your own header blocking Google's renderer?Your content security policy is enforced inside Google's renderer, because that renderer is a headless Chromium. Here is the three-step check that rules CSP in or out as the reason your page renders empty, and the point where you should stop.
- 7Mobile first indexing: how to check the version Google actually keepsMobile first indexing means Google indexes and ranks the mobile version of your page. This chapter is a two-fetch procedure — request the same address as Googlebot Smartphone and as Googlebot Desktop, then compare text volume, structured data, robots meta and image alt text — plus the seven-row parity checklist Google publishes.
- 8http-equiv meta tags: which values still do something, and what to use insteadhttp-equiv is the attribute that lets a meta element stand in for an HTTP response header, and most values people paste into it do nothing. Of the seven the HTML standard defines, four still do something and two are the ones Google reads. This chapter is the routing table.
- 9Mixed content: how to find insecure requests and fix them before the browser blocks themMixed content is an insecure request made by a page loaded over HTTPS. Browsers upgrade images, video and audio on their own and block every other insecure request, so one http:// URL can quietly stop a script or a download. Here is the audit command and the fix table.
- 10llms.txt: what it is, a template, and what it cannot doAn llms.txt lists your most useful pages so an agent can find them without crawling everything. Here is the format, a template, who actually publishes one, and the five things it does not do.
- 11How to use canonical tags: three signals, ranked by how strongly Google reads themA canonical tag names the URL you want kept when several serve the same content. Google ranks it against redirects and your sitemap, so the tag only wins when the other signals agree — here is the paste-able block, the self-check command, and the four ways it silently fails.
- 12URL structure for SEO: the six rules Google publishes, and one command to audit yoursA good url structure is readable words, hyphens, few parameters and one casing. Google publishes it as a crawling requirement, not a style preference — plus a command that audits every address in your sitemap.
- 13301 vs 302 redirect: which code Google keeps, and when each one is rightThe 301 vs 302 redirect choice decides which address stays in search results, because Google treats permanent codes as a canonical signal and temporary ones as a reason to keep the source page indexed. A decision table by situation, three curl checks, and the three ways it goes wrong.
- 14Subdomain vs subfolder: four things that split the moment you pick the subdomainSubdomain vs subfolder is a hosting decision with four documented consequences: robots.txt scope, crawl budget, Search Console property and sitemap all draw their boundary at the hostname. None of the four is a ranking claim.
- 15www vs non-www: pick one hostname, then make five places agreewww vs non-www is a decision about which hostname is your site, not a ranking question. No Google documentation says either form ranks better. What is documented: the two are separate robots.txt scopes, separate Search Console properties, and duplicate URLs until a redirect resolves them. Decision table, a four-line check and a six-row consistency table below.
- 16Schema markup for AI search: which types are still liveSchema markup for AI search is worth adding, but two of the most-recommended types no longer produce rich results. Here is what Google still supports, four blocks to paste, and the policy line that gets sites penalised.
- 17Article schema: the five properties Google actually readsArticle schema in one JSON-LD block: Google's documentation lists five supported properties and marks none of them required — here is what each one does, what the other hundred-odd schema.org properties buy you, and what shipping it will not change.
- 18FAQ schema in 2026: the rich result is gone, the markup is notFAQ schema stopped producing a rich result in Google Search in May 2026 and lost its documentation a month later. The type still validates, so the question is now whether a machine-readable question list earns its place — and this is how to decide.
- 19How to read the Rich Results TestThe Rich Results Test answers two questions: could Google fetch and render the page, and does its structured data qualify for a supported rich result type. It does not promise the result will appear. Here is how to run it, how to read each verdict, and where it stops being useful.
- 20Product structured data: two rule sets, one decision, and the price that must be above zeroProduct structured data starts with one decision: can the visitor buy on this page? Google publishes two requirement sets — merchant listings for pages that sell, product snippets for pages that review or compare — and what is optional in one is mandatory in the other. A comparison table, two pasteable JSON-LD blocks, an audit command, and the three failures that pass a validator.
- 21How to find keywords for a new website with no Search Console dataHow to find keywords for a new website before Search Console has data: three sources that need none of your own traffic, and a five-check scoring table that keeps only phrases scoring three or more.
- 22Search intent: how to tell what a search term is actually asking forSearch intent is what the person typing a term is trying to accomplish, and it decides which page type can win that term before any writing happens. Google's rater guidelines name four kinds — Know, Do, Website, Visit-in-person — and reading them off a results page takes about a minute per term.
- 23Search Console verification: five methods, and two questions that pick one for youSearch Console verification proves you control a site before Google shows you any of its data. Five methods exist, only a DNS record verifies a Domain property, and every other method covers exactly one protocol-and-hostname spelling.
- 24Read the Search Console performance report as five shapesThe Search Console performance report gives four numbers and no diagnosis. Five recognisable shapes cover almost everything a page can be doing wrong, and each one has a different next action.
- 25Why is my website not ranking: six causes, one check eachWhy is my website not ranking is six different problems, and five of them are told apart by a single Search Console reading each. Here is the check for every cause, in the order to run them, and when to stop looking.
- 26Internal linking strategy: one destination per topic, and anchor text that stands aloneAn internal linking strategy comes down to two decisions per page: which page it points at, and what words carry the link. Hub and spoke, a four-column audit sheet, a command that finds orphan pages, and Google's own test for anchor text.
- 27How to write anchor text: the href, the words, and the two fallbacksHow to write anchor text Google can actually use: a real href on an <a> element, a description of the destination between the tags, and — when the text is missing — the only two fallbacks Google's documentation names.
- 28Breadcrumb schema: four properties, and the one you should leave outBreadcrumb schema in one JSON-LD block: a BreadcrumbList with at least two ListItem entries, each carrying position, name and item — plus the one property Google tells you to omit on the last step.
- 29Nofollow vs dofollow: there is no dofollow attribute, and three values that are realNofollow vs dofollow is a choice between no rel attribute at all and one of the three values Google documents — sponsored, ugc and nofollow. There is no dofollow attribute in HTML.
- 30How to find orphan pages on your siteOrphan pages are pages no internal link points to. Google can still discover them through a sitemap, but they arrive with no anchor text and no place in your structure. Find them with one subtraction, then triage each one into link, merge, redirect or remove.
- 31How to get cited by ChatGPT, Perplexity and Google's AI surfacesHow to get cited by ChatGPT starts with what each vendor actually publishes: not ranking logic, but which crawler feeds which surface. Here is that table for four engines, plus a five-step check you can run from outside.
- 32How to rank in people also ask: collect the question, do not invent itHow to rank in people also ask starts with clerical work, not cleverness: collect the exact phrasing people type, give each question exactly one owning page, and answer it in the 80 words under the heading.
- 33How to structure content for AI: the section, not the page, is the unitHow to structure content for AI starts from what an extractor needs: where a passage begins and ends. Six checkable rules for headings, opening lines, tables and FAQ phrasing, plus a five-minute audit.
- 34Semantic HTML for search and AI: which element each part of the page should useSemantic HTML gives every region of a page a name, and a crawler reads that tree before it reads your text. This chapter gives the element-per-region table, a one-line landmark audit, and the claim we refuse to make: Google says search can rarely depend on semantic meanings, so markup is a correctness fix, not a ranking lever.
- 35H1 tag SEO: what the heading level decides, and where your h1 actually gets readH1 tag seo is one decision per page: which single heading names the whole thing. Google has published that heading order does not affect Search and that no ideal number of headings exists — while listing heading elements among the nine sources it can build a result headline from. The audit command, a seven-row reading table, and three fixes.
- 36Title tag SEO: nine sources Google can build your headline fromTitle tag seo is writing the one element Google reads first and can still overrule. Google publishes nine sources it may build a result headline from, seven situations where it replaces yours, and no length limit at all. The audit command, the seven-symptom table, and the five steps.
- 37How to write a meta description Google might actually useHow to write a meta description, given that Google says it builds snippets from page content and only falls back to your tag when it describes the page better. Six checks, each traced to a line in Google's documentation, plus what 27 large homepages actually ship.
- 38Alt text for SEO: what Google reads, and the four cases where you write nothingAlt text for SEO is a description written for someone who cannot see the image. Google treats a keyword list there as spam. Four kinds of image take an empty attribute instead, and omitting the attribute means something different again.
- 39Image seo best practices: six attributes, and the two a search engine readsImage seo best practices come down to six attributes on one tag, and a search engine reads only two of them. Alt text and the file name carry meaning; srcset, sizes, width/height and loading decide how fast the picture arrives.
- 40Open graph image size: one file that clears every surfaceOpen graph image size has one answer that satisfies every surface at once: 1200 x 630 pixels, JPG or PNG, under 5 MB. One file clears Meta's recommended dimensions and X's card limits, and works as the fallback for both.
- 41How to build a competitor comparison page that gets citedA competitor comparison page gets cited when every row is a checkable, dated, attributed fact. Here is the eight-block skeleton, the sourcing rule for each of the four claim types, and the two cases where you should not build the page at all.
- 42Entity SEO: make an answer engine say who you are, correctlyEntity SEO comes down to one structured data block, a short list of profiles you control, and writing your own name identically in seven places. Here is the block, the checklist, and how to see what an engine currently believes.
- 43Author schema: how to attach an article to a real personAuthor schema is the property that ties an article to a named person instead of to nobody. Here is the JSON-LD block, the seven places the name has to match, and what it does not buy you.
- 44IndexNow setup: key, first push, and how to read the responseIndexNow setup takes a key string, that key hosted as a text file, and one request per publish. Here is the whole path, plus the five response codes and what each one means.
- 45How to read the URL Inspection tool, line by lineThe URL Inspection tool reports what Google has on file for one page, not what your browser sees. Three of its lines decide whether the page is indexed, and this is how to read them in order — plus why a passing live test proves less than people think.
- 46Bing Webmaster Tools: setup, the AI Performance report, and what it cannot tell youBing Webmaster Tools setup takes one verification step out of five, and the report worth coming back for is AI Performance: how often your pages are cited in Microsoft Copilot and in Bing's AI answers.
- 47How to read the page indexing report, one reason at a timeThe page indexing report is the only place Google tells you why it left pages out, across the whole site at once. Read it as two questions: are your key pages in the indexed bucket, and is every reason in the not-indexed bucket one you chose. A ten-row triage table with Google's own definition of each status is below.
- 48Remove URL from Google: four situations, one tool each, and how long each one holdsRemove URL from Google is four different jobs. Your page, gone today: the Removals tool, within a day, for about six months. Your page, gone for good: 404 or 410, a password, or a noindex the crawler can read. Not your page: the Refresh Outdated Content form. A decision table, a thirty-line pre-removal check that applies Google's robots.txt precedence, and the six ways a removal quietly fails.
- 49How to create a sitemap: the four fields, and the two Google ignoresHow to create a sitemap comes down to one UTF-8 file of absolute URLs at your site root. Two of the four per-URL fields are ignored by Google outright, a third is used only if it survives being checked against the page — here is the paste-able file, the index form, and a count you can run after every deploy.
- 50When to split a sitemap, and how to write the sitemap indexA sitemap index lists other sitemaps instead of pages. Google forces one at 50,000 URLs or 50MB uncompressed, but most sites should build one earlier — split along the line you want to read a coverage number against. Template, limits table and the four quiet failure modes.
- 51UTM parameters: three you always use, four optional, and two that go nowhereUTM parameters are the labels you attach to a link so analytics can say where a visit came from. Google Analytics recognises nine, three of them are effectively required, and two are accepted and then not reported at all.
- 52GA4 direct traffic: pulling AI referrals back out of the bucketGA4 direct traffic is where a visit goes when nothing told Analytics where it came from, which is where assistant clicks often land. Here is the custom channel group that recovers part of it, and an honest account of the part it cannot.
- 53Search Console API: four services, two hostnames, and the row limit that makes it worth usingThe Search Console API is free and returns the same numbers as the interface, with one difference worth the setup: 25,000 rows per request instead of a thousand. It is four services on two hostnames, authenticated with OAuth rather than an API key, and its quotas are charged by query shape rather than by call count.
- 54PageSpeed Insights API: two sets of numbers in one response, and one of them is movingThe PageSpeed Insights API returns field data from real users and lab data from Lighthouse in the same JSON. Here is which block is which, the two parameters people miss, and where the field half is moving to.
- 55AI share of voice: measure it by hand with a fixed prompt setAI share of voice is the fraction of a fixed prompt set whose answers name you. Here is how to build the list, the seven-column record sheet, and the four things that invalidate a comparison.
- 56Content refresh SEO: sort the page into one of four verdicts firstContent refresh SEO fails at the sorting step. Refresh, merge, differentiate and delete are four different operations — here is how to tell them apart on evidence, and the three obligations each one creates.
- 57Multilingual SEO best practices: pick the languages before the tagsMultilingual SEO best practices usually start at the hreflang tag. The expensive decision is earlier: a four-row score that ranks candidate languages against data you already have, and a table of the page types not worth translating at all.
- 58Hreflang tags: how to write them, and the one rule that breaks a whole setHreflang tags only work as a set: every version must list every other version including itself, and one missing return link can make Google ignore the group. Here is how to write the codes, where the three placements go, and a command that checks reciprocity before you ship.
- 59Resource hints: how to preload, preconnect and prefetch what the first screen needsResource hints tell the browser about a resource before it discovers it. Here is which of the five types to use for what, a copyable decision table, and a command that lists every hint your page already ships.
- 60Largest contentful paint: how to measure it and fix the slowest partLargest contentful paint is the render time of the biggest element in the viewport, and the target is 2.5 seconds or less. Here is how to find the element, split the four subparts, and fix the one that dominates.
- 61Interaction to next paint: how to measure it and fix the slow interactionInteraction to next paint measures how long a page takes to respond to a click, tap, or keypress, and the target is 200 milliseconds or less at the 75th percentile. Here is how to find the slowest interaction and fix the part that causes it.
- 62Cumulative layout shift: how to find it and stop itCumulative layout shift is the Core Web Vital for visual stability, and the target is 0.1 or less at the 75th percentile. This chapter is a four-step way to find the element that moves, name it in the audit, and reserve its space — including the command we ran on BBC News.
- Level 3
Advanced
Trade-offs, failure modes, and data measured on live sites.
- 1AI Mode Search Console reporting: what the generative AI report answers, and what it doesn'tAI Mode Search Console reporting arrived as the generative AI performance report, and it counts one thing: impressions. No search terms, no clicks, no split between AI Overviews and AI Mode. Here is what it can answer, and how to rebuild the rest at page level.
- 2How to get mentioned by AI: five off-site sources, and the one you cannot buyHow to get mentioned by AI is a problem on other people's domains: an answer engine only names what it has seen named elsewhere. Five source types, what each costs, and a mention record you can keep.
- 3Keyword difficulty: three signals that say walk away, and one that says try anywayKeyword difficulty scores are vendor predictions, not Google measurements. Three things you can check yourself decide whether to abandon a search term: who owns the result page, the title-match ratio, and whether anyone is searching it at all.
- 4Keyword cannibalization: find the collision, then pick one of three fixesKeyword cannibalization has no penalty attached. Your signals split, crawling is spent twice, and Google picks the page for you. Here is how to find collisions in Search Console and decide between merging, differentiating and removing.
- 5SEO attribution: how to build one number when two systems disagreeSEO attribution breaks because Search Console and analytics do not share a key. Chain four ratios that each live inside one system, record the gap between them as a constant, and report the inputs with the number.
- 6Should I block AI crawlers? It is three decisions, not oneShould I block AI crawlers splits into three separate calls: AI search indexing, model training, and user-triggered fetches. Here is the decision tree, what each site type should pick, and why we cannot price the bandwidth side yet.
- 7When to noindex a page, and why disallow does the oppositenoindex keeps a page reachable and out of search results — but only if the crawler is allowed to fetch it. Blocking the URL in robots.txt means the rule is never read. Here is the four-tool decision table, both implementations, and the debug order that starts with robots.txt.
- 8data-nosnippet and the three page-level snippet controlsdata-nosnippet is the only one of Google's four snippet controls that works on part of a page. This is what each of the four does, which two also limit AI Overviews and AI Mode, and why most sites should use none of them.
- 9Website builder SEO: the five layers a hosted platform decides for youWebsite builder SEO comes down to five layers: you own one outright, share two, and never get the last two. This chapter names them, gives the substitute for each, and puts moving platform last.
- 10Crawl budget: Google's two size thresholds, and the report that overrules bothCrawl budget is real and Google says most sites should not think about it. Two size thresholds and one Search Console status decide whether it applies to you, and this chapter is the four-step check plus the fix list ordered by effort.
- 11Does page speed affect seo? Two documented effects, and one number to checkDoes page speed affect seo? Yes, in two documented ways that belong to two different systems: Core Web Vitals feed the ranking systems as one signal among many, and server latency changes how much of your site Google crawls.
- 12Faceted navigation: which filter URLs to let a crawler haveFaceted navigation earns crawling only where somebody searches for the thing a filter filters to. Google publishes four mechanisms; only two of them stop the fetch. Here is a decision table for the five kinds of filter URL, plus a robots.txt pattern set tested against real addresses before it ships.
- 13Pagination SEO: give every page its own URL, and never canonicalize to page onePagination SEO comes down to one rule: every page in a sequence is a separate page. Give each its own address, a canonical pointing at itself, and a real link from the page before. Five ordered steps, a head block you can copy, a five-row decision table and three curl checks.
- 14304 Not Modified: make a re-crawl cost one header instead of a whole pageA 304 Not Modified tells a crawler nothing changed, in a response with no body. Google supports two validator pairs, strongly prefers ETag, and reports that 0.017% of its fetches are cacheable. Four steps, a two-request check, and a five-row decision table.
- 15How long does it take to get indexed, and when to stop waitingHow long does it take to get indexed is four questions with one published answer. Google gives a range for crawling and none for inclusion. Here is how to tell which stage you are stuck at, and the deadline for each.
- 16Google Indexing API: who it is actually for, and what to use insteadThe Google Indexing API accepts two page types — job postings and livestream broadcast events — and nothing else. Here is the eligibility test, the quota, why plugins say otherwise, and the three routes that are open to every other kind of page.
- 17Does changing a URL affect SEO? Yes — and four quieter things a rewrite breaksDoes changing a URL affect SEO more than the rewrite itself? Almost always. Lock the URL, the canonical, your internal links, the structured data and the title before you edit anything, and know in advance which symptoms mean roll back.
- 18404 vs 410 for a page you removed, and the two cases where neither is right404 vs 410 is not a decision for Google — its documentation treats every 4xx except 429 identically. The choice that changes the outcome is between a 4xx, a 301 and a 503, and this chapter is the three questions that pick one.
