Before an engine can name a page in an answer, it has to fetch the page, decide what it is about, cut it into passages and pick one to quote. Every step can go wrong in the markup, and most of the ways it goes wrong can be seen by pressing Ctrl-U. These are the fifteen rules our site audit runs on every page it reads, in the order its queue puts them.
How the list was drawn
A rule is on this list only if two things are true of it: a person can check it from the page source, and its fix fits in a diff. That rules out a lot of real advice. "Improve the introduction" is sound and cannot be decided from a shape; "move the author bio into an aside" is a matter of taste. A finding nobody can verify costs more trust than it is worth, so neither is here.
The rules are also written to be fixed once. A finding is grouped by rule, not by page — "add alt text to the images that have none" is one job whether it touches one page or forty — because a fault that shows up on forty pages almost always lives in one template. Each rule carries a severity from one to five, and the order below is the order a queue should be worked in: a page whose content is not in its HTML has no headings worth fixing yet.
Rule 1: the content has to be in the HTML
Serve the content in the HTML. Severity 5. The check fires when a page answers with at least 20,000 bytes of markup and fewer than 120 words of text: the content exists, it is just being assembled in the browser. Most of what feeds these engines reads the HTML directly and does not run scripts, so to them the page is empty — and nothing else on this list matters for a page in that state. Check it with JavaScript off, or with curl; move the body copy into the server-rendered document. The interactive parts can stay client-rendered. The prose cannot.
Rules 2–7: one address, one subject, one shape
2. Point each canonical at the page itself. Severity 5, the joint highest. A canonical tag names the real address of a page, and everything the page earns is credited there. When it points somewhere else — often a template hard-coding the home page or a staging host — the engines quote the page under another name, sometimes on a site you do not control. The fix is to build the tag from the page's own address, then check a few sibling pages, because a wrong canonical is almost always a template.
<!-- before -->
<link rel="canonical" href="https://example.com/" />
<!-- after -->
<link rel="canonical" href="https://example.com/blog/this-post" />3. Declare a canonical on every page. Severity 3. Without one, every address that reaches the page — with a tracking parameter, with and without a trailing slash, through a staging host — is a separate page to a crawler, and the citations split across the copies. Make it absolute, not a path, and keep your analytics parameters out of it.
4. Mark where the content begins and ends. Severity 3. A main or article element tells an extractor where the page's own content stops. Without either, the boundary is guessed, and the guess usually takes the navigation, the cookie banner and the footer along — so the passage that gets quoted is padded with your menu. Either element passes; keep the navigation, sidebar and footer outside it.
5. Give every page an H1. Severity 3. With no H1, the subject is inferred from whatever the chunker reaches first — the title tag, the URL, or, on templates where the visible title is a styled div, the navigation. Mark the text a reader would call the title as an H1.
6. Cut each page down to one H1. Severity 2. Two H1s are two answers to "what is this page", and the one that wins is usually the first — on most templates, the site name. Keep the heading that names the subject and demote the rest to H2; the styling can stay identical. It is the cheapest fix on the list.
7. Step heading levels down one at a time. Severity 2, checked on pages with three or more headings. Levels are how a document says which section contains which. An H2 followed by an H4 breaks that, and a chunker will attach the orphaned section to the wrong parent or to nothing — so the quoted passage arrives without the context that made it correct. If a level was skipped to get a smaller font, change the font.
<!-- before -->
<h2>Pricing</h2>
<h4>What the trial includes</h4>
<!-- after -->
<h2>Pricing</h2>
<h3>What the trial includes</h3>Rules 8–10: dates an engine can read
8. Publish a machine-readable publish date. Severity 3, on written pages only — articles, comparisons, alternatives pages, how-tos and lists, and not reference documentation, where an undated page is not a stale one. "September 2026" in a paragraph is a date only a reader can use. When an answer has to choose between two pages that disagree, recency is one of the few tie-breakers it has, and an undated page does not get to compete on it. Wrap the date in a time element with an ISO datetime, or put datePublished in the structured data.
<time datetime="2026-09-11">11 September 2026</time>9. Add a last-updated date where there is a publish date. Severity 2. A page that declares when it was published and never says whether it has been touched since looks the same as one abandoned three years ago. Add dateModified beside datePublished and wire it to the content's real last edit — not the build date. A nightly build that stamps today on everything is worse than no date at all.
10. Keep the year in a time-sensitive title current. Severity 4, on comparisons, alternatives pages, how-tos and lists. A title that says "best tools of" a year that has passed is now a statement against you: a model choosing between two lists has been handed the reason to pick the other one. Bring the content up to date first — changing the year alone makes the title a false claim — then the title, the H1 and the description. Keep the URL; a new slug throws away the links the page has earned.
Rules 11–13: say what kind of page it is
11. Add the Open Graph type. Severity 2. og:type is the one Open Graph field that says what kind of thing a page is rather than what it looks like. Missing, it defaults to "website" everywhere the page is unfurled, so a substantial post is indexed with the same shape as your home page. Use article for posts and website for the marketing pages.
12. Mark up the FAQs you already have. Severity 3, checked when a page has three or more headings phrased as questions and no FAQPage schema. The page is already an FAQ; the schema pairs each question with its answer explicitly instead of leaving them to be paired by proximity. Mark up only questions the page genuinely answers.
13. Add a contents list to long pages. Severity 2, checked on pages over 1,200 words with four or more H2 sections and fewer than three in-page links. A long page with no contents has no addressable parts: an answer that wants to point at one section can only point at the whole page, and the reader it sends lands at the top. Give each H2 an id, link them at the top, and keep the list in the HTML rather than building it in a script — the one at the top of this post is there because this rule asked for it.
Rules 14–15: what a crawler cannot see
14. Add alt text to the images that have none. Severity 2. Nothing that reads the page can see its images, so alt text is the only description of them in the HTML. A diagram carrying the point of a section is, downstream, a blank. Write one sentence for each image that carries meaning, describing what it shows rather than naming the file, and give purely decorative images an empty alt, which is the correct answer for them.
15. Say who each quote came from. Severity 1, checked when two or more quotes on a page name nobody. A blockquote without attribution in the markup is a sentence that appears to be yours, so when it is quoted onward it is quoted as your claim — losing the credibility of a third party saying it, which is why the quote is on the page. Put a cite inside each blockquote; for a customer, name the company as well as the person.
The checklist at a glance
| Rule | Fails when | Severity |
|---|---|---|
| 1. Content in the HTML | ≥20 KB of markup, <120 words | 5 |
| 2. Canonical points at itself | It names another address | 5 |
| 3. Canonical declared | No canonical tag | 3 |
| 4. Content boundary | Neither main nor article | 3 |
| 5. An H1 | No H1 at all | 3 |
| 6. One H1 | Two or more H1s | 2 |
| 7. Heading order | A level is skipped | 2 |
| 8. Publish date | None machine-readable | 3 |
| 9. Updated date | datePublished, no dateModified | 2 |
| 10. Current year in title | Title names a past year | 4 |
| 11. og:type | Tag missing | 2 |
| 12. FAQ schema | ≥3 question headings, no FAQPage | 3 |
| 13. Contents list | >1,200 words, ≥4 H2s, <3 anchors | 2 |
| 14. Image alt text | Any image without alt | 2 |
| 15. Quote attribution | ≥2 quotes naming nobody | 1 |
Held to the same rules
It would be an odd thing to sell an audit from a site that failed it. Every page on recomma.ai that a crawler can reach is rendered in our test suite exactly as the build renders it and run through these same fifteen rules; a page that breaks one fails the suite. This post is one of those pages.
In the product, the rules run across every page in your sitemap, not only the ones somebody has already quoted — the pages with problems are usually the ones nothing has cited yet. The Site page shows what was read and which rules fail, with each failing page's markup as it is and as it should be, and a rule that touches enough of the site becomes a ticket on the same queue as everything else, re-measured fourteen days after it closes. The plans say what each tier tracks.
Find out what AI says about your brand
Add a domain and the first pass runs while you read the next one.