I have spent the last several months running the same audit, over and over, on 1.5k+ real business websites. Not a sample of ten picked because they made a good story — 1.5k+ of them, crawled and scored the same way every time.
Most of what is wrong with these sites is not exotic. It is the same handful of mistakes, repeating across businesses that have never met, built by different developers, on different platforms, in different countries. That repetition is the interesting part. When the same problem appears on hundreds of unrelated sites, it is not bad luck — it is something the tools people use, or the advice they are given, keeps producing.
This is that list, in the order I would fix it, with the real numbers behind the ones I can measure. Where I do not have a number for something, I say so rather than reaching for one.
What “1.5k+ audits” actually means
Vague authority claims are worth nothing, so here is exactly what sits behind this.
Between 13 August 2026 and 20 September 2026 I ran a crawl-based audit — same checks, same scoring, same rules — against 1.5k+ unique domains. Most are dental implant clinics: the bulk in the United States, smaller groups in France and the UAE, plus a set of hair transplant clinics in Dubai. This is not a random cross-section of every kind of business. It is a deep, repeated look at one type of local service business, audited at volume.
Why it still applies if you are not a dental clinic: nothing on this list is dental-specific. A missing canonical tag, a blocked crawler, an unlinked profile — none of that cares what industry you are in. The platforms are familiar too. More than half the corpus (55.8%) runs on WordPress, which is roughly what you would find auditing local service businesses generally.
One limitation worth stating up front: a crawl cannot see what other websites say about a business, and part of AI visibility depends entirely on that. We leave it blank rather than guess, which means the real picture is probably slightly worse than these numbers, not better.
These businesses average 83.4 out of 100 on conventional SEO and 68.6 on AI visibility. Out of all 1.5k+ sites, 1 scored an A.
That gap — roughly 14.8 points, showing up almost everywhere — is the shape of the whole problem. Sites in reasonable order for Google are frequently doing nothing at all for AI search. Several of the items below explain why.
Mistake 1: Assuming robots.txt is the only door
This is the first thing I check, and what the numbers showed surprised me enough that I went back and re-ran them.
Deliberately blocking AI answer engines in robots.txt is rare. Across the corpus, 0.1% disallow OAI-SearchBot, 0.3% disallow ChatGPT-User, 2.3% disallow PerplexityBot. Almost nobody is choosing to shut these out.
The training-only crawlers are a different story — 3.5% disallow GPTBot. That is a deliberate, defensible decision about who gets to train on your writing, and it costs you nothing in AI answers today.
Here is the part that matters. Between 3.9% and 5% of sites returned an outright HTTP error to the crawlers that answer questions live — a 403, a challenge page, a firewall rule. Almost none of that is a robots.txt decision. It is bot-protection software, set up to stop scrapers, quietly refusing the crawlers you would actually want visiting. The owner checks robots.txt, sees nothing blocking anything, and reasonably concludes they are fine.
One caveat I will state plainly, because it cuts against my own number: about one site in ten also returned an error to a request carrying Googlebot's user-agent. That figure means almost nothing. Google verifies its crawler by IP address, not by the name in the request, so a security system refusing a stranger who claims to be Googlebot is behaving correctly. Counted strictly — sites whose own robots.txt says no — exactly 1 site in 1.5k+ disallows Googlebot. Any audit that reports the first number without explaining the second is selling you a problem you do not have.
Fix check robots.txt first, because it is free. Then check what your CDN or security layer is doing, which is where the real blocks live and where nobody thinks to look.
Mistake 2: A sitemap that exists but nothing points to
17% of the sites have no discoverable sitemap at all. A separate 9.2% — 139 sites — have one that exists and works, but robots.txt never points to it. It was built. It just was not announced.
Search engines can still find pages by following links. A sitemap is the difference between eventually and reliably, particularly for anything more than a couple of clicks from the homepage.
Fix confirm your sitemap loads, then confirm robots.txt carries a Sitemap: line pointing at it. Most platforms generate one automatically — the gap is usually that nothing ever wired the two together.
Mistake 3: A noindex tag left on a page that was meant to go live
9.2% of sites had at least one noindex tag in the crawl. Some of that is intentional — a thank-you page, a legal draft, something genuinely not meant to be found. I cannot tell from a crawl which ones were deliberate, and I am not going to pretend otherwise.
What makes this one worth checking is that a noindexed page looks completely normal to a human visitor. It loads, it reads fine, nothing is visibly wrong. It simply never appears in search.
Fix search your page source for “noindex” on anything you want found, especially after a migration or a redesign.
Mistake 4: Missing canonical tags
18.5% of sites were missing a canonical tag on at least one page. A canonical tag tells Google which version of a page is the original. Without one, near-duplicate URLs — with and without a trailing slash, with a tracking parameter, with and without www — can be treated as separate competing pages rather than one.
Fix every indexable page should carry a self-referencing canonical by default. This is a platform-level setting, not a page-by-page chore.
Mistake 5: Pages with no H1 at all
39.6% of sites have at least one page with no H1 heading anywhere on it. An H1 is the cheapest, clearest signal you can give about what a page is for. A page with no H1 is a page asking Google to guess.
Fix one clear H1 per page, stating the subject in the words a customer would use. Not the business name. Not a slogan.
Mistake 6: Two H1s doing two different jobs
The opposite problem, and nearly as common: 30.5% of sites have at least one page carrying more than one H1.
The interesting part is where it clusters. Broken out by the platform each site runs on:
| Platform | Sites | Two or more H1s |
|---|---|---|
| Wix | 21 | 66.7% |
| Squarespace | 17 | 58.8% |
| Duda | 28 | 35.7% |
| Webflow | 49 | 32.7% |
| WordPress | 839 | 30.9% |
| unknown / custom | 452 | 30.3% |
| GoDaddy | 33 | 21.2% |
| SPA (Vite/React) | 25 | 8% |
| HubSpot CMS | 14 | 7.1% |
| Next.js | 24 | 4.2% |
The hand-built sites are dramatically cleaner. Next.js sites sit at 4.2%. The drag-and-drop builders are the worst: 66.7% of the Wix sites and 58.8% of the Squarespace ones. That is what you would expect when the tool decides your heading levels instead of you.
Read the small rows carefully, though. Wix is 21 sites and Squarespace is 17, so a handful of sites moves those figures several points. The direction is solid; the exact ordering between the small builders is not. Anything under ten sites is left out of the table entirely rather than published as a percentage one site could swing.
Fix one H1. Everything else that looks like a headline is an H2 or lower.
Mistake 7: No meta description, so Google writes one for you
44.3% of sites are missing a meta description somewhere — the single most common on-page gap in the whole dataset. When you do not write one, Google pulls a snippet from the page itself. Sometimes that is a reasonable summary. Often it is an awkward half-sentence with no relationship to why anyone should click.
A meta description does not move rankings directly. It is free space in the search result, and it is the easiest thing on this entire list to fix in bulk.
Mistake 8: No structured data anywhere on the site
17.4% of sites have no structured data at all. Structured data tells a machine explicitly what a page is — a business, an article, a question and its answer — rather than making it infer that from surrounding text.
Worth separating what Google actually says here from what gets sold. Google has stated structured data is not required for its AI-generated search features. It is what earns rich results in classic search, and it is the cleanest way to state plain facts about your business. Anyone selling schema markup as an AI-visibility fix is overstating it.
Fix start with Organization or LocalBusiness — whichever is actually true — before anything more elaborate. Never mark up reviews that do not exist; that is a manual-action trigger, not a grey area.
Mistake 9: Writing that never actually answers a question
64.3% of sites have no question-and-answer content anywhere.
Most business websites describe things. “We offer implant dentistry using the latest techniques.” That is not wrong, but it is not an answer to anything either. Nobody typed a question that sentence resolves. A person scanning for the thing they need, or a system assembling an answer, has nothing to grab.
Fix take the five questions customers actually ask before booking. Make each one a heading, and answer it directly in the paragraph underneath — two or three sentences that still make sense quoted on their own.
Mistake 10: Nothing on the page is shaped to be quoted
82.9% of sites have no table anywhere on them.
This matters for a boring, mechanical reason. Featured snippets and AI-generated answers pull disproportionately from content that is already structured — a table, a numbered list, a clearly separated question and answer. A paragraph has to be interpreted before it can be lifted. A table just gets copied.
Fix wherever you are already comparing options or listing steps, put it in that literal shape instead of dissolving it into prose.
Mistake 11: Broken internal links
5.8% of sites have at least one internal link that leads nowhere. Each one is a dead end for a visitor and a wasted path for whatever ranking value that link was meant to pass along.
Corrected 1 October 2026. This said 9.8% until we found the audit was also counting links a site's firewall refused to our checker, which still work for visitors. Only links that return “not found” are counted now.
Fix the easiest item on this list, because it needs no judgment — a link either resolves or it does not. Any crawl finds every instance in minutes.
Mistake 12: Redirect chains nobody noticed
4.4% of sites have a URL that redirects to another URL that redirects again before anything loads. Each hop adds delay. They accumulate quietly over a site's life: a page gets renamed, nobody updates the links pointing at the old one, a new redirect goes on top of the old redirect.
Fix point every redirect straight at its final destination. This is a one-time cleanup, not an ongoing discipline.
Mistake 13: Images carrying information with nothing describing them
Measured across the 1,470 sites with at least one informative image: 18.1% have a real gap in alt text coverage.
Worth saying how that is counted, because it is easy to get wrong in a way that flatters the finding. A decorative image — a divider, a background texture — is supposed to have empty alt text. Counting those as failures roughly doubles the number and punishes sites for doing it correctly. We had that bug in our own aggregation and reported 37% until we caught it. The figure above counts only images that carry information and have nothing describing them.
Mistake 14: The pages that make money are buried
I do not have a clean corpus-wide number for this one, so treat it as a pattern rather than a statistic. It is one of the most common internal-linking problems in practice: the page that should be earning — a specific service, a specific location — sits three or four clicks deep, reachable only through a dropdown or a footer link nobody clicks.
Fix your commercially important pages should be one or two clicks from the homepage, through a link a person would actually notice.
Mistake 15: Two of your own pages fighting for the same search
38.3% of sites have at least one pair of pages whose titles are chasing the same search. That is 564 of the 1,473 sites where the question is answerable at all — a site with only one crawled page cannot have the problem, so it is excluded rather than counted as a pass.
The usual shape is a “services” page and a “what we offer” page covering identical ground. Instead of one strong page you get two mediocre ones splitting the same signals.
How that is counted matters, because the naive version of this check flags everything. Every title on a site repeats the brand and the category, so the comparison drops any word appearing in more than half of that site's own titles before looking for an overlap. Without that filter a homepage “matches” every service page and the number is meaningless.
One honest limit: this measures that two titles target the same thing. Only Search Console can prove the two pages are actually competing — one query, two of your URLs, trading positions. Treat the figure as risk, not proof.
Fix confirm it in Search Console, then keep the stronger page, fold the other into it, and redirect.
Mistake 16: The same page with the city swapped
The one I watch most carefully, because it carries real enforcement risk. Google's scaled-content-abuse policy targets exactly this: many pages that are barely differentiated, built to catch a keyword variant rather than to say something new.
The classic version is a location-page template — same three paragraphs, same structure, city name swapped. It is tempting because it produces pages fast. It is also easy for a similarity check and a human reader to spot, and the policy judges the page, not whether a person or a model wrote it.
Fix if the argument on a page would read identically with the location swapped, it is a template, not a page. Give it something specific to that place, or merge it.
Mistake 17: Nothing on the site is signed by a person
No number here, and the reason is worth stating precisely rather than hand-waving. Our audit does check for a byline — it just started checking after this corpus was crawled. Exactly one report on disk carries the field. Producing a real percentage would mean re-crawling all 1.5k+ sites, not re-reading what we already have, so there is nothing honest to publish yet.
It is one of the most consistent gaps I see, though, and it matters more than it used to. Google's quality guidance names experience and expertise explicitly, and both attach to a person, not a company. A page with no byline and no named author gives a reader — and a system weighing what to repeat — nothing to evaluate.
Fix put a real name on anything carrying a judgment. A person, not “our team.” If they hold a relevant qualification, say so in text, not inside an image.
Mistake 18: Profiles that exist but were never linked back
47.5% of sites have no sameAs links in their structured data. Even where a business plainly has a LinkedIn page, a Facebook page, a directory listing, nothing on the website formally connects them.
sameAs is how a machine confirms your website and those profiles describe the same entity. Without it each one sits in isolation, and nothing lets a system trying to verify who you are draw a confident line between them.
Fix this only requires listing profiles you already have. Minutes of schema work, not new content.
Mistake 19: Nobody else has ever written about the business
The starkest number in the dataset: 96.7% of sites have no press mentions, no “as featured in,” nothing indicating an independent source has ever written about them.
It is also the hardest to fix and the most worth taking seriously. Everything else on this list is something you control on your own site. This one is not. Independent mentions are consistently among the strongest signals both classic search and AI systems use to decide whether to trust and repeat a claim, and you cannot generate them by editing your own pages.
Fix a programme, not a task. Local press, industry roundups, original data worth citing. Slow — and the one thing a competitor cannot copy overnight the way they can copy a meta description.
Mistake 20: A site slow enough to lose people before they read anything
Our audit marks performance as inferred rather than measured — a single request from one location is not field data — so I am not going to attach a corpus percentage to it.
What is worth saying: the direct ranking effect of speed is smaller than most people assume. The real cost is behavioural. A slow page loses the visitor before they see whether the content was any good, and no amount of good SEO elsewhere recovers that.
Fix run your own homepage through PageSpeed Insights rather than trusting a general claim. Note that INP replaced First Input Delay as the responsiveness metric in 2024, so older advice is measuring the wrong thing.
Mistake 21: Local listings that do not match the website
Last, and specific to any business serving a physical area. Your name, address and phone number need to match exactly everywhere they appear — website, Google Business Profile, every directory. A mismatch as small as a suite number in one place and not another quietly undermines the trust signal local search depends on.
This is established practice rather than something I can prove from this dataset, and I would rather label it that way than dress it up.
Fix pick the canonical version of your details and make every listing match it precisely, not approximately.
Where I would actually start
Twenty-one is too many to look at once. If I were fixing one of these sites myself, in this order:
First, the things that make you invisible outright. Check robots.txt and your security layer for blocked crawlers, confirm a sitemap exists and is declared, check for a stray noindex. An hour, combined, and it is the difference between being findable and not.
Second, the on-page fundamentals. H1s, meta descriptions, canonical tags. Boring and fast, and they compound — fixing one page teaches you the pattern for the other fifty.
Third, the structural work. Real schema, content shaped as actual answers, internal links pointing at the pages you want found. Slower, but it holds.
Last, the slow ones. Earning real mentions, building genuine attribution, fixing duplicate-content problems properly. These are not afternoon jobs, and anyone who tells you otherwise has not done one.
I have audited 1.5k+ of these sites now and still find something from this list on almost every one — including, before we fixed it, our own. The full aggregate data behind every number above is on our research page, including the parts where the sites did well.