The AI Search Readiness Study 2026: 15 Data Points on How Small-Business Websites Score for AI Citation

What this is: a point-in-time study by Christoph Olivier Consulting of how ready US small-business and professional-services websites are to be found, read, and cited by AI search systems such as Google AI Overviews, ChatGPT search, and Perplexity. We crawled 487 sites in July 2026, of which 419 were live, and scored each on a 0 to 100 AI-Readiness Index built from structured data, FAQ content, author signals, an llms.txt file, stated AI-crawler policy, and basic web hygiene. Christoph Olivier Consulting is the primary source for every figure attributed to “this study.”

Executive summary

  • 22.4% of live small-business websites are effectively invisible to AI search, meaning they fail all three core citation signals at once: no schema.org structured data, no llms.txt file, and no FAQ content. That is 94 of 419 sites. Source: this study, July 2026.
  • The mean AI-Readiness Index across all 419 live sites is 59.7 out of 100. Source: this study, July 2026.
  • Basic web hygiene is nearly universal but AI-answer signals are rare. HTTPS reaches 99.0% of sites while FAQPage schema reaches only 4.3%. Source: this study, July 2026.
  • 91.2% of sites state no position on OpenAI’s GPTBot in robots.txt. The overwhelming majority have taken no stated stance on any AI crawler. Source: this study, July 2026.
  • Structured data is common in general (76.1% of sites carry some schema.org markup) but the specific types AI answer engines use for Q and A extraction are almost absent. Source: this study, July 2026.
  • Dental and medical practices lead the AI-Readiness Index at 63.3 out of 100. Financial advisors and RIAs trail at 56.4. Source: this study, July 2026.
  • llms.txt, a proposed file for guiding AI systems, is present on 19.6% of sampled sites, which runs ahead of the roughly 10% adoption reported across a much larger general web sample. Source: this study, July 2026; SE Ranking, 2025.
  • Marketing agencies, the vertical that sells digital visibility, has the highest AI-invisibility rate in the sample at 27.4%. Source: this study, July 2026.

Key findings

The strongest findings from the July 2026 crawl of 419 live US small-business and professional-services websites. Each figure is a share of the 419 reachable sites unless stated otherwise.

  1. 22.4% are AI-invisible. 94 of 419 sites fail all three core signals (no schema, no llms.txt, no FAQ content). Source: this study, July 2026.
  2. Mean AI-Readiness Index is 59.7 out of 100. The distribution sits in the middle of the scale, not the top. Source: this study, July 2026.
  3. Reachability was 86%. Of 487 crawled sites, 419 were live and 68 were dead or unreachable. Source: this study, July 2026.
  4. 76.1% carry some schema.org structured data. Source: this study, July 2026.
  5. 52.0% carry Organization or LocalBusiness schema. Source: this study, July 2026.
  6. Only 4.3% carry FAQPage schema. That is 18 sites. Source: this study, July 2026.
  7. 21.2% have any FAQ content (FAQPage schema or question-style headings). Source: this study, July 2026.
  8. 19.6% have an llms.txt file. That is 82 sites. Source: this study, July 2026.
  9. 58.7% show an author or E-E-A-T byline. Source: this study, July 2026.
  10. 99.0% serve over HTTPS. Source: this study, July 2026.
  11. 89.0% set a mobile viewport. Source: this study, July 2026.
  12. 75.7% have a meta description. Source: this study, July 2026.
  13. 91.2% state no GPTBot policy in robots.txt. 3.1% block it and 5.7% explicitly allow it. Source: this study, July 2026.
  14. 90.2% state no Google-Extended policy. 3.1% block and 6.7% allow. Source: this study, July 2026.
  15. 98.8% state no PerplexityBot policy. 0% block and 1.2% allow. Source: this study, July 2026.
  16. Dental and medical lead by vertical at 63.3; financial advisors and RIAs trail at 56.4. Source: this study, July 2026.

Signal adoption across the sample

The table below reports the share of the 419 live sites carrying each readiness signal. Basic delivery signals are near universal. The signals that AI answer engines lean on for extraction and attribution are the ones that are scarce.

SignalCategoryShare of 419 live sitesSite count
HTTPSWeb hygiene99.0%415
Mobile viewportWeb hygiene89.0%373
Any schema.org structured dataStructured data76.1%319
Meta descriptionWeb hygiene75.7%317
Author / E-E-A-T bylineTrust signal58.7%246
Organization / LocalBusiness schemaStructured data52.0%218
Any FAQ contentAI-answer signal21.2%89
llms.txt file presentAI-answer signal19.6%82
FAQPage schemaAI-answer signal4.3%18

Site counts are computed from the reported share applied to the 419 live sites and rounded to whole sites. The share is the reported figure. Source: this study, July 2026.

Source: Christoph Olivier Consulting, AI Search Readiness Study, July 2026.

Stated AI-crawler policy in robots.txt

We read each robots.txt for named AI and AI-search user agents. The dominant pattern is silence. For every major AI crawler, more than 90% of sites take no stated position. Blocking and explicit allowing are both rare, and blocking never exceeds allowing by a wide margin.

AI crawlerOperatorNot mentionedBlockedExplicitly allowed
GPTBotOpenAI91.2%3.1%5.7%
Google-ExtendedGoogle90.2%3.1%6.7%
ClaudeBotAnthropic91.2%3.1%5.7%
OAI-SearchBotOpenAI97.9%0%2.1%
PerplexityBotPerplexity98.8%0%1.2%
CCBot / Bytespider / Applebot-ExtendedCommon Crawl / ByteDance / Apple~91%~3-4%~5%

Source: Christoph Olivier Consulting, AI Search Readiness Study, July 2026. Note: robots.txt records stated policy, not enforcement.

For external context, Cloudflare reported that as of mid-2025 only 14% of top domains used robots.txt rules to manage AI and search crawlers, and that GPTBot request volume grew 305% between May 2024 and May 2025. Our small-business sample shows an even lower rate of stated AI-crawler policy than Cloudflare’s top-domain population, which is consistent with smaller firms being slower to react. Source: Cloudflare, “From Googlebot to GPTBot: Who’s crawling your site in 2025,” July 1, 2025, https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/.

Structured data types found

Across the 319 sites carrying any schema.org markup, identity and site-structure types dominate. Types that support direct answer extraction, such as FAQPage, and content-credibility types, such as Article and Person, are far less common. The counts below are the number of sites on which each type appeared.

Schema typeSites using itPrimary function
Organization191Identity
BreadcrumbList179Site structure
LocalBusiness67Local identity
Person51Author / E-E-A-T
Service28Offering
Article24Content credibility
FAQPage18Answer extraction
Review11Reputation
Product3Offering
BlogPosting2Content credibility

Source: Christoph Olivier Consulting, AI Search Readiness Study, July 2026.

This ordering mirrors the wider web. Schema.org’s own usage dataset, released June 4, 2026, reports that the most-used terms across millions of domains are site-structure and identity types such as WebPage, WebSite, Organization, and BreadcrumbList, with JSON-LD the dominant format. Our sample reproduces that pattern inside the small-business segment: Organization and BreadcrumbList lead, and answer-focused types trail. Source: Schema.org, “Announcing the Schema.org usage statistics dataset,” June 4, 2026, https://blog.schema.org/2026/06/04/announcing-the-schema-org-usage-statistics-dataset/.

Original synthesis: three derived insights

The three analyses below combine only figures from this study, with clearly labeled external context. Each states its logic, its inputs, and its limitations.

Insight 1: The AI-Readiness Index by vertical

Logic: for each of the six verticals we report the mean AI-Readiness Index (0 to 100), the AI-invisible share, the any-schema share, and the sample size. This ranks verticals by how prepared their websites are for AI citation. Inputs: this study only. Limitation: sample sizes per vertical range from 62 to 79, so gaps of two to three points are directional rather than decisive.

RankVerticalMean AI-Readiness IndexAI-invisible %Any-schema %n
1Dental & medical63.313.9%84.8%79
2Marketing agencies63.227.4%71.2%73
3Home services (HVAC, plumbing, roofing)60.421.0%79.0%62
4Law / estate planning57.126.9%71.6%67
5Accounting / CPA57.124.6%72.3%65
6Financial advisors / RIA56.421.9%76.7%73

The spread from top to bottom is 6.9 points, from 63.3 for dental and medical to 56.4 for financial advisors and RIAs. Dental and medical also has the lowest AI-invisible share at 13.9% and the highest any-schema share at 84.8%. A notable contradiction sits inside the ranking: marketing agencies post the second-highest mean score, 63.2, yet also the highest AI-invisible share, 27.4%. That combination points to a split field where strong agency sites pull the average up while a large minority fail all three core signals. Source: this study, July 2026.

Insight 2: The hygiene gap. Sites are built for browsers, not for answer engines

Logic: we contrast adoption of long-standing web-hygiene signals against adoption of the newer signals AI answer engines use to extract and attribute content. Inputs: this study only. Limitation: JavaScript-injected schema is not counted, so answer-signal adoption is a conservative floor, and the real gap is somewhat narrower than the headline numbers.

Signal groupSignalAdoption
Web hygiene (browser era)HTTPS99.0%
Web hygiene (browser era)Mobile viewport89.0%
Web hygiene (browser era)Meta description75.7%
AI-answer signal (answer-engine era)Any FAQ content21.2%
AI-answer signal (answer-engine era)llms.txt file19.6%
AI-answer signal (answer-engine era)FAQPage schema4.3%

The single widest gap is HTTPS at 99.0% against FAQPage schema at 4.3%, a difference of 94.7 percentage points. Small businesses have completed the last visibility migration, to secure and mobile-friendly delivery, but have barely begun the next one, to machine-readable answers. The pattern holds even for the cheapest AI-answer signal: writing plain question-and-answer content, which requires no code, still reaches only 21.2% of sites. Source: this study, July 2026.

Insight 3: The silent majority on AI crawlers

Logic: we aggregate the “not mentioned” share across the two most prominent AI crawlers, OpenAI’s GPTBot and Google-Extended, to quantify how many firms have made no decision at all about AI access. Inputs: this study only. Limitation: robots.txt captures stated policy, not enforcement; a firm can block AI crawlers invisibly through a web application firewall, and such blocks do not appear in our reading.

91.2% of sites do not mention GPTBot and 90.2% do not mention Google-Extended. Across every named AI crawler we tested, more than 90% of sites are silent. Only 3.1% block GPTBot and only 5.7% explicitly allow it. The near-absence of a stated policy is the finding, not the direction of policy. Most small firms are neither courting nor refusing AI systems; they simply have not decided. For context, Cloudflare found 14% of top domains actively managing AI and search crawlers through robots.txt in mid-2025, so a stated AI-crawler policy is uncommon even among larger and better-resourced sites. Source: this study, July 2026; Cloudflare, July 1, 2025.

Methodology

Sample. We assembled 487 US small-business and professional-services websites across six verticals and ten US cities: Austin, Denver, Charlotte, Columbus, Phoenix, Tampa, Nashville, Kansas City, Portland, and Raleigh. Of the 487, 419 were reachable and live at crawl time, a reachability rate of 86%, and 68 were dead or unreachable. All figures for signal adoption are computed against the 419 live sites.

Sampling frame. Sites were sourced from public “best-of” listicles and local directories. We included independent local and regional firms only. We excluded national chains, franchises, and aggregators or directories.

Collection. Per site we made three to four HTTP GET requests: robots.txt, llms.txt, the homepage, and one interior page. Requests were made without executing JavaScript.

Timeframe. A point-in-time snapshot taken in July 2026. The design is repeatable on a quarterly basis.

Scoring rubric (0 to 100). Each site earns points on the weighted signals below. The weights sum to 100.

SignalPoints
Any schema.org structured data15
llms.txt file present15
FAQ content: FAQPage schema (15) or question-style headings (7)up to 15
AI-crawler openness (reduced if major AI search bots are blocked)15
Organization / LocalBusiness schema10
Author / E-E-A-T byline10
HTTPS8
Mobile viewport6
Meta description6

AI-invisible definition. A site is classified as AI-invisible if it fails all three core signals at once: no schema.org structured data, no llms.txt file, and no FAQ content.

Handling of conflicting or missing signals. Where a signal could be delivered by JavaScript, our no-JavaScript fetch treats it as absent, which biases those figures conservatively low. Where robots.txt is absent, the site is recorded as stating no AI-crawler policy.

Date of last update. July 23, 2026.

Data limitations

Five limitations shape how these figures should be read.

  1. Stated policy, not enforcement. robots.txt records what a site says, not what it enforces. Bots can ignore it, and web application firewall or Cloudflare blocks are invisible to our reading.
  2. Depth undercount. We sampled the homepage plus one interior page, so schema that lives only on deeper pages is undercounted.
  3. No JavaScript execution. We fetched without running JavaScript, so schema injected via JavaScript or tag managers is missed. Real schema adoption is therefore slightly higher than reported, which makes our “no schema” count a conservative overcount.
  4. Point-in-time, not a trend. This is a single July 2026 snapshot, not a time series. It is designed to be repeated quarterly.
  5. Sampling-frame skew. The frame is firms listed in public “best-of” directories with a live site, which may skew slightly more marketing-savvy than the true long tail. Real-world readiness across all small firms is likely worse than these numbers.

Source quality ranking

Tier 1, primary sources. Christoph Olivier Consulting AI Search Readiness Study, July 2026. This is the original dataset and the primary source for every figure attributed to “this study.”

Tier 2, credible research and standards bodies. Cloudflare Radar, “From Googlebot to GPTBot: Who’s crawling your site in 2025,” July 1, 2025, used for AI-crawler traffic and robots.txt context. Schema.org, “Announcing the Schema.org usage statistics dataset,” June 4, 2026, used for web-wide schema-type context. SE Ranking llms.txt study, 2025, used for web-wide llms.txt adoption context.

Tier 3, reputable journalism and expert commentary. Search Engine Journal and Search Engine Land reporting on the same llms.txt and schema.org datasets, used only as pointers to the Tier 2 primary material and not cited for any standalone figure.

Most quotable statistics

  • 22.4% of live US small-business websites are effectively invisible to AI search. Source: this study, July 2026.
  • Only 4.3% of small-business sites use FAQPage schema, the markup AI answer engines read for direct answers. Source: this study, July 2026.
  • 99.0% of small-business sites use HTTPS, but only 4.3% use FAQPage schema, a 94.7-point gap between browser-era and answer-engine-era readiness. Source: this study, July 2026.
  • 91.2% of small-business sites state no policy on OpenAI’s GPTBot. Source: this study, July 2026.
  • The average small-business website scores 59.7 out of 100 for AI search readiness. Source: this study, July 2026.
  • Marketing agencies, the firms that sell online visibility, have the highest AI-invisibility rate at 27.4%. Source: this study, July 2026.
  • Dental and medical practices are the most AI-ready vertical at 63.3 out of 100. Source: this study, July 2026.

Downloadable dataset: recommended fields

A public release of this dataset would include one row per site with the following fields.

  • site_id (anonymized)
  • vertical (one of six)
  • city (one of ten)
  • reachable (live or dead)
  • ai_readiness_index (0 to 100)
  • ai_invisible (true or false)
  • any_schema (true or false)
  • org_or_localbusiness_schema (true or false)
  • faqpage_schema (true or false)
  • any_faq_content (true or false)
  • llms_txt_present (true or false)
  • author_byline (true or false)
  • https (true or false)
  • mobile_viewport (true or false)
  • meta_description (true or false)
  • gptbot_policy (not_mentioned, blocked, allowed)
  • google_extended_policy (not_mentioned, blocked, allowed)
  • claudebot_policy (not_mentioned, blocked, allowed)
  • perplexitybot_policy (not_mentioned, blocked, allowed)
  • schema_types_found (list)
  • crawl_date

Press summary

New research from Christoph Olivier Consulting finds that many US small businesses are ready for browsers but not for AI search. In a July 2026 crawl of 487 small-business and professional-services websites, 419 of them live, 22.4% were effectively invisible to AI search, meaning they carried no structured data, no llms.txt file, and no FAQ content at all. The average site scored 59.7 out of 100 on the study’s AI-Readiness Index. Basic hygiene was near universal, with HTTPS at 99.0%, yet only 4.3% used FAQPage schema and only 19.6% had an llms.txt file. Most firms have made no decision about AI access: 91.2% state no policy on OpenAI’s GPTBot in robots.txt. Dental and medical practices led all verticals at 63.3, while financial advisors and RIAs trailed at 56.4. Marketing agencies, which sell online visibility, had the highest invisibility rate at 27.4%.

Frequently asked questions

What does AI search readiness mean?

It means how prepared a website is to be found, read, and cited by AI search systems such as Google AI Overviews, ChatGPT search, and Perplexity. In this study it is measured on a 0 to 100 index built from structured data, FAQ content, author signals, an llms.txt file, stated AI-crawler policy, and basic web hygiene. Source: this study, July 2026.

What share of small-business websites are invisible to AI search?

22.4%, or 94 of 419 live sites, failed all three core signals at once: no schema.org structured data, no llms.txt file, and no FAQ content. Source: this study, July 2026.

What is the average AI-Readiness score?

The mean AI-Readiness Index across 419 live sites is 59.7 out of 100. Source: this study, July 2026.

Which vertical is most ready for AI search?

Dental and medical practices, with a mean index of 63.3 out of 100, the lowest AI-invisible rate at 13.9%, and the highest any-schema rate at 84.8%. Source: this study, July 2026.

Which vertical is least ready?

Financial advisors and RIAs, with a mean index of 56.4 out of 100. Source: this study, July 2026.

How many small-business sites use FAQPage schema?

Only 4.3%, which is 18 of 419 live sites. FAQPage schema is the markup AI answer engines read to pull direct question-and-answer responses. Source: this study, July 2026.

How many sites have an llms.txt file?

19.6%, or 82 sites. For context, SE Ranking reported roughly 10% llms.txt adoption across a much larger general web sample of close to 300,000 domains in 2025. Source: this study, July 2026; SE Ranking, 2025.

Do small businesses block or allow AI crawlers?

Most do neither. 91.2% state no policy on OpenAI’s GPTBot in robots.txt, 3.1% block it, and 5.7% explicitly allow it. The pattern is similar for Google-Extended, ClaudeBot, and PerplexityBot. Source: this study, July 2026.

Is basic web hygiene a problem for these sites?

No. HTTPS reaches 99.0%, mobile viewport 89.0%, and meta descriptions 75.7%. The gap is in AI-answer signals, not in basic delivery. Source: this study, July 2026.

How was the study conducted and what are its limits?

We crawled 487 US small-business and professional-services websites across six verticals and ten cities in July 2026, of which 419 were live. Per site we made three to four HTTP requests without running JavaScript. The main limits are that robots.txt records stated policy rather than enforcement, that JavaScript-injected schema is undercounted, and that the sample skews toward directory-listed firms, so real-world readiness across all small firms is likely worse. Source: this study, July 2026.

Citations

Source: Christoph Olivier Consulting, AI Search Readiness Study, July 2026 (primary data).

Source: Cloudflare, “From Googlebot to GPTBot: Who’s crawling your site in 2025,” July 1, 2025. https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/

Source: Schema.org, “Announcing the Schema.org usage statistics dataset,” June 4, 2026. https://blog.schema.org/2026/06/04/announcing-the-schema-org-usage-statistics-dataset/

Source: SE Ranking, “LLMs.txt: Why Brands Rely On It and Why It Doesn’t Work,” 2025. https://seranking.com/blog/llms-txt/



About the author

Christoph Olivier Christoph Olivier is the founder of CO Consulting and a fractional CMO who has managed millions of dollars in ad spend and built a combined audience of over a million followers across social platforms. He publishes original marketing data reports here on CO Consulting.

Follow: YouTube · Instagram · LinkedIn