Research / Methodology
How AI Shelf Check works
Every check, why it matters for AI shopping answers, how it is scored, and where the guidance comes from. Last reviewed October 4, 2026.
This measures readiness, not appearance. Appearance in AI answers depends on factors no outside tool can see.
AI assistants and search engines decide what to show using signals that are not public: their own indexes, how people interact with results, other sites that mention a store, and ranking systems that change. This tool checks the parts a store controls and an outsider can verify.
What a scan does
- Validates the address. Only public http and https addresses on standard ports are accepted. The domain is resolved and refused if it points to a private, local or reserved network.
- Reads
/robots.txt. The scanner obeys it for its own requests, using the product tokenOptiVis-Labs-Shelf-Check. - At the same time, requests the homepage, the sitemap (from robots.txt, or
/sitemap.xml),/llms.txtand, to detect Shopify,/products.json. - If the sitemap is an index, opens one child sitemap, preferring one with "product" in its name.
- Picks up to 3 product pages, in this order: a product URL you entered, Shopify's public product list, product URLs in the sitemap, product links on the homepage.
- Loads those pages in parallel and runs the checks below on the HTML as delivered. No JavaScript is executed.
Limits per scan: 4 HTML pages, 5 small files, 8 seconds per request, 20 seconds in total, 3 MB per response, 4 redirects. The scanner identifies itself as OptiVis-Labs-Shelf-Check/1.0 (+https://optivisai.org/shelf-check).
When a page is larger than 3 MB, only the first 3 MB is checked. Search crawlers read more than that, so a missing item on such a page is reported as a low-confidence warning instead of a failure.
1. Crawler access (robots.txt)
If robots.txt blocks a search crawler, that service cannot read your pages for its index, and a page it cannot read is unlikely to be described or linked in its answers. Rules are matched the way Google documents: the most specific user-agent group applies, otherwise the * group; the longest matching path wins; on a tie, Allow wins. Each crawler is tested against / and against a real product path from your store.
| Crawler | Purpose | If blocked |
|---|---|---|
| OAI-SearchBot | Indexes pages for ChatGPT search results. | Fail |
| Googlebot | Crawls for Google Search, which AI Overviews and AI Mode draw on. | Fail |
| Bingbot | Crawls for Bing, which several AI assistants use for web results. | Fail |
| PerplexityBot | Indexes pages for Perplexity answers and links. | Fail |
| Claude-SearchBot | Indexes pages to improve Claude search results. | Fail |
| ChatGPT-User | Fetches a page when a ChatGPT user asks about it. | Warning |
| Claude-User | Fetches a page when a Claude user asks about it. | Warning |
| Perplexity-User | Fetches a page when a Perplexity user asks about it. | Warning |
| GPTBot | Collects content that may be used to train OpenAI models. | Not scored |
| ClaudeBot | Collects content that may be used to train Anthropic models. | Not scored |
| Google-Extended | Controls use of content for Gemini training and grounding; Google says it does not affect Search. | Not scored |
| CCBot | Builds the open Common Crawl dataset many models train on. | Not scored |
Search crawlers versus training crawlers. Training crawlers collect content that may be used to train models. They do not decide whether a store can be found in live answers today, so blocking them is a business choice, reported as a note and not scored. OpenAI states that robots.txt rules may not apply to ChatGPT-User, because it acts on a person's request. Google states that Google-Extended does not affect inclusion in Google Search. A robots.txt that answers with a server error is a failure, because Google treats that as a reason to stop crawling the site.
2. Discovery: sitemap and llms.txt
A sitemap tells crawlers which URLs exist. We pass when the sitemap (or the product sitemap it links to) lists product-looking URLs such as /products/ or /product/, warn when it lists none, and fail when there is no usable sitemap.
llms.txt is informational only. Google has said publicly that it does not use llms.txt for ranking, and its AI features guidance says you do not need new machine readable files or AI text files to appear in those features. We report whether the file exists and do not score it.
3. Product data
Structured data states product facts in a fixed vocabulary (schema.org) that machines read without guessing. Google documents Product structured data for product pages, and answer engines that quote prices and availability need those facts stated plainly. On each product page we read every JSON-LD block, including @graph and ProductGroup variants, and check:
- Product name, image and description present (fail without a name).
- An offer with price, priceCurrency and availability.
- Brand.
- Identifiers: GTIN, or MPN together with brand, passes. A SKU alone is a warning, because it is your internal code and does not identify the same item on other sites.
On the homepage we look for an Organization or store entity. Microdata is detected and mentioned, but only JSON-LD is scored.
4. Server rendering
Many AI crawlers fetch HTML and do not run JavaScript. Google can render JavaScript, but rendering can be delayed. We check whether the product name and price (from the structured data or Shopify's product list) appear in the visible text of the raw HTML. Both found is a pass, one found is a warning, neither is a failure. Confidence is medium, because prices can be formatted in ways a text match misses.
5. Page signals
- Title: present, about 15 to 70 characters.
- Meta description: present, about 50 to 170 characters (missing is a warning).
- Canonical tag: present and pointing to the same site.
- OpenGraph: og:title, og:description, og:image and og:url.
- Indexing directives:
noindexfails;nosnippetormax-snippet:0warns, because Google's AI features guidance names these as the controls that limit what is shown from a page. - Image alt text: share of images with descriptive alt text. 90% or more passes, 60% or more warns. Images with width and height of 1 are ignored as tracking pixels.
6. Shopify
A store is treated as Shopify when its HTML or headers show Shopify signals, or /products.json returns Shopify product data. That public endpoint is a Shopify default; we report it as a note. Whether a store shares its catalog through Shopify's agentic channels, such as Shopify Catalog or the ChatGPT product feed, is configured in the Shopify admin and cannot be seen from outside. We say so in every Shopify report and never claim it either way.
7. Response speed
We record time to first byte, total time and HTML size for each page, measured once from our server. Under 2 seconds and under 500 KB passes; over 5 seconds, or over the 3 MBread limit, fails. This is a rough indicator. It is not Lighthouse, PageSpeed Insights or Core Web Vitals, and it carries only 5% of the overall weight.
How the score is calculated
Each finding is pass (full credit), warning (half credit), fail (no credit) or note (not scored). Findings carry a weight inside their area; a blocked Googlebot, for example, weighs more than a short meta description. An area's score is the share of available credit, from 0 to 100. The overall score is the weighted average of the areas that had something to score:
| Area | Weight |
|---|---|
| Crawler access | 25 |
| Product data | 30 |
| Server rendering | 15 |
| Page signals | 15 |
| Discovery | 10 |
| Response speed | 5 |
85 and above reads as Ready, 60 to 84 as Partly ready, and below 60 as Needs work. The weights are our judgment of what matters most for product answers, based on the sources below. They are not published by any AI company.
Confidence
High means the observation is direct and unambiguous, such as a robots.txt rule. Medium means it is reliable but can be fooled, such as a price formatted differently. Low means treat it as a pointer, such as response time measured once.
Sources
- Google Search Central: How Google interprets the robots.txt specification
- Google: Overview of Google crawlers and fetchers (common crawlers, Google-Extended)
- Google Search Central: AI features and your website
- Google Search Central: Product structured data
- Google Search Central: Organization structured data
- Google Search Central: Understand JavaScript SEO basics
- Google Search Central: Robots meta tag, data-nosnippet and X-Robots-Tag
- Google Search Central: Build and submit a sitemap
- Google Search Central: Consolidate duplicate URLs (canonical)
- OpenAI: Overview of OpenAI crawlers (OAI-SearchBot, ChatGPT-User, GPTBot)
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Bing Webmaster Guidelines
- The Open Graph protocol
- schema.org Product
Platforms change their crawlers and guidance. If you find a check that no longer matches the current documentation, the terms page explains how to reach us.