Research / Methodology
How Fact Split and Span work
Both tools are deterministic code. No model reads the page, there is no score, and the same page at the same moment always gives the same answer.
Fact Split
We fetch one page without running JavaScript, after reading its robots.txt. For Shopify-shaped paths (/products/<handle>) we also fetch the public /products/<handle>.js document. Each surface is read for six fields.
Surfaces
- Structured data: the first Product (or product variant) in the page's JSON-LD, and its first offer.
- Visible text: the page text with scripts and styles removed. The price is the first currency amount after the main heading; availability is the first stock phrase; SKU is the text after "SKU".
- Shopify product JSON: the variant named in the address (
?variant=), or the first variant. Prices there are integer cents.
When two values count as different
- Prices are converted to cents first, so "$1,299.00", "1299" and "129900 cents" are the same. A range or a value with two amounts is not read as a price.
- Availability is mapped to in stock, out of stock, preorder or backorder. Currency symbols are compared as sets: "$" agrees with USD, CAD, AUD and other dollars.
- SKUs compare case-insensitively. GTINs compare digits without leading zeros. Variant ids compare their numeric part.
- The raw string from every surface is always shown. A field stated on fewer than two surfaces gets "Not enough surfaces", never a verdict.
Limits
- Prices that a script writes into the page after load are not seen. Only the HTML the server sends is read.
- The product feed a store sends to a shopping platform is not public, so it is never compared.
- One fetch, at the time shown. A price or stock level can change a minute later.
- The visible-text price is the first currency amount after the main heading. A page with several prices can show a different one than the buy button uses.
- The visible stock phrase is the first one on the page. A page that lists several sizes can show a phrase that belongs to a different variant than the one compared, so check a disagreement by eye before acting on it.
Span
A clause is a verbatim span when it appears in the page's visible text, or in a JSON-LD string value, character for character and case-sensitive. Only two normalizations are applied, to both sides: HTML entities are decoded, and runs of whitespace (including non-breaking spaces) become one space. Adjacent tags are separated by a space.
- Everything else counts as a difference: case, curly versus straight quotes, hyphen versus en dash, plurals, word order and every paraphrase. A paraphrase stays a miss.
- Byte offsets are UTF-8 byte positions of the first match inside the normalized visible text (or inside the named JSON-LD value). When the same characters also appear unchanged in the raw HTML source, that offset is shown too.
- Sentences are split at sentence ends and semicolons, and at ", and", ", but" and ", while" when both sides are long enough. Clauses under 2 words are not stamped. Input is limited to 600 characters and 12 clauses.
- For a miss we show the sentence with the most word overlap, if any is close, labeled as not a match. The overlap figure is a similarity, not evidence of support.
A verbatim stamp says the words are on the page. It does not say the claim is true, current, or applies to the product you are looking at.
Fetching and limits
- Our fetcher identifies itself as
OptiVis-Labs-Shelf-Check/1.0 (+https://optivisai.org/shelf-check)and follows robots.txt. If robots.txt cannot be read, only a homepage is fetched. - Private and internal addresses are refused. Each visitor can run 10 checks per hour. A fetched page is kept in memory for 10 minutes so a second check on the same page does not hit the store again.