Changelog

What changed, by date

Taken from the repository history. Each entry says whether it changed how something is measured or only how it looks, and which published figures it invalidates. So far no tool has published a figure, so nothing has been invalidated; the field is here for the day that changes.

14 entries, 3 of them measurement changes

  1. Research
    New

    Shelf Ledger, issue 0: the first published run of Fact Split on a frozen panel

    A frozen panel of 40 public Shopify product pages was read once with Fact Split. The page publishes the raw count of flagged disagreements and the count after a read of each flag, with the false alarms by kind. Domains and product names stay private; the downloads carry opaque ids. The fetch core was split out of the web tool's fetch function so the runner can pace itself without the web limiter; the web tool's behavior is unchanged.

    Invalidates: None. First published figure for Fact Split. The reader itself was not changed for the study; its price and stock rules produced most of the false alarms, which is listed as a reason to tighten them.

    Commit feat/shelf-ledger

  2. Compression Receipt
    New

    New tool: which compression rules keep the strings you marked

    Seven fixed compressors (three baselines and Context Cliff levels 1 to 4) are cut to budgets of 25%, 50% and 75% of the input's estimated tokens and scored on how many marked gold spans are still present. No model runs. The receipt records a hash of the input, not the input.

    Invalidates: None. No figure is published; the denominator to be published is stated on the tool page.

    Commit feat/compression-receipt

  3. Research
    New

    Six interactive explainers at /research/explainers; three embedded on their tool pages

    Inline SVG figures for Falsify, Fact Split, Span, the Jev distillation, the gpur router and span-verified retrieval. Each teaches one sentence, keeps its state in the address, works from the keyboard and reads completely as a still drawing. Fact Split, Span and the retrieval figure run invented pages through the tools' own code, so their stamps and verdicts are the tools' own. The Falsify, Fact Split and Span pages embed theirs.

    Invalidates: None. The explainers are illustrations. The only recorded figures they show are the Jev distillation reading already on the home ledger, unchanged.

    Commit feat/redesign

  4. Site
    Presentation only

    Redesign: paper and ink tokens, self-hosted type, a hero ledger, templates, this page and /status

    The hero gradients, smooth-scroll hijack, status pills, uppercase kickers and card grids are gone. Titles and prose are set in Source Serif 4 and every figure in IBM Plex Mono, both self-hosted. The home page is a ledger of recorded readings plus one table of work. Tool pages share an instrument template, research pages a paper template. Publication protocol v1.0 is written down and stamped in the footer.

    Invalidates: None. No figure changed. The referral study keeps its numbers; its page records a modification date of 2026-10-06 for the move onto the template.

    Commit feat/redesign

  5. Tripwire
    Measurement change

    Accept function-call syntax from small models; show raw unreadable replies

    The agent loop now reads the read_file("...") call form that small models often write, as well as the JSON form. A run that stops on two unreadable replies shows the raw replies instead of hiding them. This changes which model outputs count as a step.

    Invalidates: None. Tripwire has published no figure.

    Commit 8c484a7

  6. Context Cliff
    New

    New tool: compression against answers, with a bounded repair search

    Five fixed compression levels, tokens estimated as ceil(characters / 4), three seeded runs per question per level in the browser, and a repair search capped at 12 trials whose result is labeled the smallest tested repair.

    Invalidates: None. No figure is published; the denominators to be published are stated on the tool page.

    Commit 9c1c4e0

  7. Tripwire
    New

    New tool: scan agent instructions and replay a small model against a mocked workspace

    Deterministic scanner rules with line, column and character range per finding; four mocked tools on an in-memory workspace; one seeded run per permission configuration.

    Invalidates: None.

    Commit 6ff53da

  8. Falsify
    New

    New tool: evidence removal in the browser

    Four conditions (original, removed, replaced, control), three seeded runs each, WebLLM 0.2.85 with Qwen3-1.7B on WebGPU. Flags are string matches; the receipt records inputs, outputs, tallies, timings and settings.

    Invalidates: None.

    Commit 38d6a04

  9. Fact Split and Span
    New

    New tools: three-surface fact comparison and verbatim clause stamping

    Fact Split reads price, currency, availability, SKU, GTIN and variant id from JSON-LD, visible text and Shopify product JSON and declares a disagreement only when normalized values differ. Span stamps each clause of a sentence as a verbatim span with its UTF-8 byte offset, or as not a verbatim span. Both share the scanner's SSRF-safe fetcher and robots handling.

    Invalidates: None.

    Commits eb9616c, 809d3c3

  10. Research
    Presentation only

    Referral study: store anonymized, Google denominator shown beside 1.9%

    The auto-parts retailer is no longer named on the product and research pages. The home page shows 76 of 3,988 sessions next to the Google search rate, matching the ChatGPT figure's treatment.

    Invalidates: None. The figures (19 of 151; 76 of 3,988) are unchanged.

    Commits c014f62, 3cd0de0

  11. AI Shelf Check
    Measurement change

    Report sizes in decimal units; never soften a noindex found on a truncated page

    HTML size is reported in decimal kilobytes. When a page is cut off at the 3 MB limit, a noindex directive found in the part that was read keeps its full status instead of being downgraded. This changes a finding's status in that case.

    Invalidates: None. The scanner publishes per-store reports, not an aggregate.

    Commit dc8f2d3

  12. AI Shelf Check
    Measurement change

    Scanner identifies itself as OptiVis-Labs-Shelf-Check

    The user agent and the product token the robots.txt check looks for were renamed. A store that had allowed or blocked the old token by name is read differently from this date.

    Invalidates: None.

    Commit 4ce2a88

  13. Site
    New

    OptiVis Labs: product pages, research index and the referral study published

    Fitment Search Engine, LITIGATOR and PI ENGINE pages, the research index, and the Q3 2026 referral study with its methodology and limits. The LITIGATOR page states what the drafting gate does and names its designer as Trial Counsel to OptiVis.

    Invalidates: None.

    Commits 783181d, cdf2f95, 532c970

  14. AI Shelf Check
    New

    New tool: server-side AI shopping readiness scanner

    robots.txt, sitemap, homepage and up to three product pages fetched without JavaScript inside a 20 second budget; six weighted areas; every finding carries the URL, time, observation, confidence and fix.

    Invalidates: None.

    Commit 5b01ccc