Free toolRuns in your browser; your text stays on your deviceProtocol v1.0Updated
See which compression rules keep the facts you marked
Every way of shrinking data for a model cuts something. Paste your JSON or log, mark the exact strings that must survive, and compare compressors at the same token budget. The table names each string a rule lost, and you can export the comparison as a receipt.
- Last reading
- None published. The reading is the one you run; results are not collected into a figure.
- Protocol
- Compression Receipt 1.0 (lib/compressionreceipt.ts): seven fixed compressors written for this page or taken from Context Cliff levels 1 to 4, budgets of 25%, 50% and 75% of the input's estimated tokens, tokens estimated as ceil(characters / 4); a span is preserved when it appears after the same normalization Falsify uses; no model runs.
- Verified
- 26 of 26 unit tests passed at build . How this is verified.
- Failure state
- Empty or oversized input shows a message and no table. Spans that are not in the input are listed as not scored and excluded from the total.
- Receipt
- Export receipt (JSON) on the page, written on your device. It records a SHA-256 hash of the input, not the input.
Seven compressors, three budgets, your gold spans
0 tokens estimated (ceil(characters / 4), a rough rule for English text, not the model's tokenizer).
2. Mark up to 6 gold spans, the exact text that must survive
Choose values and words, such as an id, a name or an amount. Compression can change spacing around punctuation, so a span like "qty": 2 may not survive a minify that 2 would. Matching ignores case, curly quotes, thousands commas and runs of whitespace.
No spans yet. Without spans the table shows token counts only.
The compressors, the budgets, and what this does not show
A budget is a share of the input's estimated token count: 25%, 50%, 75%. Each compressor runs on the input and its end is cut if the result is still over the budget, so every cell in the table fits the same limit.
- Head only. Keep the start of the text and cut everything after the budget. This is the full text, cut.
- Head and tail. Keep the first half and the last half of the budget, with a one-line marker saying how many characters sit between them. The marker counts against the budget.
- Merge repeats, then head. Collapse spaces, drop blank lines and merge lines that match once a leading timestamp is ignored (the first is kept with a count such as (x4)), then keep the start. Single-line JSON has no repeated lines, so for it this is the head baseline with spaces collapsed.
- Context Cliff level 1, fitted. Whitespace removed (JSON minified), then the end is cut if the result is still over budget.
- Context Cliff level 2, fitted. Empty values dropped and repeated log lines merged, then the end is cut if the result is still over budget.
- Context Cliff level 3, fitted. Long strings, long arrays and long log lines shortened with a marker, then the end is cut if the result is still over budget.
- Context Cliff level 4, fitted. Only the top entries or the error-like lines kept, then the end is cut if the result is still over budget.
The three baselines are short rules written for this page. The four Context Cliff levels are the same rules the Context Cliff page uses. Nothing here is a library, and no other compression product is used or compared.
- This scores strings kept, not answers given. A rule can keep every span you marked and still leave a model unable to answer, or lose a span and leave the answer intact. No model runs on this page. To test answers, use Context Cliff.
- Preserved means the text is still there. A span counts as preserved when it appears in the compressed text after ignoring case, curly quotes, thousands commas and runs of whitespace. A bare number does not match inside a longer number. This is the same rule Falsify and Context Cliff use for their expected text.
- Only spans found in your input are scored. A span that is not in the full text cannot be lost, so it is listed as not scored and left out of the total. You can mark at most 6 spans of 120 characters, in an input of up to 20,000 characters.
- Tokens are estimated, not counted. The estimate is ceil(characters / 4), a rough rule for English text, not the model's tokenizer. It is the same rule for every compressor and for the budget, so the comparison is fair, but it will not match a provider's bill.
- One input and one set of spans is one observation. Which rule keeps your facts depends on where they sit in your data. A result here is not a benchmark of the rules.
Where the data goes
The comparison runs in your browser. Your pasted text and spans are not uploaded; this page has no server step for them. The exported receipt records a SHA-256 hash of your input so you can show later which input it described, and it never contains the input text. Your gold spans are written into it as you typed them. What we store. OptiVis Marketing can help when an answer engine reads your own pages wrongly.