Free toolRuns in your browser; your text stays on your deviceProtocol v1.0Updated
Find the edge where compression starts to break answers
Shrinking the data you give a model saves tokens, but a cut can remove the one fact a question needs. Paste your JSON or log, add the questions that matter, and see which level of compression still answers them. When one breaks, the tool puts omitted blocks back one at a time to show which one repaired it.
- Last reading
- None published. The reading is the one you run; results are not collected into a figure.
- Protocol
- Context Cliff 1.0 (lib/contextcliff.ts, lib/contextcliff-engine.ts): five fixed compression levels written for this page, tokens estimated as ceil(characters / 4); WebLLM 0.2.85 running Qwen3-1.7B-q4f16_1-MLC; temperature 0.7, top_p 0.9, 96 tokens, seeds 11, 22 and 33; repair search capped at 12 trials.
- Verified
- 34 of 34 unit tests passed at build . How this is verified.
- Failure state
- Without WebGPU the page says so before anything downloads. A question that never held at full text is excluded from the chart rather than counted as broken.
- Receipt
- Export receipt (JSON) on the page, written on your device.
Five levels, your questions, three runs each
Read as JSON. 307 tokens estimated at full text (ceil(characters / 4), a rough rule for English text, not the model's tokenizer).
2. Up to 3 questions, each with the text a correct answer must contain
An answer counts as correct when it contains the expected text, ignoring case and thousands commas. A bare number never matches inside a longer number.
JSON drops null, empty string, empty array and empty object values. A log merges lines that match once a leading timestamp is ignored, keeping the first with a count such as (x4).
187 tokens estimated, 39% fewer than full text, 8 cuts made.
{"service":"billing-api","region":"us-east","generatedAt":"2026-10-01T09:00:00Z","incident":{"id":"INC-2291","status":"resolved","rootCause":"Expired certificate on the payments gateway, renewed by hand after the automatic renewal job failed","resolvedAt":"2026-09-30T14:12:00Z","owner":{"name":"Priya Raman","team":"Platform"}},"orders":[{"id":1040,"status":"paid","total":58.2},{"id":1041,"status":"paid","total":91.75},{"id":1042,"status":"refunded","total":33,"note":"Customer reported a duplicate charge on the same card within one minute"},{"id":1043,"status":"paid","total":212.4},{"id":1044,"status":"paid","total":17.99},{"id":1045,"status":"pending","total":640,"note":"Awaiting manual review"},{"id":1046,"status":"paid","total":75.1}]}| Level | Rule set | Tokens (estimate) | Cuts |
|---|---|---|---|
| 0 | Full text | 307 | 0 |
| 1 | Whitespace removed | 210 | 0 |
| 2 | Empty values dropped, repeats merged | 187 | 8 |
| 3 | Long values truncated | 161 | 9 |
| 4 | Top entries kept | 125 | 8 |
5 levels x 3 questions x 3 runs = 45 model calls.
The five levels, and what this does not show
- Level 0, full text. Your input, unchanged.
- Level 1, whitespace removed. JSON is minified. A log loses repeated spaces, trailing spaces and blank lines.
- Level 2, empty values dropped, repeats merged. JSON drops null, empty string, empty array and empty object values. A log merges lines that match once a leading timestamp is ignored, keeping the first with a count such as (x4).
- Level 3, long values truncated. JSON strings longer than 80 characters and arrays longer than 5 items are cut, with a marker that says how much is missing. Log lines longer than 160 characters are cut the same way.
- Level 4, top entries kept. JSON keeps the first 6 keys of an object, the first 3 items of an array, 40 characters of a string, and nothing deeper than 4 levels. A log keeps lines that mention error, warn, fatal, fail or exception, plus the first 3 and last 3 lines, and replaces each skipped run with a marker.
Each level includes every rule below it. The compression is our own fixed rules written for this page. It is not a library, and it is not tuned to your data.
- Tokens are estimated, not counted. The estimate is ceil(characters / 4), a rough rule for English text, not the model's tokenizer. It is the same rule at every level, so the levels compare fairly with each other, but it will not match a provider's bill.
- Correct means the expected text appears in the answer. A right answer in other words counts as wrong, and a wrong answer that happens to contain the text counts as right.
- A question “holds” when most of its runs are correct. Only questions that held at full text are counted in the chart, since the others cannot break.
- The repair search is a bounded test. It orders omitted blocks by likely relevance, tries up to 8 of them alone, then growing groups, and stops after 12 trials. The result is the smallest tested repair. Combinations that were not tried may be smaller, and sampling variation can make a block look like the cause when it is not. Each trial is marked when its restored block does not contain the expected text, because such a repair says the answer is fragile, not that the block held the fact.
- One small model, one device, three samples. A model of a different size may fail at a different level. A count such as 2 of 3 is an observation on your machine, not a rate and not a benchmark. The settings and seeds are shown, and the receipt records them.
Where the data goes
The model runs on your device through WebGPU. Your pasted data and questions are not uploaded; this page has no server step for them. The model weights and its compiled library are downloaded from their public hosts the first time you consent, so those requests do happen. The model is Qwen3 1.7B (Apache 2.0) run with the open source WebLLM library. What we store. OptiVis Marketing can help when an answer engine reads your own pages wrongly.