jsonl-bench
A tool-agnostic compression benchmark for the data types of the LLM era (agent transcripts, API logs). Ranking metric: total compressed size across the corpus. A byte-exact roundtrip is required.
Source code on GitHubAbout the Project
The fairest way to compare compressors is to have everyone measure on the same data under the same rules. jsonl-bench provides a tool-agnostic, reproducible ranking on LLM-era data (agent transcripts, API logs): the winner is the tool with the smallest total output across the whole corpus — and every tool's output must decompress without a single changed byte.
The corpus is not stored in the repo; it is regenerated locally from fixed seeds, so everyone gets identical results. Note: this corpus is synthetic and the absolute ratios do not generalize to real data — stated plainly for honesty.
How It Works
Tool-agnostic
Any compressor (library or CLI) can join by being added to a single list; the benchmark isn't tied to any particular tool.
One metric: total size
Ranking is by total compressed size across the corpus. A byte-exact (lossless) roundtrip is mandatory.
Deterministic corpus
The corpus isn't stored in the repo; it is regenerated locally from fixed seeds (verified with SHA256SUMS), so everyone measures on the same data.
Add a tool with a PR
A new tool joins with a single PR that adds a command/call to the TOOLS list in run_bench.py. The tool must be installable by anyone.
Tech Stack
Fully open source
jsonl-bench is published as open source on GitHub. Read the rules, add your own tool with a PR and reproduce the results locally.