Open SourcePythonbyte-exact roundtrip

jsonl-bench

A tool-agnostic compression benchmark for the data types of the LLM era (agent transcripts, API logs). Ranking metric: total compressed size across the corpus. A byte-exact roundtrip is required.

Source code on GitHub

About the Project

The fairest way to compare compressors is to have everyone measure on the same data under the same rules. jsonl-bench provides a tool-agnostic, reproducible ranking on LLM-era data (agent transcripts, API logs): the winner is the tool with the smallest total output across the whole corpus — and every tool's output must decompress without a single changed byte.

The corpus is not stored in the repo; it is regenerated locally from fixed seeds, so everyone gets identical results. Note: this corpus is synthetic and the absolute ratios do not generalize to real data — stated plainly for honesty.

How It Works

$ pip install zstandard
$ python gen_corpus.py # generate the corpus (deterministic, SHA256SUMS)
$ python run_bench.py # leaderboard → RESULTS.md

Tool-agnostic

Any compressor (library or CLI) can join by being added to a single list; the benchmark isn't tied to any particular tool.

One metric: total size

Ranking is by total compressed size across the corpus. A byte-exact (lossless) roundtrip is mandatory.

Deterministic corpus

The corpus isn't stored in the repo; it is regenerated locally from fixed seeds (verified with SHA256SUMS), so everyone measures on the same data.

Add a tool with a PR

A new tool joins with a single PR that adds a command/call to the TOOLS list in run_bench.py. The tool must be installable by anyone.

Tech Stack

PythonzstandardBenchmark designDeterministic corpusSHA256 verification
Published as open source

Fully open source

jsonl-bench is published as open source on GitHub. Read the rules, add your own tool with a PR and reproduce the results locally.