Hi Nikhil,
> v4 attached, rebased on e27f3b2cad7 and renumbered now that the first
> patch is in:
I ran some checks and numbers on v4, on master 3c5d9d914fa. Scripts
and the full tables are attached.
On-disk format. The cover letter says pglz and lz4 values keep their
representation bit for bit. Writing the same data with master and
with v4, the raw bytes of the inline compressed datums (pageinspect)
and of the TOAST chunks hash the same, for pglz and lz4. pg_upgrade
from master to v4 keeps every value, and verify_heapam(check_toast)
finds nothing. It is also clean on v4 for a column holding pglz, lz4
and zstd values at once, and after VACUUM FULL, UPDATE and DELETE on
a zstd table.
No slowdown for the existing methods. v4 against master with lz4,
paired within each of 6 rounds: the median ratio is between 0.97 and
1.06 for every point, and no point is slower in all rounds. The two
rounds with pglz look the same.
zstd itself, on -O2 builds, hot cache, 200 MB per case, sizes as a
percentage of the raw data (pglz / lz4 / zstd):
- SGML docs in 32 kB values: 35.7% / 41.0% / 26.9%
- JSON documents in 1 MB values: 12.1% / 18.0% / 4.6%
- base64 text in 32 kB values: 102.9% / 102.9% / 77.0%
The last one is worth noting: pglz and lz4 get nothing out of base64
or hex text, and zstd saves about a quarter.
The cost is on reads of small values. For 50,000 JSON values of
4 kB, reading them all takes 140 ms with zstd against 53 ms with lz4,
and a 100-byte prefix costs the same as the whole value (140 ms,
against 23 ms with lz4).
One thing in 0004 that is easy to improve:
zstd_decompress_datum_slice() creates and frees a ZSTD_DCtx for every
value. The attached diff keeps one per backend, reset on each use,
and uses it in both decompression paths. zstd, 4 rounds, ratio of
the diff to v4:
- 32 kB values, slice and prefix: 0.66 to 0.77, faster in every round
- 32 kB values, full read: 0.87 to 0.89 for docs and base64
- 1 MB base64 values, slice: 0.70
- 4 kB values: 0.94 to 0.97
make check passes with it. At 4 kB the prefix still costs as much as
the whole value, so what is left there is per frame, not the context.
Regards,
Manu
diff --git a/src/backend/access/common/toast_compression.c
b/src/backend/access/common/toast_compression.c
index 849ec174539..ef7c65e1276 100644
--- a/src/backend/access/common/toast_compression.c
+++ b/src/backend/access/common/toast_compression.c
@@ -29,6 +29,30 @@
/* GUC */
int default_toast_compression = DEFAULT_TOAST_COMPRESSION;
+#ifdef USE_ZSTD
+/*
+ * zstd decompression context, created on first use and kept for the life of
+ * the backend. Creating one per value costs more than decompressing a small
+ * value. It is reset at the start of each use, so an error in the middle of
+ * a decompression leaves nothing behind that matters.
+ */
+static ZSTD_DCtx *zstd_dctx = NULL;
+
+static ZSTD_DCtx *
+zstd_get_dctx(void)
+{
+ if (zstd_dctx == NULL)
+ {
+ zstd_dctx = ZSTD_createDCtx();
+ if (zstd_dctx == NULL)
+ elog(ERROR, "could not create zstd decompression
context");
+ }
+ else
+ ZSTD_DCtx_reset(zstd_dctx, ZSTD_reset_session_only);
+ return zstd_dctx;
+}
+#endif
+
#define NO_COMPRESSION_SUPPORT(method) \
ereport(ERROR, \
(errcode(ERRCODE_FEATURE_NOT_SUPPORTED), \
@@ -327,7 +351,8 @@ zstd_decompress_datum(const varlena *value)
result = (varlena *) palloc(VARDATA_COMPRESSED_GET_EXTSIZE(value) +
VARHDRSZ);
/* decompress the data */
- rawsize = ZSTD_decompress(VARDATA(result),
+ rawsize = ZSTD_decompressDCtx(zstd_get_dctx(),
+ VARDATA(result),
VARDATA_COMPRESSED_GET_EXTSIZE(value),
(const char *) value
+ VARHDRSZ_COMPRESSED_LONG,
VARSIZE(value) -
VARHDRSZ_COMPRESSED_LONG);
@@ -374,9 +399,7 @@ zstd_decompress_datum_slice(const varlena *value, int32
slicelength)
*/
result = (varlena *) palloc(slicelength + VARHDRSZ);
- dctx = ZSTD_createDCtx();
- if (dctx == NULL)
- elog(ERROR, "could not create zstd decompression context");
+ dctx = zstd_get_dctx();
inbuf.src = (const char *) value + VARHDRSZ_COMPRESSED_LONG;
inbuf.size = VARSIZE(value) - VARHDRSZ_COMPRESSED_LONG;
@@ -409,9 +432,6 @@ zstd_decompress_datum_slice(const varlena *value, int32
slicelength)
}
}
- /* release the context before any possible error is thrown */
- ZSTD_freeDCtx(dctx);
-
if (failed)
ereport(ERROR,
(errcode(ERRCODE_DATA_CORRUPTED),
#!/bin/bash
# Benchmarks TOAST compression methods, master against v4 of "ZSTD TOAST
# compression, and an extensible compression method encoding".
# bench.sh BUILD METHOD [SIZES]
# BUILD base | v4 (see build.sh)
# METHOD pglz | lz4 | zstd (zstd only exists in v4)
# SIZES value sizes in bytes (default: 4000 32000 1000000)
# One fresh cluster per run. Three corpora, the same bytes for every build
# and method:
# docs the PostgreSQL SGML documentation, cut into values of SIZE bytes
# json generated JSON documents with realistic, repetitive keys
# random base64 of random bytes (setseed), nearly incompressible
# Storage is the default (EXTENDED), so small values are compressed inline
# and large ones compressed and moved out of line. Prints one "key=value"
# line per measurement.
set -eu
BUILD=$1 METHOD=$2
SIZES=${3:-4000 32000 1000000}
REPS=${REPS:-5}
W=$HOME/pgzstd; B=$W/i-$BUILD/bin
DOCS=$W/src-base/doc/src/sgml
D=$(mktemp -d /tmp/claude-1000/zs.XXXX); P=$((56300 + RANDOM % 100))
# SQL_ASCII and C: substr() then works on bytes and fetches only the chunks
# it needs from the 10 MB source; compression sees the same bytes either way.
"$B/initdb" -D $D -A trust --no-sync -U postgres -E SQL_ASCII --locale=C
>/dev/null
cat >> $D/postgresql.conf <<EOF
port = $P
unix_socket_directories = '/tmp'
shared_buffers = 2GB
max_wal_size = 32GB
checkpoint_timeout = 30min
autovacuum = off
jit = off
default_toast_compression = $METHOD
EOF
"$B/pg_ctl" -D $D -l $D/log -w start >/dev/null
q() { "$B/psql" -X -qAt -h /tmp -p $P -U postgres -d postgres -v
ON_ERROR_STOP=1 "$@"; }
ms() { python3 -c "import time; print(int(time.time()*1000))"; }
q -c "CREATE EXTENSION pg_prewarm"
# The documentation as one text, about 10 MB (file order is fixed).
cat $(ls $DOCS/*.sgml $DOCS/ref/*.sgml | sort) > $D/docs.txt
# Stored uncompressed, so substr() fetches only the chunks it needs instead
# of decompressing all 10 MB for every value.
q -c "CREATE TABLE docs_src (t text)"
q -c "ALTER TABLE docs_src ALTER COLUMN t SET STORAGE EXTERNAL"
q -c "INSERT INTO docs_src SELECT pg_read_file('$D/docs.txt')"
DOCS_LEN=$(q -c "SELECT length(t) FROM docs_src")
out() { echo "build=$BUILD method=$METHOD corpus=$CORPUS size=$SIZE $*"; }
for SIZE in $SIZES; do
# About 200 MB of raw data per corpus and size.
N=$(( 200000000 / SIZE )); [ $N -gt 50000 ] && N=50000
for CORPUS in docs json random; do
case $CORPUS in
docs) src="SELECT g, substr((SELECT t FROM docs_src), 1 + (g * 7919) %
$((DOCS_LEN - SIZE)), $SIZE)
FROM generate_series(1, $N) g" ;;
json) src="SELECT g, left(string_agg(format(
'{\"id\": %s, \"customer\": {\"name\": \"customer %s\",
\"tier\": \"%s\"}, \"status\": \"%s\", \"items\": [{\"sku\": \"SKU-%s\",
\"qty\": %s, \"price\": %s}], \"tags\": [\"%s\", \"%s\"]}',
g * 1000 + i, (g * 31 + i) % 5000,
(ARRAY['gold','silver','bronze'])[1 + i % 3],
(ARRAY['paid','shipped','pending','refunded'])[1 + (g +
i) % 4],
(g * 17 + i) % 100000, 1 + i % 9, round(((g * 13 + i) %
100000) / 100.0, 2),
(ARRAY['web','store','app'])[1 + g % 3],
(ARRAY['promo','none'])[1 + i % 2]), ','), $SIZE)
FROM generate_series(1, $N) g,
generate_series(1, $SIZE / 150 + 1) i
GROUP BY g" ;;
random) src="SELECT g, left(string_agg(md5(random()::text) ||
encode(sha256(random()::text::bytea), 'base64'), ''), $SIZE)
FROM generate_series(1, $N) g, generate_series(1, $SIZE /
76 + 1) i
GROUP BY g" ;;
esac
q -c "DROP TABLE IF EXISTS t, raw"
q -c "SELECT setseed(0.42)"
q -c "CREATE UNLOGGED TABLE raw (id int, v text)"
q -c "ALTER TABLE raw ALTER COLUMN v SET STORAGE EXTERNAL"
q -c "INSERT INTO raw $src"
q -c "CREATE TABLE t (id int PRIMARY KEY, v text)"
q -c "SELECT pg_prewarm('raw'), pg_prewarm((SELECT reltoastrelid FROM
pg_class WHERE relname = 'raw'))" >/dev/null
q -c "CHECKPOINT"
# 1. Load: compression happens here.
lsn0=$(q -c "SELECT pg_current_wal_insert_lsn()")
t0=$(ms); q -c "INSERT INTO t SELECT id, v FROM raw"; t1=$(ms)
wal=$(q -c "SELECT pg_wal_lsn_diff(pg_current_wal_insert_lsn(), '$lsn0')")
sizes=$(q -c "SELECT format('rows=%s raw_mb=%s heap_kb=%s toast_kb=%s
compressed_pct=%s',
count(*), round(sum(octet_length(v)) / 1e6, 1),
pg_relation_size('t') / 1024,
pg_relation_size((SELECT reltoastrelid FROM pg_class WHERE
relname = 't')) / 1024,
round(100.0 * count(*) FILTER (WHERE
pg_column_compression(v) IS NOT NULL) / count(*), 1))
FROM t")
out "op=load ms=$((t1 - t0)) wal_bytes=$wal $sizes"
q -c "VACUUM ANALYZE t"; q -c "CHECKPOINT"
q -c "SELECT pg_prewarm('t'), pg_prewarm((SELECT reltoastrelid FROM
pg_class WHERE relname = 't'))" >/dev/null
# 2. Full read, decompressing every value (hot).
for rep in $(seq $REPS); do
t0=$(ms); q -c "SELECT sum(length(v || 'x')) FROM t" >/dev/null; t1=$(ms)
out "op=read_all rep=$rep ms=$((t1 - t0))"
done
# 3. A 100-byte slice from the middle of every value (hot).
for rep in $(seq $REPS); do
t0=$(ms); q -c "SELECT sum(length(substr(v, $SIZE / 2, 100))) FROM t"
>/dev/null; t1=$(ms)
out "op=slice rep=$rep ms=$((t1 - t0))"
done
# 4. A 100-byte prefix of every value (hot): pglz can stop early here.
for rep in $(seq $REPS); do
t0=$(ms); q -c "SELECT sum(length(substr(v, 1, 100))) FROM t" >/dev/null;
t1=$(ms)
out "op=prefix rep=$rep ms=$((t1 - t0))"
done
done
done
"$B/pg_ctl" -D $D -m fast -w stop >/dev/null
rm -rf $D
#!/bin/bash
# Checks the claim in the v4 cover letter: "Existing on-disk data is not
# affected by any of them; pglz and lz4 values keep their current
# representation, bit for bit."
# 1. The same pglz and lz4 data written by master and by v4: the raw bytes
# of the inline compressed datums (pageinspect) and of the TOAST chunks
# must be identical.
# 2. pg_upgrade of the master cluster to v4: every value reads back the
# same, and amcheck (verify_heapam with check_toast) finds nothing.
# 3. zstd on v4: amcheck after ALTER ... SET COMPRESSION, VACUUM FULL, and
# copying values between pglz, lz4 and zstd columns.
# Run through the build lane: compila -n zstd-format -- bash format_check.sh
set -u
W=$HOME/pgzstd
T=$(mktemp -d /tmp/claude-1000/zfmt.XXXX)
DOCS=$W/src-base/doc/src/sgml
cat $(ls $DOCS/*.sgml $DOCS/ref/*.sgml | sort) > $T/docs.txt
setup="
CREATE EXTENSION pageinspect; CREATE EXTENSION amcheck;
CREATE TABLE docs_src AS SELECT pg_read_file('$T/docs.txt') AS t;
CREATE TABLE t_pglz (id int, v text COMPRESSION pglz);
CREATE TABLE t_lz4 (id int, v text COMPRESSION lz4);
-- Each table from the source text: copying compressed values between
-- tables keeps them as they are, whatever the target column says.
INSERT INTO t_pglz SELECT g, substr(t, 1 + (g * 7919) % 5000000, CASE WHEN g %
2 = 0 THEN 4000 ELSE 40000 END)
FROM docs_src, generate_series(1, 2000) g;
INSERT INTO t_lz4 SELECT g, substr(t, 1 + (g * 7919) % 5000000, CASE WHEN g % 2
= 0 THEN 4000 ELSE 40000 END)
FROM docs_src, generate_series(1, 2000) g;
CHECKPOINT;"
# One line per table: inline compressed datums, TOAST chunks, all values.
digest="
SELECT r, 'inline=' || md5(string_agg(md5(a.t_attrs[2]), '' ORDER BY blk, a.lp))
|| ' chunks=' || (SELECT md5(string_agg(md5(chunk_data), '' ORDER BY
chunk_id, chunk_seq))
FROM REL)
|| ' values=' || (SELECT md5(string_agg(md5(v), '' ORDER BY id)) FROM
TBL)
FROM (SELECT 'TBL'::text r) x,
generate_series(0, pg_relation_size('TBL') / 8192 - 1) blk,
heap_page_item_attrs(get_raw_page('TBL', blk::int), 'TBL'::regclass) a
WHERE a.t_attrs IS NOT NULL
GROUP BY r"
start() { "$1/bin/pg_ctl" -D "$2" -o "-p $3 -c unix_socket_directories=$T" -l
"$2.log" -w start > /dev/null; }
stop() { "$1/bin/pg_ctl" -D "$2" -m fast -w stop > /dev/null; }
q() { local I=$1 P=$2; shift 2; "$I/bin/psql" -X -qAt -h "$T" -p $P -U
postgres -d postgres "$@"; }
digests() { # install port
for tbl in t_pglz t_lz4; do
rel=$(q $1 $2 -c "SELECT reltoastrelid::regclass FROM pg_class WHERE
relname = '$tbl'")
q $1 $2 -v ON_ERROR_STOP=1 -c "$(echo "$digest" | sed "s/TBL/$tbl/g;
s/REL/$rel/g")" \
|| echo "$tbl: DIGEST FAILED"
done
}
echo "== 1. same data written by master and by v4"
for b in base v4c; do
I=$W/i-$b
"$I/bin/initdb" -D "$T/d-$b" -U postgres --no-sync -A trust > /dev/null 2>&1
start $I "$T/d-$b" 55461
q $I 55461 -c "$setup" > /dev/null
digests $I 55461 > "$T/digest-$b.txt"
sed "s/^/ $b: /" "$T/digest-$b.txt"
stop $I "$T/d-$b"
done
if [ "$(grep -c 'inline=' "$T/digest-base.txt")" != 2 ]; then echo " NO
DIGESTS, check failed"
elif cmp -s "$T/digest-base.txt" "$T/digest-v4c.txt"; then echo " identical"
else echo " DIFFERENT"; fi
echo "== 2. pg_upgrade master -> v4"
"$W/i-v4c/bin/initdb" -D "$T/d-up" -U postgres --no-sync -A trust > /dev/null
2>&1
( cd "$T" && "$W/i-v4c/bin/pg_upgrade" -b "$W/i-base/bin" -B "$W/i-v4c/bin" -d
"$T/d-base" -D "$T/d-up" \
-U postgres -s "$T" -p 55461 -P 55462 > "$T/pg_upgrade.log" 2>&1 )
echo " pg_upgrade rc=$?"
start $W/i-v4c "$T/d-up" 55462
digests $W/i-v4c 55462 > "$T/digest-up.txt"
cmp -s "$T/digest-base.txt" "$T/digest-up.txt" && echo " digests after
upgrade: identical to master" || { echo " digests after upgrade: DIFFERENT";
cat "$T/digest-up.txt"; }
for tbl in t_pglz t_lz4; do
echo " amcheck $tbl: $(q $W/i-v4c 55462 -c "SELECT count(*) FROM
verify_heapam('$tbl', check_toast => true)") problems,"\
"methods: $(q $W/i-v4c 55462 -c "SELECT string_agg(DISTINCT
coalesce(pg_column_compression(v), 'none'), ',') FROM $tbl")"
done
echo "== 3. zstd on v4"
q $W/i-v4c 55462 -v ON_ERROR_STOP=1 > /dev/null <<'SQL'
CREATE TABLE t_zstd (id int, v text COMPRESSION zstd);
INSERT INTO t_zstd SELECT id, v || '' FROM t_pglz;
-- A column that ends up holding pglz, lz4 and zstd values at once.
CREATE TABLE t_mix (id int, v text COMPRESSION pglz);
INSERT INTO t_mix SELECT id, v || '' FROM t_pglz;
ALTER TABLE t_mix ALTER COLUMN v SET COMPRESSION zstd;
INSERT INTO t_mix SELECT id + 100000, v || '' FROM t_lz4;
INSERT INTO t_mix SELECT id + 200000, v FROM t_lz4;
VACUUM FULL t_zstd;
CREATE TABLE t_back (id int, v text COMPRESSION lz4);
INSERT INTO t_back SELECT id, v || '' FROM t_zstd;
UPDATE t_zstd SET v = v || 'x' WHERE id % 10 = 0;
DELETE FROM t_zstd WHERE id % 7 = 0;
VACUUM t_zstd;
SQL
for tbl in t_zstd t_mix t_back; do
echo " $tbl: amcheck problems $(q $W/i-v4c 55462 -c "SELECT count(*) FROM
verify_heapam('$tbl', check_toast => true)"),"\
"methods $(q $W/i-v4c 55462 -c "SELECT string_agg(m || ':' || n, ' ')
FROM (SELECT coalesce(pg_column_compression(v), 'none') m, count(*) n FROM $tbl
GROUP BY 1 ORDER BY 1) s")"
done
echo " values equal to the source: $(q $W/i-v4c 55462 -c "SELECT count(*)
FROM t_pglz a JOIN t_back b USING (id) WHERE a.v = b.v") of $(q $W/i-v4c 55462
-c "SELECT count(*) FROM t_pglz")"
stop $W/i-v4c "$T/d-up"
echo "== logs in $T"
1. v4 / master, same method (median time ratio; 1.00 = no change)
corpus size op | pglz | lz4 | round-to-round spread of master (pglz, lz4)
docs 4000 load | 1.00 | 0.99 | 1.00, 1.00
docs 4000 read_all | 1.00 | 1.01 | 1.00, 1.00
docs 4000 slice | 1.00 | 1.02 | 1.01, 1.01
docs 4000 prefix | 1.02 | 1.03 | 1.00, 1.00
docs 32000 load | 1.00 | 1.03 | 1.00, 1.01
docs 32000 read_all | 0.99 | 1.03 | 1.03, 1.00
docs 32000 slice | 1.00 | 1.04 | 1.00, 1.02
docs 32000 prefix | 1.05 | 1.03 | 1.05, 1.00
docs 1000000 load | 1.00 | 1.01 | 1.00, 1.01
docs 1000000 read_all | 1.01 | 1.00 | 1.01, 1.00
docs 1000000 slice | 1.00 | 1.01 | 1.01, 1.03
docs 1000000 prefix | 1.00 | 0.96 | 1.00, 1.04
json 4000 load | 1.01 | 1.02 | 1.01, 1.02
json 4000 read_all | 1.02 | 1.02 | 1.02, 1.02
json 4000 slice | 1.00 | 1.02 | 1.00, 1.00
json 4000 prefix | 1.04 | 1.00 | 1.04, 1.00
json 32000 load | 1.01 | 1.05 | 1.00, 1.03
json 32000 read_all | 1.01 | 1.04 | 1.00, 1.00
json 32000 slice | 1.03 | 1.06 | 1.00, 1.00
json 32000 prefix | 1.00 | 1.08 | 1.05, 1.00
json 1000000 load | 1.00 | 0.99 | 1.00, 1.01
json 1000000 read_all | 1.00 | 1.00 | 1.00, 1.01
json 1000000 slice | 0.98 | 0.98 | 1.04, 1.00
json 1000000 prefix | 1.00 | 1.00 | 1.00, 1.00
random 4000 load | 1.02 | 1.02 | 1.02, 1.03
random 4000 read_all | 1.03 | 1.12 | 1.01, 1.02
random 4000 slice | 1.02 | 1.06 | 1.01, 1.02
random 4000 prefix | 1.03 | 1.02 | 1.01, 1.02
random 32000 load | 1.01 | 1.11 | 1.00, 1.02
random 32000 read_all | 1.00 | 0.99 | 1.02, 1.02
random 32000 slice | 1.00 | 1.00 | 1.00, 1.04
random 32000 prefix | 1.00 | 1.02 | 1.04, 1.08
random 1000000 load | 1.05 | 1.00 | 1.03, 1.02
random 1000000 read_all | 0.99 | 0.98 | 1.01, 1.02
random 1000000 slice | 1.00 | 1.00 | 1.00, 1.00
random 1000000 prefix | 1.05 | 1.00 | 1.10, 1.00
2. Methods on v4: stored size (% of raw), then median ms per operation
corpus size | pglz | lz4 | zstd
docs 4000 stored | 48.2% | 57.4% | 36.3%
docs 4000 load ms | 1683 | 540 | 854
docs 4000 read_all ms | 224 | 105 | 228
docs 4000 slice ms | 140 | 88 | 227
docs 4000 prefix ms | 42 | 62 | 227
docs 32000 stored | 35.7% | 41.0% | 26.9%
docs 32000 load ms | 2030 | 390 | 592
docs 32000 read_all ms | 196 | 76 | 146
docs 32000 slice ms | 120 | 57 | 172
docs 32000 prefix ms | 22 | 34 | 168
docs 1000000 stored | 30.8% | 32.8% | 21.0%
docs 1000000 load ms | 2028 | 384 | 556
docs 1000000 read_all ms | 234 | 101 | 142
docs 1000000 slice ms | 103 | 38 | 59
docs 1000000 prefix ms | 11 | 23 | 29
json 4000 stored | 18.6% | 29.3% | 13.7%
json 4000 load ms | 798 | 278 | 445
json 4000 read_all ms | 57 | 53 | 140
json 4000 slice ms | 44 | 43 | 142
json 4000 prefix ms | 24 | 23 | 140
json 32000 stored | 13.0% | 25.8% | 8.7%
json 32000 load ms | 814 | 200 | 248
json 32000 read_all ms | 46 | 49 | 81
json 32000 slice ms | 34 | 38 | 122
json 32000 prefix ms | 19 | 27 | 111
json 1000000 stored | 12.1% | 18.0% | 4.6%
json 1000000 load ms | 792 | 192 | 213
json 1000000 read_all ms | 78 | 78 | 94
json 1000000 slice ms | 26 | 26 | 35
json 1000000 prefix ms | 10 | 17 | 18
random 4000 stored | 137.8% | 137.8% | 78.1%
random 4000 load ms | 1026 | 510 | 711
random 4000 read_all ms | 105 | 116 | 223
random 4000 slice ms | 86 | 91 | 225
random 4000 prefix ms | 84 | 86 | 223
random 32000 stored | 102.9% | 102.9% | 77.0%
random 32000 load ms | 1394 | 366 | 370
random 32000 read_all ms | 56 | 56 | 140
random 32000 slice ms | 25 | 26 | 167
random 32000 prefix ms | 25 | 26 | 165
random 1000000 stored | 102.8% | 102.8% | 73.7%
random 1000000 load ms | 1504 | 290 | 499
random 1000000 read_all ms | 89 | 88 | 168
random 1000000 slice ms | 11 | 11 | 113
random 1000000 prefix ms | 11 | 11 | 48
v4/lz4 / base/lz4: median of per-round ratios [min, max] over 6 rounds
docs 4000 load: 0.99 [0.89, 1.04]
docs 4000 prefix: 0.99 [0.90, 1.02]
docs 4000 read_all: 0.99 [0.90, 1.02]
docs 4000 slice: 0.98 [0.92, 1.01]
docs 32000 load: 1.00 [0.91, 1.02]
docs 32000 prefix: 1.00 [0.85, 1.09]
docs 32000 read_all: 0.99 [0.91, 1.01]
docs 32000 slice: 0.99 [0.90, 1.02]
docs 1000000 load: 1.00 [0.97, 1.04]
docs 1000000 prefix: 1.00 [0.92, 1.08]
docs 1000000 read_all: 0.99 [0.83, 1.28]
docs 1000000 slice: 0.99 [0.93, 1.08]
json 4000 load: 1.01 [0.92, 1.06]
json 4000 prefix: 1.00 [0.85, 1.04]
json 4000 read_all: 0.97 [0.90, 1.04]
json 4000 slice: 0.99 [0.92, 1.02]
json 32000 load: 0.98 [0.94, 1.11]
json 32000 prefix: 0.98 [0.93, 1.08]
json 32000 read_all: 0.99 [0.94, 1.06]
json 32000 slice: 1.03 [0.95, 1.05]
json 1000000 load: 1.00 [0.97, 1.13]
json 1000000 prefix: 1.06 [0.94, 1.11]
json 1000000 read_all: 0.98 [0.91, 1.07]
json 1000000 slice: 1.00 [0.90, 1.04]
random 4000 load: 1.01 [0.92, 1.13]
random 4000 prefix: 0.99 [0.89, 1.04]
random 4000 read_all: 1.01 [0.93, 1.07]
random 4000 slice: 0.99 [0.91, 1.01]
random 32000 load: 0.99 [0.96, 1.02]
random 32000 prefix: 1.02 [0.90, 1.19]
random 32000 read_all: 1.00 [0.91, 1.00]
random 32000 slice: 0.98 [0.90, 1.04]
random 1000000 load: 1.05 [0.78, 1.17]
random 1000000 prefix: 1.09 [1.00, 1.27]
random 1000000 read_all: 0.97 [0.93, 1.24]
random 1000000 slice: 1.09 [0.92, 1.27]
v4r/zstd / v4/zstd: median of per-round ratios [min, max] over 4 rounds
docs 4000 load: 0.99 [0.96, 1.03]
docs 4000 prefix: 0.97 [0.89, 0.99] <-- all rounds faster
docs 4000 read_all: 0.96 [0.95, 0.97] <-- all rounds faster
docs 4000 slice: 0.97 [0.94, 0.98] <-- all rounds faster
docs 32000 load: 0.98 [0.97, 1.01]
docs 32000 prefix: 0.77 [0.73, 0.78] <-- all rounds faster
docs 32000 read_all: 0.89 [0.86, 0.90] <-- all rounds faster
docs 32000 slice: 0.76 [0.71, 0.77] <-- all rounds faster
docs 1000000 load: 0.98 [0.97, 1.01]
docs 1000000 prefix: 0.95 [0.90, 1.00]
docs 1000000 read_all: 0.99 [0.95, 1.04]
docs 1000000 slice: 0.99 [0.97, 1.00]
json 4000 load: 0.98 [0.92, 1.00]
json 4000 prefix: 0.95 [0.88, 0.97] <-- all rounds faster
json 4000 read_all: 0.94 [0.88, 0.98] <-- all rounds faster
json 4000 slice: 0.94 [0.91, 0.96] <-- all rounds faster
json 32000 load: 0.97 [0.96, 1.00] <-- all rounds faster
json 32000 prefix: 0.69 [0.65, 0.70] <-- all rounds faster
json 32000 read_all: 0.99 [0.98, 1.01]
json 32000 slice: 0.66 [0.64, 0.67] <-- all rounds faster
json 1000000 load: 0.99 [0.97, 1.04]
json 1000000 prefix: 1.00 [0.94, 1.17]
json 1000000 read_all: 0.98 [0.92, 1.00]
json 1000000 slice: 1.01 [0.94, 1.03]
random 4000 load: 0.99 [0.93, 1.00] <-- all rounds faster
random 4000 prefix: 0.96 [0.94, 0.97] <-- all rounds faster
random 4000 read_all: 0.97 [0.93, 0.98] <-- all rounds faster
random 4000 slice: 0.97 [0.95, 0.99] <-- all rounds faster
random 32000 load: 0.96 [0.85, 1.00]
random 32000 prefix: 0.77 [0.73, 0.78] <-- all rounds faster
random 32000 read_all: 0.87 [0.84, 0.88] <-- all rounds faster
random 32000 slice: 0.72 [0.71, 0.74] <-- all rounds faster
random 1000000 load: 1.02 [0.96, 1.06]
random 1000000 prefix: 0.99 [0.98, 1.02]
random 1000000 read_all: 1.01 [0.97, 1.03]
random 1000000 slice: 0.70 [0.68, 0.71] <-- all rounds faster