Methodology / version 1

Know what was measured.

A benchmark describes a particular experiment. Here are the conditions, selection rules, and limitations behind every number on CacheArena.

01 / Disclosure

Owned by Lavik. Scrutiny welcome.

CacheArena is owned and maintained by Lavik. Version 1 republishes benchmark measurements authored by Lavik / EloqData. We have not independently rerun these experiments and do not claim that this is an independent testing organization.

Fairness here means publishing the source, keeping different experiments separate, exposing material configuration differences, and applying the same selection rules to every system. It does not mean that every configuration has identical durability, memory use, or operational cost.

What these results cannot establish

A universal winner, production capacity, equal crash durability, a full cost comparison, or a statistically significant performance advantage. The 10M- and 1B-key report has one run per point; the 200M-key summary does not publish repeat-trial uncertainty.

Corrections and additional evidence are welcome through the Lavik repository’s issue tracker. The first release does not contain third-party submissions or independent validation.

02 / Display rules

One measurement, all of its metrics.

  1. Choose a cohort and workload. No cross-cohort aggregate score is calculated. Missing measurements are never treated as zero.
  2. Choose the connection count. The default 10M- and 1B-key views select each system’s peak measured QPS. A fixed connection count lets you compare at the same client concurrency. The 200M-key cohort uses 80 connections throughout.
  3. Choose thread settings. Redis and Valkey default to their highest-throughput I/O-thread configuration for the selected workload across the full sweep. You can instead select 1, 2, 4, 8, or 16 I/O threads. Thread settings are not selected to minimize latency.
  4. Sort by your chosen metric. QPS sorts descending; latency sorts ascending. Changing the metric preserves the selected measurement, including its connections and server threads. Lower latency at a peak-QPS point can still reflect lower concurrency.
  5. Inspect the evidence. Charts start at zero. Displayed QPS is rounded to whole operations per second; downloads retain source precision. The full connection sweep and source reports remain downloadable.

For example, the 1B-key SET peaks report Lavik at 764,939 QPS / 10.879 ms p99 and Garnet at 751,057 QPS / 5.535 ms p99. Throughput and tail latency answer different questions; both stay visible.

03 / 1B keys

1 billion keys. Storage-tier controls.

1,024-byte values, uniform random GET or overwriting SET, memtier 2.5.1, pipeline 1, 16 client threads, and no rate limit. Each point runs for 60 seconds at 80, 160, 320, 640, 1,280, or 2,560 connections with distinct client seeds.

ServerAMD EPYC 9V748 physical cores / 16 logical CPUs · ~126 GiB RAM
ClientAMD EPYC 9V4516 cores · ~31 GiB RAM
Storage6 × ~1.92 TB NVMeSame benchmark drives, different access paths
CPU placementCPUs 0–15One product at a time; sequential acquisition
  • Lavik 0.1.0-beta.1: downloaded standard release, 16 workers, kernel TCP, raw SPDK devices, 8 GiB of hugepages in addition to other memory use. Fresh dataset, no additional 1B GET warmup. Completion cap 16; Tomb Raider tombstone sweep disabled.
  • Dragonfly 1.40.2: 16 proactors, 96 GiB maxmemory, Tiered Storage on RAID0/XFS. Receives a 180-second GET warmup.
  • Garnet 2.1.5: 64 GiB hybrid log, 32 GiB read cache, 16 GiB index, Libaio storage on RAID0/XFS; 180-second GET warmup. Minimum pool threads are 16, not a fixed total thread count.
Reused controls and incomplete post-test validation

The report contains 22 new Lavik points and 124 peer points reused from an earlier sweep on the same hosts. Products were not interleaved. Garnet’s pre-test exact count passed, but its post-test full scan was stopped and its post-test count and value-length check remain unverified. All reported GETs had zero misses and all 146 points had zero connection errors; those checks do not replace the omitted validation.

Each point is measured once. Host timing, compiler/build changes, cache history, memory allocations, and persistence differences limit causal claims. Read the full method and scope.

04 / 10M keys

10 million keys. NVMe SSD versus DRAM.

The same hosts as the 1B-key cohort and CPU affinity are used with 1,024-byte values. Points run for 30 seconds at 80–1,280 connections, with a 10-second GET warmup. The client uses the historical scripts’ default correlated seeds, unlike the distinct seeds in the 1B cohort.

Redis 8.8.0 and Valkey 9.1.0 test 1, 2, 4, 8, and 16 I/O threads, with AOF and automatic snapshots disabled. Each thread configuration starts from a freshly generated, hash-verified 10M-key RDB. Lavik stores values on NVMe SSD through its SPDK backend and keeps a key index in DRAM. Redis and Valkey store their values in DRAM. The shared host does not imply equal memory usage or persistence guarantees.

The source report’s peak selections are Redis GET/SET at 16 I/O threads, Valkey GET at 16 and SET at 8. The comparison reproduces those defaults and exposes the other settings.

05 / 200M keys

200 million keys. Nine backend configurations.

Server and client are separate Azure Standard_L16s_v3 VMs. The local systems use two NVMe drives, uniformly random 1,000–4,000-byte values, eight memtier threads with ten connections each, and unlimited-rate 300-second windows. Each backend loads once, then runs GET, 1:1 mixed, and SET in that order without a database restart or page-cache reset.

Lavik uses 16 workers with defragmentation enabled across SPDK, raw-device io_uring, and XFS-file io_uring. Only Dragonfly receives an additional 180-second GET warmup. Other systems retain the state from loading and earlier workloads; background work is included according to the configurations below.

SystemVersion and configuration differences
LavikThree backends. Source report build, distinct from the beta-release test. Defragmentation enabled with workload-specific budgets. About 4 GiB of registered buffers across 16 workers.
Dragonfly1.40.1 · 16 proactors · RAID0/XFS · buffered I/O and normal page cache · cooling disabled.
Garnet2.1.3 · 64 GiB hybrid log + 32 GiB read cache + 4 GiB index · AOF/checkpoints disabled · Lookup compaction every 300 seconds. DBSIZE failure with --no-obj; validation uses successful loads and INFO store.
Apache Kvrocks2.16.0 · 16 workers · 80 GiB block cache · WAL, sync, compression and Blob GC disabled; automatic compaction enabled.
Pika4.0.3 tag, binary reports 4.0.2 · 16 network / 32 request threads · 24 GiB block cache + 32 GiB RTC cache · WAL/binlog and compression disabled; compaction enabled.
Tendis2.8.4-rocksdb-v8.5.3 · 16 executor threads · ten stores, 72 GiB shared cache · WAL/binlog and compression disabled; compaction and Blob GC enabled.
KeyDB On Flash6.3.4 · beta Flash feature · four server threads · 64 GiB hot tier · RocksDB WAL retained. Uniform access differs from the recommended hot-key distribution.

These are configuration-specific observations. The systems do not have equal memory budgets or crash-recovery guarantees. Read every fairness note in the report.

06 / Separate experiment

Azure Managed Redis: context, without a rank.

The managed-service addendum tests a 480 GB / 16 vCPU managed instance, Redis 7.4.3, with 154,088,000 keys and approximately 400 GB of data. It uses 80 connections over a Private Endpoint, TLS disabled, and 60-second windows. Dataset size, network path, duration, and service model differ from the local comparison.

WorkloadQPSp99 (ms)p99.9 (ms)
Read-only GET44,480.155.85511.839
Write-only SET84,863.933.1997.839
1:1 read / write55,351.995.31111.839

The report records 100% GET hits and zero errors in these formal windows. These rows are excluded from the local ranking and no managed-versus-self-managed performance ratio is presented.

The source report’s historical estimates were $2,414.28/month for the managed instance and $1,152.67/month for the Lavik server VM. Both exclude the benchmark client and have different service boundaries. They are not current quotes or full TCO; CacheArena does not derive a price-performance ranking from them.

Read the separate capacity and performance test

07 / Evidence

Follow a number back to its source.

Source content is pinned to Lavik repository revision d88a5e8fdcc5702272b26f9853f4baad7a146988. Downloads preserve the original reports and their copyright notices. The published data is transcribed or parsed from those sources; no performance values are estimated.

200M keys: storage-tier comparison

27 local measurements · 3 separate Azure measurements · reproduction protocol

Pinned reportReport snapshot

Normalized dataset (JSON) · Source checksums and revision · Source Apache 2.0 license

Reproduction requires dedicated test hardware

The original acquisition scripts include host-specific device allowlists and destructive scratch-storage preparation. Review and adapt the original instructions for verified disposable devices; do not execute them against an existing database.

To propose a correction, include the source row, product version, workload, settings, and supporting evidence in a repository issue.