LSM tuning knobs include memtable size, bloom bits per key, compaction style, and block cache. RocksDB benchmarks report write amplification, read amplification, and p99 latency together because optimizing one metric often hurts another.
Workload Dimensions
Sequential write, random read, mixed read/write, and range scan. Track bytes written per user byte inserted.
c
Loading…
Warm up before timing steady-state compaction
Compare size-tiered vs leveled compaction policies
Plot latency histograms, not just averages
Document hardware: SSD vs HDD changes winners
Working Through the Exercise
Keep the relevant documentation open while you implement. When your output disagrees with the reference, trace one failing case by hand before changing random lines.
Trace first: walk one failing input step by step
Read the manual: skim the official doc for the mechanism named in this task
Compare baselines: save measurements from the prior step so speedups are honest
Why for this exercise
You will run a modeled or implemented KV workload and print throughput plus amplification metrics. This exercise requires interpreting one tradeoff between write amplification and read latency.
Document one invariant you will assert in tests and how you would detect its violation from observable symptoms.
Write a C program that models an LSM benchmark from workload parameters: amplification metrics, disk bytes, and sustainable write rate.
Input (stdin, whitespace-separated):
ops: number of operations in the run
read_cost_us: cost of one disk read (microseconds, may be fractional)
write_amp: write amplification factor: disk bytes written per byte the user wrote (RocksDB's leveled compaction typically lands at 10-30x)
bytes_per_op: average bytes a user operation writes
Model:
avg_put_us = read_cost_us * write_amp (a write pays compaction re-reads/re-writes; read amplification multiplies the per-read cost)