SysDeck/klanker-gate/docs/benchmark-report.md

19 KiB
Executable File

Benchmark and Performance Matrix

Measured 2026-07-29 · Gateway v0.9.0 · Deno 2.9.3 (V8 14.9.207.2-rusty) · PostgreSQL 18 (pgvector/pgvector:0.8.5-pg18) · Windows 10 host, 8 logical CPUs / 16 GB · Docker Desktop 29.6.2 (Linux containers, 8 CPUs visible).

Historical runs from 2026-07-13 and 2026-07-22 are preserved in the appendix, which is where the pre-PostgreSQL numbers now live. Everything above the appendix was measured on the current tree.

Two harnesses, two questions

scripts/load-bench.ts (deno task test:load) measures end-to-end request cost and is the source of sections 1-5. deno task bench measures single pure functions on the hot path, which is the one thing the load harness cannot isolate; it is the source of section 6 (decision-log 78, superseding 76). Never carry a number from one into a sentence about the other.

Method

Two harness shapes for the end-to-end numbers in sections 1-5, because one shape cannot answer both questions. Section 6 is a third shape and is described there.

In-process. Client, gateway, and mock upstream share a single Deno event loop. This isolates gateway overhead from provider latency and is the right shape for comparing one commit against another, which is what the measure-first rule in decision-log 45 needs. It cannot measure process count: a multi-process gateway cannot be hosted inside the load generator's own event loop.

External target. The load generator and the mock upstream run in a container on the same Docker network as a real, already-running gateway, so the gateway is a separate process tree with a real PostgreSQL behind it. Keeping the generator inside the container network matters: driving it from the Windows host instead put Docker Desktop's NAT hop in the measurement path and capped throughput before the gateway did.

Latency is per request including full body drain. Unless stated otherwise N=3000 with a discarded warmup, and one request in five uses the SSE streaming path.

Never compare a number from one shape against a number from the other. The container rows do strictly more work than the in-process rows (a real PostgreSQL round trip, a real network hop to the upstream) and are only meaningful against each other.

Reproduce

# In-process, concurrency sweep
deno task test:load                                            # 500 / 50
deno run --allow-net --allow-env scripts/load-bench.ts 3000 100

# Streaming share: 0 = never, 1 = always, N = every Nth (default 5)
FROSTY_BENCH_STREAM_EVERY=0 deno run --allow-net --allow-env \
  scripts/load-bench.ts 3000 50

# External target: pin the mock upstream, point the gateway at it, then drive
FROSTY_BENCH_TARGET=http://frosty-bench:8080 \
FROSTY_BENCH_UPSTREAM_PORT=9099 \
  deno run --unstable-net --allow-net --allow-env scripts/load-bench.ts 3000 100

FROSTY_BENCH_TARGET, FROSTY_BENCH_UPSTREAM_PORT, FROSTY_BENCH_STREAM_EVERY, FROSTY_BENCH_MODEL, and FROSTY_BENCH_KEY are documented in the header of scripts/load-bench.ts. The positional load-bench.ts [requests] [concurrency] contract is unchanged. The mock upstream is served in both modes, and an external gateway only dials it per request, so there is no start-order dependency.

Matrix

1. Concurrency (in-process, mixed streaming)

Concurrency RPS p50 p95 p99 max Failures
1 670.5 1.14 ms 3.08 ms 4.91 ms 31.81 ms 0
10 669.6 12.68 ms 25.87 ms 38.58 ms 76.44 ms 0
50 1069.2 43.21 ms 66.25 ms 125.97 ms 134.06 ms 0
100 999.2 90.85 ms 139.86 ms 258.55 ms 288.32 ms 0
200 821.4 200.98 ms 575.61 ms 727.06 ms 803.17 ms 0

Peak is at 50 concurrent. Past that, p50 grows almost linearly while throughput falls: queue wait, not processing cost.

2. Streaming share (in-process)

Concurrency Non-streaming RPS All-streaming RPS Cost of streaming
10 1035.0 953.1 -8%
50 1005.0 818.1 -19%
100 931.5 831.8 -11%

p50 non-streaming / all-streaming: 8.40 / 9.28 ms at c=10, 43.26 / 52.76 ms at c=50, 86.59 / 115.51 ms at c=100. Zero failures throughout. The passive tee (withStreamCompletion + StreamAccumulator) is the cost here, and it is the price of reconstructing the assistant message for plugins, cost, and logging without buffering the client stream.

3. Process count (container, PostgreSQL state, log store ON)

This is the shipped default shape: FROSTY_LOG_STORE=pg. N=3000, c=100. Process counts were confirmed inside the container: 1, 3, 5, 9 running apps/gateway/main.ts = supervisor + N children.

FROSTY_WORKERS Serving procs RPS p50 p95 p99 max Failures
unset / 1 1 553.6 159.98 ms 356.39 ms 437.89 ms 533.65 ms 0
2 1 + 2 732.0 121.15 ms 268.59 ms 390.99 ms 490.25 ms 0
4 1 + 4 680.2 108.34 ms 359.35 ms 433.95 ms 514.75 ms 0
8 1 + 8 419.1 169.69 ms 602.83 ms 792.08 ms 1280.88 ms 0

4. Durable log store on vs off (container)

Identical to section 3 with FROSTY_LOG_STORE=off. This isolates the cost of the durable request-log write, which is a PostgreSQL round trip per request.

FROSTY_WORKERS Log store ON Log store OFF Delta
1 553.6 rps 902.4 rps +63%
2 732.0 rps 1158.7 rps +58%
4 680.2 rps 1153.7 rps +70%

p50 with the store off: 95.62 ms (1 worker), 77.19 ms (2), 68.92 ms (4). Zero failures.

5. Boot

FROSTY_WORKERS Container start to /healthz 200
1 1745 ms
4 2469 ms

Fan-out costs roughly 240 ms per worker at boot: each child opens its own PostgreSQL pool and its own LISTEN connection.

6. Micro-benchmarks (deno task bench)

Six *_bench.ts files sit next to the source they measure. They are pure, need no network, database, or gateway, and run in about 40 seconds. Run them with:

deno task bench                                     # everything
deno bench --unstable-net --allow-net --allow-env --allow-read \
  packages/core/src/translate_bench.ts              # one file

Numbers below are one run on the reference box (Deno 2.9.4, Intel i7-6700HQ, Windows 10). They are a comparison instrument, not a spec. Re-run the suite on the machine you are optimizing on, before and after the change, and apply decision-log 45: a win inside run-to-run variance is rejected.

Cache key projection (packages/cache/src/semantic_bench.ts, store_bench.ts). keyFor runs on every cacheable request; digestKey runs only when a shared L2 tier is attached.

Benchmark time/iter iter/s
keyFor, 2-message request 625.0 ns 1,600,000
keyFor, 21-message thread + 2 tools 9.9 µs 101,200
keyFor, same thread, excludeSystemPrompt 10.5 µs 95,220
promptText, 21-message thread 1.4 µs 718,500
digestKey, 120-byte key 54.9 µs 18,220
digestKey, 37 KB key 265.7 µs 3,764

Governance admission (packages/governance/src/virtual_keys_bench.ts). check is the synchronous half of the admission path, once per governed request.

Benchmark time/iter iter/s
check, admit (hash, lookup, budgets, windows) 3.2 µs 316,300
check, reject unknown token 2.4 µs 409,400
hashVirtualKeyToken alone 2.1 µs 474,000
chars/4 token estimate, 6 KB body 8.2 ns 122,000,000
JSON.parse model extraction, same 6 KB body 5.7 µs 176,900

Stream translation and accumulation (packages/core/src/translate_bench.ts, accumulate_bench.ts). One iteration processes a whole 74-chunk response, not one chunk.

Benchmark time/iter iter/s
identity passthrough, 74 chunks (Web Streams floor) 119.6 µs 8,359
Anthropic Messages, 74 chunks 644.5 µs 1,551
OpenAI Responses, 74 chunks 629.2 µs 1,589
Google GenAI, 74 chunks 340.1 µs 2,941
Cohere v2, 74 chunks 369.2 µs 2,709
legacy completions, 74 chunks 372.5 µs 2,684
StreamAccumulator, 67 text chunks 2.0 µs 493,400
StreamAccumulator, 15 tool-call chunks 1.5 µs 668,300

Thinking-block translation. Decision-log 79 gave AnthropicStreamTranslator a thinking-block lifecycle for canonical reasoning deltas, and item 78 requires a changed translator to carry a bench row. The base fixture has no reasoning deltas, so without this case the path would ship unmeasured. Re-run on a quieter machine than the table above, so compare the two rows to each other rather than to the numbers above them.

Benchmark time/iter iter/s
identity passthrough, 74 chunks (floor) 93.3 µs 10,720
Anthropic Messages, 74 chunks 293.4 µs 3,408
Anthropic Messages + thinking blocks, 106 chunks 473.8 µs 2,111

Per chunk that is 3.96 µs without reasoning deltas and 4.47 µs with them: the extra wall-clock is the extra 32 chunks, not a per-chunk regression, so the thinking-block lifecycle costs about what an equivalent text block costs. Well inside the noise band item 45 requires a win to clear, which is the point - this row exists to catch a future regression, not to claim a gain.

State key encoding (packages/config/src/store_bench.ts). keyId runs on every durable read, write, and counter operation; hasPrefix and compareKeys run per candidate row inside a prefix listing.

Benchmark time/iter iter/s
keyId, 3-part path 213.2 ns 4,690,000
keyId, 5-part path 288.7 ns 3,464,000
hasPrefix, 5-part vs 3-part 48.0 ns 20,850,000
compareKeys, two 5-part 48.9 ns 20,460,000

What section 6 says:

  • Nothing on the pure hot path is worth optimizing. The most expensive per-request pure function measured is keyFor on a 21-message thread at ~10 µs, against a p50 of 1.14 ms for a whole in-process request (section 1) and 100 ms to 10 s for a real upstream call. Governance admission is 3.2 µs, which independently reproduces the "~0.1 ms" end-to-end result from the 2026-07-22 addendum below.
  • The translators are dominated by Web Streams plumbing, not by translation. The identity-passthrough baseline is 119.6 µs of every translator's 340-645 µs over the same 74 chunks. The spread between translators moved by more than 2x across repeat runs on this box, so treat it as noise; the plumbing floor is the finding. Anyone proposing a faster translator should first check whether they are proposing a faster TransformStream.
  • The chars/4 estimate is free; the parse next to it is not. 8.2 ns against 5.7 µs for the JSON.parse the same middleware does to recover the requested model, roughly 700x. Replacing the estimate with a real tokenizer would be a correctness argument, never a performance one.
  • Token hashing is most of admission. 2.1 µs of the 3.2 µs check is the SHA-256 of the bearer token. That is the deliberate cost of never storing a raw token, and it is not a candidate for removal.
  • digestKey is the one L2 cost worth knowing. 54.9 µs for a small key, and it is paid twice per request (read and write) when a shared tier is attached. This is why L1 keeps using the raw string and never hashes.

What the numbers mean

The durable log store is the throughput ceiling, not the gateway. Turning it off is worth 58-70% across every process count. It is on by default (FROSTY_LOG_STORE=pg), which is the right default - the dashboard trail is the product - but an operator who needs throughput more than history has one knob that moves more than anything else in this document.

Fan-out buys about one doubling on this box, and the knee is early. One to two workers is +32% with the log store on and +28% with it off. Two to four is flat. Eight workers is worse than one. With 8 logical CPUs shared between the gateway, PostgreSQL, and the in-network load generator, 8 workers oversubscribes the machine, and the p99 and max columns show it: 792 ms and 1281 ms against 438 ms and 534 ms at a single worker. The transferable result is the shape, not the peak: set FROSTY_WORKERS well below core count when the database and the load source are co-resident, and measure rather than assuming N cores means N workers.

Latency under load is queuing, not work. p50 tracks concurrency almost linearly while throughput plateaus, in both harness shapes. That is Little's Law queue wait, and it reproduces the conclusion the 2026-07-22 addendum reached on the pre-PostgreSQL backend: there is no per-request hot-path bottleneck.

Streaming costs 8-19%, and the cost grows with concurrency.

Nothing failed. Zero failed requests across every run in every section, at up to 200 concurrent connections and up to 9 processes.

Not measured, and why

  • Real provider latency and throughput. Network-dominated and provider-specific. A real LLM call is 100 ms to 10 s upstream, which is one to four orders of magnitude above anything in this document.
  • Separate replicas behind a load balancer, as opposed to workers sharing a port. Note this is the topology where FROSTY_SHARED_RATE_LIMIT=auto cannot help: auto keys off FROSTY_WORKERS, which a second machine does not see. Set it to on explicitly. See multi-process.md.
  • Long-haul soak and memory growth over hours.
  • L1/L2 cache hit-path throughput.
  • Host-native multi-process. Windows has no SO_REUSEPORT, so planCluster() serves single-process by design and says so on the boot line. Every process-count row here is containerized.
  • PgBouncer under the pgbouncer profile. The connection budget is workers x FROSTY_PG_POOL_SIZE + workers; the pooler is the escape hatch past max_connections=200, and it is untested for throughput here.

Appendix: historical waves

Retained for the trend line. These predate the PostgreSQL consolidation (decision-log 57, 61) and were measured against the retired Deno KV backend on Deno 2.9.x. Their storage claims no longer describe the system; their latency conclusions do.

Waves 1-2, 2026-07-13

Run Requests Concurrency RPS p50 p95 p99 max Failures
Baseline 500 50 691.9 58.6 ms 180.3 ms 185.1 ms 190.8 ms 0
Stress 2000 100 679.3 117.5 ms 224.3 ms 627.7 ms 672.4 ms 0
Wave-2 re-check 500 50 - 54.1 ms 150.4 ms 156.0 ms 163.8 ms 0

Wave-2 re-check (post streaming/MCP/vector changes): latency held or improved at every percentile with zero failures.

Addendum 2026-07-22: governed hot-path profile

Differential profiling of the fully governed request path after the hash-token/reserve-before-admit hardening, attributing latency by isolating one factor at a time. In-process harness, mock upstream, N=3000 per stage.

Per-request cost, p50, single request in flight:

Stage End-to-end Handler-direct
A · direct mock (HTTP floor) 0.18 ms -
B · gateway, ungoverned, tiny 1.03 ms 0.51 ms
C · gateway, governed, tiny 1.06 ms 0.61 ms
D · gateway, governed, 200 KB 4.06 ms 2.37 ms
Ds · gateway, governed, streaming 1.23 ms -

Conclusions that still hold on the current backend:

  • The full governance admission pipeline (hash lookup, hierarchy, scope, reserve, epoch admit, cost accounting) adds ~0.1 ms (C minus B). That is under 1% of any real LLM call, so no hot-path optimization is warranted.
  • The only input-scaling cost is body size, roughly 1.8 ms server-side per 200 KB, dominated by the unavoidable read plus re-serialization. A measured parse-once optimization returned a win inside noise and was rejected under decision-log 45.

Conclusions that have since been overtaken:

  • "The lever is horizontal scaling, which first requires externalizing the single-node KV state." That work shipped: state is in PostgreSQL, budgets and rate-limit windows are fleet-wide, and section 3 above measures the result.

Packaging verification, 2026-07-13 (not re-run)

  • Docker: image built, container booted, /healthz 200, schema validation live inside the container. The KV volume persistence noted at the time no longer applies - durable state is in the postgres-data volume.
  • deno compile: 85.5 MB standalone executable, booted and served (/healthz, /v1/models). Caveats recorded then: the unstable flags and permission set are baked at compile time, and the same-origin UI needs apps/control-ui/dist next to the binary or falls back to API-only mode. Not re-verified on the current tree, and the flag set has changed (--unstable-net replaced --unstable-kv), so treat the size and the verdict as historical. A compiled binary now also needs a reachable FROSTY_PG_URL like every other run mode.
  • Running multiple processes - what FROSTY_WORKERS does, where it does not work, and the connection budget.
  • TODO.md - the register that replaced the retired decision log. The numbers this report cites as provenance are items 45 (measure-first), 57 and 61 (PostgreSQL), 62 (multi-process), 71 (shared rate limits) and 78 (the micro-benchmark suite, superseding 76); they no longer resolve to a file.