SysDeck/klanker-gate/docs/benchmark-report.md

365 lines
19 KiB
Markdown
Executable File

# Benchmark and Performance Matrix
Measured 2026-07-29 · Gateway v0.9.0 · Deno 2.9.3 (V8 14.9.207.2-rusty) ·
PostgreSQL 18 (`pgvector/pgvector:0.8.5-pg18`) · Windows 10 host, 8 logical CPUs
/ 16 GB · Docker Desktop 29.6.2 (Linux containers, 8 CPUs visible).
Historical runs from 2026-07-13 and 2026-07-22 are preserved in the
[appendix](#appendix-historical-waves), which is where the pre-PostgreSQL
numbers now live. Everything above the appendix was measured on the current
tree.
## Two harnesses, two questions
`scripts/load-bench.ts` (`deno task test:load`) measures **end-to-end request
cost** and is the source of sections 1-5. `deno task bench` measures **single
pure functions** on the hot path, which is the one thing the load harness cannot
isolate; it is the source of section 6 (decision-log 78, superseding 76). Never
carry a number from one into a sentence about the other.
## Method
Two harness shapes for the end-to-end numbers in sections 1-5, because one shape
cannot answer both questions. Section 6 is a third shape and is described there.
**In-process.** Client, gateway, and mock upstream share a single Deno event
loop. This isolates gateway overhead from provider latency and is the right
shape for comparing one commit against another, which is what the measure-first
rule in decision-log 45 needs. It cannot measure process count: a multi-process
gateway cannot be hosted inside the load generator's own event loop.
**External target.** The load generator and the mock upstream run in a container
on the same Docker network as a real, already-running gateway, so the gateway is
a separate process tree with a real PostgreSQL behind it. Keeping the generator
inside the container network matters: driving it from the Windows host instead
put Docker Desktop's NAT hop in the measurement path and capped throughput
before the gateway did.
Latency is per request including full body drain. Unless stated otherwise N=3000
with a discarded warmup, and one request in five uses the SSE streaming path.
**Never compare a number from one shape against a number from the other.** The
container rows do strictly more work than the in-process rows (a real PostgreSQL
round trip, a real network hop to the upstream) and are only meaningful against
each other.
### Reproduce
```bash
# In-process, concurrency sweep
deno task test:load # 500 / 50
deno run --allow-net --allow-env scripts/load-bench.ts 3000 100
# Streaming share: 0 = never, 1 = always, N = every Nth (default 5)
FROSTY_BENCH_STREAM_EVERY=0 deno run --allow-net --allow-env \
scripts/load-bench.ts 3000 50
# External target: pin the mock upstream, point the gateway at it, then drive
FROSTY_BENCH_TARGET=http://frosty-bench:8080 \
FROSTY_BENCH_UPSTREAM_PORT=9099 \
deno run --unstable-net --allow-net --allow-env scripts/load-bench.ts 3000 100
```
`FROSTY_BENCH_TARGET`, `FROSTY_BENCH_UPSTREAM_PORT`,
`FROSTY_BENCH_STREAM_EVERY`, `FROSTY_BENCH_MODEL`, and `FROSTY_BENCH_KEY` are
documented in the header of [`scripts/load-bench.ts`](../scripts/load-bench.ts).
The positional `load-bench.ts [requests] [concurrency]` contract is unchanged.
The mock upstream is served in both modes, and an external gateway only dials it
per request, so there is no start-order dependency.
## Matrix
### 1. Concurrency (in-process, mixed streaming)
| Concurrency | RPS | p50 | p95 | p99 | max | Failures |
| ----------- | ---------- | --------- | --------- | --------- | --------- | -------- |
| 1 | 670.5 | 1.14 ms | 3.08 ms | 4.91 ms | 31.81 ms | 0 |
| 10 | 669.6 | 12.68 ms | 25.87 ms | 38.58 ms | 76.44 ms | 0 |
| 50 | **1069.2** | 43.21 ms | 66.25 ms | 125.97 ms | 134.06 ms | 0 |
| 100 | 999.2 | 90.85 ms | 139.86 ms | 258.55 ms | 288.32 ms | 0 |
| 200 | 821.4 | 200.98 ms | 575.61 ms | 727.06 ms | 803.17 ms | 0 |
Peak is at 50 concurrent. Past that, p50 grows almost linearly while throughput
falls: queue wait, not processing cost.
### 2. Streaming share (in-process)
| Concurrency | Non-streaming RPS | All-streaming RPS | Cost of streaming |
| ----------- | ----------------- | ----------------- | ----------------- |
| 10 | 1035.0 | 953.1 | -8% |
| 50 | 1005.0 | 818.1 | -19% |
| 100 | 931.5 | 831.8 | -11% |
p50 non-streaming / all-streaming: 8.40 / 9.28 ms at c=10, 43.26 / 52.76 ms at
c=50, 86.59 / 115.51 ms at c=100. Zero failures throughout. The passive tee
(`withStreamCompletion` + `StreamAccumulator`) is the cost here, and it is the
price of reconstructing the assistant message for plugins, cost, and logging
without buffering the client stream.
### 3. Process count (container, PostgreSQL state, log store ON)
This is the shipped default shape: `FROSTY_LOG_STORE=pg`. N=3000, c=100. Process
counts were confirmed inside the container: 1, 3, 5, 9 running
`apps/gateway/main.ts` = supervisor + N children.
| `FROSTY_WORKERS` | Serving procs | RPS | p50 | p95 | p99 | max | Failures |
| ---------------- | ------------- | --------- | --------- | --------- | --------- | ---------- | -------- |
| unset / 1 | 1 | 553.6 | 159.98 ms | 356.39 ms | 437.89 ms | 533.65 ms | 0 |
| 2 | 1 + 2 | **732.0** | 121.15 ms | 268.59 ms | 390.99 ms | 490.25 ms | 0 |
| 4 | 1 + 4 | 680.2 | 108.34 ms | 359.35 ms | 433.95 ms | 514.75 ms | 0 |
| 8 | 1 + 8 | 419.1 | 169.69 ms | 602.83 ms | 792.08 ms | 1280.88 ms | 0 |
### 4. Durable log store on vs off (container)
Identical to section 3 with `FROSTY_LOG_STORE=off`. This isolates the cost of
the durable request-log write, which is a PostgreSQL round trip per request.
| `FROSTY_WORKERS` | Log store ON | Log store OFF | Delta |
| ---------------- | ------------ | -------------- | -------- |
| 1 | 553.6 rps | **902.4 rps** | **+63%** |
| 2 | 732.0 rps | **1158.7 rps** | **+58%** |
| 4 | 680.2 rps | **1153.7 rps** | **+70%** |
p50 with the store off: 95.62 ms (1 worker), 77.19 ms (2), 68.92 ms (4). Zero
failures.
### 5. Boot
| `FROSTY_WORKERS` | Container start to `/healthz` 200 |
| ---------------- | --------------------------------- |
| 1 | 1745 ms |
| 4 | 2469 ms |
Fan-out costs roughly 240 ms per worker at boot: each child opens its own
PostgreSQL pool and its own LISTEN connection.
### 6. Micro-benchmarks (`deno task bench`)
Six `*_bench.ts` files sit next to the source they measure. They are pure, need
no network, database, or gateway, and run in about 40 seconds. Run them with:
```bash
deno task bench # everything
deno bench --unstable-net --allow-net --allow-env --allow-read \
packages/core/src/translate_bench.ts # one file
```
Numbers below are one run on the reference box (Deno 2.9.4, Intel i7-6700HQ,
Windows 10). **They are a comparison instrument, not a spec.** Re-run the suite
on the machine you are optimizing on, before and after the change, and apply
decision-log 45: a win inside run-to-run variance is rejected.
**Cache key projection** (`packages/cache/src/semantic_bench.ts`,
`store_bench.ts`). `keyFor` runs on every cacheable request; `digestKey` runs
only when a shared L2 tier is attached.
| Benchmark | time/iter | iter/s |
| ------------------------------------------ | --------- | --------- |
| `keyFor`, 2-message request | 625.0 ns | 1,600,000 |
| `keyFor`, 21-message thread + 2 tools | 9.9 µs | 101,200 |
| `keyFor`, same thread, excludeSystemPrompt | 10.5 µs | 95,220 |
| `promptText`, 21-message thread | 1.4 µs | 718,500 |
| `digestKey`, 120-byte key | 54.9 µs | 18,220 |
| `digestKey`, 37 KB key | 265.7 µs | 3,764 |
**Governance admission** (`packages/governance/src/virtual_keys_bench.ts`).
`check` is the synchronous half of the admission path, once per governed
request.
| Benchmark | time/iter | iter/s |
| ----------------------------------------------- | --------- | ----------- |
| `check`, admit (hash, lookup, budgets, windows) | 3.2 µs | 316,300 |
| `check`, reject unknown token | 2.4 µs | 409,400 |
| `hashVirtualKeyToken` alone | 2.1 µs | 474,000 |
| chars/4 token estimate, 6 KB body | 8.2 ns | 122,000,000 |
| `JSON.parse` model extraction, same 6 KB body | 5.7 µs | 176,900 |
**Stream translation and accumulation** (`packages/core/src/translate_bench.ts`,
`accumulate_bench.ts`). One iteration processes a whole 74-chunk response, not
one chunk.
| Benchmark | time/iter | iter/s |
| --------------------------------------------------- | --------- | ------- |
| identity passthrough, 74 chunks (Web Streams floor) | 119.6 µs | 8,359 |
| Anthropic Messages, 74 chunks | 644.5 µs | 1,551 |
| OpenAI Responses, 74 chunks | 629.2 µs | 1,589 |
| Google GenAI, 74 chunks | 340.1 µs | 2,941 |
| Cohere v2, 74 chunks | 369.2 µs | 2,709 |
| legacy completions, 74 chunks | 372.5 µs | 2,684 |
| `StreamAccumulator`, 67 text chunks | 2.0 µs | 493,400 |
| `StreamAccumulator`, 15 tool-call chunks | 1.5 µs | 668,300 |
**Thinking-block translation.** Decision-log 79 gave `AnthropicStreamTranslator`
a thinking-block lifecycle for canonical reasoning deltas, and item 78 requires
a changed translator to carry a bench row. The base fixture has no `reasoning`
deltas, so without this case the path would ship unmeasured. Re-run on a quieter
machine than the table above, so compare the two rows to each other rather than
to the numbers above them.
| Benchmark | time/iter | iter/s |
| ------------------------------------------------ | --------- | ------ |
| identity passthrough, 74 chunks (floor) | 93.3 µs | 10,720 |
| Anthropic Messages, 74 chunks | 293.4 µs | 3,408 |
| Anthropic Messages + thinking blocks, 106 chunks | 473.8 µs | 2,111 |
Per chunk that is 3.96 µs without reasoning deltas and 4.47 µs with them: the
extra wall-clock is the extra 32 chunks, not a per-chunk regression, so the
thinking-block lifecycle costs about what an equivalent text block costs. Well
inside the noise band item 45 requires a win to clear, which is the point - this
row exists to catch a future regression, not to claim a gain.
**State key encoding** (`packages/config/src/store_bench.ts`). `keyId` runs on
every durable read, write, and counter operation; `hasPrefix` and `compareKeys`
run per candidate row inside a prefix listing.
| Benchmark | time/iter | iter/s |
| ----------------------------- | --------- | ---------- |
| `keyId`, 3-part path | 213.2 ns | 4,690,000 |
| `keyId`, 5-part path | 288.7 ns | 3,464,000 |
| `hasPrefix`, 5-part vs 3-part | 48.0 ns | 20,850,000 |
| `compareKeys`, two 5-part | 48.9 ns | 20,460,000 |
What section 6 says:
- **Nothing on the pure hot path is worth optimizing.** The most expensive
per-request pure function measured is `keyFor` on a 21-message thread at ~10
µs, against a p50 of 1.14 ms for a whole in-process request (section 1) and
100 ms to 10 s for a real upstream call. Governance admission is 3.2 µs, which
independently reproduces the "~0.1 ms" end-to-end result from the 2026-07-22
addendum below.
- **The translators are dominated by Web Streams plumbing, not by translation.**
The identity-passthrough baseline is 119.6 µs of every translator's 340-645 µs
over the same 74 chunks. The spread _between_ translators moved by more than
2x across repeat runs on this box, so treat it as noise; the plumbing floor is
the finding. Anyone proposing a faster translator should first check whether
they are proposing a faster `TransformStream`.
- **The `chars/4` estimate is free; the parse next to it is not.** 8.2 ns
against 5.7 µs for the `JSON.parse` the same middleware does to recover the
requested model, roughly 700x. Replacing the estimate with a real tokenizer
would be a correctness argument, never a performance one.
- **Token hashing is most of admission.** 2.1 µs of the 3.2 µs `check` is the
SHA-256 of the bearer token. That is the deliberate cost of never storing a
raw token, and it is not a candidate for removal.
- **`digestKey` is the one L2 cost worth knowing.** 54.9 µs for a small key, and
it is paid twice per request (read and write) when a shared tier is attached.
This is why L1 keeps using the raw string and never hashes.
## What the numbers mean
**The durable log store is the throughput ceiling, not the gateway.** Turning it
off is worth 58-70% across every process count. It is on by default
(`FROSTY_LOG_STORE=pg`), which is the right default - the dashboard trail is the
product - but an operator who needs throughput more than history has one knob
that moves more than anything else in this document.
**Fan-out buys about one doubling on this box, and the knee is early.** One to
two workers is +32% with the log store on and +28% with it off. Two to four is
flat. Eight workers is _worse than one_. With 8 logical CPUs shared between the
gateway, PostgreSQL, and the in-network load generator, 8 workers oversubscribes
the machine, and the p99 and max columns show it: 792 ms and 1281 ms against 438
ms and 534 ms at a single worker. The transferable result is the shape, not the
peak: set `FROSTY_WORKERS` well below core count when the database and the load
source are co-resident, and measure rather than assuming N cores means N
workers.
**Latency under load is queuing, not work.** p50 tracks concurrency almost
linearly while throughput plateaus, in both harness shapes. That is Little's Law
queue wait, and it reproduces the conclusion the 2026-07-22 addendum reached on
the pre-PostgreSQL backend: there is no per-request hot-path bottleneck.
**Streaming costs 8-19%,** and the cost grows with concurrency.
**Nothing failed.** Zero failed requests across every run in every section, at
up to 200 concurrent connections and up to 9 processes.
## Not measured, and why
- **Real provider latency and throughput.** Network-dominated and
provider-specific. A real LLM call is 100 ms to 10 s upstream, which is one to
four orders of magnitude above anything in this document.
- **Separate replicas behind a load balancer,** as opposed to workers sharing a
port. Note this is the topology where `FROSTY_SHARED_RATE_LIMIT=auto` cannot
help: `auto` keys off `FROSTY_WORKERS`, which a second machine does not see.
Set it to `on` explicitly. See [multi-process.md](guides/multi-process.md).
- **Long-haul soak and memory growth** over hours.
- **L1/L2 cache hit-path throughput.**
- **Host-native multi-process.** Windows has no SO_REUSEPORT, so `planCluster()`
serves single-process by design and says so on the boot line. Every
process-count row here is containerized.
- **PgBouncer under the `pgbouncer` profile.** The connection budget is
`workers x FROSTY_PG_POOL_SIZE + workers`; the pooler is the escape hatch past
`max_connections=200`, and it is untested for throughput here.
## Appendix: historical waves
Retained for the trend line. **These predate the PostgreSQL consolidation**
(decision-log 57, 61) and were measured against the retired Deno KV backend on
Deno 2.9.x. Their storage claims no longer describe the system; their latency
conclusions do.
### Waves 1-2, 2026-07-13
| Run | Requests | Concurrency | RPS | p50 | p95 | p99 | max | Failures |
| --------------- | -------- | ----------- | ----- | -------- | -------- | -------- | -------- | -------- |
| Baseline | 500 | 50 | 691.9 | 58.6 ms | 180.3 ms | 185.1 ms | 190.8 ms | 0 |
| Stress | 2000 | 100 | 679.3 | 117.5 ms | 224.3 ms | 627.7 ms | 672.4 ms | 0 |
| Wave-2 re-check | 500 | 50 | - | 54.1 ms | 150.4 ms | 156.0 ms | 163.8 ms | 0 |
Wave-2 re-check (post streaming/MCP/vector changes): latency held or improved at
every percentile with zero failures.
### Addendum 2026-07-22: governed hot-path profile
Differential profiling of the fully governed request path after the
hash-token/reserve-before-admit hardening, attributing latency by isolating one
factor at a time. In-process harness, mock upstream, N=3000 per stage.
Per-request cost, p50, single request in flight:
| Stage | End-to-end | Handler-direct |
| --------------------------------- | ---------- | -------------- |
| A · direct mock (HTTP floor) | 0.18 ms | - |
| B · gateway, ungoverned, tiny | 1.03 ms | 0.51 ms |
| C · gateway, **governed**, tiny | 1.06 ms | 0.61 ms |
| D · gateway, governed, 200 KB | 4.06 ms | 2.37 ms |
| Ds · gateway, governed, streaming | 1.23 ms | - |
Conclusions that still hold on the current backend:
- The full governance admission pipeline (hash lookup, hierarchy, scope,
reserve, epoch admit, cost accounting) adds **~0.1 ms** (C minus B). That is
under 1% of any real LLM call, so no hot-path optimization is warranted.
- The only input-scaling cost is body size, roughly 1.8 ms server-side per 200
KB, dominated by the unavoidable read plus re-serialization. A measured
parse-once optimization returned a win inside noise and was rejected under
decision-log 45.
Conclusions that have since been overtaken:
- "The lever is horizontal scaling, which first requires externalizing the
single-node KV state." That work shipped: state is in PostgreSQL, budgets and
rate-limit windows are fleet-wide, and section 3 above measures the result.
### Packaging verification, 2026-07-13 (not re-run)
- **Docker**: image built, container booted, `/healthz` 200, schema validation
live inside the container. The KV volume persistence noted at the time no
longer applies - durable state is in the `postgres-data` volume.
- **`deno compile`**: 85.5 MB standalone executable, booted and served
(`/healthz`, `/v1/models`). Caveats recorded then: the unstable flags and
permission set are baked at compile time, and the same-origin UI needs
`apps/control-ui/dist` next to the binary or falls back to API-only mode.
**Not re-verified on the current tree**, and the flag set has changed
(`--unstable-net` replaced `--unstable-kv`), so treat the size and the verdict
as historical. A compiled binary now also needs a reachable `FROSTY_PG_URL`
like every other run mode.
## Related
- [Running multiple processes](guides/multi-process.md) - what `FROSTY_WORKERS`
does, where it does not work, and the connection budget.
- [TODO.md](../TODO.md) - the register that replaced the retired decision log.
The numbers this report cites as provenance are items 45 (measure-first), 57
and 61 (PostgreSQL), 62 (multi-process), 71 (shared rate limits) and 78 (the
micro-benchmark suite, superseding 76); they no longer resolve to a file.