Files
vdm/core/docs/m7-baseline.md
T
samiandClaude Sonnet 5 b60d4e6f5b core: add tools/bench (throughput/load/alloc-check) and record the M7 baseline
Three subcommands in one binary, driving vdm::Engine directly (docs/04 §8):

- throughput: a single download against a fast local origin
  (support/local_server.hpp, busybox httpd), reporting Mbps/CPU%/RSS. Gates
  on --require-mbps/--max-cpu-pct only when passed, so the ctest smoke
  registration stays a correctness check, not a hardware-dependent
  perf gate -- the real 1-Gbit-link sign-off is a manual/CI job (see the
  file's header comment).
- load: N concurrent tasks against tools/testserver's `throttled` mode
  (support/testserver_client.hpp), reporting peak RSS via getrusage(). Paced
  externally rather than through the engine's own rate::RateLimiter or
  busybox: the limiter's pause/resume path allocates on every throttle event
  (would contaminate alloc-check's measurement) and under heavy segment
  contention was found to starve individual tasks indefinitely (see
  docs/adr/0016, added here); busybox couldn't sustain the DoD's ~160
  concurrent connections (20 tasks * default_segments=8) reliably. The
  ctest registration runs at reduced concurrency under sanitizer presets --
  see the CMakeLists.txt comment and the ADR's postscript.
- alloc-check: operator new/delete overridden process-wide, sampling the
  allocation count across a steady mid-transfer window against a paced
  tools/testserver origin. Caught a real bug in the same change (see the
  http_client.cpp commit) and, by dropping its Engine mid-download to end
  cleanly, also surfaced the quiesce() use-after-free (see that commit).

core/docs/m7-baseline.md records actual measured numbers against the M1/M7
DoD lines, including where they don't clear yet (RSS ~70 MB vs a 60 MB
target; throughput/CPU only measured on loopback, no 1 Gbit link available
here) rather than rounding them away.

docs/adr/0016 documents a rate::RateLimiter fairness gap found building the
load subcommand: a single shared TokenBucket under heavy segment contention
has no fairness ordering across its peek/commit race and can starve a
waiter well past what its configured rate implies. Filed as a follow-up
(it's a core/src/rate design question, not a tools/bench one) rather than
fixed here, along with a related TSan-only load-test straggler that could
not be root-caused in this environment.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01Q3QrF7rCt21bkAjt9BCDFQ
2026-09-11 13:12:57 +04:00

92 lines
5.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M7 performance baseline
Measured against `docs/04-engine-design.md` §8's targets, via `tools/bench/vdm_bench`
(see that file's header comment for the exact commands — reproduced below with their
actual output). `--preset release`, this machine, 2026-09-11. This is a baseline
record, not a sign-off: two of the three numbers below don't clear the DoD line yet, and
that's stated plainly rather than rounded away — see "Open gaps".
## Commands and results
```
$ cmake --preset release && cmake --build --preset release
$ bin/vdm_bench throughput --size 5G --require-mbps 940 --max-cpu-pct 8
throughput: 5368709120 bytes in 2.99s
throughput 14371.00 Mbps
cpu 196.52 % of one core
peak RSS 22.81 MiB
```
Against `tools/bench/support/local_server.hpp`'s busybox loopback server, not a real 1
Gbit link — no such link was available to test against in this environment, so
`--require-mbps`/`--max-cpu-pct` weren't meaningfully exercised here (loopback trivially
clears 940 Mbps; the 196% CPU figure reflects driving a link far faster than 1 Gbit, not
the 1-Gbit-saturated cost the target is about). This needs re-running against a real
1 Gbit peer before it can stand as the actual M1/M7 sign-off number.
```
$ bin/vdm_bench load --tasks 20 --require-rss-kb 61440
load: 20 tasks, 0 failed, 17.06s wall
peak RSS 69.77 MiB
FAIL: peak RSS 71448 KiB > allowed 61440 KiB
```
Default `Config` (`default_segments=8`, `max_active_segments=32`, `default_buffer_bytes=1
MiB`), default `--task-size 4M`. Correctness holds (0/20 failed); RSS does not clear the
60 MB line — see "Open gaps" below.
```
$ bin/vdm_bench alloc-check --size 512M --window-s 2
alloc-check: 4 allocations in 2.00s (budget 15)
```
Clears the no-allocation-on-the-hot-path bar (docs/agents/AGENT-CORE.md) comfortably.
This number is *after* a real fix landed in the same change:
`net::HttpClient::Impl::drain_commands` was constructing an (always-allocating, in
libstdc++) `std::deque` on every worker-loop iteration regardless of whether any command
was actually pending — once per curl_multi_poll wake, i.e. on the transfer hot path. Fixed
by checking `w.queue.empty()` under the lock before touching `local` at all. Before the
fix this bench reported thousands of allocations/sec under any sustained transfer.
## ASan / UBSan / TSan (M1 DoD: "20-task load test... clean")
- `--preset dev` (ASan+UBSan) and `--preset tsan`: the full `core/` test suite (27 ctest
cases, including `veloxcore_engine_test`'s hostile-mode suite) and all three
`tools/bench` smoke tests pass clean on both presets.
- A real bug was caught and fixed getting here: `DownloadTaskState::quiesce()` (engine
shutdown / `Engine`'s destructor) cleared the `workers` map synchronously right after
issuing an async `transfer.cancel()`, racing the HttpClient worker thread's still-in-flight
write callback into a heap-use-after-free on the segment's ring buffer — ASan-caught via
`alloc-check`, which (by design) drops its `Engine` while a download is still active.
Fixed by having `quiesce()` wait for each worker to drain itself through the same
`seg_finished` path every other exit uses, instead of tearing the map down itself.
- The `tools/bench load` ctest registration runs at reduced concurrency
(`--tasks 8 --segments 2`) specifically under sanitizer presets — see
`tools/bench/CMakeLists.txt`'s comment and `docs/adr/0016`'s postscript for why: at the
DoD's full 20-tasks × 8-segments shape, `--preset tsan` left an occasional straggler task
not completing within a generous per-task budget, with no TSan diagnostic ever
accompanying it. Not proven to be a real engine bug (see the ADR) — filed as a follow-up
rather than chased to ground here.
## Open gaps
1. **RSS is ~70 MB against a 60 MB target (~10 MB over, ~18%).** `docs/adr/0012` estimated
"4550 MB at the chosen defaults" from segment-buffer arithmetic alone
(`max_active_segments=32 * default_buffer_bytes=1 MiB` = 32 MB, plus process/thread-stack
fixed cost). A minimal single-tiny-task run here measured that fixed cost at ~14.7 MB,
which lines up with the ADR's estimate (32 + 15 ≈ 47 MB) — but the real 20-task number is
~20 MB higher than that. Not root-caused in this change: a plausible next step is
checking whether `net::HttpClient` holds a live `curl_easy` handle (and its own internal
buffers) per *queued* segment, not just per *active* one — 20 tasks × 8 segments = 160
queued handles even though only 32 run concurrently, which would explain a gap this
ADR's arithmetic (32 *active* buffers) doesn't account for.
2. **Throughput/CPU numbers are loopback-only.** No 1 Gbit link was available to test
against; re-run `throughput --size 5G --require-mbps 940 --max-cpu-pct 8` against a real
one before treating this as signed off.
3. **`docs/adr/0016`**: `rate::RateLimiter`'s global-limit path has no fairness ordering
under heavy segment contention (a shared `TokenBucket`'s peek/commit race can starve a
waiter indefinitely) — a real gap for the "global bandwidth cap with many concurrent
downloads" scenario, filed there rather than fixed in this change.
4. **The TSan-only load-test straggler** noted above (`docs/adr/0016`'s postscript) —
not root-caused; needs reproducing outside a shared/virtualized sandbox to tell "TSan is
just slow here" apart from a real timing-sensitive bug (a plausible candidate named in
the ADR: `CURLOPT_LOW_SPEED_TIME` false-tripping under TSan's slowdown).