# tests/integration — cross-lane runtime gates Owned by PKG/QA. Real binaries against `tools/mockd` / `tools/testserver`, headless. Unit tests live in each lane; this tree is for behaviour that only shows up when the pieces run together (throughput, memory over time, reconnect). ## Nightly integration run `tests/integration/nightly_run.py`, wired as the `nightly-integration` job in `ci.yml` (`schedule: '17 3 * * *'`, plus `workflow_dispatch` for an on-demand run). Real `veloxd` + `tools/testserver`, 50 concurrent downloads mixing hostile testserver modes (`flaky-reset`, `throttled`, `no-range`, mostly `plain`), asserting: 1. every completed file's SHA-256 matches testserver's own `//sha256/` route (computed by this script re-hashing the file on disk — never trusting veloxd's own claim of success); 2. every task reaches a terminal state inside the timeout (a task stuck retrying forever is a failure, not a hang for CI's `timeout(1)` to paper over); 3. `veloxd`'s own open-FD count (`/proc//fd`) returns to within a small fixed tolerance of its pre-run baseline, checked after a settle window. `veloxd` runs isolated: `XDG_RUNTIME_DIR`, `XDG_DATA_HOME` and `XDG_CONFIG_HOME` all point into a fresh `mkdtemp()` (not this session's scratch dir — its path is long enough to overflow `AF_UNIX`'s ~108-byte `sun_path`; verified live via "File name too long" before switching to `tempfile.mkdtemp()`). `saveTo.allowedRoots` is seeded straight into `velox.db` after a migrations-only warm-up start: `settings.set` returns `-32603 "not implemented in this build"` on the `veloxd` this job builds — verified live, not assumed — so direct DB seeding is the only entry point that currently exists, not a workaround for a wrong contract. Every task also gets its own `filename` override on `download.add`: tasks share `(mode, size)` pairs by design (testserver's content is a pure function of path, not of who's asking), and without distinct filenames they raced each other for the same destination path — this was caught live on the first real run of this script (48/50 "passed" with io_errors and checksum mismatches on the collided tasks) before the `filename` override was added. ### What makes each check go red, proven once Per the standing note in this repo's history (four green checks that didn't look where the bug was — an always-false guard, a `ctest` label matching zero tests, `--check` validating pkg-config instead of apt names, conformance validating only fixtures): every assertion here was forced red once, on purpose, before being trusted. | Assertion | Forced via | Observed | |---|---|---| | Checksum match | `VDM_NIGHTLY_FORCE_BAD_HASH=1` (substitutes a wrong hash for task 0's comparison only, after the real download and hash succeed) | `FAIL: task 0 (plain, 256K) checksum mismatch: got 6f4c254c…, testserver says 0000…0000`, exit 1 | | FD-leak tolerance | `VDM_NIGHTLY_FORCE_FD_LEAK=1` (adds 25 to the post-run FD count) | `FAIL: veloxd leaked file descriptors: 19 -> 49 (tolerance 10)`, exit 1 | | A task that can't succeed still fails the run | `VDM_NIGHTLY_FORCE_HOSTILE_STALL=1` (task 0's URL points at 192.0.2.1, TEST-NET-1 — unroutable, so the connection just hangs) | `FAIL: task 0 (plain, 256K) ended in state 'failed', error={'code': 'timeout', ...}`, exit 1 | The first version of the `HOSTILE_STALL` hook used testserver's `416-always` mode, expecting it to never yield a 2xx to a Range probe — but `veloxd` correctly falls back to a plain full GET on a 416 and the task completed fine, so that attempt proved nothing (a real finding in itself: worth knowing the daemon handles this correctly). Re-run any of the three whenever the corresponding assertion changes, to re-prove it still catches what it claims to — that's the point of the hooks living in the script rather than being one-off manual edits. ### Proven: the harness cannot leak its daemon A run of this script was itself killed hard (its own harness process, not a graceful stop) mid-download and left `veloxd` running under `systemd --user` for 4h40m — the `finally:` teardown never got to run, because nothing runs after `SIGKILL`. Fixed by giving every child (`veloxd` and `testserver.py`) `PR_SET_PDEATHSIG` (via a `preexec_fn` calling `prctl` through `ctypes`) plus its own process group (`start_new_session=True`): the kernel now delivers `SIGKILL` to a child the instant its parent dies, by any means, without the harness needing to run any code at all. `finally:` still does the graceful SIGTERM-then-SIGKILL `killpg` for the normal-exit path; `PR_SET_PDEATHSIG` is what covers the path `finally:` cannot reach. Proven live: started the harness, waited for its "baseline FDs" log line (proof the real `veloxd` and `testserver.py` were both already up as separate process groups under it), then `kill -9`'d the harness itself and confirmed both children were gone within 1.5 s — nothing left running, nothing to clean up by hand. ### Known limitation: single-instance lock is per-euid, not per-XDG-tree `veloxd`'s single-instance guard binds an abstract-namespace socket keyed only by `geteuid()` (`daemon/src/main.cpp`), so `XDG_RUNTIME_DIR` isolation does not let two `veloxd` processes for the same Unix user run side by side — confirmed live: a second instance exits with "another instance is already running for this user" even with fully distinct `XDG_*` dirs. Harmless on a real CI runner (one job, one user, one `veloxd`) but means this script cannot run concurrently with another `veloxd` on the same machine — worth knowing before parallelizing this job or running it by hand next to another lane's manual testing. ## GUI M1 definition-of-done gates (R3) `gui/docs/pkg-qa-requests-m1.md` R3: three GUI DoD items are not unit tests and have nowhere to run. GUI owns the harness; PKG/QA owns the CI job. This is the wiring contract so the two halves meet without another round trip. ### What GUI built, and the invocation contract `gui/tests/dod/run.sh [--json ]` (`gate` one of `scroll-60fps` / `rss-flat` / `unhappy-path`), backed by `gui/tests/dod/dod_harness.cpp` (target `gui-dod-harness`). Headless by default — `run.sh` sets `QT_QPA_PLATFORM=offscreen` itself, so the CI jobs below need no Xvfb. | `` | Pass / fail condition | Budget | |---|---|---| | `scroll-60fps` | `mockd --tasks 10000`, scripted fling scroll; **fail** if p99 frame > 16.6 ms (×4 under an ASan/UBSan build — `VELOX_DOD_FRAME_BUDGET_MS` overrides outright) | per-PR | | `rss-flat` | `mockd --tasks 10000` + progress events, 10 min; **fail** if RSS growth past warm-up exceeds a 20 MiB slack (`VELOX_DOD_RSS_SLACK_KIB` overrides) | nightly | | `unhappy-path` | Three phases (`--slow` / `--flaky ` / `--drop-connection `), one `mockd` restart each; **fail** on crash, on the harness's own watchdog firing (75 s), or if connection state never (re)reaches `Connected` | per-PR | Contract, as built: exit `0` pass, non-zero fail; a hang is the harness's own watchdog converting itself into a non-zero exit (`3`), never something CI needs `timeout(1)` around; `--json` writes one result object per gate (`unhappy-path` merges its three phases into one file); no network, no writes outside a tempdir, no leaked child process on any exit path (`run.sh`'s `trap cleanup EXIT INT TERM`). ### CI jobs — wired `gui-dod` (per-PR: `scroll-60fps` + `unhappy-path`) and `gui-dod-nightly` (`rss-flat`, `schedule`/`workflow_dispatch` only) are live in `ci.yml`. Both bootstrap, build `gui-dod-harness`, `npm ci` in `tools/mockd`, run `gui/tests/dod/run.sh`, and upload the `--json` output as an artifact — no Xvfb step, since `run.sh` already runs offscreen. ### Forced red, once per gate, before wiring it required | Gate | Forced via | Observed | |---|---|---| | `scroll-60fps` | `VELOX_DOD_FRAME_BUDGET_MS=0.01` | `FAIL: p99=50.73 ms mean=31.34 ms max=52.88 ms budget=0.01 ms over 240 steps, 10000 rows`, exit 1 | | `rss-flat` | `VELOX_DOD_RSS_SLACK_KIB=-999999999` (guarantees `growthKiB > slack` regardless of actual RSS behavior that run) | `FAIL: growth=13644 KiB slack=-999999999 KiB over 15s (warmup 1s), 10000 rows`, exit 1 | | `unhappy-path` | ran `gui-dod-harness` directly (bypassing `run.sh`, which always starts a working `mockd`) against a `--sock` path with nothing listening | 75 s of `Connecting`/`Reconnecting`, then `FAIL: dod_harness watchdog fired — hung past 75s`, exit 3 | The real (non-forced) runs all pass live: `scroll-60fps` p99 51.28 ms against a 66.4 ms budget (ASan/UBSan build), `rss-flat` growth 13504 KiB against the 20480 KiB slack over a 15 s smoke duration, `unhappy-path` all three phases `PASS` with `finalConnected=true`. ### Known gap: `unhappy-path`'s drop-connection phase doesn't drop anything Filed by GUI in `gui/docs/proto-requests-m1.md`: `mockd --drop-connection` only works over WebSocket — `startUds()` never wires the periodic-drop timer `startWs()` has, so the UDS transport (GUI/CLI/nmhost's only one) never sees a connection actually die. GUI's own live verification: 45 s observing a `--drop-connection 5` mockd over UDS, `stateChanged` never fires. The phase still runs and its JSON honestly records `sawDisruption: false` rather than silently passing as if it proved something — that's what the real run above shows, `pass: true` alongside `sawDisruption: false`, so a reviewer reading the JSON sees exactly how much this phase currently covers. This is PROTO's fix, not PKG/QA's or GUI's to route around locally — coordinating on it rather than patching mockd from this lane. Once PROTO lands `dropEverySec` on `startUds`, `unhappy-path`'s drop-connection phase starts exercising a real drop and this note comes out; until then `gui-dod` stays required as specified (crash/hang/never-reconnects still catch real regressions), just not yet catching a swallowed real disconnect. ### Status Wired and required: `gui-dod` runs per-PR, `gui-dod-nightly` on schedule. Revisit once PROTO's UDS `--drop-connection` fix lands (see above).