Files
vdm/tests/conformance/README.md
T
samiandClaude Sonnet 5 5b03eff926 proto: wire veloxd into the conformance suite's canonical entry point
run.sh's TS runner only ever started mockd; veloxd existed but nothing in
the ctest -L conformance path touched it, so mockd's always-valid fixtures
were the only thing ts/replay.ts ever saw. That let a real bug through:
veloxd's download.get can return segments: 0, which TaskSummary.segments
forbids (minimum 1, required) — nothing caught it.

Add a 3b step to run.sh (still the one canonical entry, per ADR 0014):
builds veloxd, starts it isolated (its own XDG_RUNTIME_DIR/XDG_DATA_HOME/
XDG_CONFIG_HOME), seeds saveTo.allowedRoots/defaultDir directly into the
isolated velox.db (settings.set is itself a stub, and the default
~/Downloads root doesn't isolate download.add's writes), then replays
every fixture against it over both transports.

Most handlers are still stubs (daemon/docs/deferrals.md D1-D4b). Fixtures
that hit them get an expected-failure entry in the new veloxd-xfail.json,
loaded by replay.ts's new --xfail flag. This is a maintained allowlist,
not a snapshot: a listed fixture that unexpectedly *passes* is flipped
back to a failure (applyXfail), so the list can only shrink as DAEMON
lands handlers, never rot into a list nobody rechecks. download.get's
segments: 0 is deliberately *not* on it — that's the regression this
step exists to catch.

Also hardened setupBindings: a server that can't even complete fixture
binding used to take the whole runner down with an uncaught exception
before a single fixture was checked. It's now a reported Outcome instead,
so the run still produces a coherent report. That robustness fix earned
its keep immediately: veloxd's download.add crashes on startMode
"later" (a valid, documented StartMode — "the File Info dialog's
Download Later button") with a SQLite CHECK constraint violation, because
migrations/0001_initial.sql's start_mode CHECK never had 'later' added to
it (and includes 'manual'/'auto', neither a contract value). That's a
second, more severe bug this wiring found, unrelated to segments: 0 and
currently blocking most of the veloxd run — filed for DAEMON in
tests/conformance/README.md, not fixed here (out of lane). capture.offer
and capture.getRules are also stubs but missing from deferrals.md's
D-list; xfailed with a note asking DAEMON to add the row.

Verified live once against a real, isolated veloxd before this session's
sandbox became persistently contended for veloxd's single-instance lock
(UID-scoped, not namespaced by XDG_RUNTIME_DIR — daemon/src/main.cpp;
documented as a caveat in the README): it built, started isolated, seeded
settings, connected over both transports, and surfaced the startMode bug
above as a real, non-xfailed failure — confirming the whole pipeline
including --xfail end to end. segments: 0 is confirmed by direct reading
of daemon/src/store/tasks.{hpp,cpp} (TaskRow::eff_segments defaults to 0,
copied verbatim into TaskSummary.segments) rather than by a second live
run reaching that specific fixture, since setup itself fails first on the
startMode bug above. mockd path re-verified green after these changes
(200/200, up from 196/196 — the new setup/$taskId outcomes are visible
and passing).

Recommendation for PKG: don't flip this required yet. The existing
`conformance` ctest entry is already a required check, and right now the
startMode bug fails most of the veloxd run, not just the one expected
segments: 0 case — merging as-is would block every lane's PRs on two
DAEMON bugs at once, one of them unrelated to what this task set out to
catch. Required once DAEMON lands a fix for startMode "later" at minimum;
segments: 0 can stay red for a while by design, same as any other tracked
regression.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01SFeUKLbdHizrJjLBeK7ffz
2026-09-11 12:52:13 +04:00

95 lines
6.2 KiB
Markdown

# tests/conformance — one suite, four runners
**This is a required check on every lane's PR.** It is the mechanism that makes four
parallel lanes safe: the C++ daemon and the TypeScript extension are proved compatible
without either having run against the other.
```sh
./tests/conformance/run.sh # starts its own mockd + veloxd
./tests/conformance/run.sh --uds /run/user/1000/velox/velox.sock --ws-port 52000
```
## The runners
| Runner | Needs | Asserts |
|---|---|---|
| `check_contract.py` | python3, jsonschema | schemas parse and resolve; the documented surface matches the schema surface both ways; every method has a success fixture; every fixture validates; `SettingKey` and `Settings` agree; **committed generated code is not stale** |
| `cpp/` | a C++23 compiler, nlohmann | every golden payload parses into the generated structs, serialises back stably, and goes through the real `dispatch()`; privileged methods are refused `-32003` over the WebSocket |
| `ts/replay.ts` against mockd | node ≥ 20 | the TS client and the fixtures agree with each other — mockd always answers every fixture correctly by construction, so this cannot catch veloxd disagreeing with the contract |
| `ts/replay.ts` against veloxd | the above, plus a C++23 toolchain (`daemon/CMakeLists.txt`'s deps) | the **real** daemon it builds and starts, isolated (its own `XDG_RUNTIME_DIR`/`XDG_DATA_HOME`/`XDG_CONFIG_HOME`), answers every fixture — except ones hitting a still-stubbed handler, excused by `veloxd-xfail.json` (see below) |
### `veloxd-xfail.json`
`daemon/docs/deferrals.md` (D1-D4b) lists the handlers still stubbed out (`-32603 not
implemented`); the fixtures that hit them can't pass against veloxd yet and are listed
here with why, keyed by fixture path. This is a maintained allowlist, not a snapshot: a
listed fixture that unexpectedly *passes* is flipped back to a failure by `replay.ts`
(`applyXfail`) rather than silently accepted, so an entry has to be deleted the same PR
that closes the handler — the list can only shrink, never rot into "things nobody checks."
This is also where a real veloxd bug shows up before it reaches anyone else: an
implemented handler returning something the schema forbids (e.g. `download.get` with
`segments: 0`, which `TaskSummary.segments` requires >= 1) is **not** in the allowlist, so
it fails the run for real. That is the point of running against veloxd at all, not just
mockd.
### Current status against veloxd: red, for two reasons — not just the one
As of this wiring, the veloxd runner does not pass, and shouldn't yet:
1. **`download.get` / `download.list` can return `segments: 0`.** `store::TaskRow::eff_segments`
defaults to `0` and `to_summary` copies it straight into `TaskSummary.segments`
(`daemon/src/store/tasks.{hpp,cpp}`), which the schema forbids (minimum 1, required).
This isn't only a post-completion thing — it's any task the engine hasn't segmented yet,
which includes a task the instant it's added. This is the regression this runner exists
to catch, and it is deliberately **not** in `veloxd-xfail.json`.
2. **New finding: `download.add` with `startMode: "later"` always fails.** `later` is a
valid, documented `StartMode` (`contracts/schema/types/StartMode.schema.json`:
`["now", "later", "queue"]` — "'later' is the File Info dialog's Download Later
button"), and `on_download_add` stores it verbatim as `tasks.start_mode`
(`daemon/src/rpc/dispatcher.cpp`). But the `start_mode` column's `CHECK` constraint
(`daemon/src/store/migrations/0001_initial.sql`) only allows
`'auto','now','queue','manual'` — no `'later'`, and `'manual'`/`'auto'` aren't contract
values at all. Every `download.add` with `startMode: "later"` — including this runner's
own fixture-binding setup, which needs one to exist before it can replay any
`$taskId`-referencing fixture — fails `-32603` on a SQLite `CHECK` violation. This is not
in `veloxd-xfail.json` either: it's not a stub (D-list), it's a real, currently-shipping
bug, and it's why the veloxd run is red across most of the suite right now, not just on
`download.get`. Filed to DAEMON; not fixed here (out of lane).
`run.sh` also runs one scenario that cannot be shown against a healthy server: with the
daemon answering slower than `capture.offer`'s 750 ms deadline, the client must give up and
let Firefox take the download. **That is the fail-open guarantee, and it is checked here.**
## What "passing" means
The runners check the contract, not the implementation's opinions. Results are compared by
shape and validated against the generated validators; error codes are compared exactly.
Byte-equality with a golden file is deliberately *not* asserted, because a live daemon
returns its own ids and its own clock — see `contracts/fixtures/README.md`.
Adding a method without a fixture fails `check_contract.py`. Regenerating and forgetting to
commit the output fails it too.
## CI
Wired in as `ctest -L conformance` (`.github/workflows/ci.yml`'s `conformance` job, PKG/QA;
see `docs/adr/0014-conformance-runs-through-ctest.md`), required on every branch. It starts
and stops its own `mockd` and its own `veloxd`; nothing else needs to be running.
Needs: `python3` with `jsonschema`, a C++23 compiler, the same deps `daemon/CMakeLists.txt`
needs (SQLite3, nlohmann-json, libsecret — `tools/bootstrap.sh` installs all of it), and
Node ≥ 20.
### One caveat: veloxd's single-instance lock is not isolation-aware
veloxd refuses to start a second copy for the same user — an abstract-namespace socket
keyed by UID only, not by `$XDG_RUNTIME_DIR` (`daemon/src/main.cpp`,
`acquire_single_instance_lock`). This run's isolated veloxd collides with that lock exactly
like any other copy would: if a real veloxd (or another worktree's integration run) is
already up for this user when `run.sh` starts, the veloxd step fails fast with "another
instance is already running for this user" rather than silently testing the wrong daemon.
A fresh CI runner never hits this — only concurrent local runs can. If that turns out to
bite CI in practice (two conformance jobs sharing a runner user, say), the fix belongs in
`daemon/` (scope the lock name to the runtime dir), not here.