daemon: fold in CORE's sign-off on ADR 0013

CORE accepted ADR 0013 as written, no amendments
(core/docs/adr-0013-core-response.md, lane/core@c65e664). Folds in:

- pause()/resume() idempotency contract, precisely: no-op success on
  an already-paused task, ALSO on a terminal task (pause racing
  completion isn't an error), resume() no-op on a non-paused task, the
  only error is task_not_found, and no state-change event fires for a
  no-op call.
- tasks_starved pinned as {connecting, downloading} AND
  segments_active == 0 — a structural exclusion of retry_wait and
  auto-paused tasks rather than a special case, with the full state
  table CORE gave.
- restart handling confirmed fully; two non-blocking notes from CORE
  about work-interruption during verifying/assembling.
- "auto-pause" adopted as the term, no new wire/API surface.

Status updated: accepted by CORE; PROTO's item 3 (permit `error` on
event.task.state when state=="paused") is the one remaining blocker
before daemon/src/sched/'s pause/resume logic can be written
correctness-preservingly.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01Upd9WhG9oppieig5nRDLig
This commit is contained in:
2026-09-10 00:05:11 +04:00
co-authored by Claude Sonnet 5
parent 19d3cd2b1d
commit 62cda074c1
+76 -35
View File
@@ -1,9 +1,13 @@
# ADR 0013 — Task state-machine ownership: CORE, DAEMON, and the shared `paused` # ADR 0013 — Task state-machine ownership: CORE, DAEMON, and the shared `paused`
**Status:** proposed · **Date:** 2026-09-09 · **Lane:** DAEMON, drafted at PROTO's request **Status:** accepted by CORE, PROTO's open item 3 outstanding · **Date:** 2026-09-09
(D1 in `core/docs/proto-requests-m1.md`), for CORE and PROTO to review and sign off. · **Lane:** DAEMON, drafted at PROTO's request (D1 in `core/docs/proto-requests-m1.md`).
**Needs:** explicit accept from CORE and PROTO before `daemon/src/sched/` calls into CORE's **CORE's response:** `core/docs/adr-0013-core-response.md` (`lane/core`, commit `c65e664`)
task-start API, since the calling convention below is derived from this split. — accepted as written, no amendments, plus the pinned `tasks_starved` definition and
`pause()`/`resume()` idempotency contract folded in below.
**Remaining blocker:** PROTO landing the `error`-on-`paused` widening (open item 3) —
`daemon/src/sched/`'s pause/resume logic is unblocked from CORE's side already, but not
buildable correctness-preservingly until that lands.
## Context ## Context
@@ -77,19 +81,24 @@ Both sides can put a task in `paused`, for disjoint reasons:
state-change callback — the same path every other CORE-driven transition uses. This is state-change callback — the same path every other CORE-driven transition uses. This is
not a new mechanism. not a new mechanism.
The contract that makes this safe: The contract that makes this safe, per CORE's sign-off (`core/docs/adr-0013-core-response.md`):
- **`pause()` is idempotent.** DAEMON calling `pause()` on a task CORE already - **`pause(TaskId)` is idempotent: a no-op success** on a task already `paused` (however
auto-paused is a no-op, not an error — DAEMON does not need to know CORE got there it got there), and **also a no-op success on a terminal task** (`complete`/`failed`/
first. Symmetrically, if CORE were ever mid-transition to `paused` when DAEMON's pause `cancelled`) — a pause racing a completion is not an error. **`resume(TaskId)` is a
request arrives, the result is still `paused`, once. no-op success on a task that is not paused.** The only error either returns is
`task_not_found`. CORE emits no state-change event for a no-op call — DAEMON reads the
resulting state from the normal state-change callback / `TaskDetail`, never from
`pause()`'s or `resume()`'s return value. If CORE is mid-transition into `paused` (its
own auto-pause) when DAEMON's `pause()` arrives, the task ends up `paused` once, with
exactly one state-change event.
- **DAEMON persists *why* a task is paused**, in the `tasks` table, not in `TaskState` - **DAEMON persists *why* a task is paused**, in the `tasks` table, not in `TaskState`
itself (the wire type stays a flat enum — this is DAEMON-local bookkeeping, not a itself (the wire type stays a flat enum — this is DAEMON-local bookkeeping, not a
contract change). A `pauseReason` distinguishing at least `user`, `schedule`, contract change). A `pauseReason` distinguishing at least `user`, `schedule`,
`queue_stopped`, `admission_reconcile`, and CORE's `error.code` when auto-paused. This `queue_stopped`, `admission_reconcile`, and CORE's `error.code` when auto-paused. This
is what makes §3's resume rule possible — but it needs a contract fix first: see the is what makes §3's resume rule possible — but it needs a contract fix first: see "A
gap in open item 3 below. `error.code` is not currently carried on a transition into contract gap this ADR surfaces" below. `error.code` is not currently carried on a
`paused` at all. transition into `paused` at all.
- **CORE does not need to track why it's paused past the current occurrence.** Once - **CORE does not need to track why it's paused past the current occurrence.** Once
paused, CORE's job is done; DAEMON is the only side that later decides whether to paused, CORE's job is done; DAEMON is the only side that later decides whether to
resume, and DAEMON is also the only side with persistent storage to remember the resume, and DAEMON is also the only side with persistent storage to remember the
@@ -118,14 +127,28 @@ and asserts `tasks_starved == 0` in steady state past a bounded delay. A task in
`retry_wait` also holds zero segments, deliberately, for up to 60 s (`docs/04` §7's backoff `retry_wait` also holds zero segments, deliberately, for up to 60 s (`docs/04` §7's backoff
cap) — and it is **not** admission-starved, it is CORE's own policy holding it idle. cap) — and it is **not** admission-starved, it is CORE's own policy holding it idle.
Resolution: `retry_wait` must be distinguishable from true starvation without DAEMON **Pinned by CORE's sign-off**, tightening ADR 0011's definition rather than special-casing
guessing from timing. CORE's `starved_tasks()` (ADR 0011) excludes any task CORE itself it: `tasks_starved` counts only tasks whose `TaskState` is `connecting` or `downloading`
holds at zero by policy — `retry_wait` and CORE-initiated `paused` both fall outside **and** `segments_active == 0` — precisely "the allocator has not granted a slot to a task
`tasks_starved` by construction, because that accessor is about the segment allocator that is asking for one."
failing to grant a slot to a task that wants one, not about a task that isn't asking. A
task in `retry_wait` still counts against `connection.maxConcurrentDownloads` from | Task state | In `tasks_starved`? | Why |
DAEMON's side (it is running, not requeued) but contributes nothing to the segment |---|---|---|
allocator's guarantee pass until it re-enters `connecting`. | `probing` | no | uses the probe pool (ADR 0011 §5), not the segment budget |
| `connecting`, 0 segments | **yes** | admitted + probed, waiting on the allocator's first grant — the real starvation case |
| `connecting`/`downloading`, ≥1 segment | no | a segment in its own `connecting` sub-state counts as held (ADR 0011 amendment A3) |
| `downloading`, 0 segments | **yes** | held slots and lost them all (e.g. every segment failed and is being re-requested) — transient, still real |
| `retry_wait` | no | CORE's backoff timer holds it at zero *deliberately*; not asking the allocator for anything until it re-enters `connecting` |
| `paused` (either-initiated) | no | not asking for a slot |
| `new`, `queued`, `assembling`, `verifying`, terminal | no | outside the counted state set |
So `retry_wait` and auto-paused tasks are excluded **structurally, by not being in the
`{connecting, downloading}` state set** — not via a special case that could rot as the
engine evolves. `starved_tasks()` returns exactly the TaskIds in this count;
`starved_since(id)` is defined only for them. A task in `retry_wait` still counts against
`connection.maxConcurrentDownloads` from DAEMON's side (it is running, not requeued) but
contributes nothing to the segment allocator's guarantee pass until it re-enters
`connecting`.
### 5. Restart — CORE holds no persistent state, DAEMON reloads to `queued` ### 5. Restart — CORE holds no persistent state, DAEMON reloads to `queued`
@@ -143,6 +166,18 @@ where to actually resume from `.veloxpart.meta` and re-validates with `If-Range`
`paused` tasks reload as `paused`, with their `pauseReason` intact, and are not `paused` tasks reload as `paused`, with their `pauseReason` intact, and are not
auto-admitted — §3 applies identically after a restart as it does live. auto-admitted — §3 applies identically after a restart as it does live.
**Confirmed by CORE, fully.** `start(TaskId)` transparently checks for a valid
`.veloxpart.meta` sidecar (CORE's stage 5), re-validates with `If-Range` (`docs/04` §5),
and either resumes from the recorded offsets or restarts if the sidecar is missing or
fails validation. DAEMON does nothing special beyond rewriting CORE-owned states to
`queued` and re-admitting, as this section already said.
**Two notes from CORE, not objections to the ADR:** the *transition* into `paused` is
honoured from any CORE state per §1's table, but the *work interruption* is best-effort —
a pause during `verifying` discards the in-progress hash and re-hashes from the start on
resume (cheap, bounded); a pause during `assembling` (HLS/DASH mux) is an M4 concern and
may not be cleanly interruptible mid-mux. Neither changes this ADR's API shape.
## Consequences ## Consequences
- `daemon/src/sched/` calls `start(TaskId)` exactly once per admission (`queued → - `daemon/src/sched/` calls `start(TaskId)` exactly once per admission (`queued →
@@ -151,10 +186,10 @@ auto-admitted — §3 applies identically after a restart as it does live.
- The `tasks` table needs a `pauseReason` column (DAEMON-local; not a wire type) before - The `tasks` table needs a `pauseReason` column (DAEMON-local; not a wire type) before
`sched/`'s pause/resume logic can be written correctly — flagging as a concrete `sched/`'s pause/resume logic can be written correctly — flagging as a concrete
follow-up, not blocking this ADR's acceptance. follow-up, not blocking this ADR's acceptance.
- CORE's `starved_tasks()` (ADR 0011) must exclude `retry_wait` and CORE-auto-paused tasks - CORE's `starved_tasks()` (ADR 0011) excludes `retry_wait` and CORE-auto-paused tasks by
by construction; if that isn't already true in CORE's accessor, it needs to be before construction — pinned in CORE's sign-off as the `{connecting, downloading} ∧
DAEMON relies on the starvation invariant, since otherwise every backoff cycle would segments_active == 0` definition in §4's table — so DAEMON's governor-invariant warning
trip DAEMON's governor-invariant warning as a false positive. will not false-positive on a normal backoff cycle.
- No `contracts/` change. `TaskState`, `event.task.state`, and `previousState` are already - No `contracts/` change. `TaskState`, `event.task.state`, and `previousState` are already
sufficient; this ADR is entirely about which process calls which function when. sufficient; this ADR is entirely about which process calls which function when.
@@ -190,15 +225,21 @@ follow-up (minor: widening an existing field's presence condition, per
tell CORE-paused from user-paused on the wire in general — `error: null` on a tell CORE-paused from user-paused on the wire in general — `error: null` on a
DAEMON-initiated pause is sufficient, since DAEMON already knows it just did that. DAEMON-initiated pause is sufficient, since DAEMON already knows it just did that.
## Open, for CORE and PROTO to confirm or amend ## Resolution of the four open items
1. Does CORE's `starved_tasks()` already exclude `retry_wait` and auto-paused tasks, or 1. **`starved_tasks()` excludes `retry_wait`/auto-paused by construction — confirmed.**
does this ADR ask for a behavior change there (§4)? Pinned as a design commitment in CORE's sign-off (§4's table above): the counted set is
2. Is `pause()` idempotent today, or does calling it on an already-paused task currently `{connecting, downloading} ∧ segments_active == 0`, so exclusion is structural, not a
return an error CORE needs to relax? special case.
3. PROTO: land the `error`-on-`paused` widening above before DAEMON writes the resume 2. **`pause()`/`resume()` idempotency — confirmed as specified in §2 above**, plus two
logic in §3 — otherwise DAEMON has no correctness-preserving way to implement it and details this ADR hadn't anticipated: a no-op call also succeeds against a *terminal*
would have to guess from timing, which is the exact failure mode ADR 0011 was written task (pause racing completion isn't an error), and no state-change event fires for a
to rule out for a different pair of governors. no-op — DAEMON reads resulting state from the callback/`TaskDetail`, never from the
4. Naming: is "auto-pause" the term CORE already uses internally, or does CORE have an call's return value.
existing name for this that this ADR should adopt instead of introducing a new one? 3. **Still open — PROTO.** CORE's half is ready (`ErrorInfo` is populated on every
auto-pause transition today's callback path would carry, using the existing B1
taxonomy); the wire only needs `error` permitted when `state == "paused"`. This is the
one remaining blocker on `daemon/src/sched/`'s pause/resume logic.
4. **CORE adopts "auto-pause".** No new wire or API term — the discriminator stays
`state == paused` plus the presence of an `Error` (present ⇒ CORE-initiated, absent ⇒
DAEMON-initiated), exactly as §2 already specified.