core: sign off on ADR 0013 (task state-machine ownership)

Accept as written — no amendments. Answers to the four open items:

1. tasks_starved / starved_tasks() count ONLY tasks in {connecting,
   downloading} with segments_active == 0 (the allocator owes a slot to a
   task that is asking). retry_wait and paused (either-initiated) are
   outside that state set, so they are excluded by construction, not by a
   special case. Design commitment; the accessor is stage 6/8.
2. pause() is idempotent: no-op success on an already-paused or terminal
   task; resume() no-op success on a non-paused task; only task_not_found
   errors. No state-change event for a no-op.
3. PROTO's item — CORE confirms its half: auto-pause reports the
   transition with ErrorInfo populated (auth_required / server_file_changed
   / disk_full / path_rejected), already in util/error.hpp. Ready once
   PROTO permits error on state=="paused".
4. CORE adopts "auto-pause"; the wire/API discriminator stays
   state==paused + presence of the Error code.

Notes back: pause during verifying re-hashes from scratch on resume;
pause during assembling is M4; restart handling in §5 agreed —
start(TaskId) re-derives resume position from .veloxpart.meta + If-Range.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01HPPSGhiArbvQgwC2DNiURS
This commit is contained in:
2026-09-10 00:02:02 +04:00
co-authored by Claude Sonnet 5
parent 201ebc55d4
commit c65e664c29
+97
View File
@@ -0,0 +1,97 @@
# CORE response to ADR 0013 (task state-machine ownership)
**Verdict: accept as written.** No amendments to the decision. `daemon/src/sched/`'s
pause/resume logic is unblocked from CORE's side once PROTO lands open item 3 (the
`error`-on-`paused` widening — see below).
The split in §1, the shared-`paused` + idempotency contract in §2, the cross-reason resume
rule in §3, the `retry_wait ≠ starvation` resolution in §4, and the restart handling in §5
all match what CORE agreed in `contracts/proto-answers-m1.md` D1 and the ADR 0011 response.
Everything below is a **design commitment** — CORE's state machine is stage 8 and
`starved_tasks()` is stage 6/8; none of it is built yet. So "does CORE already do X" is
answered as "CORE will do X, and here is the spec it will be built to", not as a report on
existing code.
---
## Answers to the four open items
### 1 — `starved_tasks()` / `tasks_starved` exclude `retry_wait` and auto-paused, by construction
Not a behavior change (nothing to change yet); a commitment baked into the accessor's
definition. Pinning ADR 0011's "running task with `segments_active == 0`":
`tasks_starved` counts **only tasks whose task-state is `connecting` or `downloading` and
whose `segments_active == 0`.** That is precisely "the segment allocator has not granted a
slot to a task that is asking for one." Consequences:
| Task state | In `tasks_starved`? | Why |
|---|---|---|
| `probing` | no | uses the probe pool (ADR 0011 §5), not the segment budget |
| `connecting`, `segments_active == 0` | **yes** | admitted + probed, waiting on the allocator's first grant — the real starvation case |
| `connecting`/`downloading`, `segments_active ≥ 1` | no | a segment in its own `connecting` sub-state counts as a held slot (ADR 0011 A3) |
| `downloading`, `segments_active == 0` | **yes** | held slots, lost them all (e.g. every segment failed and is being re-requested) — transient, still a real "allocator owes this task a slot" |
| `retry_wait` | no | CORE's backoff timer holds it at zero *deliberately*; it is not asking the allocator for anything until it re-enters `connecting` |
| `paused` (DAEMON- **or** CORE-initiated) | no | a paused task is not asking for a slot |
| `new`, `queued`, `assembling`, `verifying`, `complete`, `failed`, `cancelled` | no | not in the `{connecting, downloading}` set |
So `retry_wait` and auto-paused tasks fall outside `tasks_starved` **because they are not
in the counted state set**, not because of a special case that could rot. `starved_tasks()`
returns exactly the TaskIds in that count; `starved_since(id)` is defined only for them.
A `retry_wait` task still counts against `connection.maxConcurrentDownloads` from DAEMON's
side (it is running, not requeued) and contributes nothing to the allocator's guarantee
pass until it re-enters `connecting` — matches §4 exactly.
### 2 — `pause()` is idempotent (committed)
- `pause(TaskId)` on a task already in `paused` (however it got there) → **no-op, returns
success**. DAEMON never needs to know CORE auto-paused first.
- `pause(TaskId)` on a task in a terminal state (`complete`/`failed`/`cancelled`) →
**no-op, returns success** — a pause racing a completion is not an error.
- `resume(TaskId)` on a task that is not paused → **no-op, returns success**.
- The only error `pause()`/`resume()` return is `task_not_found`.
- CORE does not report a state-change event for a no-op pause (nothing changed). DAEMON
reads the resulting state via the normal state-change callback / `TaskDetail`, not from
`pause()`'s return.
- If CORE is mid-transition to `paused` (its own auto-pause) when DAEMON's `pause()`
arrives, the task ends up `paused` once, with one state-change event.
### 3 — PROTO: `error`-on-`paused` widening
Not CORE's to land, but CORE confirms its half: when CORE auto-pauses, it reports the
transition through the same state-change callback every CORE transition uses, with the
`ErrorInfo` populated — `auth_required` (401/407), `server_file_changed`, `disk_full`,
`path_rejected` for a mid-run destination failure. Those are the `vdm::Error` values from
the B1 taxonomy (`core/docs/proto-requests-m1.md`), already implemented in
`core/include/vdm/util/error.hpp`. **CORE is ready**; PROTO only needs to permit `error` to
be present on `event.task.state` when `state == "paused"`, and DAEMON to project CORE's
`ErrorInfo` onto it. `error: null` on a DAEMON-initiated pause is correct and sufficient.
Until PROTO lands it, DAEMON cannot implement §3's cross-reason resume rule
correctness-preservingly — agree it should not guess from timing.
### 4 — "auto-pause" naming
CORE has no established internal term (the state machine is unbuilt). **CORE adopts
"auto-pause"** for the informal concept. On the wire and in the API there is no new term:
the discriminator is `state == paused` plus the `Error` code (present ⇒ CORE-initiated,
absent ⇒ DAEMON-initiated), exactly as the ADR's §2 and the rejected-alternatives section
describe.
---
## Two notes back to DAEMON (not objections)
- **§1 table, `paused` from `assembling`/`verifying`:** the *state transition* is honoured
from any CORE state as the table says. The *work* interruption is best-effort: a pause
during `verifying` discards the in-progress hash and re-hashes from the start of the
file on resume (cheap, bounded); a pause during `assembling` (HLS/DASH mux) is an M4
concern and may not be cleanly interruptible mid-mux. Neither affects M1 or the API
shape.
- **§5 restart:** CORE agrees fully. CORE holds no persistent state; `start(TaskId)`
transparently checks for a valid `.veloxpart.meta` sidecar (stage 5), re-validates with
`If-Range` (`docs/04` §5), and either resumes from the recorded offsets or restarts if
the validator failed / the sidecar is corrupt. DAEMON does nothing special on restart
beyond rewriting CORE-owned states to `queued` and re-admitting — which is what §5 says.