core: answer PROTO's bufferBytes range question; accept inclusive endByte
buffer-sizing.md: the frozen 4 KiB–8 MiB and docs/04's 64 KiB–64 MiB both miss. Recommend 64 KiB – 16 MiB, default 2 MiB, max_total_buffer_bytes unchanged at 256 MiB: - 4 KiB floor is smaller than one libcurl write callback -> a syscall per chunk; 64 KiB is the smallest floor that coalesces. - throughput vs write size is flat past ~8 MiB on NVMe; 8–16 MiB is disk-stall absorption headroom for the fast-pipe/slow-disk case; 64 MiB is cache pressure for zero gain. - 32 segments x 64 MiB = 2 GiB vs the 256 MiB cap means the docs/04 max is unreachable past 4 total active segments — a misleading Options value. 16 MiB is reachable for single-/light-multitask and clamps to 8 MiB under heavy parallelism, which is correct. - default 4 MiB x 20 downloads = 80 MiB, busting the "<=60 MB RSS / 20 downloads" DoD; 2 MiB fits. Filed as request B4. proto-requests-m1.md: B3 endByte accepted as inclusive (HTTP Range semantics, no curl-boundary off-by-one); [start,end) ask withdrawn; stage 6 designed against inclusive. New B3a: the Content-Length: 0 whole-file case needs a representable zero-length segment — min_segment_bytes means CORE never makes empty segments mid-download, so it's only the degenerate case; mild preference for startByte+length over an endByte=startByte-1 sentinel. Co-Authored-By: Claude Sonnet 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01HPPSGhiArbvQgwC2DNiURS
This commit is contained in:
@@ -0,0 +1,116 @@
|
||||
# CORE → PROTO — `bufferBytes` range (answer to the freeze question)
|
||||
|
||||
`contracts/` froze `bufferBytes` at **4 KiB – 8 MiB**; `docs/04` §4 line 70 says
|
||||
**64 KiB – 64 MiB, default 4 MiB**. They disagree on the floor (16×), the ceiling (8×),
|
||||
and the RSS budget can't hold the default. CORE owns the ring buffer, the 32-segment
|
||||
ceiling and `max_total_buffer_bytes`, so here is the range that is actually right and why.
|
||||
PROTO to land schema + `docs/04` §4 + an ADR together.
|
||||
|
||||
## Recommendation
|
||||
|
||||
| Field | Value | |
|
||||
|---|---|---|
|
||||
| `bufferBytes` minimum | **65536** (64 KiB) | from `docs/04`; the frozen 4 KiB is wrong — see below |
|
||||
| `bufferBytes` maximum | **16777216** (16 MiB) | 8 MiB is defensible if simplicity wins; 64 MiB is not |
|
||||
| `bufferBytes` default | **2097152** (2 MiB) | 4 MiB busts the RSS DoD — see below |
|
||||
| `max_total_buffer_bytes` default | **268435456** (256 MiB), unchanged | the real ceiling; the backstop for everything |
|
||||
|
||||
## Why the floor is 64 KiB, not 4 KiB
|
||||
|
||||
The buffer's first job is to coalesce libcurl write-callback deliveries into one `pwrite`.
|
||||
Over HTTP/2 a single write callback is routinely 16–64 KiB, and can be up to ~256 KiB. A
|
||||
4 KiB ring buffer is **smaller than one callback**: every callback would have to flush
|
||||
mid-call (or loop), so the "one `pwrite` per fill" design degrades to a syscall per curl
|
||||
chunk — the exact thing the buffer exists to avoid. 64 KiB (≈4 typical chunks) is the
|
||||
smallest floor that still buys anything. 4 KiB is a page size that wandered into a
|
||||
throughput knob.
|
||||
|
||||
## Why the ceiling is ~16 MiB, not 64 MiB
|
||||
|
||||
Two things the buffer buys, past coalescing:
|
||||
|
||||
1. **Large sequential writes.** On NVMe, write throughput as a function of write size is
|
||||
flat by ~1–4 MiB. From 4 MiB to 8 MiB you gain a little on syscall overhead at
|
||||
multi-Gbit; past 8 MiB there is **no throughput left to get** — you are only adding
|
||||
`fdatasync` latency and page-cache pressure (a 40 GB ISO must not evict the user's
|
||||
working set — `docs/04` §4).
|
||||
2. **Absorbing a disk stall without stalling the socket.** This is the only reason to go
|
||||
above 8 MiB. Fast link + bursty storage (HDD, SMR, USB, a network mount): 1 Gbit is
|
||||
~125 MB/s, so 16 MiB per segment ≈ 130 ms of write-stall cover; across 4–8 segments,
|
||||
~0.5–1 s aggregate — enough to ride out a seek storm. Beyond 16 MiB the marginal cover
|
||||
isn't worth the cache footprint.
|
||||
|
||||
So: **8 MiB captures all the throughput; 8–16 MiB is stall-absorption headroom for the
|
||||
fast-pipe/slow-disk power user; >16 MiB is waste.**
|
||||
|
||||
## The arithmetic you asked about: 32 × 64 MiB vs a 256 MiB cap
|
||||
|
||||
`max_total_buffer_bytes` (256 MiB) caps the sum of every active segment's buffer across
|
||||
every active task. The effective per-segment buffer is
|
||||
|
||||
```
|
||||
effective = clamp( requested,
|
||||
64 KiB,
|
||||
floor(max_total_buffer_bytes / active_segment_count) )
|
||||
```
|
||||
|
||||
Reachable `bufferBytes` before the cap clamps it, by workload:
|
||||
|
||||
| Active work | segments | cap ÷ segments | 16 MiB reachable? |
|
||||
|---|---|---|---|
|
||||
| 1 task, 1 seg (small file / non-resumable) | 1 | 256 MiB | yes |
|
||||
| 1 task, 8 seg (typical big download) | 8 | 32 MiB | yes, unclamped |
|
||||
| 1 task, 16 seg | 16 | 16 MiB | exactly at the cap |
|
||||
| 1 task, 32 seg | 32 | 8 MiB | **clamps to 8 MiB** |
|
||||
| 4 tasks × 8 seg | 32 | 8 MiB | clamps to 8 MiB |
|
||||
|
||||
Now the same table with a **64 MiB** ceiling: it is unreachable the moment total active
|
||||
segments exceed **4** (256 / 64). At the default 8 segments the user asks for 64 MiB and
|
||||
silently gets 32 MiB; at 32 segments they get 8 MiB. A maximum that no realistic
|
||||
configuration can actually use is a misleading number in the Options dialog. 16 MiB is
|
||||
reachable for the single-task and light-multitask cases — the cases where deep buffering
|
||||
is the point — and degrades predictably (to 8 MiB) exactly when per-segment buffering
|
||||
stops mattering because each segment is only getting 1/32 of the link.
|
||||
|
||||
32 segments is already past the point of diminishing returns on segment count itself
|
||||
(`docs/04` §3: "more segments than [1 MiB each] is pure overhead and gets you
|
||||
rate-limited"); clamping their buffers to 8 MiB is the right behaviour, not a regression.
|
||||
|
||||
## The default: 4 MiB fails the RSS DoD
|
||||
|
||||
`docs/04` §8: "≤ 60 MB RSS with 20 active downloads at default buffers." Twenty active
|
||||
downloads, each with at least one segment:
|
||||
|
||||
- default **4 MiB** → 20 × 4 = **80 MiB in buffers alone**, before curl handles, TLS
|
||||
buffers, thread stacks, and the task table. Busts 60 MB outright. The 256 MiB global
|
||||
cap does not save you — 80 < 256, so nothing clamps.
|
||||
- default **2 MiB** → 20 × 2 = 40 MiB, leaving ~20 MiB for everything else. Fits.
|
||||
- A single 8-segment download at 2 MiB is 16 MiB of buffer — already ample for line rate
|
||||
on NVMe (see the throughput-plateau point above).
|
||||
|
||||
So the default has to be **2 MiB** for the RSS target and the buffer default to be
|
||||
consistent, or `docs/04` §8 has to be renegotiated. 2 MiB is the cheaper fix and is not a
|
||||
throughput compromise.
|
||||
|
||||
## Clamp behaviour CORE will implement
|
||||
|
||||
- Per **task start** and on any change to the active-segment count (new task, task
|
||||
finishing, a steal), recompute `effective` for every live segment by the formula above.
|
||||
- Never clamp below the 64 KiB floor. If `max_total_buffer_bytes / active_segment_count`
|
||||
is itself below 64 KiB (would need >4096 concurrent segments — not reachable at the
|
||||
32-per-task ceiling and a sane concurrent-task limit), CORE admits fewer concurrent
|
||||
segments rather than shipping a sub-floor buffer.
|
||||
- The resulting `effective` value is what CORE reports upward for the readback field
|
||||
(request **B2a**). `requested` is echoed back too so the GUI can show "8 MiB (using
|
||||
2 MiB)".
|
||||
|
||||
## Net contract delta for PROTO
|
||||
|
||||
- `DownloadSpec.bufferBytes`, `download.update` patch: `minimum: 65536`, `maximum:
|
||||
16777216`, `default: 2097152`.
|
||||
- `docs/04` §4 line ~70: "Default 2 MiB, range 64 KiB – 16 MiB. Silently reduced to fit
|
||||
`max_total_buffer_bytes` (256 MiB) across all active segments; the effective value is
|
||||
reported back."
|
||||
- `docs/04` §8: keep "≤ 60 MB RSS / 20 downloads" — it now holds at the 2 MiB default.
|
||||
- ADR: record the throughput-plateau + global-cap-arithmetic reasoning; note 8 MiB was
|
||||
considered for the ceiling and 16 MiB chosen for stall absorption.
|
||||
@@ -93,16 +93,47 @@ CORE will emit per segment:
|
||||
| CORE field | Type | Note |
|
||||
|---|---|---|
|
||||
| `index` | int ≥ 0 | **`event.task.progress` currently says `i`.** Pick one name for both. |
|
||||
| `start` | int ≥ 0 | absolute byte offset, inclusive |
|
||||
| `end` | int ≥ 0 | absolute byte offset, **exclusive** — range is `[start, end)` |
|
||||
| `startByte` | int ≥ 0 | absolute byte offset, inclusive |
|
||||
| `endByte` | int | absolute byte offset, **inclusive** — range is `[startByte, endByte]` (see below) |
|
||||
| `completed` | int ≥ 0 | bytes written in this range so far |
|
||||
| `speedBps` | int ≥ 0 | current per-segment rate |
|
||||
| `state` | enum | `connecting` \| `downloading` \| `stalled` \| `complete` \| `failed` |
|
||||
|
||||
Asks: (a) reconcile `i` vs `index` — one spelling in both the `Segment` type and the
|
||||
`event.task.progress` payload; (b) confirm the half-open `[start, end)` convention in the
|
||||
schema `description` so DAEMON and GUI don't off-by-one the last byte; (c) confirm the
|
||||
segment `state` enum values.
|
||||
**Resolved by PROTO at freeze:** `endByte` is **inclusive** (matches HTTP `Range`
|
||||
semantics — `Range: bytes=start-end` is inclusive — and removes an off-by-one at the curl
|
||||
boundary). CORE designs stage 6 (segmenter/stealer) against inclusive. The earlier
|
||||
`[start, end)` ask is withdrawn.
|
||||
|
||||
Still open: (a) reconcile `i` vs `index`; (b) confirm the segment `state` enum values;
|
||||
(c) **empty-segment representation** — see B3a.
|
||||
|
||||
### B3a. Zero-length segment must be representable — *PROTO is fixing; CORE's requirement*
|
||||
|
||||
With `endByte` inclusive and `minimum: 0`, a zero-length segment (`endByte = startByte -
|
||||
1`) at offset 0 is `endByte = -1`, which the schema forbids. The one case CORE actually
|
||||
needs: a **whole-file zero-length download** (`Content-Length: 0`) — one segment, length
|
||||
0. It is a valid HTTP response and the daemon/GUI must be able to hold it.
|
||||
|
||||
CORE will **not** produce empty segments mid-download: the `min_segment_bytes` floor
|
||||
(1 MiB, `docs/04` §3) means the segmenter never splits below 1 MiB and the stealer only
|
||||
takes a half-range if it is ≥ that floor. So B3a is purely about the degenerate
|
||||
whole-file case.
|
||||
|
||||
Preference: encode segments as `startByte` + `length` (+ `completed`) rather than an
|
||||
inclusive `endByte` with a `startByte - 1` sentinel — `length: 0` is then the natural
|
||||
representation and there is no negative value to allow. If `endByte` inclusive stays,
|
||||
then a 0-byte task needs an explicit encoding (an `empty`/`length` field, or permitting
|
||||
`endByte = startByte - 1` with `minimum: -1`) — any of those work for CORE as long as
|
||||
total length 0 round-trips. Flag back if the chosen fix needs anything else from CORE.
|
||||
|
||||
### B4. `bufferBytes` range is wrong in the frozen schema — *see `buffer-sizing.md`*
|
||||
|
||||
`contracts/` froze `bufferBytes` at 4 KiB – 8 MiB; `docs/04` §4 says 64 KiB – 64 MiB
|
||||
default 4 MiB; the RSS DoD (`docs/04` §8) can't hold either default. CORE's analysis and
|
||||
the recommended range (**64 KiB – 16 MiB, default 2 MiB**, `max_total_buffer_bytes`
|
||||
unchanged at 256 MiB) with the global-cap arithmetic is in
|
||||
[`core/docs/buffer-sizing.md`](buffer-sizing.md). PROTO to land schema + `docs/04` §4 +
|
||||
ADR together.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user