First pass counted one buffer per download; it is one per segment. 20 active downloads at the default 8 segments = 160 buffers, so at 20 tasks the binding constraint is the global cap, not the per-segment default — 256 MiB and the "<=60 MB RSS / 20 downloads" target (line 125) cannot both hold whatever the default is. Floor (64 KiB) and ceiling (16 MiB) unchanged — the 256/64 unreachability argument is stronger under per-segment accounting. Changes: - default 1 MiB (was 2): with the cap below, 32 live segments x 1 MiB = 32 MiB buffers -> ~45-50 MiB RSS, line 125 holds with margin. - NEW maxActiveSegments (default 32): a global concurrent-segment cap is the actual mechanism that bounds "20 active downloads"; docs/01 §2 implies it, docs/04 never states it. Without it no buffer policy hits 60 MB. - maxTotalBufferBytes 128 MiB (was 256) and it must be ADDED to the contract — currently absent, so the clamp CORE implements has no wire representation and Options can't show/set it. Folded into B2a. - line 125: keep 60 MB "given maxActiveSegments=32 and default buffers", or explicitly raise to 120 MB — ADR records which. Flagged that changing it is a defensible outcome CORE owns, not a number that quietly loses. - bufferBytes bounds are in FOUR schema files, not three: Settings.schema.json connection.bufferBytes also has 4096-8388608. Co-Authored-By: Claude Sonnet 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01HPPSGhiArbvQgwC2DNiURS
8.1 KiB
CORE → PROTO — bufferBytes range + the RSS-budget reconciliation
contracts/ froze bufferBytes at 4 KiB – 8 MiB in four places; docs/04 §4 line 70
says 64 KiB – 64 MiB, default 4 MiB. The floor and ceiling are settled below (64 KiB /
16 MiB — my earlier reasoning holds). The default and max_total_buffer_bytes are
not — my first pass counted one buffer per download; it is one per segment
("Per-segment ring buffer", line 68; "Write buffer per connection", line 70), which
changes the whole RSS picture. This version fixes that.
PROTO to land schema + docs/04 §4 + §8 + an ADR together.
Recommendation
| Knob | Value | |
|---|---|---|
bufferBytes minimum |
65536 (64 KiB) | frozen 4 KiB is below one libcurl callback — see §Floor |
bufferBytes maximum |
16777216 (16 MiB) | see §Ceiling; the 256/64 unreachability argument is stronger per-segment |
bufferBytes default |
1048576 (1 MiB) | see §Default+RSS; 2 MiB and 4 MiB both bust line 125 |
maxActiveSegments (new) |
32 | global cap on concurrently-transferring segments — the actual bound on "20 active downloads" |
maxTotalBufferBytes |
134217728 (128 MiB), down from 256 | backstop for the buffer knob; RSS-safe given the 1 MiB default + the segment cap |
Floor — 64 KiB, not 4 KiB
The buffer's first job is to coalesce libcurl write-callback deliveries into one pwrite.
Over HTTP/2 a single write callback is routinely 16–64 KiB, up to ~256 KiB. A 4 KiB ring
buffer is smaller than one callback: the "one pwrite per fill" design degrades to a
syscall per curl chunk — the exact thing the buffer exists to prevent. 64 KiB (≈4 typical
chunks) is the smallest floor that buys anything. 4 KiB is a page size that wandered into
a throughput knob.
Ceiling — 16 MiB, not 64 MiB
Two things the buffer buys past coalescing:
- Large sequential writes. On NVMe, throughput vs. write size is flat by ~1–4 MiB.
4→8 MiB gains a little on syscall overhead at multi-Gbit; past 8 MiB there is no
throughput left to get, only
fdatasynclatency and page-cache pressure (a 40 GB ISO must not evict the user's working set —docs/04§4). - Absorbing a disk stall without stalling the socket — the only reason to exceed 8 MiB. Fast link + bursty storage (HDD, SMR, USB, network mount): 1 Gbit ≈ 125 MB/s, so 16 MiB/segment ≈ 130 ms of write-stall cover, ~0.5 s across 4 segments — enough to ride out a seek storm. Beyond 16 MiB the marginal cover isn't worth the footprint.
8 MiB captures all throughput; 8–16 MiB is stall-absorption headroom for the fast-pipe/slow-disk case; >16 MiB is waste.
maxTotalBufferBytes / segments unreachability — worse per-segment
effective = clamp(requested, 64 KiB, floor(maxTotalBufferBytes / live_segment_count)).
With a 64 MiB ceiling and the old 256 MiB cap, 64 MiB is unreachable once total live segments exceed 4 (256/64). Under correct per-segment accounting "4 total segments" is half of one default download (8 segments) — so the frozen-doc ceiling is unreachable in essentially every real configuration. 16 MiB with a 128 MiB cap is reachable up to 8 live segments (128/16) — i.e. exactly the single-download slow-disk case the ceiling exists for — and clamps predictably beyond that, where per-segment buffering has stopped mattering because each segment holds a small fraction of the link.
Default + RSS — the part that didn't land
docs/04 line 125: "≤ 60 MB RSS with 20 active downloads at default buffers."
Per-segment accounting: 20 downloads × the default 8 segments = 160 buffers, not 20.
| default | 160 buffers | after a 256 MiB cap | vs 60 MB RSS |
|---|---|---|---|
| 4 MiB | 640 MiB | 256 MiB | ~4× over |
| 2 MiB | 320 MiB | 256 MiB | ~4× over |
| 1 MiB | 160 MiB | 160 MiB (cap not binding) | still ~3× over |
| any | — | cap is the binding constraint at 20 tasks, not the default | — |
So the per-segment default is not what decides the 20-task case — the global cap is, whatever the default. 256 MiB and 60 MB RSS cannot both hold. Hitting 60 MB across 160 segments is ~48 MiB of buffers total (≈300 KiB each), leaving ~12 MB for 160 curl handles + TLS + the daemon — which is not achievable; 160 live TLS connections alone are ~10–15 MB.
The real fix: cap concurrent segments, not just total buffer bytes
No download manager runs 160 simultaneous connections for 20 downloads. docs/01 §2
already implies this ("1 curl-multi transfer thread per ~8 active segments (capped)");
docs/04 never states the cap. Add maxActiveSegments (default 32) — a global
ceiling on segments actually transferring at once. 20 "active" downloads then means ~32
live connections with the rest of each download's segments queued, not 160.
RSS with maxActiveSegments = 32, default bufferBytes = 1 MiB:
- buffers: 32 × 1 MiB = 32 MiB (the 128 MiB cap isn't even engaged at the default)
- 32 curl/HTTP2/TLS connections: ~2–3 MiB
- daemon base (RPC loop, event batching, SQLite cache + WAL, ~8 thread stacks resident, task table): ~8–12 MiB
- total ≈ 45–50 MiB RSS — line 125 holds, at the default, with margin.
docs/04 §8 keeps the 60 MB number but must state it now depends on
maxActiveSegments = 32 and default buffers. A power user who overrides every download to
16 MiB gets clamped by maxTotalBufferBytes to 128 / 32 = 4 MiB per live segment →
128 MiB of buffers, ~140 MiB RSS — deliberately, and outside "at default buffers", so
line 125 is unaffected.
If 60 MB is the wrong target
It's defensible to raise it instead. 160 (or even 32) live TLS connections have an
irreducible cost, and IDM itself uses more. If PROTO/docs prefer, change line 125 to
"≤ 120 MB RSS with 20 active downloads at default buffers" and keep
maxActiveSegments higher (64). CORE's recommendation is the 60 MB + cap-at-32 route
because "lean daemon" is in the brief and 45–50 MiB is comfortably achievable — but it
must be an explicit decision in the ADR, not the number that quietly loses.
Clamp behaviour CORE will implement
- On task start and on any change to
live_segment_count(new task, task finishing, a steal), recomputeeffectivefor every live segment by the formula above. - Never below the 64 KiB floor. If
maxTotalBufferBytes / live_segment_count< 64 KiB (needs >2048 live segments at a 128 MiB cap — not reachable undermaxActiveSegments = 32), CORE admits fewer concurrent segments rather than shipping a sub-floor buffer. effectiveandrequestedboth reported upward so the GUI shows "16 MiB (using 4 MiB)" — request B2a.
Net contract delta for PROTO
bufferBytesbounds in FOUR schema files (not three):DownloadSpec,TaskDetail,download.updatepatch, andSettings.schema.jsonconnection.bufferBytes— all currently4096–8388608. Change every one tominimum: 65536,maximum: 16777216,default: 1048576.maxTotalBufferBytesis absent from the contract entirely. The clamp CORE implements has no wire representation, so the Options dialog can neither show nor set it. AddSettings.schema.json connection.maxTotalBufferBytes(default134217728) — fold into the B2a follow-up alongsideeffectiveBufferBytes.maxActiveSegmentsis new. AddSettings.schema.json connection.maxActiveSegments(default32). CORE enforces it; DAEMON's scheduler needs to know it to decide what to start.docs/04§4 line ~70: "Default 1 MiB, range 64 KiB – 16 MiB. Reduced to fitmaxTotalBufferBytes(128 MiB) across all live segments; effective value reported back."docs/04§8 line 125: keep "≤ 60 MB RSS / 20 downloads" but add "givenmaxActiveSegments = 32and default buffers" — or raise to 120 MB (see above). ADR records which and why.- ADR: throughput-plateau reasoning for the ceiling; the per-segment RSS arithmetic;
the
maxActiveSegmentsaddition as the mechanism that makes line 125 hold; 8 MiB considered and rejected for the ceiling in favour of 16 MiB stall absorption.