Files
vdm/core/include/vdm
samiandClaude Sonnet 5 322a20efa5 core: fix SegmentBudget over-admission and add a wait-list wakeup
Root cause of the M7 RSS gap (core/docs/m7-baseline.md: ~70 MB measured
against a 60 MB target): SegmentBudget::confirm_slot() only checked a task's
own held count against its own target -- never the engine-wide active_ sum.
reallocate_locked()'s two-pass fairness allocation does bound
sum(target) <= max_active_ at the moment it computes a plan, but that bound
says nothing about sum(held): a task can be legitimately holding more than
its own just-lowered target for a while (yield is deferred to a segment
boundary, never mid-segment -- ADR 0011 A1), and another task's target can
correctly rise to claim that capacity before the first task has physically
released it. Both confirm_slot() calls could then succeed against their own,
individually-correct targets while sum(held) exceeded max_active_ --
tools/bench heap-profile caught this directly: budget.active reading 56-86
against a total of 32.

confirm_slot() now also checks active_ < max_active_, unconditionally, as a
backstop that doesn't depend on any task's target bookkeeping being in sync
with what every other task holds. That creates a liveness question the
original design never answered: a task denied only by this new check has a
target that's already correct, so it never changes again and
reallocate_locked()'s plain "fire a callback when a task's target changes"
mechanism never revisits it. Task gained a waiting_for_slot flag, set on
exactly this denial; release_slot()/deregister_task() (the only two places
that free real capacity) now hand a freed slot directly to the
highest-priority waiting task via wake_one_waiter_locked(), if
reallocate_locked()'s own plan didn't already produce a callback for anyone.

download_task.cpp's fill_slots_locked() needed a matching fix: a woken
task's stalled segments (SegState::stalled -- backed off mid-retry, its own
release_slot() already called) have no live worker and never surface
through Segmenter::assign_slot(), which only hands out unassigned or fresh
ranges. fill_slots_locked() now restarts any stalled segment with no live
worker directly (bounded by slot_target, same as its assign_slot() loop)
before looking for new work; a segment it doesn't get to keeps its own
scheduled retry_worker() timer as a second chance.

Also fixes a real TSan-caught data race this work surfaced: SegWorker::
speed_bps was written only by its own segment's curl callback and, before
Progress.speed_bps's polled-path fix, only ever read from that same thread
-- safe without synchronization. snapshot_progress() reading it from
whatever thread calls DownloadHandle::progress() broke that invariant
(workers_mu's shared_lock protects the workers map's structure, not an
individual SegWorker's fields). Now std::atomic<double> with relaxed
ordering -- an informational EMA, nothing synchronizes real state on it --
rather than adding a lock to the write side.

core/tests/segment/budget_test.cpp adds two tests reproducing the actual
gap (budget_active_never_exceeds_max_active_segments_under_concurrent_load,
budget_wait_list_wakes_a_task_whose_target_never_changed) plus a sanity
baseline (budget_release_wakes_a_denied_waiter), and introduces AsyncFakeTask
+ TestTimer for the one existing test that drives the budget from multiple
concurrent threads -- mirroring production's real dispatch (register_task()'s
on_target lambda posts through host.schedule(), download_task.cpp, never a
synchronous call) rather than adding reentrancy-guarding machinery to
SegmentBudget itself to compensate for a synchronous test double being
unlike production. See docs/adr/0017 for the full writeup, including what an
earlier version of this fix got wrong chasing a same-thread reentrancy
hazard that doesn't actually exist in production.

core/docs/m7-baseline.md updated: the RSS number now clears the DoD line
(45.41 MiB via heap-profile), root-caused rather than just re-measured.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01Q3QrF7rCt21bkAjt9BCDFQ
2026-09-12 10:58:14 +04:00
..

libveloxcore — public API

Status: M1 in progress. util/, net/ (http_client, probe, url, content_disposition), io/ (sparse_file, write_buffer), meta/veloxpart, segment/ (segmenter, budget), rate/ (token_bucket), task/+engine.hpp (the download engine itself — vdm::Engine, DownloadSpec/DownloadHandle/DownloadCallbacks), and rules/ (filename sanitization, collision policy, rule-table matching) are landed. media/ is M4, not started — see core/docs/engine-api-m1.md for the engine API's own DAEMON-review history.

Layering (CLAUDE.md §3): this library knows nothing about JSON, SQL, Qt, or RPC. Input is a spec value; output is bytes on disk plus typed callbacks. DAEMON projects engine state onto the wire contract's TaskSummary / TaskDetail / events — see core/docs/proto-requests-m1.md for the shapes that projection needs frozen.

Every header under core/include/vdm/ compiles standalone (-Wall -Wextra -Wpedantic -Werror, C++23). Clean under ASan/UBSan and TSan.


util/ — foundations

vdm/util/error.hpp

enum class Error — the engine-wide failure taxonomy (network / HTTP / content / local I/O / metadata / probe / retry / internal). This is CORE's own vocabulary; it is not a wire type. error_name(Error) gives a stable snake_case string; is_retryable(Error) is the advisory retry hint the task policy consults.

struct ErrorInfo { Error code; std::string context; int http_status; bool retryable; Error cause; } — the payload carried by every failed Result. .to_string() renders "<name>: <context> (HTTP <n>)".

vdm/util/result.hpp

Result<T> — return-based error channel, a thin wrapper over std::expected<T, ErrorInfo>. Errors are returned, never thrown, on anything that runs during a transfer.

  • Result<int> r = 42; / Result<int> r = Err{Error::timeout, "..."}; / Result<T> r = Error::not_found;
  • r.has_value(), explicit operator bool, r.value() / *r / r->, r.error(), r.code(), r.value_or(x)
  • monadic and_then / transform / transform_error (forward to std::expected)
  • Result<void> specialization; vdm::ok() success sentinel
  • VDM_TRY(expr) — return the error if expr failed
  • VDM_TRY_ASSIGN(auto x, expr) — bind the value or return the error

vdm/util/bytes.hpp

Byte / ByteSpan / ConstByteSpan aliases; as_bytes(string_view) / as_chars(span). Little-endian fixed-width codec load_le<T> / store_le<T> and a bounds-checked sequential ByteReader (.u8/.u16/.u32/.u64, .raw(n), .lp_string(), .overran()). Built for the .veloxpart.meta reader and the 4-byte NM framing; every read is bounds-checked and latches on overrun (reader-first, fuzz-ready).

vdm/util/event_bus.hpp

EventBus — typed, thread-safe in-process pub/sub. subscribe<E>(fn) -> Token, publish<E>(ev) (synchronous, calling thread, registration order), unsubscribe(Token), and RAII subscribe_scoped<E> returning a Subscription. Handlers may (un)subscribe or publish during dispatch. Handlers must not throw. Not a hot-path structure — progress is coalesced to ≤4 Hz upstream.

vdm/util/thread_pool.hpp

ThreadPool — fixed-size std::jthread pool for bounded off-loop work (hashing, fsync batches, DNS pre-resolve). submit(fn, args...) -> std::future<R>; propagates exceptions through the future; drains already-queued tasks on destruction. Not the transfer loop — net/ will own one curl_multi per dedicated worker.

vdm/util/log.hpp

Sink interface — core does no I/O itself. LogSink abstract base; DAEMON installs one via set_log_sink(), default discards. CallbackSink adapter (with a min-level filter). VDM_LOG_{TRACE,DEBUG,INFO,WARN,ERROR}(category, fmt, args...)std::format syntax, only formatted when a sink is installed and wants the level.


rules/ — filename sanitization, collision policy, rule-table matching

Pure functions only: no I/O, no filesystem access, no notion of the wire Rule type or its JSON/SQL representation. DAEMON owns the rule table (storage, rules.upsert, the generated Rule type) and decodes it into the plain structs below before calling in.

vdm/rules/filename.hpp

sanitize_filename(raw, max_bytes = 255) — turns a raw candidate (from net::parse_content_disposition or net::url_filename, neither of which is filesystem-safe by design — see their own headers) into one safe to create on ext4, APFS, and NTFS alike: strips separators/control bytes, folds NTFS-illegal characters to _, neutralizes reserved Windows device names (CON, COM1, ...), and clamps length on a UTF-8 boundary. Total: never empty, never throws. Not the path-traversal security boundary — that's DAEMON's fs/safepath, which runs after this and is the one that matters adversarially.

vdm/rules/collision.hpp

resolve_collision(desired, exists, policy, max_attempts = 1000) — given an existence predicate (DAEMON supplies a real one; tests supply an in-memory set), finds the next free name Explorer/Finder-style ("name (1).ext", "name (2).ext", ...) under CollisionPolicy::rename, or returns desired unchanged under ::overwrite. Never fabricates a guaranteed-unique name past max_attempts — returns the last candidate tried and leaves "still colliding" for the caller to treat as a real error.

vdm/rules/match.hpp

match_rules(rules, input) -> optional<RuleAction> — the evaluation half of contracts/schema/types/Rule.schema.json: tries rules in ascending priority order (ties keep table order), skips disabled rows, returns the first whose every present match clause (extensions, mime_types, host_pattern, url_pattern, min_size_bytes/max_size_bytes) is satisfied — an absent clause is not a constraint, and a size clause never matches speculatively when MatchInput::size_bytes is still unknown (pre-probe). std::nullopt means no rule matched; the caller's own default category applies. glob_match(pattern, text) — the */? matcher host_pattern and url_pattern both use, case-insensitive, bounded work even on a pathological all-* pattern (iterative, not recursive).