From 201ebc55d4fff7312608ee59f4ee51b5f248c1f3 Mon Sep 17 00:00:00 2001 From: sami Date: Wed, 9 Sep 2026 23:53:55 +0400 Subject: [PATCH 1/2] core: net/probe + Content-Disposition parser + URL splitter (stage 3) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit net/content_disposition — total parser for the mojibake-prone header: RFC 6266 filename (quoted/token), RFC 5987 filename* ext-values (charset'lang'pct-encoded, incl. RFC 2231 continuations), legacy RFC 2047 encoded-words (=?UTF-8?B?..?= / ?Q?), and raw Latin-1 bytes; prefers filename* over filename; strips path components AFTER decoding (a base64 payload can hold '/'). 22-case test table. net/text_codec (internal) — percent-decode, UTF-8 validation, Latin-1-> UTF-8, base64, RFC 2047 — shared by the CD parser and the URL splitter. net/url — a small total URL splitter (scheme/userinfo/host/port/path/ query/fragment, http(s) validity) and url_filename() for the last path segment; used for the filename fallback. net/probe — HEAD then a ranged GET bytes=0-0 that PROVES resumability (206 + matching Content-Range + a validator), rather than trusting Accept-Ranges which servers lie about; the ranged GET is also the HEAD- refused (403/405/501) fallback. 401/407 -> success result with requires_auth, not an error. Runs on its own pool (max_concurrent, default 4) outside the segment budget per ADR 0011 §5. suggest_filename() does the resolution order (explicit -> disposition -> URL -> download.bin) with a light strip; rules/ (stage 9) owns the authoritative sanitize. tools/fuzz — libFuzzer targets for the CD parser and the URL splitter, compiling the parser sources directly so they're fully instrumented; self-guards on VELOX_BUILD_FUZZ + Clang (the top-level CMake adds every tools/* unconditionally). Seed corpora included. Fixed on the way: a p -> Transfer -> State -> cbs -> p reference cycle in Prober that leaked every probe (drop the stored Transfer; the worker keeps State alive). Tests green under ASan/UBSan and TSan. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01HPPSGhiArbvQgwC2DNiURS --- core/CMakeLists.txt | 18 +- core/include/vdm/net/content_disposition.hpp | 42 +++ core/include/vdm/net/probe.hpp | 90 +++++ core/include/vdm/net/url.hpp | 40 ++ core/src/net/content_disposition.cpp | 279 ++++++++++++++ core/src/net/probe.cpp | 348 ++++++++++++++++++ core/src/net/text_codec.cpp | 262 +++++++++++++ core/src/net/text_codec.hpp | 40 ++ core/src/net/url.cpp | 126 +++++++ core/tests/CMakeLists.txt | 25 +- core/tests/net/content_disposition_test.cpp | 148 ++++++++ core/tests/net/probe_test.cpp | 163 ++++++++ core/tests/net/url_test.cpp | 71 ++++ tools/fuzz/CMakeLists.txt | 42 +++ tools/fuzz/content_disposition_fuzz.cpp | 23 ++ tools/fuzz/corpus/content_disposition/ext5987 | 1 + tools/fuzz/corpus/content_disposition/inline | 1 + tools/fuzz/corpus/content_disposition/plain | 1 + tools/fuzz/corpus/content_disposition/rfc2047 | 1 + .../fuzz/corpus/content_disposition/traversal | 1 + tools/fuzz/corpus/url/basic | 1 + tools/fuzz/corpus/url/encoded | 1 + tools/fuzz/corpus/url/ftp | 1 + tools/fuzz/corpus/url/ipv6 | 1 + tools/fuzz/url_fuzz.cpp | 24 ++ 25 files changed, 1731 insertions(+), 19 deletions(-) create mode 100644 core/include/vdm/net/content_disposition.hpp create mode 100644 core/include/vdm/net/probe.hpp create mode 100644 core/include/vdm/net/url.hpp create mode 100644 core/src/net/content_disposition.cpp create mode 100644 core/src/net/probe.cpp create mode 100644 core/src/net/text_codec.cpp create mode 100644 core/src/net/text_codec.hpp create mode 100644 core/src/net/url.cpp create mode 100644 core/tests/net/content_disposition_test.cpp create mode 100644 core/tests/net/probe_test.cpp create mode 100644 core/tests/net/url_test.cpp create mode 100644 tools/fuzz/CMakeLists.txt create mode 100644 tools/fuzz/content_disposition_fuzz.cpp create mode 100644 tools/fuzz/corpus/content_disposition/ext5987 create mode 100644 tools/fuzz/corpus/content_disposition/inline create mode 100644 tools/fuzz/corpus/content_disposition/plain create mode 100644 tools/fuzz/corpus/content_disposition/rfc2047 create mode 100644 tools/fuzz/corpus/content_disposition/traversal create mode 100644 tools/fuzz/corpus/url/basic create mode 100644 tools/fuzz/corpus/url/encoded create mode 100644 tools/fuzz/corpus/url/ftp create mode 100644 tools/fuzz/corpus/url/ipv6 create mode 100644 tools/fuzz/url_fuzz.cpp diff --git a/core/CMakeLists.txt b/core/CMakeLists.txt index 299a4f2..f5cbdb1 100644 --- a/core/CMakeLists.txt +++ b/core/CMakeLists.txt @@ -13,6 +13,10 @@ add_library(veloxcore STATIC src/util/thread_pool.cpp src/net/curl_error.cpp src/net/http_client.cpp + src/net/text_codec.cpp + src/net/content_disposition.cpp + src/net/url.cpp + src/net/probe.cpp ) add_library(velox::core ALIAS veloxcore) @@ -37,13 +41,7 @@ if(VELOX_BUILD_TESTS) add_subdirectory(tests) endif() -# libFuzzer is clang-only; a GCC configure with -DVELOX_BUILD_FUZZ=ON (the `ci` preset) -# must not hard-fail. No fuzz targets exist yet — they arrive with net/probe (stage 3) -# and meta/veloxpart (stage 5). -if(VELOX_BUILD_FUZZ AND NOT CMAKE_CXX_COMPILER_ID MATCHES "Clang") - message(STATUS "veloxcore: VELOX_BUILD_FUZZ set but compiler is " - "${CMAKE_CXX_COMPILER_ID}; fuzz targets need Clang and will be skipped.") -endif() -# When fuzz targets land (stage 3: Content-Disposition, URL; stage 5: .veloxpart.meta), -# they are added here under `if(VELOX_BUILD_FUZZ AND CMAKE_CXX_COMPILER_ID MATCHES "Clang")` -# and live in ${CMAKE_SOURCE_DIR}/tools/fuzz (owned by lane CORE). +# Fuzz targets live in ${CMAKE_SOURCE_DIR}/tools/fuzz (lane CORE) and are wired in by the +# top-level CMakeLists.txt, which add_subdirectory()s every tools/* with a CMakeLists. +# tools/fuzz/CMakeLists.txt self-guards on VELOX_BUILD_FUZZ + a Clang compiler. +# Present: Content-Disposition, URL (stage 3). Coming: .veloxpart.meta (stage 5). diff --git a/core/include/vdm/net/content_disposition.hpp b/core/include/vdm/net/content_disposition.hpp new file mode 100644 index 0000000..7696d15 --- /dev/null +++ b/core/include/vdm/net/content_disposition.hpp @@ -0,0 +1,42 @@ +// vdm/net/content_disposition.hpp — parse a Content-Disposition header into a filename. +// +// This is a classic mojibake source (AGENT-CORE §3): RFC 6266 `filename`, RFC 5987 +// `filename*` ext-values, RFC 2047 encoded-words in the legacy quoted form, and raw +// Latin-1 bytes all show up in the wild. The parser is total — hostile input yields a +// best-effort or empty result, never a throw or a crash — and has its own test table +// (content_disposition_test.cpp) and a fuzz target (tools/fuzz). +// +// This header compiles standalone. + +#ifndef VDM_NET_CONTENT_DISPOSITION_HPP +#define VDM_NET_CONTENT_DISPOSITION_HPP + +#include +#include + +namespace vdm::net { + +struct ContentDisposition { + enum class Type { none, inline_, attachment, form_data, other }; + + Type type = Type::none; + + // Best-effort UTF-8 filename, path components stripped, or empty when the header + // carries none. NOT sanitized for the filesystem — that is rules/ (stage 9); this + // only decodes and de-mojibakes. `..` and control characters may still be present. + std::string filename; + + // The filename came from an RFC 5987 `filename*` ext-value (preferred over a plain + // `filename` per RFC 6266 §4.3 when both are present). + bool filename_from_ext = false; + + [[nodiscard]] bool is_attachment() const noexcept { return type == Type::attachment; } + [[nodiscard]] bool has_filename() const noexcept { return !filename.empty(); } +}; + +// Parse the value of a Content-Disposition header (everything after the colon). +[[nodiscard]] ContentDisposition parse_content_disposition(std::string_view header_value); + +} // namespace vdm::net + +#endif // VDM_NET_CONTENT_DISPOSITION_HPP diff --git a/core/include/vdm/net/probe.hpp b/core/include/vdm/net/probe.hpp new file mode 100644 index 0000000..2122190 --- /dev/null +++ b/core/include/vdm/net/probe.hpp @@ -0,0 +1,90 @@ +// vdm/net/probe.hpp — "what is at this URL?" without downloading it. +// +// Feeds the File Info dialog (docs/04 §2). HEAD first; a ranged GET `bytes=0-0` follows to +// PROVE resumability (a 206 with a matching Content-Range) rather than trust +// `Accept-Ranges`, which servers lie about (docs/06 R4). The ranged GET is also the +// fallback when HEAD is refused (403/405/501). +// +// Runs on its own small worker pool, sized outside the segment budget (ADR 0011 §5) so a +// burst of probes can't starve transfers and capture.offer's 750 ms path never waits on +// one. +// +// This header compiles standalone. + +#ifndef VDM_NET_PROBE_HPP +#define VDM_NET_PROBE_HPP + +#include +#include +#include +#include +#include +#include +#include + +#include "vdm/net/content_disposition.hpp" +#include "vdm/net/http_types.hpp" +#include "vdm/util/result.hpp" + +namespace vdm::net { + +struct ProbeRequest { + std::string url; + std::vector headers; // browser headers, verbatim + std::vector cookies; + std::string user_agent; + std::string referrer; + ProxyConfig proxy; + + long connect_timeout_ms = 15000; + long overall_timeout_ms = 25000; // download.probe deadline is 30 s +}; + +struct ProbeResult { + std::string effective_url; + std::vector redirect_chain; // requested URL first, effective_url last + long http_status = 0; + + std::optional total_size; // full-resource size, if known + std::string mime; // Content-Type value (params kept) + std::string etag; + std::string last_modified; + + bool accept_ranges = false; // server advertised Accept-Ranges: bytes + bool resumable = false; // PROVEN: ranged GET -> 206 + matching Content-Range, + // and a validator (ETag or Last-Modified) is present + bool requires_auth = false; // a 401/407 was seen + + // Decoded, path-stripped; NOT filesystem-sanitized (rules/ owns that, stage 9). + std::string filename_from_disposition; + std::string filename_from_url; + ContentDisposition::Type disposition_type = ContentDisposition::Type::none; +}; + +// Resolution order (docs/04 §2.5): explicit user name -> Content-Disposition -> URL path +// segment -> "download.bin". Only a light path/control strip here; rules/ does the +// authoritative sanitize, byte cap, and collision handling. +[[nodiscard]] std::string suggest_filename(const ProbeResult &r, + std::string_view explicit_name = {}); + +class Prober { + public: + // max_concurrent bounds outstanding probe transfers; the rest queue. + explicit Prober(unsigned max_concurrent = 4); + ~Prober(); + Prober(const Prober &) = delete; + Prober &operator=(const Prober &) = delete; + + // Async. `done` runs on an internal worker thread, exactly once. A 401/407 is a + // SUCCESS result with requires_auth = true (the GUI collects credentials), not an + // error; transport failures and hard HTTP errors (404/410/5xx) are errors. + void probe(ProbeRequest req, std::function)> done); + + private: + struct Impl; + std::unique_ptr impl_; +}; + +} // namespace vdm::net + +#endif // VDM_NET_PROBE_HPP diff --git a/core/include/vdm/net/url.hpp b/core/include/vdm/net/url.hpp new file mode 100644 index 0000000..b9dc2d3 --- /dev/null +++ b/core/include/vdm/net/url.hpp @@ -0,0 +1,40 @@ +// vdm/net/url.hpp — a small, total URL splitter. +// +// Not a full RFC 3986 parser (libcurl does the real fetching); just enough to pull a +// filename out of a path and to sanity-check a scheme. Total on hostile input — it has a +// fuzz target (tools/fuzz) — never throws, never asserts. +// +// This header compiles standalone. + +#ifndef VDM_NET_URL_HPP +#define VDM_NET_URL_HPP + +#include +#include +#include +#include + +namespace vdm::net { + +struct SplitUrl { + std::string scheme; // lowercased, without "://" + std::string userinfo; // before '@', if any + std::string host; // lowercased; bracketed IPv6 keeps its brackets stripped + std::optional port; + std::string path; // includes the leading '/', or empty + std::string query; // without the '?' + std::string fragment; // without the '#' + bool valid = false; // scheme + host present and scheme is http/https + + [[nodiscard]] bool is_http() const noexcept { return scheme == "http" || scheme == "https"; } +}; + +[[nodiscard]] SplitUrl split_url(std::string_view url); + +// The last non-empty path segment, percent-decoded, path components stripped. Empty when +// the path has no usable segment (ends in '/', is empty, or is only "/"). +[[nodiscard]] std::string url_filename(std::string_view url); + +} // namespace vdm::net + +#endif // VDM_NET_URL_HPP diff --git a/core/src/net/content_disposition.cpp b/core/src/net/content_disposition.cpp new file mode 100644 index 0000000..d86aab6 --- /dev/null +++ b/core/src/net/content_disposition.cpp @@ -0,0 +1,279 @@ +// vdm/net/content_disposition.cpp + +#include "vdm/net/content_disposition.hpp" + +#include +#include +#include +#include + +#include "net/text_codec.hpp" + +namespace vdm::net { +namespace { + +using detail::decode_rfc2047; +using detail::latin1_to_utf8; +using detail::percent_decode; +using detail::to_utf8_best_effort; + +std::string_view trim(std::string_view s) { + while (!s.empty() && (s.front() == ' ' || s.front() == '\t')) + s.remove_prefix(1); + while (!s.empty() && + (s.back() == ' ' || s.back() == '\t' || s.back() == '\r' || s.back() == '\n')) + s.remove_suffix(1); + return s; +} + +std::string ascii_lower(std::string_view s) { + std::string r(s); + for (char &c : r) + if (c >= 'A' && c <= 'Z') + c = static_cast(c - 'A' + 'a'); + return r; +} + +// Split "type; a=b; c*=d; e=\"f;g\"" into the type and a param list, honouring quoted +// strings (a ';' inside quotes is not a separator). For quoted values the value stored is +// the *raw* inner text (quotes removed, escapes NOT yet resolved) with `quoted = true`; +// callers resolve escapes and strip path components together (order matters — see +// unquote_strip). Keys are lowercased. +struct Param { + std::string key; + std::string value; + bool quoted = false; +}; +struct Params { + std::string type; + std::vector kv; + + [[nodiscard]] const Param *find(std::string_view key) const { + for (const auto &p : kv) + if (p.key == key) + return &p; + return nullptr; + } +}; + +Params tokenize(std::string_view h) { + Params out; + std::size_t i = 0; + const std::size_t n = h.size(); + + auto read_segment = [&]() -> std::string_view { + std::size_t start = i; + bool in_q = false; + for (; i < n; ++i) { + char c = h[i]; + if (c == '"') { + in_q = !in_q; + } else if (c == '\\' && in_q && i + 1 < n) { + ++i; // skip escaped char + } else if (c == ';' && !in_q) { + break; + } + } + std::string_view seg = h.substr(start, i - start); + if (i < n) + ++i; // consume ';' + return seg; + }; + + out.type = ascii_lower(trim(read_segment())); + + while (i < n) { + std::string_view seg = trim(read_segment()); + if (seg.empty()) + continue; + auto eq = seg.find('='); + if (eq == std::string_view::npos) { + out.kv.emplace_back(ascii_lower(seg), std::string{}); + continue; + } + std::string key = ascii_lower(trim(seg.substr(0, eq))); + std::string_view rawval = trim(seg.substr(eq + 1)); + + Param param; + param.key = std::move(key); + if (rawval.size() >= 2 && rawval.front() == '"') { + std::string_view inner = rawval.substr(1); + auto close = inner.rfind('"'); + if (close != std::string_view::npos) + inner = inner.substr(0, close); + param.value.assign(inner); // raw, escapes unresolved + param.quoted = true; + } else { + param.value.assign(rawval); + } + out.kv.push_back(std::move(param)); + } + return out; +} + +std::string strip_path(std::string s) { + auto slash = s.find_last_of("/\\"); + if (slash != std::string::npos) + s.erase(0, slash + 1); + return s; +} + +// Resolve ONLY the `\"` escape (needed so a quote can appear mid-name). Every other +// backslash is kept literal and later treated as a path separator by strip_path — real +// Windows paths in the wild use `\` unescaped, and path-traversal defence matters more +// than supporting the vanishingly rare filename with a literal backslash. +std::string unescape_dquote(std::string_view raw) { + std::string out; + out.reserve(raw.size()); + for (std::size_t i = 0; i < raw.size(); ++i) { + if (raw[i] == '\\' && i + 1 < raw.size() && raw[i + 1] == '"') { + out.push_back('"'); + ++i; + } else { + out.push_back(raw[i]); + } + } + return out; +} + +// Decode an RFC 5987 ext-value: charset'lang'pct-encoded-octets +std::string decode_ext_value(std::string_view v) { + auto q1 = v.find('\''); + if (q1 == std::string_view::npos) + return percent_decode(v); // malformed: best effort + auto q2 = v.find('\'', q1 + 1); + if (q2 == std::string_view::npos) + return percent_decode(v.substr(q1 + 1)); + + std::string_view charset = v.substr(0, q1); + std::string_view enc = v.substr(q2 + 1); + std::string bytes = percent_decode(enc); + + std::string cs = ascii_lower(charset); + if (cs == "iso-8859-1" || cs == "latin1") + return latin1_to_utf8(bytes); + return to_utf8_best_effort(bytes); // utf-8 or unknown -> best effort +} + +// Reassemble RFC 2231 continuations: name*0*, name*1, name*2* ... in order. +std::string join_continuations(const Params &p, std::string_view base, bool &is_ext) { + std::vector> parts; + is_ext = false; + for (const auto ¶m : p.kv) { + const std::string &k = param.key; + const std::string &val = param.value; + if (k.size() <= base.size() + 1 || k.compare(0, base.size(), base) != 0) + continue; + if (k[base.size()] != '*') + continue; + std::string_view rest(k); + rest.remove_prefix(base.size() + 1); // after "base*" + bool star = false; + if (!rest.empty() && rest.back() == '*') { + star = true; + rest.remove_suffix(1); + } + int idx = 0; + for (char c : rest) { + if (c < '0' || c > '9') { + idx = -1; + break; + } + idx = idx * 10 + (c - '0'); + } + if (idx < 0) + continue; + if (star) + is_ext = true; + parts.emplace_back(idx, val); + } + if (parts.empty()) + return {}; + std::sort(parts.begin(), parts.end(), + [](const auto &a, const auto &b) { return a.first < b.first; }); + + // Piece 0 (if star-form) carries charset'lang' prefix; later pieces are raw + // percent-encoded. Concatenate the percent-encoded text then decode once. + std::string charset_prefix; + std::string enc; + bool first = true; + for (auto &[idx, val] : parts) { + if (first && is_ext) { + auto q1 = val.find('\''); + auto q2 = (q1 == std::string::npos) ? std::string::npos : val.find('\'', q1 + 1); + if (q2 != std::string::npos) { + charset_prefix = val.substr(0, q2 + 1); + enc += val.substr(q2 + 1); + } else { + enc += val; + } + } else { + enc += val; + } + first = false; + } + if (is_ext) + return decode_ext_value(charset_prefix + enc); + return to_utf8_best_effort(percent_decode(enc)); +} + +ContentDisposition::Type classify(std::string_view t) { + if (t == "inline") + return ContentDisposition::Type::inline_; + if (t == "attachment") + return ContentDisposition::Type::attachment; + if (t == "form-data") + return ContentDisposition::Type::form_data; + if (t.empty()) + return ContentDisposition::Type::none; + return ContentDisposition::Type::other; +} + +} // namespace + +ContentDisposition parse_content_disposition(std::string_view header_value) { + ContentDisposition cd; + header_value = trim(header_value); + if (header_value.empty()) + return cd; + + Params p = tokenize(header_value); + cd.type = classify(p.type); + + // RFC 6266 §4.3: prefer filename* over filename. + std::string ext_name; + bool ext_is_ext = false; + if (const Param *fstar = p.find("filename*")) { + ext_name = decode_ext_value(fstar->value); + ext_is_ext = true; + } else { + std::string joined = join_continuations(p, "filename", ext_is_ext); + if (!joined.empty()) + ext_name = std::move(joined); + } + + std::string plain_name; + if (const Param *f = p.find("filename")) { + // Resolve `\"`, then decode legacy encoded-words if present — BEFORE stripping + // path components, since a base64 payload can legitimately contain '/'. + std::string raw = f->quoted ? unescape_dquote(f->value) : f->value; + bool had_ew = false; + std::string decoded = decode_rfc2047(raw, &had_ew); + plain_name = had_ew ? std::move(decoded) : to_utf8_best_effort(raw); + } + + if (!ext_name.empty()) { + cd.filename = strip_path(std::move(ext_name)); + cd.filename_from_ext = ext_is_ext; + } else if (!plain_name.empty()) { + cd.filename = strip_path(std::move(plain_name)); + cd.filename_from_ext = false; + } + + // Trim ASCII whitespace the decoders may have left at the edges. + std::string_view fv = trim(cd.filename); + cd.filename.assign(fv); + return cd; +} + +} // namespace vdm::net diff --git a/core/src/net/probe.cpp b/core/src/net/probe.cpp new file mode 100644 index 0000000..608e02e --- /dev/null +++ b/core/src/net/probe.cpp @@ -0,0 +1,348 @@ +// vdm/net/probe.cpp + +#include "vdm/net/probe.hpp" + +#include +#include +#include + +#include "net/curl_error.hpp" +#include "vdm/net/http_client.hpp" +#include "vdm/net/url.hpp" + +namespace vdm::net { +namespace { + +std::string_view trim(std::string_view s) { + while (!s.empty() && (s.front() == ' ' || s.front() == '\t')) + s.remove_prefix(1); + while (!s.empty() && (s.back() == ' ' || s.back() == '\t')) + s.remove_suffix(1); + return s; +} + +bool iequals(std::string_view a, std::string_view b) { + return HeaderList::iequals(a, b); +} + +// "bytes 0-0/12345" -> 12345 ; "bytes 0-0/*" or malformed -> nullopt +std::optional total_from_content_range(std::string_view v) { + auto slash = v.find('/'); + if (slash == std::string_view::npos) + return std::nullopt; + std::string_view tail = trim(v.substr(slash + 1)); + if (tail.empty() || tail == "*") + return std::nullopt; + std::uint64_t n = 0; + auto [p, ec] = std::from_chars(tail.data(), tail.data() + tail.size(), n); + (void)p; + if (ec != std::errc{}) + return std::nullopt; + return n; +} + +std::string light_sanitize(std::string_view in) { + std::string out; + out.reserve(in.size()); + for (unsigned char c : in) { + if (c == '/' || c == '\\' || c == 0) { + out.push_back('_'); + } else if (c >= 0x20 || (c & 0x80)) { // keep printable ASCII + all UTF-8 bytes + out.push_back(static_cast(c)); + } + } + while (!out.empty() && (out.back() == '.' || out.back() == ' ')) + out.pop_back(); + if (out == "." || out == "..") + out.clear(); + return out; +} + +} // namespace + +std::string suggest_filename(const ProbeResult &r, std::string_view explicit_name) { + std::string cand; + if (!explicit_name.empty()) + cand = light_sanitize(explicit_name); + if (cand.empty() && !r.filename_from_disposition.empty()) + cand = light_sanitize(r.filename_from_disposition); + if (cand.empty() && !r.filename_from_url.empty()) + cand = light_sanitize(r.filename_from_url); + if (cand.empty()) + cand = "download.bin"; + return cand; +} + +// --- Prober::Impl --------------------------------------------------------------------- +struct Prober::Impl { + struct Job { + ProbeRequest req; + std::function)> done; + }; + + struct P { + Impl *self = nullptr; + Job job; + ProbeResult result; + bool had_head_ok = false; + bool delivered = false; + }; + + explicit Impl(unsigned max_concurrent) + : max_(max_concurrent ? max_concurrent : 1), client_(HttpClient::Options{.workers = 2}) {} + + unsigned max_; + HttpClient client_; + std::mutex mu_; + unsigned inflight_ = 0; + std::deque pending_; + + void submit(Job j) { + { + std::lock_guard lk(mu_); + if (inflight_ >= max_) { + pending_.push_back(std::move(j)); + return; + } + ++inflight_; + } + start(std::move(j)); + } + + void finish_one() { + Job next; + bool have_next = false; + { + std::lock_guard lk(mu_); + --inflight_; + if (!pending_.empty()) { + next = std::move(pending_.front()); + pending_.pop_front(); + ++inflight_; + have_next = true; + } + } + if (have_next) + start(std::move(next)); + } + + Request base_request(const ProbeRequest &pr) { + Request r; + r.url = pr.url; + r.headers = pr.headers; + r.cookies = pr.cookies; + r.user_agent = pr.user_agent; + r.referrer = pr.referrer; + r.proxy = pr.proxy; + r.follow_redirects = true; + r.accept_encoding = false; + r.connect_timeout_ms = pr.connect_timeout_ms; + r.overall_timeout_ms = pr.overall_timeout_ms; + r.low_speed_bytes_per_sec = 0; // probes are tiny; no stall detector + r.low_speed_secs = 0; + return r; + } + + void start(Job j) { + auto p = std::make_shared

(); + p->self = this; + p->job = std::move(j); + p->result.effective_url = p->job.req.url; + + Request req = base_request(p->job.req); + req.method = Method::head; + + TransferCallbacks cbs; + cbs.on_head = [p](const ResponseHead &h) { + absorb_head(p->result, h); + return DataAction::abort; + }; + cbs.on_data = [](ConstByteSpan) { return DataAction::abort; }; + cbs.on_finished = [p](Result r) { on_head_done(p, std::move(r)); }; + + // The worker keeps Transfer::State alive; the callbacks keep `p` alive. Storing + // the Transfer in `p` would make a p -> Transfer -> State -> cbs -> p cycle. + client_.start(std::move(req), std::move(cbs)); + } + + static void absorb_head(ProbeResult &res, const ResponseHead &h) { + if (h.status) + res.http_status = h.status; + if (!h.effective_url.empty()) + res.effective_url = h.effective_url; + refresh_common(res, h); + if (h.content_length && !res.total_size) + res.total_size = h.content_length; + } + + static void refresh_common(ProbeResult &res, const ResponseHead &h) { + if (auto v = h.headers.get("Content-Type"); v && res.mime.empty()) { + std::string_view mv = *v; + mv = mv.substr(0, mv.find(';')); + res.mime.assign(trim(mv)); + } + if (auto v = h.headers.get("ETag"); v && res.etag.empty()) + res.etag.assign(*v); + if (auto v = h.headers.get("Last-Modified"); v && res.last_modified.empty()) + res.last_modified.assign(*v); + if (auto v = h.headers.get("Accept-Ranges")) { + if (iequals(trim(*v), "bytes")) + res.accept_ranges = true; + } + if (auto v = h.headers.get("Content-Disposition"); + v && res.filename_from_disposition.empty()) { + ContentDisposition cd = parse_content_disposition(*v); + res.disposition_type = cd.type; + if (cd.has_filename()) + res.filename_from_disposition = cd.filename; + } + } + + static bool has_validator(const ProbeResult &r) { + return !r.etag.empty() || !r.last_modified.empty(); + } + + static void on_head_done(std::shared_ptr

p, Result r) { + const long status = p->result.http_status; + + if (status == 0) { // never got headers -> transport failure + deliver(p, r.has_value() ? Result(ErrorInfo(Error::probe_failed)) + : Result(std::move(r).error())); + return; + } + p->had_head_ok = (status >= 200 && status < 300); + + if (status == 401 || status == 407) { + p->result.requires_auth = true; + finalize_and_deliver(p); + return; + } + if (status == 403 || status == 405 || status == 501) { + start_range_get(p); // HEAD refused; the ranged GET is now the primary probe + return; + } + if (status >= 400) { + deliver(p, ErrorInfo(detail::error_from_curl(CURLE_OK, status), "probe HEAD", + static_cast(status))); + return; + } + // 2xx. Prove resumability with a ranged GET unless the server said "none". + if (auto ar = p->result.accept_ranges; !ar) { + // Accept-Ranges absent or not "bytes": still try one ranged GET — servers + // that support ranges without advertising are common (docs/06 R4). + } + start_range_get(p); + } + + static void start_range_get(const std::shared_ptr

&p) { + Request req = p->self->base_request(p->job.req); + req.method = Method::get; + req.range = ByteRange{0, 0}; + + TransferCallbacks cbs; + cbs.on_head = [p](const ResponseHead &h) { + absorb_range_head(p->result, h); + return DataAction::abort; // we don't need the one body byte + }; + cbs.on_data = [](ConstByteSpan) { return DataAction::abort; }; + cbs.on_finished = [p](Result r) { on_range_done(p, std::move(r)); }; + + p->self->client_.start(std::move(req), std::move(cbs)); + } + + static void absorb_range_head(ProbeResult &res, const ResponseHead &h) { + if (h.status) + res.http_status = h.status; + if (!h.effective_url.empty()) + res.effective_url = h.effective_url; + refresh_common(res, h); + + if (h.status == 206) { + res.accept_ranges = true; // proven, not just advertised + if (auto cr = h.headers.get("Content-Range")) { + if (auto total = total_from_content_range(*cr)) + res.total_size = total; + } + } else if (h.status == 200) { + if (h.content_length) + res.total_size = h.content_length; + } + } + + static void on_range_done(std::shared_ptr

p, Result r) { + const long status = p->result.http_status; + + if (status == 0) { // range GET died at transport level + if (p->had_head_ok) { + p->result.resumable = false; + finalize_and_deliver(p); + } else { + deliver(p, r.has_value() ? Result(ErrorInfo(Error::probe_failed)) + : Result(std::move(r).error())); + } + return; + } + if (status == 401 || status == 407) { + p->result.requires_auth = true; + finalize_and_deliver(p); + return; + } + if (status == 206) { + p->result.resumable = has_validator(p->result); + finalize_and_deliver(p); + return; + } + if (status == 200 || status == 416) { + p->result.resumable = false; + finalize_and_deliver(p); + return; + } + if (status >= 400) { + if (p->had_head_ok) { + p->result.resumable = false; + finalize_and_deliver(p); + } else { + deliver(p, ErrorInfo(detail::error_from_curl(CURLE_OK, status), "probe ranged GET", + static_cast(status))); + } + return; + } + p->result.resumable = false; + finalize_and_deliver(p); + } + + static void finalize_and_deliver(const std::shared_ptr

&p) { + ProbeResult &res = p->result; + if (res.effective_url.empty()) + res.effective_url = p->job.req.url; + res.filename_from_url = url_filename(res.effective_url); + if (res.filename_from_url.empty() && res.effective_url != p->job.req.url) + res.filename_from_url = url_filename(p->job.req.url); + + res.redirect_chain.clear(); + res.redirect_chain.push_back(p->job.req.url); + if (res.effective_url != p->job.req.url) + res.redirect_chain.push_back(res.effective_url); + + deliver(p, ProbeResult(res)); + } + + static void deliver(const std::shared_ptr

&p, Result out) { + if (p->delivered) + return; + p->delivered = true; + Impl *self = p->self; + p->job.done(std::move(out)); + self->finish_one(); + } +}; + +// --- Prober ----------------------------------------------------------------------- +Prober::Prober(unsigned max_concurrent) : impl_(std::make_unique(max_concurrent)) {} +Prober::~Prober() = default; + +void Prober::probe(ProbeRequest req, std::function)> done) { + impl_->submit(Impl::Job{std::move(req), std::move(done)}); +} + +} // namespace vdm::net diff --git a/core/src/net/text_codec.cpp b/core/src/net/text_codec.cpp new file mode 100644 index 0000000..ccc6f38 --- /dev/null +++ b/core/src/net/text_codec.cpp @@ -0,0 +1,262 @@ +// vdm/net/text_codec.cpp + +#include "net/text_codec.hpp" + +#include +#include + +namespace vdm::net::detail { +namespace { + +int hex_val(char c) noexcept { + if (c >= '0' && c <= '9') + return c - '0'; + if (c >= 'a' && c <= 'f') + return c - 'a' + 10; + if (c >= 'A' && c <= 'F') + return c - 'A' + 10; + return -1; +} + +void append_utf8(std::string &out, std::uint32_t cp) { + if (cp <= 0x7F) { + out.push_back(static_cast(cp)); + } else if (cp <= 0x7FF) { + out.push_back(static_cast(0xC0 | (cp >> 6))); + out.push_back(static_cast(0x80 | (cp & 0x3F))); + } else if (cp <= 0xFFFF) { + out.push_back(static_cast(0xE0 | (cp >> 12))); + out.push_back(static_cast(0x80 | ((cp >> 6) & 0x3F))); + out.push_back(static_cast(0x80 | (cp & 0x3F))); + } else { + out.push_back(static_cast(0xF0 | (cp >> 18))); + out.push_back(static_cast(0x80 | ((cp >> 12) & 0x3F))); + out.push_back(static_cast(0x80 | ((cp >> 6) & 0x3F))); + out.push_back(static_cast(0x80 | (cp & 0x3F))); + } +} + +bool charset_is(std::string_view cs, std::string_view want) { + if (cs.size() != want.size()) + return false; + for (std::size_t i = 0; i < cs.size(); ++i) { + char a = cs[i], b = want[i]; + if (a >= 'A' && a <= 'Z') + a = static_cast(a - 'A' + 'a'); + if (b >= 'A' && b <= 'Z') + b = static_cast(b - 'A' + 'a'); + if (a != b) + return false; + } + return true; +} + +bool is_utf8(std::string_view cs) { + return charset_is(cs, "utf-8") || charset_is(cs, "utf8"); +} +bool is_latin1(std::string_view cs) { + return charset_is(cs, "iso-8859-1") || charset_is(cs, "latin1") || + charset_is(cs, "iso8859-1") || charset_is(cs, "windows-1252"); +} + +} // namespace + +std::string percent_decode(std::string_view in, bool plus_as_space) { + std::string out; + out.reserve(in.size()); + for (std::size_t i = 0; i < in.size(); ++i) { + char c = in[i]; + if (c == '%' && i + 2 < in.size()) { + int hi = hex_val(in[i + 1]); + int lo = hex_val(in[i + 2]); + if (hi >= 0 && lo >= 0) { + out.push_back(static_cast((hi << 4) | lo)); + i += 2; + continue; + } + } + if (c == '+' && plus_as_space) { + out.push_back(' '); + continue; + } + out.push_back(c); + } + return out; +} + +bool is_valid_utf8(std::string_view s) noexcept { + std::size_t i = 0; + const std::size_t n = s.size(); + auto cont = [&](std::size_t k) { + return k < n && (static_cast(s[k]) & 0xC0) == 0x80; + }; + while (i < n) { + unsigned char c = static_cast(s[i]); + if (c < 0x80) { + ++i; + } else if ((c & 0xE0) == 0xC0) { + if (!cont(i + 1)) + return false; + std::uint32_t cp = (c & 0x1F) << 6 | (static_cast(s[i + 1]) & 0x3F); + if (cp < 0x80) + return false; // overlong + i += 2; + } else if ((c & 0xF0) == 0xE0) { + if (!cont(i + 1) || !cont(i + 2)) + return false; + std::uint32_t cp = (c & 0x0F) << 12 | + (static_cast(s[i + 1]) & 0x3F) << 6 | + (static_cast(s[i + 2]) & 0x3F); + if (cp < 0x800 || (cp >= 0xD800 && cp <= 0xDFFF)) + return false; + i += 3; + } else if ((c & 0xF8) == 0xF0) { + if (!cont(i + 1) || !cont(i + 2) || !cont(i + 3)) + return false; + std::uint32_t cp = (c & 0x07) << 18 | + (static_cast(s[i + 1]) & 0x3F) << 12 | + (static_cast(s[i + 2]) & 0x3F) << 6 | + (static_cast(s[i + 3]) & 0x3F); + if (cp < 0x10000 || cp > 0x10FFFF) + return false; + i += 4; + } else { + return false; + } + } + return true; +} + +std::string latin1_to_utf8(std::string_view s) { + std::string out; + out.reserve(s.size() + s.size() / 2); + for (char ch : s) + append_utf8(out, static_cast(ch)); + return out; +} + +std::string to_utf8_best_effort(std::string_view s) { + return is_valid_utf8(s) ? std::string(s) : latin1_to_utf8(s); +} + +std::string base64_decode(std::string_view in) { + auto val = [](char c) -> int { + if (c >= 'A' && c <= 'Z') + return c - 'A'; + if (c >= 'a' && c <= 'z') + return c - 'a' + 26; + if (c >= '0' && c <= '9') + return c - '0' + 52; + if (c == '+') + return 62; + if (c == '/') + return 63; + return -1; + }; + std::string out; + out.reserve(in.size() / 4 * 3 + 3); + std::uint32_t acc = 0; + int bits = 0; + for (char c : in) { + if (c == '=' || c == '\r' || c == '\n' || c == ' ' || c == '\t') + continue; + int v = val(c); + if (v < 0) + continue; // skip stray bytes + acc = (acc << 6) | static_cast(v); + bits += 6; + if (bits >= 8) { + bits -= 8; + out.push_back(static_cast((acc >> bits) & 0xFF)); + } + } + return out; +} + +std::string decode_rfc2047(std::string_view in, bool *had_encoded_word) { + if (had_encoded_word) + *had_encoded_word = false; + std::string out; + out.reserve(in.size()); + std::size_t i = 0; + const std::size_t n = in.size(); + + while (i < n) { + auto start = in.find("=?", i); + if (start == std::string_view::npos) { + out.append(in.substr(i)); + break; + } + out.append(in.substr(i, start - i)); + + // =?charset?enc?text?= + auto q1 = in.find('?', start + 2); + if (q1 == std::string_view::npos) { + out.append(in.substr(start)); + break; + } + auto q2 = in.find('?', q1 + 1); + if (q2 == std::string_view::npos || q2 != q1 + 2) { + out.append("=?"); + i = start + 2; + continue; + } + auto end = in.find("?=", q2 + 1); + if (end == std::string_view::npos) { + out.append(in.substr(start)); + break; + } + + std::string_view charset = in.substr(start + 2, q1 - (start + 2)); + char enc = in[q1 + 1]; + std::string_view text = in.substr(q2 + 1, end - (q2 + 1)); + + std::string bytes; + if (enc == 'B' || enc == 'b') { + bytes = base64_decode(text); + } else if (enc == 'Q' || enc == 'q') { + for (std::size_t k = 0; k < text.size(); ++k) { + char c = text[k]; + if (c == '_') { + bytes.push_back(' '); + } else if (c == '=' && k + 2 < text.size()) { + int hi = hex_val(text[k + 1]); + int lo = hex_val(text[k + 2]); + if (hi >= 0 && lo >= 0) { + bytes.push_back(static_cast((hi << 4) | lo)); + k += 2; + } else { + bytes.push_back(c); + } + } else { + bytes.push_back(c); + } + } + } else { + out.append(in.substr(start, end + 2 - start)); // unknown encoding: verbatim + i = end + 2; + continue; + } + + if (is_utf8(charset)) + out.append(to_utf8_best_effort(bytes)); + else if (is_latin1(charset)) + out.append(latin1_to_utf8(bytes)); + else + out.append(to_utf8_best_effort(bytes)); + + if (had_encoded_word) + *had_encoded_word = true; + i = end + 2; + + // RFC 2047: whitespace between adjacent encoded words is elided. + std::size_t j = i; + while (j < n && (in[j] == ' ' || in[j] == '\t')) + ++j; + if (j < n && in.compare(j, 2, "=?") == 0) + i = j; + } + return out; +} + +} // namespace vdm::net::detail diff --git a/core/src/net/text_codec.hpp b/core/src/net/text_codec.hpp new file mode 100644 index 0000000..42ba97a --- /dev/null +++ b/core/src/net/text_codec.hpp @@ -0,0 +1,40 @@ +// vdm/net/text_codec.hpp — internal: small byte/text codecs for header parsing. +// +// Not a public header. Everything here is pure, allocation-bounded, and total (no throw, +// no assert on input): the inputs come off the wire from untrusted servers. + +#ifndef VDM_NET_TEXT_CODEC_HPP +#define VDM_NET_TEXT_CODEC_HPP + +#include +#include +#include + +namespace vdm::net::detail { + +// Percent-decode ("%XX"). A stray '%' or a non-hex digit after it is emitted literally. +// `plus_as_space` handles application/x-www-form-urlencoded style; off for URL paths. +[[nodiscard]] std::string percent_decode(std::string_view in, bool plus_as_space = false); + +// True if `s` is well-formed UTF-8 (no overlong forms, no surrogates, no > U+10FFFF). +[[nodiscard]] bool is_valid_utf8(std::string_view s) noexcept; + +// Reinterpret each byte as a Latin-1 (ISO-8859-1) code point and re-encode as UTF-8. +[[nodiscard]] std::string latin1_to_utf8(std::string_view s); + +// Decode standard base64 (RFC 4648, '+' '/', optional '=' padding). Whitespace is +// skipped. Invalid trailing bits are dropped. Returns the decoded bytes. +[[nodiscard]] std::string base64_decode(std::string_view in); + +// Decode RFC 2047 "encoded-word" runs: =?charset?B?..?= / =?charset?Q?..?=. Text outside +// encoded words is passed through. Only UTF-8 and ISO-8859-1/Latin-1 charsets are +// transcoded; anything else is passed through as-is (best effort). `had_encoded_word` +// reports whether at least one well-formed word was found. +[[nodiscard]] std::string decode_rfc2047(std::string_view in, bool *had_encoded_word = nullptr); + +// If `s` is valid UTF-8, return it unchanged; otherwise treat it as Latin-1 and transcode. +[[nodiscard]] std::string to_utf8_best_effort(std::string_view s); + +} // namespace vdm::net::detail + +#endif // VDM_NET_TEXT_CODEC_HPP diff --git a/core/src/net/url.cpp b/core/src/net/url.cpp new file mode 100644 index 0000000..7054245 --- /dev/null +++ b/core/src/net/url.cpp @@ -0,0 +1,126 @@ +// vdm/net/url.cpp + +#include "vdm/net/url.hpp" + +#include + +#include "net/text_codec.hpp" + +namespace vdm::net { +namespace { + +std::string ascii_lower(std::string_view s) { + std::string r(s); + for (char &c : r) + if (c >= 'A' && c <= 'Z') + c = static_cast(c - 'A' + 'a'); + return r; +} + +bool valid_scheme(std::string_view s) { + if (s.empty()) + return false; + if (!((s[0] >= 'a' && s[0] <= 'z') || (s[0] >= 'A' && s[0] <= 'Z'))) + return false; + for (char c : s) { + bool ok = (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9') || + c == '+' || c == '-' || c == '.'; + if (!ok) + return false; + } + return true; +} + +} // namespace + +SplitUrl split_url(std::string_view url) { + SplitUrl out; + + auto scheme_end = url.find("://"); + if (scheme_end == std::string_view::npos) + return out; + std::string_view scheme = url.substr(0, scheme_end); + if (!valid_scheme(scheme)) + return out; + out.scheme = ascii_lower(scheme); + + std::string_view rest = url.substr(scheme_end + 3); + + // authority ends at the first '/', '?' or '#' + std::size_t auth_end = rest.size(); + for (std::size_t i = 0; i < rest.size(); ++i) { + char c = rest[i]; + if (c == '/' || c == '?' || c == '#') { + auth_end = i; + break; + } + } + std::string_view authority = rest.substr(0, auth_end); + std::string_view tail = rest.substr(auth_end); + + if (auto at = authority.rfind('@'); at != std::string_view::npos) { + out.userinfo.assign(authority.substr(0, at)); + authority = authority.substr(at + 1); + } + + std::string_view host = authority; + std::string_view port; + if (!authority.empty() && authority.front() == '[') { + auto close = authority.find(']'); + if (close != std::string_view::npos) { + host = authority.substr(1, close - 1); // strip brackets + if (close + 1 < authority.size() && authority[close + 1] == ':') + port = authority.substr(close + 2); + } + } else if (auto colon = authority.rfind(':'); colon != std::string_view::npos) { + host = authority.substr(0, colon); + port = authority.substr(colon + 1); + } + out.host = ascii_lower(host); + + if (!port.empty()) { + unsigned v = 0; + auto [p, ec] = std::from_chars(port.data(), port.data() + port.size(), v); + (void)p; + if (ec == std::errc{} && v > 0 && v <= 65535) + out.port = static_cast(v); + } + + // tail = path [ '?' query ] [ '#' fragment ] + std::string_view path_and_rest = tail; + if (auto hash = path_and_rest.find('#'); hash != std::string_view::npos) { + out.fragment.assign(path_and_rest.substr(hash + 1)); + path_and_rest = path_and_rest.substr(0, hash); + } + if (auto q = path_and_rest.find('?'); q != std::string_view::npos) { + out.query.assign(path_and_rest.substr(q + 1)); + path_and_rest = path_and_rest.substr(0, q); + } + out.path.assign(path_and_rest); + + out.valid = !out.host.empty() && out.is_http(); + return out; +} + +std::string url_filename(std::string_view url) { + SplitUrl u = split_url(url); + std::string_view path = u.path; + if (path.empty()) + return {}; + + auto slash = path.find_last_of('/'); + std::string_view seg = (slash == std::string_view::npos) ? path : path.substr(slash + 1); + if (seg.empty()) + return {}; + + std::string name = detail::percent_decode(seg); + // Guard against a decoded segment that reintroduces a separator or NUL. + for (char &c : name) + if (c == '/' || c == '\\' || c == '\0') + c = '_'; + if (name == "." || name == "..") + return {}; + return name; +} + +} // namespace vdm::net diff --git a/core/tests/CMakeLists.txt b/core/tests/CMakeLists.txt index 9f83bde..6a465ac 100644 --- a/core/tests/CMakeLists.txt +++ b/core/tests/CMakeLists.txt @@ -23,14 +23,21 @@ vdm_add_test(veloxcore_log_test util/log_test.cpp) # net/ integration tests drive tools/testserver (lane PKG/QA). Skip cleanly if it isn't # in the tree yet (lanes merge independently). +vdm_add_test(veloxcore_content_disposition_test net/content_disposition_test.cpp) +vdm_add_test(veloxcore_url_test net/url_test.cpp) + set(_testserver ${CMAKE_SOURCE_DIR}/tools/testserver/testserver.py) -vdm_add_test(veloxcore_http_client_test net/http_client_test.cpp) -target_include_directories(veloxcore_http_client_test PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/net) -if(EXISTS ${_testserver}) - target_compile_definitions(veloxcore_http_client_test - PRIVATE VDM_TESTSERVER_PY="${_testserver}") - set_tests_properties(veloxcore_http_client_test PROPERTIES TIMEOUT 120) -else() - message(STATUS "veloxcore: tools/testserver not present; http_client_test will skip " - "its server-backed cases.") +foreach(net_it http_client probe) + vdm_add_test(veloxcore_${net_it}_test net/${net_it}_test.cpp) + target_include_directories(veloxcore_${net_it}_test + PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/net) + if(EXISTS ${_testserver}) + target_compile_definitions(veloxcore_${net_it}_test + PRIVATE VDM_TESTSERVER_PY="${_testserver}") + set_tests_properties(veloxcore_${net_it}_test PROPERTIES TIMEOUT 120) + endif() +endforeach() +if(NOT EXISTS ${_testserver}) + message(STATUS "veloxcore: tools/testserver not present; net integration tests will " + "skip their server-backed cases.") endif() diff --git a/core/tests/net/content_disposition_test.cpp b/core/tests/net/content_disposition_test.cpp new file mode 100644 index 0000000..61057db --- /dev/null +++ b/core/tests/net/content_disposition_test.cpp @@ -0,0 +1,148 @@ +#include "vdm/net/content_disposition.hpp" + +#include +#include + +#include "vtest.hpp" + +using vdm::net::ContentDisposition; +using vdm::net::parse_content_disposition; +using Type = vdm::net::ContentDisposition::Type; + +namespace { +// "€" is U+20AC -> UTF-8 E2 82 AC. "£" is U+00A3 -> UTF-8 C2 A3. +const std::string kEuro = "\xE2\x82\xAC"; +const std::string kPound = "\xC2\xA3"; +const std::string kEAcute = "\xC3\xA9"; // é U+00E9 +} // namespace + +VT_TEST(cd_plain_quoted) { + auto cd = parse_content_disposition(R"(attachment; filename="report.pdf")"); + VT_CHECK(cd.type == Type::attachment); + VT_CHECK_EQ(cd.filename, std::string("report.pdf")); + VT_CHECK(!cd.filename_from_ext); +} + +VT_TEST(cd_plain_token_unquoted) { + auto cd = parse_content_disposition("attachment; filename=report.pdf"); + VT_CHECK_EQ(cd.filename, std::string("report.pdf")); +} + +VT_TEST(cd_inline_no_filename) { + auto cd = parse_content_disposition("inline"); + VT_CHECK(cd.type == Type::inline_); + VT_CHECK(!cd.has_filename()); +} + +VT_TEST(cd_rfc5987_utf8_ext_value) { + auto cd = parse_content_disposition("attachment; filename*=UTF-8''%e2%82%ac%20rates.pdf"); + VT_CHECK_EQ(cd.filename, kEuro + " rates.pdf"); + VT_CHECK(cd.filename_from_ext); +} + +VT_TEST(cd_prefers_ext_over_plain) { + auto cd = parse_content_disposition( + R"(attachment; filename="EURO rates.pdf"; filename*=UTF-8''%e2%82%ac%20rates.pdf)"); + VT_CHECK_EQ(cd.filename, kEuro + " rates.pdf"); + VT_CHECK(cd.filename_from_ext); +} + +VT_TEST(cd_rfc5987_latin1_ext_value) { + // %A3 = £ in ISO-8859-1 + auto cd = parse_content_disposition("attachment; filename*=ISO-8859-1''%A3rates.pdf"); + VT_CHECK_EQ(cd.filename, kPound + "rates.pdf"); +} + +VT_TEST(cd_legacy_rfc2047_base64) { + // base64(" rates.pdf") with euro as UTF-8 + // "€ rates.pdf" -> bytes E2 82 AC 20 72 61 74 65 73 2E 70 64 66 -> base64: + auto cd = + parse_content_disposition(R"(attachment; filename="=?UTF-8?B?4oKsIHJhdGVzLnBkZg==?=")"); + VT_CHECK_EQ(cd.filename, kEuro + " rates.pdf"); +} + +VT_TEST(cd_legacy_rfc2047_qencoded_latin1) { + // =?ISO-8859-1?Q?=A3rates.pdf?= -> £rates.pdf + auto cd = parse_content_disposition(R"(attachment; filename="=?ISO-8859-1?Q?=A3rates.pdf?=")"); + VT_CHECK_EQ(cd.filename, kPound + "rates.pdf"); +} + +VT_TEST(cd_raw_utf8_bytes_in_quotes) { + std::string h = "attachment; filename=\"caf" + kEAcute + ".txt\""; + auto cd = parse_content_disposition(h); + VT_CHECK_EQ(cd.filename, "caf" + kEAcute + ".txt"); +} + +VT_TEST(cd_raw_latin1_byte_in_quotes) { + // 0xE9 is 'é' in Latin-1; not valid UTF-8 alone -> transcoded + std::string h = "attachment; filename=\"caf\xE9.txt\""; + auto cd = parse_content_disposition(h); + VT_CHECK_EQ(cd.filename, "caf" + kEAcute + ".txt"); +} + +VT_TEST(cd_strips_path_components) { + VT_CHECK_EQ(parse_content_disposition(R"(attachment; filename="../../etc/passwd")").filename, + std::string("passwd")); + // real Windows paths in the wild use unescaped backslashes as separators + VT_CHECK_EQ(parse_content_disposition(R"(attachment; filename="C:\Windows\evil.exe")").filename, + std::string("evil.exe")); + // a base64 payload can contain '/', so decode must happen before path stripping + VT_CHECK_EQ(parse_content_disposition(R"(attachment; filename="=?UTF-8?B?Li4vLi4vc2VjcmV0?=")") + .filename, + std::string("secret")); // decodes to "../../secret", then stripped +} + +VT_TEST(cd_quoted_dquote_escape) { + // \" is the one escape we resolve, so a quote can appear mid-name + auto cd = parse_content_disposition(R"(attachment; filename="quote\"here.txt")"); + VT_CHECK_EQ(cd.filename, std::string("quote\"here.txt")); +} + +VT_TEST(cd_semicolon_inside_quotes_is_not_a_separator) { + auto cd = parse_content_disposition(R"(attachment; filename="a;b;c.txt")"); + VT_CHECK_EQ(cd.filename, std::string("a;b;c.txt")); +} + +VT_TEST(cd_form_data) { + auto cd = parse_content_disposition(R"(form-data; name="file"; filename="upload.bin")"); + VT_CHECK(cd.type == Type::form_data); + VT_CHECK_EQ(cd.filename, std::string("upload.bin")); +} + +VT_TEST(cd_rfc2231_continuations) { + auto cd = parse_content_disposition( + "attachment; filename*0*=UTF-8''%e2%82%ac; filename*1*=%20rates; filename*2=.pdf"); + VT_CHECK_EQ(cd.filename, kEuro + " rates.pdf"); +} + +VT_TEST(cd_empty_and_garbage_do_not_crash) { + VT_CHECK(!parse_content_disposition("").has_filename()); + VT_CHECK(!parse_content_disposition(";;;;").has_filename()); + VT_CHECK(!parse_content_disposition("attachment;").has_filename()); + VT_CHECK(!parse_content_disposition(R"(attachment; filename=)").has_filename()); + VT_CHECK(!parse_content_disposition(R"(attachment; filename="")").has_filename()); + // truncated ext-value + auto cd = parse_content_disposition("attachment; filename*=UTF-8''%e2%82"); + VT_CHECK(cd.type == Type::attachment); // no crash; filename is whatever fell out + // truncated encoded-word + auto trunc = parse_content_disposition(R"(attachment; filename="=?UTF-8?B?4oKs")"); + (void)trunc; +} + +VT_TEST(cd_bad_percent_escapes_in_ext_value) { + // stray % and non-hex digits are emitted literally, no crash + auto cd = parse_content_disposition("attachment; filename*=UTF-8''%ZZ%%file%2"); + VT_CHECK(cd.type == Type::attachment); +} + +VT_TEST(cd_case_insensitive_keys_and_type) { + auto cd = parse_content_disposition(R"(ATTACHMENT; FileName="x.txt")"); + VT_CHECK(cd.type == Type::attachment); + VT_CHECK_EQ(cd.filename, std::string("x.txt")); +} + +VT_TEST(cd_unknown_type_is_other) { + auto cd = parse_content_disposition(R"(signal; filename="x.txt")"); + VT_CHECK(cd.type == Type::other); + VT_CHECK_EQ(cd.filename, std::string("x.txt")); +} diff --git a/core/tests/net/probe_test.cpp b/core/tests/net/probe_test.cpp new file mode 100644 index 0000000..9b680b6 --- /dev/null +++ b/core/tests/net/probe_test.cpp @@ -0,0 +1,163 @@ +#include "vdm/net/probe.hpp" + +#include +#include +#include + +#include "testserver_fixture.hpp" +#include "vtest.hpp" + +using namespace vdm; +using namespace vdm::net; +using vdm::testing::TestServer; + +namespace { +Result run_probe(Prober &p, const std::string &url) { + std::promise> prom; + auto fut = prom.get_future(); + ProbeRequest req; + req.url = url; + req.overall_timeout_ms = 15000; + p.probe(std::move(req), [&](Result r) { prom.set_value(std::move(r)); }); + if (fut.wait_for(std::chrono::seconds(25)) != std::future_status::ready) + return Err{Error::timeout, "probe test wait"}; + return fut.get(); +} +const std::string kEuro = "\xE2\x82\xAC"; +} // namespace + +VT_TEST(probe_plain_resumable_file) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + auto r = run_probe(p, srv.url("/plain/file/2M")); + VT_REQUIRE(r.has_value()); + const auto &pr = r.value(); + VT_CHECK_EQ(pr.http_status, 206); // proven via the ranged GET + VT_CHECK(pr.accept_ranges); + VT_CHECK(pr.resumable); // 206 + ETag validator + VT_REQUIRE(pr.total_size.has_value()); + VT_CHECK_EQ(*pr.total_size, 2u * 1024 * 1024); + VT_CHECK(!pr.etag.empty()); + VT_CHECK_EQ(pr.filename_from_url, std::string("2M")); + VT_CHECK(!pr.requires_auth); +} + +VT_TEST(probe_no_range_server_is_not_resumable) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + auto r = run_probe(p, srv.url("/no-range/file/1M")); + VT_REQUIRE(r.has_value()); + const auto &pr = r.value(); + VT_CHECK(!pr.resumable); // no 206 ever + VT_CHECK(!pr.accept_ranges); + VT_REQUIRE(pr.total_size.has_value()); + VT_CHECK_EQ(*pr.total_size, 1u * 1024 * 1024); +} + +VT_TEST(probe_lying_accept_ranges_still_not_resumable) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + // advertises Accept-Ranges: bytes but never returns 206 + auto r = run_probe(p, srv.url("/lies-about-accept-ranges/file/1M")); + VT_REQUIRE(r.has_value()); + const auto &pr = r.value(); + VT_CHECK(pr.accept_ranges); // it advertised + VT_CHECK(!pr.resumable); // ...but never proved it -> resume is off +} + +VT_TEST(probe_reads_utf8_content_disposition) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + auto r = run_probe(p, srv.url("/utf8-content-disposition/file/8K")); + VT_REQUIRE(r.has_value()); + VT_CHECK_EQ(r.value().filename_from_disposition, kEuro + " rates.pdf"); + VT_CHECK_EQ(suggest_filename(r.value()), kEuro + " rates.pdf"); +} + +VT_TEST(probe_reads_legacy_content_disposition) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + auto r = run_probe(p, srv.url("/legacy-content-disposition/file/8K")); + VT_REQUIRE(r.has_value()); + VT_CHECK(!r.value().filename_from_disposition.empty()); // decoded, not the raw =?...?= + VT_CHECK(r.value().filename_from_disposition.find("=?") == std::string::npos); +} + +VT_TEST(probe_follows_redirect_and_reports_effective_url) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + auto r = run_probe(p, srv.url("/redirect-chain/file/64K")); + VT_REQUIRE(r.has_value()); + const auto &pr = r.value(); + VT_CHECK(pr.effective_url != srv.url("/redirect-chain/file/64K")); + VT_REQUIRE(pr.redirect_chain.size() >= 2); + VT_CHECK_EQ(pr.redirect_chain.front(), srv.url("/redirect-chain/file/64K")); + VT_CHECK_EQ(pr.redirect_chain.back(), pr.effective_url); +} + +VT_TEST(probe_401_is_requires_auth_not_error) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + auto r = run_probe(p, srv.url("/401-basic/file/64K")); + VT_REQUIRE(r.has_value()); // success result... + VT_CHECK(r.value().requires_auth); // ...flagged for the credential dialog +} + +VT_TEST(probe_404_is_an_error) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(4); + + auto r = run_probe(p, srv.url("/plain/nope")); + VT_REQUIRE(!r.has_value()); + VT_CHECK_EQ(r.error().code, Error::not_found); +} + +VT_TEST(probe_dead_host_is_connect_failed) { + Prober p(4); + std::promise> prom; + auto fut = prom.get_future(); + ProbeRequest req; + req.url = "http://127.0.0.1:1/x"; + req.connect_timeout_ms = 2000; + p.probe(std::move(req), [&](Result r) { prom.set_value(std::move(r)); }); + VT_REQUIRE(fut.wait_for(std::chrono::seconds(10)) == std::future_status::ready); + auto r = fut.get(); + VT_REQUIRE(!r.has_value()); + VT_CHECK_EQ(r.error().code, Error::connect_failed); +} + +VT_TEST(probe_pool_serialises_excess_requests) { + TestServer srv; + VT_REQUIRE(srv.available()); + Prober p(2); // only 2 at a time + + constexpr int kN = 8; + std::vector>> futs; + std::vector>> proms(kN); + for (int i = 0; i < kN; ++i) { + futs.push_back(proms[i].get_future()); + ProbeRequest req; + req.url = srv.url("/plain/file/4K"); + p.probe(std::move(req), + [pr = &proms[i]](Result r) { pr->set_value(std::move(r)); }); + } + for (auto &f : futs) { + VT_REQUIRE(f.wait_for(std::chrono::seconds(30)) == std::future_status::ready); + VT_CHECK(f.get().has_value()); + } +} diff --git a/core/tests/net/url_test.cpp b/core/tests/net/url_test.cpp new file mode 100644 index 0000000..4678184 --- /dev/null +++ b/core/tests/net/url_test.cpp @@ -0,0 +1,71 @@ +#include "vdm/net/url.hpp" + +#include + +#include "vtest.hpp" + +using vdm::net::split_url; +using vdm::net::SplitUrl; +using vdm::net::url_filename; + +VT_TEST(url_basic_https) { + auto u = split_url("https://example.com/path/to/file.iso?x=1#frag"); + VT_CHECK(u.valid); + VT_CHECK_EQ(u.scheme, std::string("https")); + VT_CHECK_EQ(u.host, std::string("example.com")); + VT_CHECK(!u.port.has_value()); + VT_CHECK_EQ(u.path, std::string("/path/to/file.iso")); + VT_CHECK_EQ(u.query, std::string("x=1")); + VT_CHECK_EQ(u.fragment, std::string("frag")); +} + +VT_TEST(url_port_userinfo_lowercasing) { + auto u = split_url("HTTP://User:pw@Host.EXAMPLE.com:8080/a"); + VT_CHECK_EQ(u.scheme, std::string("http")); + VT_CHECK_EQ(u.userinfo, std::string("User:pw")); + VT_CHECK_EQ(u.host, std::string("host.example.com")); + VT_REQUIRE(u.port.has_value()); + VT_CHECK_EQ(*u.port, 8080); +} + +VT_TEST(url_ipv6_host) { + auto u = split_url("http://[2001:db8::1]:9000/file"); + VT_CHECK_EQ(u.host, std::string("2001:db8::1")); + VT_REQUIRE(u.port.has_value()); + VT_CHECK_EQ(*u.port, 9000); +} + +VT_TEST(url_no_path) { + auto u = split_url("https://example.com"); + VT_CHECK(u.valid); + VT_CHECK_EQ(u.path, std::string("")); +} + +VT_TEST(url_non_http_is_invalid_but_parsed) { + auto u = split_url("ftp://host/file"); + VT_CHECK(!u.valid); // not http/https + VT_CHECK_EQ(u.scheme, std::string("ftp")); + VT_CHECK(!u.is_http()); +} + +VT_TEST(url_garbage_does_not_crash) { + VT_CHECK(!split_url("").valid); + VT_CHECK(!split_url("not a url").valid); + VT_CHECK(!split_url("://noscheme/x").valid); + VT_CHECK(!split_url("http://").valid); + auto a = split_url("http://////"); + auto b = split_url("https://h/%%%/%"); + (void)a; + (void)b; +} + +VT_TEST(url_filename_extraction) { + VT_CHECK_EQ(url_filename("https://x.com/a/b/report%20final.pdf"), + std::string("report final.pdf")); + VT_CHECK_EQ(url_filename("https://x.com/a/b/file.iso?sig=abc"), std::string("file.iso")); + VT_CHECK_EQ(url_filename("https://x.com/dir/"), std::string("")); + VT_CHECK_EQ(url_filename("https://x.com"), std::string("")); + VT_CHECK_EQ(url_filename("https://x.com/%2e%2e"), std::string("")); // ".." rejected + VT_CHECK_EQ(url_filename("https://x.com/a%2Fb"), + std::string("a_b")); // decoded '/' neutralised +} diff --git a/tools/fuzz/CMakeLists.txt b/tools/fuzz/CMakeLists.txt new file mode 100644 index 0000000..5f9f9d7 --- /dev/null +++ b/tools/fuzz/CMakeLists.txt @@ -0,0 +1,42 @@ +# libFuzzer targets for CORE parsers. Lane CORE owns tools/fuzz. +# +# Self-guarding: the top-level CMakeLists.txt add_subdirectory()s every tools/* that has a +# CMakeLists, unconditionally, so this file must opt out on its own when fuzzing isn't +# wanted or the compiler can't do libFuzzer. +# +# Each target compiles the parser sources directly (not the whole libveloxcore) so the +# code under test is fully fuzzer-instrumented without a second build of the library. + +if(NOT VELOX_BUILD_FUZZ) + return() +endif() +if(NOT CMAKE_CXX_COMPILER_ID MATCHES "Clang") + message(STATUS "tools/fuzz: libFuzzer needs Clang (have ${CMAKE_CXX_COMPILER_ID}); " + "skipping fuzz targets.") + return() +endif() + +set(_core ${CMAKE_SOURCE_DIR}/core) +set(_fuzz_flags -g -O1 -fsanitize=fuzzer,address,undefined -fno-omit-frame-pointer) + +function(vdm_add_fuzzer name) + add_executable(${name} ${ARGN}) + target_include_directories(${name} PRIVATE ${_core}/include ${_core}/src) + target_compile_features(${name} PRIVATE cxx_std_23) + target_compile_options(${name} PRIVATE ${_fuzz_flags}) + target_link_options(${name} PRIVATE ${_fuzz_flags}) +endfunction() + +vdm_add_fuzzer(fuzz_content_disposition + content_disposition_fuzz.cpp + ${_core}/src/net/content_disposition.cpp + ${_core}/src/net/text_codec.cpp) + +vdm_add_fuzzer(fuzz_url + url_fuzz.cpp + ${_core}/src/net/url.cpp + ${_core}/src/net/text_codec.cpp) + +# Seed corpora live next to the harnesses. +file(GLOB _cd_seeds ${CMAKE_CURRENT_SOURCE_DIR}/corpus/content_disposition/*) +file(GLOB _url_seeds ${CMAKE_CURRENT_SOURCE_DIR}/corpus/url/*) diff --git a/tools/fuzz/content_disposition_fuzz.cpp b/tools/fuzz/content_disposition_fuzz.cpp new file mode 100644 index 0000000..65e2cc3 --- /dev/null +++ b/tools/fuzz/content_disposition_fuzz.cpp @@ -0,0 +1,23 @@ +// Fuzz target for the Content-Disposition parser (AGENT-CORE §3: mojibake source, needs a +// fuzz target). The parser must be total on any input — no crash, no UB, bounded work. +// +// clang++ -std=c++23 -fsanitize=fuzzer,address,undefined ... (see CMakeLists.txt) +// ./fuzz_content_disposition -max_len=4096 corpus/content_disposition/ + +#include +#include +#include + +#include "vdm/net/content_disposition.hpp" + +extern "C" int LLVMFuzzerTestOneInput(const std::uint8_t *data, std::size_t size) { + std::string_view header(reinterpret_cast(data), size); + auto cd = vdm::net::parse_content_disposition(header); + + // Light invariants: a returned filename never contains a path separator or NUL. + for (char c : cd.filename) + if (c == '/' || c == '\\' || c == '\0') + __builtin_trap(); + + return 0; +} diff --git a/tools/fuzz/corpus/content_disposition/ext5987 b/tools/fuzz/corpus/content_disposition/ext5987 new file mode 100644 index 0000000..244462a --- /dev/null +++ b/tools/fuzz/corpus/content_disposition/ext5987 @@ -0,0 +1 @@ +attachment; filename*=UTF-8''%e2%82%ac%20rates.pdf \ No newline at end of file diff --git a/tools/fuzz/corpus/content_disposition/inline b/tools/fuzz/corpus/content_disposition/inline new file mode 100644 index 0000000..621650e --- /dev/null +++ b/tools/fuzz/corpus/content_disposition/inline @@ -0,0 +1 @@ +inline \ No newline at end of file diff --git a/tools/fuzz/corpus/content_disposition/plain b/tools/fuzz/corpus/content_disposition/plain new file mode 100644 index 0000000..2d3e658 --- /dev/null +++ b/tools/fuzz/corpus/content_disposition/plain @@ -0,0 +1 @@ +attachment; filename="report.pdf" \ No newline at end of file diff --git a/tools/fuzz/corpus/content_disposition/rfc2047 b/tools/fuzz/corpus/content_disposition/rfc2047 new file mode 100644 index 0000000..f968ec5 --- /dev/null +++ b/tools/fuzz/corpus/content_disposition/rfc2047 @@ -0,0 +1 @@ +attachment; filename="=?UTF-8?B?4oKsIHJhdGVzLnBkZg==?=" \ No newline at end of file diff --git a/tools/fuzz/corpus/content_disposition/traversal b/tools/fuzz/corpus/content_disposition/traversal new file mode 100644 index 0000000..57d8c93 --- /dev/null +++ b/tools/fuzz/corpus/content_disposition/traversal @@ -0,0 +1 @@ +attachment; filename="../../etc/passwd" \ No newline at end of file diff --git a/tools/fuzz/corpus/url/basic b/tools/fuzz/corpus/url/basic new file mode 100644 index 0000000..921da60 --- /dev/null +++ b/tools/fuzz/corpus/url/basic @@ -0,0 +1 @@ +https://example.com/path/to/file.iso?sig=abc#f \ No newline at end of file diff --git a/tools/fuzz/corpus/url/encoded b/tools/fuzz/corpus/url/encoded new file mode 100644 index 0000000..6cb07bf --- /dev/null +++ b/tools/fuzz/corpus/url/encoded @@ -0,0 +1 @@ +https://x.com/%2e%2e/%2f \ No newline at end of file diff --git a/tools/fuzz/corpus/url/ftp b/tools/fuzz/corpus/url/ftp new file mode 100644 index 0000000..9a6e5be --- /dev/null +++ b/tools/fuzz/corpus/url/ftp @@ -0,0 +1 @@ +ftp://host/x \ No newline at end of file diff --git a/tools/fuzz/corpus/url/ipv6 b/tools/fuzz/corpus/url/ipv6 new file mode 100644 index 0000000..b397676 --- /dev/null +++ b/tools/fuzz/corpus/url/ipv6 @@ -0,0 +1 @@ +http://[2001:db8::1]:9000/a%20b.bin \ No newline at end of file diff --git a/tools/fuzz/url_fuzz.cpp b/tools/fuzz/url_fuzz.cpp new file mode 100644 index 0000000..6e7f6f0 --- /dev/null +++ b/tools/fuzz/url_fuzz.cpp @@ -0,0 +1,24 @@ +// Fuzz target for the URL splitter (AGENT-CORE stage 3: "a fuzz target ... URL parsing"). +// split_url / url_filename must be total on any input. + +#include +#include +#include + +#include "vdm/net/url.hpp" + +extern "C" int LLVMFuzzerTestOneInput(const std::uint8_t *data, std::size_t size) { + std::string_view url(reinterpret_cast(data), size); + + auto u = vdm::net::split_url(url); + (void)u; + + std::string name = vdm::net::url_filename(url); + for (char c : name) + if (c == '/' || c == '\\' || c == '\0') + __builtin_trap(); + if (name == "." || name == "..") + __builtin_trap(); + + return 0; +} From c65e664c2913f6bcffa2c1192fee290372599a2e Mon Sep 17 00:00:00 2001 From: sami Date: Thu, 10 Sep 2026 00:02:02 +0400 Subject: [PATCH 2/2] core: sign off on ADR 0013 (task state-machine ownership) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Accept as written — no amendments. Answers to the four open items: 1. tasks_starved / starved_tasks() count ONLY tasks in {connecting, downloading} with segments_active == 0 (the allocator owes a slot to a task that is asking). retry_wait and paused (either-initiated) are outside that state set, so they are excluded by construction, not by a special case. Design commitment; the accessor is stage 6/8. 2. pause() is idempotent: no-op success on an already-paused or terminal task; resume() no-op success on a non-paused task; only task_not_found errors. No state-change event for a no-op. 3. PROTO's item — CORE confirms its half: auto-pause reports the transition with ErrorInfo populated (auth_required / server_file_changed / disk_full / path_rejected), already in util/error.hpp. Ready once PROTO permits error on state=="paused". 4. CORE adopts "auto-pause"; the wire/API discriminator stays state==paused + presence of the Error code. Notes back: pause during verifying re-hashes from scratch on resume; pause during assembling is M4; restart handling in §5 agreed — start(TaskId) re-derives resume position from .veloxpart.meta + If-Range. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01HPPSGhiArbvQgwC2DNiURS --- core/docs/adr-0013-core-response.md | 97 +++++++++++++++++++++++++++++ 1 file changed, 97 insertions(+) create mode 100644 core/docs/adr-0013-core-response.md diff --git a/core/docs/adr-0013-core-response.md b/core/docs/adr-0013-core-response.md new file mode 100644 index 0000000..791ac75 --- /dev/null +++ b/core/docs/adr-0013-core-response.md @@ -0,0 +1,97 @@ +# CORE response to ADR 0013 (task state-machine ownership) + +**Verdict: accept as written.** No amendments to the decision. `daemon/src/sched/`'s +pause/resume logic is unblocked from CORE's side once PROTO lands open item 3 (the +`error`-on-`paused` widening — see below). + +The split in §1, the shared-`paused` + idempotency contract in §2, the cross-reason resume +rule in §3, the `retry_wait ≠ starvation` resolution in §4, and the restart handling in §5 +all match what CORE agreed in `contracts/proto-answers-m1.md` D1 and the ADR 0011 response. + +Everything below is a **design commitment** — CORE's state machine is stage 8 and +`starved_tasks()` is stage 6/8; none of it is built yet. So "does CORE already do X" is +answered as "CORE will do X, and here is the spec it will be built to", not as a report on +existing code. + +--- + +## Answers to the four open items + +### 1 — `starved_tasks()` / `tasks_starved` exclude `retry_wait` and auto-paused, by construction + +Not a behavior change (nothing to change yet); a commitment baked into the accessor's +definition. Pinning ADR 0011's "running task with `segments_active == 0`": + +`tasks_starved` counts **only tasks whose task-state is `connecting` or `downloading` and +whose `segments_active == 0`.** That is precisely "the segment allocator has not granted a +slot to a task that is asking for one." Consequences: + +| Task state | In `tasks_starved`? | Why | +|---|---|---| +| `probing` | no | uses the probe pool (ADR 0011 §5), not the segment budget | +| `connecting`, `segments_active == 0` | **yes** | admitted + probed, waiting on the allocator's first grant — the real starvation case | +| `connecting`/`downloading`, `segments_active ≥ 1` | no | a segment in its own `connecting` sub-state counts as a held slot (ADR 0011 A3) | +| `downloading`, `segments_active == 0` | **yes** | held slots, lost them all (e.g. every segment failed and is being re-requested) — transient, still a real "allocator owes this task a slot" | +| `retry_wait` | no | CORE's backoff timer holds it at zero *deliberately*; it is not asking the allocator for anything until it re-enters `connecting` | +| `paused` (DAEMON- **or** CORE-initiated) | no | a paused task is not asking for a slot | +| `new`, `queued`, `assembling`, `verifying`, `complete`, `failed`, `cancelled` | no | not in the `{connecting, downloading}` set | + +So `retry_wait` and auto-paused tasks fall outside `tasks_starved` **because they are not +in the counted state set**, not because of a special case that could rot. `starved_tasks()` +returns exactly the TaskIds in that count; `starved_since(id)` is defined only for them. + +A `retry_wait` task still counts against `connection.maxConcurrentDownloads` from DAEMON's +side (it is running, not requeued) and contributes nothing to the allocator's guarantee +pass until it re-enters `connecting` — matches §4 exactly. + +### 2 — `pause()` is idempotent (committed) + +- `pause(TaskId)` on a task already in `paused` (however it got there) → **no-op, returns + success**. DAEMON never needs to know CORE auto-paused first. +- `pause(TaskId)` on a task in a terminal state (`complete`/`failed`/`cancelled`) → + **no-op, returns success** — a pause racing a completion is not an error. +- `resume(TaskId)` on a task that is not paused → **no-op, returns success**. +- The only error `pause()`/`resume()` return is `task_not_found`. +- CORE does not report a state-change event for a no-op pause (nothing changed). DAEMON + reads the resulting state via the normal state-change callback / `TaskDetail`, not from + `pause()`'s return. +- If CORE is mid-transition to `paused` (its own auto-pause) when DAEMON's `pause()` + arrives, the task ends up `paused` once, with one state-change event. + +### 3 — PROTO: `error`-on-`paused` widening + +Not CORE's to land, but CORE confirms its half: when CORE auto-pauses, it reports the +transition through the same state-change callback every CORE transition uses, with the +`ErrorInfo` populated — `auth_required` (401/407), `server_file_changed`, `disk_full`, +`path_rejected` for a mid-run destination failure. Those are the `vdm::Error` values from +the B1 taxonomy (`core/docs/proto-requests-m1.md`), already implemented in +`core/include/vdm/util/error.hpp`. **CORE is ready**; PROTO only needs to permit `error` to +be present on `event.task.state` when `state == "paused"`, and DAEMON to project CORE's +`ErrorInfo` onto it. `error: null` on a DAEMON-initiated pause is correct and sufficient. + +Until PROTO lands it, DAEMON cannot implement §3's cross-reason resume rule +correctness-preservingly — agree it should not guess from timing. + +### 4 — "auto-pause" naming + +CORE has no established internal term (the state machine is unbuilt). **CORE adopts +"auto-pause"** for the informal concept. On the wire and in the API there is no new term: +the discriminator is `state == paused` plus the `Error` code (present ⇒ CORE-initiated, +absent ⇒ DAEMON-initiated), exactly as the ADR's §2 and the rejected-alternatives section +describe. + +--- + +## Two notes back to DAEMON (not objections) + +- **§1 table, `paused` from `assembling`/`verifying`:** the *state transition* is honoured + from any CORE state as the table says. The *work* interruption is best-effort: a pause + during `verifying` discards the in-progress hash and re-hashes from the start of the + file on resume (cheap, bounded); a pause during `assembling` (HLS/DASH mux) is an M4 + concern and may not be cleanly interruptible mid-mux. Neither affects M1 or the API + shape. +- **§5 restart:** CORE agrees fully. CORE holds no persistent state; `start(TaskId)` + transparently checks for a valid `.veloxpart.meta` sidecar (stage 5), re-validates with + `If-Range` (`docs/04` §5), and either resumes from the recorded offsets or restarts if + the validator failed / the sidecar is corrupt. DAEMON does nothing special on restart + beyond rewriting CORE-owned states to `queued` and re-admitting — which is what §5 says.