daemon: fix startMode 'later' -32603 and isolate the single-instance lock by runtime dir

Two bugs blocking PROTO's live-veloxd conformance check.

1. tasks.start_mode's CHECK was ('auto','now','queue','manual') — not the
   contract's StartMode enum (['now','later','queue']) at all. 'later', a real
   documented value (the File Info dialog's Download Later button), hit the CHECK
   on every insert and surfaced as an unhandled -32603; 'auto'/'manual' were
   never contract values to begin with.

   Migration 0003 rebuilds tasks (SQLite can't ALTER a CHECK) with the contract's
   values, remapping existing rows by what they actually meant: 'auto' -> 'now'
   (eligible for the scheduler immediately), 'manual' -> 'later' (parked, matching
   StartMode's own "lands the task in paused" description). store_migrations_test
   covers the remap and that 'later' inserts clean while the retired spellings
   are rejected.

   dispatcher.cpp's on_download_add matched: default (absent startMode) is now
   'now' instead of the invented 'auto'; 'later' actually lands the task in
   `paused` (pause_reason 'user') instead of a dead 'manual' -> `new` branch that
   spec.startMode (typed as the 3-value enum) could never even reach.
   TaskRow::start_mode's in-memory default followed suit ('now').

2. main.cpp's single-instance guard bound an abstract socket named
   "velox-daemon-<euid>" — one name per user, system-wide. XDG_RUNTIME_DIR
   isolation never reached it: a leaked test veloxd held the lock for 4h40m and
   locked out every other isolated instance with the same euid (PROTO, EXT, the
   orchestrator), real daemon included.

   Extracted rpc/single_instance.{hpp,cpp} (was a static in main.cpp, untestable)
   and derived the abstract-socket name from a hash of the resolved runtime dir
   path instead of euid alone. The real per-user daemon is still unique (its
   runtime dir is unique to it); isolated instances pointed at their own runtime
   dirs now coexist. main() resolves the runtime dir before acquiring the lock
   (was the other way around). New single_instance_test covers same-dir refusal,
   different-dir coexistence, and release-on-close.

Verified against real veloxd binaries, not just unit tests: startMode: "later"
via a live download.add lands in `paused`; two veloxd with different runtime
dirs run concurrently, two with the same one and the second refuses with the
runtime dir named in the error. Full ctest: 39/39.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GRDjHGgpYmMoPE2UFbe7pP
This commit is contained in:
2026-09-11 17:04:02 +04:00
co-authored by Claude Sonnet 5
parent a967eca669
commit de748cc2fc
10 changed files with 270 additions and 35 deletions
+1
View File
@@ -23,3 +23,4 @@ veloxd_test(store_tasks LIBS veloxd_store)
veloxd_test(sched_scheduler LIBS veloxd_sched veloxd_rpc)
veloxd_test(event_hub LIBS veloxd_rpc)
veloxd_test(store_categories_queues LIBS veloxd_store)
veloxd_test(single_instance LIBS veloxd_rpc)
+40
View File
@@ -0,0 +1,40 @@
// The single-instance lock: isolated by runtime dir, not just by euid. This is the bug a
// leaked test veloxd exploited — one abstract-socket name per user meant every isolated
// instance (real daemon, tests, other lanes) fought over the same lock.
#include <unistd.h>
#include "check.hpp"
#include "rpc/single_instance.hpp"
using namespace velox::daemon::rpc;
void run() {
// Two different runtime dirs: both acquire the lock independently.
{
const int a = acquire_single_instance_lock("/run/user/1000/velox-test-a");
const int b = acquire_single_instance_lock("/run/user/1000/velox-test-b");
CHECK(a >= 0);
CHECK(b >= 0);
if (a >= 0) ::close(a);
if (b >= 0) ::close(b);
}
// Same runtime dir: the second attempt is refused while the first still holds it.
{
const int first = acquire_single_instance_lock("/run/user/1000/velox-test-shared");
CHECK(first >= 0);
const int second = acquire_single_instance_lock("/run/user/1000/velox-test-shared");
CHECK(second < 0);
if (first >= 0) ::close(first);
if (second >= 0) ::close(second);
// Releasing (closing) the fd frees the abstract-namespace name immediately — a
// third attempt at the same dir succeeds once the first is gone.
const int third = acquire_single_instance_lock("/run/user/1000/velox-test-shared");
CHECK(third >= 0);
if (third >= 0) ::close(third);
}
}
TEST_MAIN()
+50
View File
@@ -86,6 +86,56 @@ void run() {
.has_value());
}
// --- 0003: start_mode is rebuilt to the contract's values, existing rows mapped -
{
auto db = Db::open(":memory:");
CHECK(db.has_value());
if (!db) return;
// Build a real pre-0003 db (schema 1..2) with rows in the old, non-contract
// start_mode spelling, the way an actually-released daemon would have them.
for (const auto& m : embedded_migrations()) {
if (m.version > 2) break;
CHECK(db->exec(m.sql).has_value());
CHECK(db->set_user_version(m.version).has_value());
}
CHECK(db->exec("INSERT INTO tasks(task_id,url,save_dir,created_at,start_mode) "
"VALUES('auto1','http://x','/tmp','2026-09-10T00:00:00Z','auto')")
.has_value());
CHECK(db->exec("INSERT INTO tasks(task_id,url,save_dir,created_at,start_mode) "
"VALUES('man1','http://x','/tmp','2026-09-10T00:00:00Z','manual')")
.has_value());
CHECK(db->exec("INSERT INTO tasks(task_id,url,save_dir,created_at,start_mode) "
"VALUES('q1','http://x','/tmp','2026-09-10T00:00:00Z','queue')")
.has_value());
CHECK(migrate_to_head(*db).has_value());
CHECK_EQ(db->user_version(), head);
auto start_mode_of = [&](const char* id) -> std::string {
auto st = db->prepare("SELECT start_mode FROM tasks WHERE task_id=?1");
if (!st || !st->bind(1, std::string_view(id))) return "";
auto row = st->step();
if (!row || !*row) return "";
return std::string(st->column_text(0));
};
CHECK_EQ(start_mode_of("auto1"), std::string("now"));
CHECK_EQ(start_mode_of("man1"), std::string("later"));
CHECK_EQ(start_mode_of("q1"), std::string("queue")); // passes through unchanged
// The bug this migration closes: 'later' — a real, documented StartMode value —
// used to hit the old CHECK and fail every insert. It's accepted now, and the two
// retired spellings are gone for good.
CHECK(db->exec("INSERT INTO tasks(task_id,url,save_dir,created_at,start_mode) "
"VALUES('later1','http://x','/tmp','2026-09-10T00:00:00Z','later')")
.has_value());
CHECK(db->exec("INSERT INTO tasks(task_id,url,save_dir,created_at,start_mode) "
"VALUES('bad1','http://x','/tmp','2026-09-10T00:00:00Z','auto')")
.has_value() == false);
CHECK(db->exec("INSERT INTO tasks(task_id,url,save_dir,created_at,start_mode) "
"VALUES('bad2','http://x','/tmp','2026-09-10T00:00:00Z','manual')")
.has_value() == false);
}
// --- re-running the migrator on an at-head DB is a no-op ------------------------
{
auto db = Db::open(":memory:");