Skip to content

fix(daemon): close stores in one teardown, drop capacity retry - #2751

Merged
ScriptedAlchemy merged 1 commit into
masterfrom
fleet/teardown-lease-release
Sep 30, 2026
Merged

ScriptedAlchemy merged 1 commit into
masterfrom
fleet/teardown-lease-release

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

Root cause

#2476 (lease outlives teardown). The in-process production harness forked the daemon's terminal store close. Its shutdown skipped five steps:

  • the store-telemetry sampler's retained handle release (release_retained_handles_for_shutdown);
  • the Git-transaction store-actor join;
  • the native-integration store-actor join;
  • the profile session refresh services release;
  • the retirement-reaper join.

So every store kept a client lease past teardown. The first shutdown logged store runtime still held at shutdown … blockers=[ClientLeases { count: 1 }] for global.db, user-sessions.db and tracedecay.db, and { count: 2 } for sessions.db. No WAL was truncated, and a same-process reopen raced the old owners, which is the delivery_settlement_recorder_already_running path in #2476. The portable daemon shutdown had the same fork and also missed the telemetry release.

#2585 retry. #2585 retried a capacity-retirement store release that ClientLeases refused. The owner-scoped blocker it was built for was the automation scheduler's session lease, and #2516 already moved that into teardown. Capacity teardown now joins every owner-scoped store client. A 12-project journey where every project uses its session stores releases each retired owner on its first attempt. So the retry, the stored release closure and the per-open retry pass only covered a lease held outside the owner.

Change

  • One terminal store close. StoreAdministration::close_stores_for_shutdown runs the ordered sequence: close admission, cancel and join reconciliation, release telemetry handles, drain retained owners, close idle stores. The Unix engine, the portable daemon (bootstrap.rs) and the production harness all call it; the three forks are deleted.
  • Harness joins the daemon's owners. It now also joins the Git-transaction and native-integration actors, the profile session refresh services, and the retirement reapers (Unix, as in the engine).
  • Capacity retirement runs once. Teardown and store release run in one tracked task (spawn_and_track_fallible(owner, retirement)).
  • Retry deleted. Removed CapacityRetirementRelease, capacity_release, failed_capacity_release, retry_failed_capacity_releases in both files, the per-open retry call, the Clone on CapacityRetirementStores, and the two retry tracker tests.
  • Foreign leases. A lease held outside the owner still gets the typed, retryable project_server_capacity_reached refusal that names the store blocker, exactly once. The owner's next retirement supersedes the failed receipt. prior_completions_for_owner waits only on pending receipts, so a reported failure is never replayed (the daemon: a failed capacity retirement refuses every later project open until restart #2547 symptom).

Evidence

Fail-before

New test: daemon::production_harness::project_server_capacity_journey_test::shut_down_composition_releases_its_session_stores_for_an_immediate_reopen. It mounts the session stores through tracedecay_hook_runtime ingest_transcript, shuts the harness down, requires every store WAL to be truncated, then reopens and requires the session stores to mount again. On master production sources:

assertion `left == right` failed: shutdown releases every store lease, so each writer truncates its WAL
  left: [("global.db", 3007632), ("user-sessions.db", 2970552), ("sessions.db", 3650352), ("tracedecay.db", 753992)]
 right: [("global.db", 0), ("user-sessions.db", 0), ("sessions.db", 0), ("tracedecay.db", 0)]
test result: FAILED. 0 passed; 1 failed

On this branch it passes, with no still held at shutdown lines.

Retry deletion

  • capacity_retirement_releases_session_stores_its_owner_used_on_the_first_attempt passes with the retry gone. It opens 12 projects, each ingesting through its session stores, past the 8-server cache, and every retirement releases on its only attempt.
  • capacity_retirement_blocked_by_a_foreign_store_lease_is_typed_and_not_replayed replaces fix(daemon): retry a refused capacity retirement #2585's journey. A foreign mounted_project_sessions lease gets the typed refusal naming ClientLeases; every later open serves while the lease is still held; the retired project reopens.
  • The tracker test failed_capacity_retirement_is_reported_once_and_superseded_by_the_next asserts the literal refusal text, prior_completions_for_owner(..).len() == 0 after a failure, and supersession.

Focused runs (non-zero counts)

  • cargo test -p tracedecay --lib --features test-helpers -- daemon::branch_admin daemon::production_harness daemon::project_composition daemon::tests::lifecycle: 72 passed.
  • cargo test -p tracedecay --features test-helpers --test daemon_suite -- store_shutdown_checkpoint_test project_capacity_reuse: 4 passed. This is the real engine shutdown through the shared close, plus external capacity reuse.
  • cargo test -p tracedecay --features test-helpers,test-transport --test mcp_suite -- mcp_handler_test::work_test mcp_handler_test::index_path_settings_test: 4 passed (the existing harness reopen() users).

Lint and cross-check

  • cargo clippy -p tracedecay --all-targets --features test-helpers,test-transport -- -D warnings: clean.
  • cargo fmt --all -- --check: clean.
  • cargo check --workspace --all-targets --target x86_64-pc-windows-gnu --features tracedecay/test-transport,tracedecay/test-helpers,tracedecay-cli/test-transport: clean.
  • ripwire --quality-delta=<merge-base>..HEAD gates on one pair, close_stores_for_shutdown against SessionStoreAccess::set_parse_offset. The two only share the .await.map_err(|error| error.to_string())? shape; no logic or helper is shared. The dead-code rows are #[tokio::test] functions.
  • ripwire --edit-check on close_stores_for_shutdown and spawn_and_track_fallible: 0 incompatible callers.

Built-CLI journey

Setup: debug CLI from this branch, isolated HOME/XDG profile, one daemon under systemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G. Ten git projects are initialized one after another, and each ingests through its session stores. That is past the 8-server cache, so capacity retires idle owners. Then projects 1–3 are revisited, and the daemon gets SIGTERM by exact PID.

proj01 init rc=0
proj01 ingest rc=0 status:** accepted_for_replay
…
proj10 init rc=0
proj10 ingest rc=0 status:** accepted_for_replay
revisit proj01 rc=0 count:** 1      (first call answered typed retryable code-graph-unavailable while the retired owner re-warmed)
revisit proj02 rc=0 count:** 1
revisit proj03 rc=0 count:** 1
proj01..proj03 full_published=2 each (retired under capacity, then reopened); proj04 full_published=1
daemon log lines matching `still held at shutdown|retirement failed|capacity_reached|is blocked`: 0
after graceful stop:
      1 global.db wal=0
     10 projects/<id>/sessions.db wal=0
     10 projects/<id>/tracedecay.db wal=0
      1 user-sessions.db wal=0

Fixes #2476

Harness shutdown forked the daemon's terminal store close and skipped the
store-telemetry handle release plus the Git-transaction, native-integration,
profile-refresh and retirement-reaper joins, so every store kept a client
lease past teardown, its WAL was never truncated and a same-process reopen
raced the old owners. Daemon (Unix and portable) and harness shutdown now
share StoreAdministration::close_stores_for_shutdown, and the harness joins
the same owners the daemon does.

Capacity retirement teardown already joins every owner-scoped store client,
so its store release runs once inside the retirement task. The stored
release closure and the per-open retry pass are deleted; a foreign lease
still gets the typed capacity refusal once, and the owner's next retirement
supersedes the failed receipt instead of replaying it.

Fixes #2476
@changeset-bot

changeset-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: d972f1b

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@ScriptedAlchemy
ScriptedAlchemy merged commit 54fb091 into master Sep 30, 2026
4 of 7 checks passed
@ScriptedAlchemy
ScriptedAlchemy deleted the fleet/teardown-lease-release branch September 30, 2026 06:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Harness reopen never re-mounts session stores: delivery recorder spool lease outlives composition shutdown

1 participant