fix(daemon): close stores in one teardown, drop capacity retry - #2751
Merged
Merged
Conversation
Harness shutdown forked the daemon's terminal store close and skipped the store-telemetry handle release plus the Git-transaction, native-integration, profile-refresh and retirement-reaper joins, so every store kept a client lease past teardown, its WAL was never truncated and a same-process reopen raced the old owners. Daemon (Unix and portable) and harness shutdown now share StoreAdministration::close_stores_for_shutdown, and the harness joins the same owners the daemon does. Capacity retirement teardown already joins every owner-scoped store client, so its store release runs once inside the retirement task. The stored release closure and the per-open retry pass are deleted; a foreign lease still gets the typed capacity refusal once, and the owner's next retirement supersedes the failed receipt instead of replaying it. Fixes #2476
|
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root cause
#2476 (lease outlives teardown). The in-process production harness forked the daemon's terminal store close. Its shutdown skipped five steps:
release_retained_handles_for_shutdown);So every store kept a client lease past teardown. The first shutdown logged
store runtime still held at shutdown … blockers=[ClientLeases { count: 1 }]for global.db, user-sessions.db and tracedecay.db, and{ count: 2 }for sessions.db. No WAL was truncated, and a same-process reopen raced the old owners, which is thedelivery_settlement_recorder_already_runningpath in #2476. The portable daemon shutdown had the same fork and also missed the telemetry release.#2585 retry. #2585 retried a capacity-retirement store release that
ClientLeasesrefused. The owner-scoped blocker it was built for was the automation scheduler's session lease, and #2516 already moved that into teardown. Capacity teardown now joins every owner-scoped store client. A 12-project journey where every project uses its session stores releases each retired owner on its first attempt. So the retry, the stored release closure and the per-open retry pass only covered a lease held outside the owner.Change
StoreAdministration::close_stores_for_shutdownruns the ordered sequence: close admission, cancel and join reconciliation, release telemetry handles, drain retained owners, close idle stores. The Unix engine, the portable daemon (bootstrap.rs) and the production harness all call it; the three forks are deleted.spawn_and_track_fallible(owner, retirement)).CapacityRetirementRelease,capacity_release,failed_capacity_release,retry_failed_capacity_releasesin both files, the per-open retry call, theCloneonCapacityRetirementStores, and the two retry tracker tests.project_server_capacity_reachedrefusal that names the store blocker, exactly once. The owner's next retirement supersedes the failed receipt.prior_completions_for_ownerwaits only on pending receipts, so a reported failure is never replayed (the daemon: a failed capacity retirement refuses every later project open until restart #2547 symptom).Evidence
Fail-before
New test:
daemon::production_harness::project_server_capacity_journey_test::shut_down_composition_releases_its_session_stores_for_an_immediate_reopen. It mounts the session stores throughtracedecay_hook_runtime ingest_transcript, shuts the harness down, requires every store WAL to be truncated, then reopens and requires the session stores to mount again. On master production sources:On this branch it passes, with no
still held at shutdownlines.Retry deletion
capacity_retirement_releases_session_stores_its_owner_used_on_the_first_attemptpasses with the retry gone. It opens 12 projects, each ingesting through its session stores, past the 8-server cache, and every retirement releases on its only attempt.capacity_retirement_blocked_by_a_foreign_store_lease_is_typed_and_not_replayedreplaces fix(daemon): retry a refused capacity retirement #2585's journey. A foreignmounted_project_sessionslease gets the typed refusal namingClientLeases; every later open serves while the lease is still held; the retired project reopens.failed_capacity_retirement_is_reported_once_and_superseded_by_the_nextasserts the literal refusal text,prior_completions_for_owner(..).len() == 0after a failure, and supersession.Focused runs (non-zero counts)
cargo test -p tracedecay --lib --features test-helpers -- daemon::branch_admin daemon::production_harness daemon::project_composition daemon::tests::lifecycle: 72 passed.cargo test -p tracedecay --features test-helpers --test daemon_suite -- store_shutdown_checkpoint_test project_capacity_reuse: 4 passed. This is the real engine shutdown through the shared close, plus external capacity reuse.cargo test -p tracedecay --features test-helpers,test-transport --test mcp_suite -- mcp_handler_test::work_test mcp_handler_test::index_path_settings_test: 4 passed (the existing harnessreopen()users).Lint and cross-check
cargo clippy -p tracedecay --all-targets --features test-helpers,test-transport -- -D warnings: clean.cargo fmt --all -- --check: clean.cargo check --workspace --all-targets --target x86_64-pc-windows-gnu --features tracedecay/test-transport,tracedecay/test-helpers,tracedecay-cli/test-transport: clean.ripwire --quality-delta=<merge-base>..HEADgates on one pair,close_stores_for_shutdownagainstSessionStoreAccess::set_parse_offset. The two only share the.await.map_err(|error| error.to_string())?shape; no logic or helper is shared. The dead-code rows are#[tokio::test]functions.ripwire --edit-checkonclose_stores_for_shutdownandspawn_and_track_fallible: 0 incompatible callers.Built-CLI journey
Setup: debug CLI from this branch, isolated HOME/XDG profile, one daemon under
systemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G. Ten git projects are initialized one after another, and each ingests through its session stores. That is past the 8-server cache, so capacity retires idle owners. Then projects 1–3 are revisited, and the daemon gets SIGTERM by exact PID.Fixes #2476