Skip to content

Separate integration checks from full release verification - #808

Merged
justin808 merged 10 commits into
mainfrom
jg-codex/speed-delivery-policy
Sep 10, 2026
Merged

justin808 merged 10 commits into
mainfrom
jg-codex/speed-delivery-policy

Conversation

@justin808

@justin808 justin808 commented Sep 9, 2026 •

Copy link
Copy Markdown
Member

Why

Repositories can already customize their validation commands, but verification, review, and closeout need a consistent way to distinguish selected integration checks from complete promotion evidence. This makes those choices explicit while preserving required failures, trusted policy, and separate release authority.

Closes #514. Closes #787.

What changed

  • Define repository-owned integration and promotion checks, with coverage reports consumed by verify, run-ci, autoreview preparation, and close-session.
  • Add executable low-impact-tool and critical-service examples: the same overview edit runs selected checks in one and the complete gate in the other.
  • Block both phases when HEAD, index, or working validator policy differs from the executed trusted policy; an older checklist cannot establish completeness for newly required checks.
  • Replay trusted-base selection, risky and unknown changes, legacy behavior, promotion, persistent required failures, and candidate changes during validation.
  • Require execution through a trusted selection entry point. The teaching example recognizes only two reviewed overview contents; arbitrary operational English and invalid bytes retain full coverage.
  • Use the captured candidate head for selection and reporting. Compare tracked bytes, Git executable modes, and symlink targets with HEAD independently of Git status hints; hidden edits and check-induced mutations cannot qualify clean promotion.
  • Explicitly block submodules and unsupported candidate entries in the examples; these require adopter-owned capture. Preserve required results for every repository and coverage fields only when its wrapper reports them.

This change includes the delivery contract and executable replay fixtures. The design could have shipped separately from its executable adoption examples; combining them makes the documented contract verifiable in the same PR. New YAML settings, a shared resolver, executable fast/balanced/strict presets, and additional effort counters remain deferred.

How to review and verify

Start with docs/delivery-policy.md, then compare the two example wrappers and their replay tests. Candidate policy edits must not reduce their own checks; selected success must not qualify omitted promotion checks. Existing repositories retain their current behavior unless they explicitly customize their wrappers.

Test plan

  • Delivery-policy replay: 13 tests and 779 assertions, including candidate-added required checks, hidden staged/committed validator changes in both examples/phases, and the concurrent-code-commit race.
  • Independent feature, corrective, and prior integration reviews passed.
  • Current policy-boundary correction passed full precommit validation, including installer and stack suites.
  • Current clean committed local full validation, including installer and stack suites: PASS.
  • Current-head full hosted validation and completed review cohort: PASS; all 46 threads resolved.
  • Changelog: deferred_to_update_changelog.
Agent details

Candidate and verification evidence

The current Codex review completed with no new findings. Claude published a substantive review and two optional consolidation suggestions; both were dispositioned without further source changes. Its launcher reported seven permission denials, so launcher success alone is not treated as review evidence. All 46 review threads are resolved.

Clean committed local full validation passed. Current-head full hosted validation passed, including installer and stack suites, on tested integration 479aa94e2061408a7697b6d4a5142ffa1e1974a7 against base 48ed4f81b725c63808d78e99e12d05c920e6223a. An earlier candidate failed an unchanged batch-status timing assertion in run34426073347. The cause remains unknown and is recorded in #260; this policy correction does not claim to fix that failure.

Candidate fc2c5d61c72a964394dfd63478c6d19d9a768f2c integrates main 48ed4f81b725c63808d78e99e12d05c920e6223a. Selection/reporting use the captured candidate head. Both wrappers verify captured HEAD and working validator source identity against the trusted wrapper, and reject staged validator changes before reporting coverage. The added replay covers staged and committed validator changes hidden by restoring trusted working bytes, across both examples and phases. Independent review checked the complete HEAD/index/worktree boundary; dirty promotion already failed before this correction, so no release bypass is claimed.

The prior integration proof preserves the actual base. This candidate adds only the independently reviewed three-file policy-identity correction and replay; all other feature bytes and current-main CHANGELOG.md are preserved. The preceding full hosted run passed on a2c52cae0ed2659ac47a0c0c1bd335485c61f209, tested integration 7b76b63fcd06b187cbe86b9528ba3ccd5553ad97, and the same base. It is historical evidence, not qualification of this corrected candidate.

The approved 3eabd95d9cf7fc9d95ff663fb12288f746fc828c candidate passed clean local full validation and full hosted validation. Its independent local integration review was clean. The completed hosted Codex review and resolved prior findings remain recorded; the Claude launcher reported nine permission denials and no published review artifact, so its success was not counted as a clean review. These are historical results; the current candidate's gates are listed above.

The regressions reproduce hidden tracked edits, executable-mode changes, candidate mutations, dirty gitlinks, unsupported nested entries, and operational prose that must retain full coverage. The teaching wrappers compare complete candidate bytes, with a per-file process cost. They are examples; no production adoption or measured performance improvement is claimed.

QA Evidence

  • QA lane: independent review and executable behavioral replay.
  • Scope checked: repository-owned integration and promotion policy, reporting, example wrappers, and failure preservation.
  • Tested at: fc2c5d61c72a964394dfd63478c6d19d9a768f2c; clean committed local full validation passed.
  • Automated checks: reviewed feature replays and lint passed; clean committed local full validation passed; current-head full hosted validation passed; completed review cohort has no unresolved findings.
  • Manual checks: independent complete example identity/selection-path and reporting-contract review; exact correction/current-base integration proof.
  • User-visible UI change: no.
  • Visual evidence: not applicable: workflow instructions and command-line examples.
  • Interaction change: no application interaction change.
  • Interaction evidence: not applicable: no application interaction.
  • Visual fix: no.
  • Negative control: not applicable: no visual fix.
  • Performance evidence: not applicable: no measured performance claim.
  • Findings: prior confirmed candidate-capture and operational-prose findings fixed; candidate-added-check omission, mismatched selection/capture heads, and hidden HEAD/index policy changes now block.
  • QA required: yes.
  • QA required rationale: verification choices must preserve required coverage, failures, and promotion boundaries.
  • QA lane status: satisfied.
  • Release-blocking status: clear; current-head full hosted validation and review complete.
  • Process-gap disposition: checklist+replay.

Decision log

Use existing repository wrappers and verification/review stopping rules. Shared configuration, a resolver, additional loop counters, broader presets, and automatic consumer adoption remain deferred. This PR does not claim the remaining outcomes of #392.

@github-actions github-actions Bot added the coderabbit:first-pass Triggers CodeRabbit's automatic first-pass pull-request review. label Sep 9, 2026
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-10T06:33:11.765883Z fc2c5d6 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 58 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 4ec629f0-0689-4669-bf70-9e865addeccb

📥 Commits

Reviewing files that changed from the base of the PR and between ad8175a and 79bd7e5.

📒 Files selected for processing (15)
  • CONTEXT.md
  • bin/delivery-policy-replay-test.rb
  • bin/validate
  • docs/README.md
  • docs/adoption.md
  • docs/delivery-policy.md
  • docs/seam-design.md
  • examples/delivery-policy/README.md
  • examples/delivery-policy/critical/.agents/bin/validate
  • examples/delivery-policy/low-impact/.agents/bin/validate
  • skills/autoreview/SKILL.md
  • skills/close-session/SKILL.md
  • skills/run-ci/SKILL.md
  • skills/verify/SKILL.md
  • skills/verify/references/verification-evidence.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e3cb357cd9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread skills/run-ci/SKILL.md Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
@justin808

Copy link
Copy Markdown
Member Author

Review disposition at aec8735cab837183d449439fe9eae0b2120105fd: all original-wave findings have a verified response and their conversations are resolved. Updated-head hosted CI and review remain separate gates.

Review decisions and evidence

Confirmed dirty-state mutation and hidden-untracked-file defects are fixed, and reporting order is clarified. Eight replay tests/353 assertions, complete lint, the explicitly partial pre-commit validate run, and independent corrective review pass. Unchanged installer/stack evidence comes from the original clean committed full validation. SHA-256 selection was declined as an optional expansion of the narrow example; unknown formats keep full coverage. No follow-up issue was created.

CodeRabbit explicitly skipped review and is not counted as a completed clean review. Automated summary comments are status artifacts; substantive inline findings are accounted for.

Future full-PR scans should start after this comment unless check all reviews is requested.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: aec8735cab

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread skills/run-ci/SKILL.md Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate
Comment thread skills/verify/SKILL.md
Comment thread skills/run-ci/SKILL.md
Comment thread examples/delivery-policy/critical/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread bin/delivery-policy-replay-test.rb
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate
Comment thread docs/delivery-policy.md
@claude

claude Bot commented Sep 9, 2026

Copy link
Copy Markdown

Review summary

The Ruby example wrappers (examples/delivery-policy/{low-impact,critical}/.agents/bin/validate) and the replay test are logically sound — I traced the trusted-base diff check, the dirty/untracked-file tamper detection, and the promotion-blocking logic against all 8 replay scenarios and didn't find a correctness bug. No security issues in the example scripts (no untrusted input reaches system/Open3).

Main concern is scope/proportionality, and one behavioral change that reaches beyond this PR's stated goal — see inline comments.

Scope: bin/delivery-policy-replay-test.rb (230 lines) + the two example wrappers (126 lines) + their README (56 lines) = ~412 of the ~656 added lines, and their only "caller" is each other (the test replays the examples; nothing else in the skill logic invokes them). The PR description itself notes "the design could have shipped separately from its executable adoption examples." Given that observation is already made by the author, I'd suggest actually splitting it: land docs/delivery-policy.md + the skill/doc wording updates now, and ship the executable examples + replay harness as a follow-up once the contract text is settled (wording changes to the doc would otherwise require re-touching the examples anyway).

Comment thread skills/verify/SKILL.md Outdated
Comment thread skills/close-session/SKILL.md Outdated
Comment thread bin/delivery-policy-replay-test.rb

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 96523f1baf

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread CONTEXT.md
@claude

claude Bot commented Sep 9, 2026

Copy link
Copy Markdown

Reviewed the diff (15 files, +703/-21). One inline comment posted on a concrete content bug in CONTEXT.md. Summary of the rest:

Scope vs. stated Why

The PR's own body says: "The design could have shipped separately from its executable adoption examples." I'd take that at face value. Concretely:

  • docs/delivery-policy.md (128 lines) + the CONTEXT.md/docs/adoption.md/docs/seam-design.md/docs/README.md edits are the actual policy contract — that's the part the stated Why ("a consistent way to distinguish selected integration checks from complete promotion evidence") needs.
  • examples/delivery-policy/{low-impact,critical}/.agents/bin/validate (140 lines) + examples/delivery-policy/README.md (60 lines) + bin/delivery-policy-replay-test.rb (252 lines) — roughly 450 of the 700 added lines — exist solely to demonstrate and test each other. Nothing in this repo's real .agents/bin/validate, CI, or any other caller uses these wrappers; the replay test's only purpose is to exercise the example wrappers. This repo itself doesn't adopt "selected coverage" anywhere.

Given the PR already acknowledges the split is viable, I'd suggest actually splitting it: land the contract/doc/skill changes (verify, run-ci, autoreview, close-session, verification-evidence.md) now, and defer the toy example repos + their dedicated 252-line test suite to a follow-up once a real adopting repo needs the worked example. That also shrinks the review surface for the part that actually changes agent behavior.

Duplication

The "selected coverage does not prove complete promotion coverage" point is restated near-verbatim in CONTEXT.md, docs/adoption.md, docs/delivery-policy.md, docs/seam-design.md, examples/delivery-policy/README.md, and across verify/run-ci/autoreview/close-session SKILL.md files plus verification-evidence.md. Most of the skill files correctly link to verification-evidence.md#delivery-coverage instead of repeating the prose, but autoreview/SKILL.md's new paragraph doesn't — worth linking there too for consistency and less to keep in sync later.

No functional bugs found in the example wrapper scripts themselves (git argv usage avoids shell injection; the state-diffing logic for detecting mid-validation mutation looks correct on inspection). I was not able to execute bin/delivery-policy-replay-test.rb in this review sandbox to independently confirm the claimed 8 tests / 377 assertions.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 20305c82fb

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread examples/delivery-policy/critical/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/critical/.agents/bin/validate
Comment thread examples/delivery-policy/critical/.agents/bin/validate
Comment thread CONTEXT.md
Comment thread skills/verify/SKILL.md
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate
Comment thread CONTEXT.md
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread skills/autoreview/SKILL.md
Comment thread skills/verify/SKILL.md
@justin808 justin808 added complexity:neutral Bounded repair, tests, docs or evidence with little net structural change. Not merge approval. P2 Useful follow-up: schedule after higher-impact work triage:drain-first Resolve the existing PR or concrete blocker before starting more work. Not merge approval. labels Sep 10, 2026
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate
@claude

claude Bot commented Sep 10, 2026

Copy link
Copy Markdown

Review summary

I read through the full diff (docs, skill instructions, the two example validate wrappers, and the replay test). Overall the design docs and skill-file guidance are coherent and consistent with each other (cross-references/anchors all resolve correctly, e.g. verification-evidence.md#delivery-coverage), and the example wrappers' candidate-tamper detection (comparing tracked bytes/modes/symlink targets against HEAD independently of git status) is genuinely careful for what it's trying to demonstrate.

One scope/proportion observation, per this review's focus:

Diff is dominated by illustrative code that explicitly disclaims production use. Of the 846 added lines, bin/delivery-policy-replay-test.rb (334 lines, 11 tests / 619 assertions) plus the two examples/delivery-policy/*/.agents/bin/validate wrappers (78 + 112 = 190 lines) total ~524 lines — about 62% of the PR — and all of it is teaching material the PR body itself says has "no production adoption or measured performance improvement... claimed." The core, adopted-by-everyone-using-this-repo change is the doc + skill-instruction updates (docs/delivery-policy.md, docs/adoption.md, docs/seam-design.md, the skills/*/SKILL.md edits, verification-evidence.md) — roughly 280 lines, and that content stands on its own.

The PR description anticipates this ("The design could have shipped separately from its executable adoption examples; combining them makes the documented contract verifiable in the same PR"), so this may be a deliberate tradeoff already made. But given the size and independence of the two halves, splitting the examples + replay test into a follow-up PR would let the policy contract land and be reviewed at its own (much smaller) size, while the illustrative wrappers get reviewed on their own terms without inflating the core change's diff.

Also left an inline note on the duplicated ~40-line candidate-capture helper block that's identical between the low-impact and critical example wrappers.

No correctness, security, or performance issues found beyond what's already disclosed in the PR body (per-file git cat-file process cost in the examples).

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a2c52cae0e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread examples/delivery-policy/low-impact/.agents/bin/validate Outdated
Comment thread examples/delivery-policy/low-impact/.agents/bin/validate
Comment thread docs/adoption.md
@claude

claude Bot commented Sep 10, 2026

Copy link
Copy Markdown

Review summary

I read the full diff (docs/delivery-policy.md, the two example .agents/bin/validate wrappers, bin/delivery-policy-replay-test.rb, and the six skill/doc files touched). I didn't find functional bugs in the executable code — the replay tests exercise the two example wrappers' policy-integrity checks (hidden tracked edits, staged/committed validator-policy tampering, dirty/nested-repo candidates, stale evidence reuse) and the logic they assert on checks out against a manual read of candidate_state/candidate_file in both wrappers.

Two scope/simplicity findings posted inline:

  1. Duplicated ~50-line block between the two new example wrappers (examples/delivery-policy/low-impact/.agents/bin/validate vs .../critical/.agents/bin/validate) — candidate_file, candidate_state, and the trusted-policy verification preamble are byte-identical. Since both files are new in this PR, this is a plain DRY opportunity (shared file the two wrappers load), not the "shared resolver" the PR explicitly (and correctly) defers.
  2. The same normative policy sentences are restated near-verbatim across 6+ files (docs/delivery-policy.md, docs/adoption.md, docs/seam-design.md, skills/autoreview/SKILL.md, skills/close-session/SKILL.md, skills/run-ci/SKILL.md, skills/verify/SKILL.md, skills/verify/references/verification-evidence.md) — e.g. "selected success does not prove/qualify omitted checks" and "merge authority is separate from release/deployment permission" each appear multiple times almost word-for-word. Given docs/delivery-policy.md is introduced in this same PR as the canonical contract doc, the other files could link to it instead of re-deriving the rules, which would substantially cut this diff's size and avoid future drift.

No security or performance issues stood out in the example wrappers beyond their own acknowledged scope (they're explicitly documented as teaching examples, not a production-hardened path, and the per-file git cat-file calls in candidate_state are fine at example scale).

@justin808
justin808 merged commit f26ed1e into main Sep 10, 2026
14 checks passed
@justin808
justin808 deleted the jg-codex/speed-delivery-policy branch September 10, 2026 06:57
justin808 added a commit that referenced this pull request Sep 10, 2026
…usted-base-provenance

* origin/main:
  Separate integration checks from full release verification (#808)
justin808 added a commit that referenced this pull request Sep 10, 2026
…ical-token-budgets

* origin/main:
  Separate integration checks from full release verification (#808)
  Make address-review summaries human-ready (#795)
  Speed up PR validation for ordinary documentation (#806)
  Batch installer copies to reduce validation overhead (#807)
  Report oversized PR diffs as blocked preflight coverage (#748)
justin808 added a commit that referenced this pull request Sep 11, 2026
…iet-draft-reviews

* origin/main:
  Accept trusted-base provenance for merged changes (#523)
  Limit optional review nits to the initial fix pass (#811)
  Make merge approval receipts human-first (#820)
  Simplify human-facing writing guidance (#821)
  Make ordinary PR closeout lightweight (#818)
  Separate integration checks from full release verification (#808)
  Make address-review summaries human-ready (#795)
justin808 added a commit that referenced this pull request Sep 11, 2026
…data-trust-boundary

* origin/main:
  Accept trusted-base provenance for merged changes (#523)
  Limit optional review nits to the initial fix pass (#811)
  Make merge approval receipts human-first (#820)
  Simplify human-facing writing guidance (#821)
  Make ordinary PR closeout lightweight (#818)
  Separate integration checks from full release verification (#808)
  Make address-review summaries human-ready (#795)

# Conflicts:
#	skills/pr-batch/bin/pr-security-preflight
#	skills/pr-batch/bin/pr-security-preflight-test.rb
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

coderabbit:first-pass Triggers CodeRabbit's automatic first-pass pull-request review. complexity:neutral Bounded repair, tests, docs or evidence with little net structural change. Not merge approval. P2 Useful follow-up: schedule after higher-impact work triage:drain-first Resolve the existing PR or concrete blocker before starting more work. Not merge approval.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Implement repo-owned verification tuning through the seam after delivery-policy design Define repository delivery policy and task merge authority

1 participant