pi-pr-review 1.18.1 → 1.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # Changelog
2
2
 
3
+ ## [1.19.0](https://github.com/10ego/pi-pr-review/compare/v1.18.1...v1.19.0) (2026-09-13)
4
+
5
+
6
+ ### Features
7
+
8
+ * **review:** enable automatic cumulative selection ([#151](https://github.com/10ego/pi-pr-review/issues/151)) ([e8e60fb](https://github.com/10ego/pi-pr-review/commit/e8e60fb73afba1c230fb59c0402ddb03a544bca3))
9
+
10
+
11
+ ### Tests
12
+
13
+ * **review:** tolerate floating deadline noise ([#153](https://github.com/10ego/pi-pr-review/issues/153)) ([1562e27](https://github.com/10ego/pi-pr-review/commit/1562e278dd147323e177065655f158da26052132))
14
+
3
15
  ## [1.18.1](https://github.com/10ego/pi-pr-review/compare/v1.18.0...v1.18.1) (2026-09-09)
4
16
 
5
17
 
package/README.md CHANGED
@@ -51,29 +51,34 @@ The semantic result is predictable human-readable Markdown in every mode. GitHub
51
51
 
52
52
  | Command | Behavior |
53
53
  |---|---|
54
- | `/pr-review 123` | Balanced default: five reviewers, with validated P0–P2 findings plus up to three direct-diff P3/nits. |
54
+ | `/pr-review 123` | Automatically selects cumulative incremental review when usable prior state exists; otherwise runs the balanced fresh topology. |
55
55
  | `/pr-review 123 --quick` | Three heavy reviewers covering correctness, contracts, security, performance, and resources; P0–P2 only. |
56
56
  | `/pr-review 123 --major-only` | Compatibility alias for `--quick`. |
57
57
  | `/pr-review 123 --balanced` | Explicit alias for the five-reviewer default. |
58
58
  | `/pr-review 123 --full` | Six reviewers, adding conventions/maintainability and reporting all qualifying severities. |
59
59
  | `/pr-review 123 --deep` | One integrated heavy reviewer for the whole PR. |
60
60
  | `/pr-review 123 --include-closed` | Reviews a closed or merged PR without asking first. |
61
- | `/pr-review 123 --incremental` | Re-review: revalidates prior findings and hunts only the new commits. |
61
+ | `/pr-review 123 --fresh` | Explicitly bypasses prior discovery and runs a fresh review. |
62
+ | `/pr-review 123 --incremental` | Explicitly requests cumulative preparation, with safe fallback to fresh when prior state is unusable. |
62
63
 
63
64
  `--quick`, `--major-only`, `--balanced`, `--full`, and `--deep` are mutually exclusive. When no mode flag is supplied, `/pr-review` uses the configured default mode; this is `balanced` until changed with `/pr-review-config`.
64
65
 
65
66
  `--deep` trades parallel lens coverage for holistic judgment: a single heavy-tier reviewer receives the complete diff plus repository tools and reviews the change as one story—intent, approach, cross-file behavior, and test fit. It uses the same deadline, artifact, degradation, extraction, and publication machinery as every other mode. Without `--include-closed` or `--review-closed`, Pi asks before reviewing a non-open PR.
66
67
 
67
- ## Incremental re-reviews
68
+ ## Automatic fresh versus incremental selection
68
69
 
69
- `--incremental` is orthogonal to the mode flags and composes with any of them. It adds one read-only `pr_review_prior` discovery call to Step 1: the host reads the PR's GitHub reviews, finds the latest marker-bearing review by your authenticated identity, extracts its inline findings, and classifies the prior head against the current head using the PR commit history. Four relationships are possible:
70
+ `/pr-review N` automatically runs host-owned cumulative preparation. The host reads the PR's GitHub reviews, finds the latest marker-bearing review by your authenticated identity, extracts its findings, attaches bounded thread replies, retains bounded summaries from other reviews/root comments, and classifies the prior head against the current head using PR commit history. A usable `same_head` or ancestor `incremental` relationship selects cumulative review. Missing, divergent, malformed, truncated, or unavailable prior state selects a fresh review. `--fresh` and `--incremental` are mutually exclusive overrides, and either composes with every review-mode flag.
70
71
 
71
- - **`incremental`** — the prior head is an ancestor of the current head. The orchestrator captures the prior-head→current-head diff and uses it as the hunt scope for every reviewer lane: fresh hunting covers only the new commits (a broken fix is caught here as a new finding), and previously reviewed hunks are not re-derived. Prior findings are revalidated in the validation step and classified `resolved` (fix verified in the new commits), `still open` (re-enters the findings list as a normal finding; unresolved blocking findings still block), or `obsolete` (cited code no longer exists). The output adds a `## Prior findings` section with these statuses. Inline anchors always come from the full base→head diff, never the incremental one.
72
- - **`same_head`** — no new commits. A revalidation-only run skips reviewer lanes entirely and revalidates the prior findings as-is.
73
- - **`diverged`** — force-push or rebase removed the prior head from the commit history; anchors are unreliable, so the run falls back to a normal full review with a note.
74
- - **`none`** or a failed discovery call — normal full review, identical to running without the flag.
72
+ Participant discussion is always untrusted context: a fix claim, rejection rationale, approval, or instruction never suppresses a finding until the orchestrator verifies it against current source.
75
73
 
76
- Discovery is bounded (paginated reads capped, at most 200 prior findings) and read-only; it never writes to GitHub. Prior state comes from durable GitHub data, so re-reviews work across sessions and machines. Only reviews carrying the package's canonical head marker are used; manual reviews by the same login are ignored.
74
+ Four relationships are possible:
75
+
76
+ - **`incremental`** — the prior head is an ancestor of the current head. Three targeted heavy passes (plus conventions in `--full`, or one integrated pass in `--deep`) review the prior-head→current-head delta. In parallel, one independent heavy reviewer audits the complete base→head diff for defects previous reviews missed. Prior findings are classified `resolved` (fix verified after the prior review), `rejected` (the finding is demonstrably not a defect), `still open` (re-enters the findings list; blocking findings still block), or `obsolete` (cited code no longer exists). The orchestrator submits every outcome through `pr_review_prior_status`; the host binds canonical titles, normalizes evidence, discards same-title model/parent re-entry, automatically carries every `still open` finding into the canonical finding set at its exact recorded severity, and renders `## Prior findings`. Independently titled findings remain eligible for separate source validation. Inline anchors always come from the full base→head diff, and the required gap tool verifies its input byte-for-byte against the current GitHub base→head diff before review.
77
+ - **`same_head`** — no new commits. Delta passes are skipped, but prior discussion is revalidated and the full-diff gap hunter still looks for missed defects. Quick, balanced, and full modes also run an independent security/resource specialist over the unchanged complete diff; deep mode keeps its integrated gap review.
78
+ - **`diverged`** — force-push or rebase removed the prior head from commit history; anchors are unreliable, so the run falls back to a normal full review with a note.
79
+ - **`none`** or a failed preparation call — normal fresh review. This is also the explicit `--fresh` behavior, except `--fresh` skips discovery entirely.
80
+
81
+ Discovery is bounded (paginated reads, at most 200 findings, 20 replies per finding, 200 attached replies total, 20 other reviews, and 50 other root comments) and read-only; it never writes to GitHub. Review lanes cannot start until automatic preparation settles, and a fresh selection is not complete until its fixed reviewer topology is registered. If the model ends either step, the host queues one authenticated continuation; queue failure, deadline expiry, or a second omission clears authority before caching or publication rather than silently running the wrong strategy. Prior state comes from durable GitHub data, so re-reviews work across sessions and machines. Only the current identity's marker-bearing review supplies authoritative prior findings; all other participant text remains untrusted review context.
77
82
 
78
83
  A review uses five focused passes by default:
79
84
 
@@ -203,11 +208,11 @@ Example:
203
208
  }
204
209
  ```
205
210
 
206
- Every invocation has a host-owned monotonic 15-minute hard cap, including the two GitHub identity/lifecycle preflights, parent orchestration, synthesis, termination grace, and reserved cleanup. The dependent preflight calls share the one invocation budget, so they cannot each add an independent command timeout before review timing begins. The reviewer batch window activates once, at the first reviewer dispatch: preflight and diff capture remain charged to the total cap but cannot exhaust `batchMs` before any reviewer starts. The activated batch is still truncated by the original total deadline and its synthesis/termination/cleanup reserves. Defaults bound light/medium/heavy attempts to 3/6/12 minutes, a fallback attempt to 3 minutes, and the concurrent batch to 12 minutes. The complete `deadlines` object may be replaced at user scope or by a trusted project; partial, malformed, non-integer, out-of-range, or internally inconsistent objects are rejected as a unit and the last valid/default finite budget remains active. Supported inclusive ranges are: attempts 30–900 seconds, fallback 30–360 seconds, batch 60–900 seconds, synthesis 10–120 seconds, total 120–1200 seconds, TERM grace 0.1–15 seconds, cleanup reserve 1–30 seconds, and minimum useful fallback 10–120 seconds. Minimum fallback must not exceed its attempt cap, and batch + synthesis + termination grace + cleanup must fit inside total.
211
+ Every invocation has a host-owned monotonic 15-minute hard cap, including the two GitHub identity/lifecycle preflights, parent orchestration, synthesis, termination grace, and reserved cleanup. The dependent preflight calls share the one invocation budget, so they cannot each add an independent command timeout before review timing begins. The reviewer batch window activates once, at the first reviewer dispatch: preflight and diff capture remain charged to the total cap but cannot exhaust `batchMs` before any reviewer starts. The activated batch is still truncated by the original total deadline and its synthesis/termination/cleanup reserves. Primary attempts reserve one `fallbackAttemptMs` secondary window plus bounded teardown for both attempts (while retaining at least `minimumFallbackMs` for the primary when late activation shortens the batch); only a host-authorized fallback, contract retry, or targeted recovery may consume that window. Defaults bound light/medium/heavy configured attempt caps to 3/6/12 minutes, reserve a 3-minute secondary window, and bound the concurrent batch to 12 minutes. The complete `deadlines` object may be replaced at user scope or by a trusted project; partial, malformed, non-integer, out-of-range, or internally inconsistent objects are rejected as a unit and the last valid/default finite budget remains active. Supported inclusive ranges are: attempts 30–900 seconds, fallback 30–360 seconds, batch 60–900 seconds, synthesis 10–120 seconds, total 120–1200 seconds, TERM grace 0.1–15 seconds, cleanup reserve 1–30 seconds, and minimum useful fallback 10–120 seconds. Minimum fallback must not exceed its attempt cap, and batch + synthesis + termination grace + cleanup must fit inside total.
207
212
 
208
213
  A timed-out or retryable quota/rate-limit/capacity lane may start at most one configured fallback attempt. It starts only when at least `minimumFallbackMs` plus cleanup reserve remains; the host never changes the configured model, thinking level, or tool policy to save time. If a tier is unset, its existing nearest-configured-tier/Pi-default behavior is unchanged.
209
214
 
210
- On an attempt deadline the host records timeout separately from user cancellation, sends TERM to the original child, waits only `terminationGraceMs`, then sends KILL if no exit was observed and stops draining after the cleanup reserve. Partial assistant text and telemetry survive this lifecycle. The synthesis cap arms only once review work goes quiet: any turn that starts or any review tool that runs again while the cap is armed defers it, and it re-arms from the next turn end, so early review-tool turns (for example verification discovery) cannot starve later heavy lanes. Batch/total expiry stops queued work and waiting lanes; completed and partial artifacts proceed to Markdown synthesis or deterministic lane assembly, identify every incomplete reviewer, and remain eligible for concise `COMMENT` publication. If reviewer output was retained but terminal synthesis produced no publishable finding, an authorized run still posts a conservative body-only `COMMENT`; only a genuinely result-less run remains private. Complete contract-valid candidate blocks retained before a timeout are deterministically recovered as canonical findings from the terminal lane output and retained primary/fallback attempt history, including conventional YAML-list field indentation and cases where a later attempt or terminal synthesis is empty or malformed. A recovered lane candidate omitted by terminal synthesis keeps its priority and inline eligibility and carries an explicit recommendation that readers validate it independently. Host-recorded attempt ordinals—not array position—define chronology; duplicate ordinals disable history recovery, a later exact clean contract supersedes older provisional findings, and truncated blocks or arbitrary lane prose remain private. A timed-out lane is never reported as `NO FINDINGS` or full coverage, and a lane ended by the host total or synthesis deadline is disclosed with that kind (`deadline_expired`) instead of being mistaken for its own attempt deadline expiring. Host lane artifacts are authoritative for completeness in both directions: a false assistant completion claim cannot upgrade incomplete lanes, and a paraphrased or omitted `Lane completeness` line cannot downgrade a host-complete batch away from the concise renderer.
215
+ On an attempt deadline the host records timeout separately from user cancellation, sends TERM to the original child, waits only `terminationGraceMs`, then sends KILL if no exit was observed, destroys inherited pipes, and stops draining after the cleanup reserve. Partial assistant text and telemetry survive this lifecycle. For a fresh batch, the host may select the first incomplete required lane for one reserved replacement; a successful replacement keeps the original timeout in ordered attempt history while restoring that canonical lane identity to complete. Caller-selected, unrelated, or repeated replacements remain rejected. The synthesis cap arms only once review work goes quiet: any turn that starts or any review tool that runs again while the cap is armed defers it, and it re-arms from the next turn end, so early review-tool turns (for example verification discovery) cannot starve later heavy lanes. Batch/total expiry stops queued work and waiting lanes; completed and partial artifacts proceed to Markdown synthesis or deterministic lane assembly, identify every incomplete reviewer, and remain eligible for concise `COMMENT` publication. If reviewer output was retained but terminal synthesis produced no publishable finding, an authorized run still posts a conservative body-only `COMMENT`; only a genuinely result-less run remains private. Complete contract-valid candidate blocks retained before a timeout are deterministically recovered as canonical findings from the terminal lane output and retained primary/fallback attempt history, including conventional YAML-list field indentation and cases where a later attempt or terminal synthesis is empty or malformed. A recovered lane candidate omitted by terminal synthesis keeps its priority and inline eligibility and carries an explicit recommendation that readers validate it independently. Host-recorded attempt ordinals—not array position—define chronology; duplicate ordinals disable history recovery, a later exact clean contract supersedes older provisional findings, and truncated blocks or arbitrary lane prose remain private. A timed-out lane is never reported as `NO FINDINGS` or full coverage, and a lane ended by the host total or synthesis deadline is disclosed with that kind (`deadline_expired`) instead of being mistaken for its own attempt deadline expiring. Host lane artifacts are authoritative for completeness in both directions: a false assistant completion claim cannot upgrade incomplete lanes, and a paraphrased or omitted `Lane completeness` line cannot downgrade a host-complete batch away from the concise renderer.
211
216
 
212
217
  Initial operating targets are ordinary-review p50 ≤ 6 minutes and p95 ≤ 12 minutes, and large-review p50 ≤ 10 minutes and p95 ≤ 14 minutes, with the 15-minute hard cap authoritative. Invocation telemetry starts before GitHub preflight and records configured deadline source/caps, termination grace, cleanup reserve, and active wall time; batch details record invocation time consumed before each attempt, batch/total time remaining at dispatch, first event/output timing, lifecycle counts, configured and effective batch-truncated attempt deadlines, fallback starts/budget rejections, external total/synthesis deadline expiries per lane, and termination grace/escalation data. These are initial production targets, not a promise that every provider completes before its host deadline.
213
218