pi-pr-review 1.17.2 → 1.17.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,19 @@
1
1
  # Changelog
2
2
 
3
+ ## [1.17.4](https://github.com/10ego/pi-pr-review/compare/v1.17.3...v1.17.4) (2026-09-01)
4
+
5
+
6
+ ### Bug Fixes
7
+
8
+ * **review:** default heavy attempts to twelve minutes ([#130](https://github.com/10ego/pi-pr-review/issues/130)) ([7a3af3b](https://github.com/10ego/pi-pr-review/commit/7a3af3b3bf9f5beb0a8c309c3e7524f782a1d88f))
9
+
10
+ ## [1.17.3](https://github.com/10ego/pi-pr-review/compare/v1.17.2...v1.17.3) (2026-08-31)
11
+
12
+
13
+ ### Bug Fixes
14
+
15
+ * **review:** preserve validated confidence ratings ([#128](https://github.com/10ego/pi-pr-review/issues/128)) ([27843b7](https://github.com/10ego/pi-pr-review/commit/27843b7c6469827e65fc5d7c69511b20be3615ba))
16
+
3
17
  ## [1.17.2](https://github.com/10ego/pi-pr-review/compare/v1.17.1...v1.17.2) (2026-08-31)
4
18
 
5
19
 
package/README.md CHANGED
@@ -175,7 +175,7 @@ Example:
175
175
  },
176
176
  "tools": ["read", "bash", "grep", "find", "ls"],
177
177
  "deadlines": {
178
- "attemptMs": { "light": 180000, "medium": 360000, "heavy": 480000 },
178
+ "attemptMs": { "light": 180000, "medium": 360000, "heavy": 720000 },
179
179
  "fallbackAttemptMs": 180000,
180
180
  "batchMs": 720000,
181
181
  "synthesisMs": 60000,
@@ -191,7 +191,7 @@ Example:
191
191
  }
192
192
  ```
193
193
 
194
- Every invocation has a host-owned monotonic 15-minute hard cap, including the two GitHub identity/lifecycle preflights, parent orchestration, synthesis, termination grace, and reserved cleanup. The dependent preflight calls share the one invocation budget, so they cannot each add an independent command timeout before review timing begins. The reviewer batch window activates once, at the first reviewer dispatch: preflight and diff capture remain charged to the total cap but cannot exhaust `batchMs` before any reviewer starts. The activated batch is still truncated by the original total deadline and its synthesis/termination/cleanup reserves. Defaults bound light/medium/heavy attempts to 3/6/8 minutes, a fallback attempt to 3 minutes, and the concurrent batch to 12 minutes. The complete `deadlines` object may be replaced at user scope or by a trusted project; partial, malformed, non-integer, out-of-range, or internally inconsistent objects are rejected as a unit and the last valid/default finite budget remains active. Supported inclusive ranges are: attempts 30–720 seconds, fallback 30–360 seconds, batch 60–840 seconds, synthesis 10–120 seconds, total 120–1200 seconds, TERM grace 0.1–15 seconds, cleanup reserve 1–30 seconds, and minimum useful fallback 10–120 seconds. Minimum fallback must not exceed its attempt cap, and batch + synthesis + termination grace + cleanup must fit inside total.
194
+ Every invocation has a host-owned monotonic 15-minute hard cap, including the two GitHub identity/lifecycle preflights, parent orchestration, synthesis, termination grace, and reserved cleanup. The dependent preflight calls share the one invocation budget, so they cannot each add an independent command timeout before review timing begins. The reviewer batch window activates once, at the first reviewer dispatch: preflight and diff capture remain charged to the total cap but cannot exhaust `batchMs` before any reviewer starts. The activated batch is still truncated by the original total deadline and its synthesis/termination/cleanup reserves. Defaults bound light/medium/heavy attempts to 3/6/12 minutes, a fallback attempt to 3 minutes, and the concurrent batch to 12 minutes. The complete `deadlines` object may be replaced at user scope or by a trusted project; partial, malformed, non-integer, out-of-range, or internally inconsistent objects are rejected as a unit and the last valid/default finite budget remains active. Supported inclusive ranges are: attempts 30–720 seconds, fallback 30–360 seconds, batch 60–840 seconds, synthesis 10–120 seconds, total 120–1200 seconds, TERM grace 0.1–15 seconds, cleanup reserve 1–30 seconds, and minimum useful fallback 10–120 seconds. Minimum fallback must not exceed its attempt cap, and batch + synthesis + termination grace + cleanup must fit inside total.
195
195
 
196
196
  A timed-out or retryable quota/rate-limit/capacity lane may start at most one configured fallback attempt. It starts only when at least `minimumFallbackMs` plus cleanup reserve remains; the host never changes the configured model, thinking level, or tool policy to save time. If a tier is unset, its existing nearest-configured-tier/Pi-default behavior is unchanged.
197
197
 
@@ -293,7 +293,7 @@ Each finding includes:
293
293
 
294
294
  - severity: `P0`, `P1`, `P2`, `P3`, or `nit`;
295
295
  - whether it blocks the verdict;
296
- - an explanation and confidence score;
296
+ - an explanation and independently validated confidence score (the host preserves the supplied 0.0–1.0 value instead of manufacturing `1.00`; scoreless or malformed current Markdown degrades fail-closed, and pre-fix cached completion records are invalidated);
297
297
  - a diff-anchored file and line range when available.
298
298
 
299
299
  | Severity | Meaning |
@@ -18,7 +18,7 @@ export interface ReviewDeadlineConfig {
18
18
 
19
19
  /** Conservative initial caps; the 15 minute invocation cap makes 30 minute reviews impossible. */
20
20
  export const DEFAULT_REVIEW_DEADLINES: Readonly<ReviewDeadlineConfig> = Object.freeze({
21
- attemptMs: Object.freeze({ light: 180_000, medium: 360_000, heavy: 480_000 }),
21
+ attemptMs: Object.freeze({ light: 180_000, medium: 360_000, heavy: 720_000 }),
22
22
  fallbackAttemptMs: 180_000,
23
23
  batchMs: 720_000,
24
24
  synthesisMs: 60_000,
@@ -384,14 +384,18 @@ function parseFindings(text: string): { findings: ReviewFindingLike[]; count: nu
384
384
  const severity = explicitSeverity?.toLowerCase() === "nit" ? "nit" : explicitSeverity?.toUpperCase();
385
385
  const normalizedTag = tagged?.toLowerCase() === "nit" ? "nit" : tagged?.toUpperCase();
386
386
  const rationale = field(block, "Rationale") ?? field(block, "Why");
387
+ const confidenceText = field(block, "Confidence");
388
+ const confidence = confidenceText === undefined ? undefined : Number(confidenceText.trim());
389
+ const validConfidence = confidenceText !== undefined &&
390
+ /^(?:0(?:\.\d+)?|1(?:\.0+)?)$/.test(confidenceText.trim()) && Number.isFinite(confidence);
387
391
  const location = parseLocation(field(block, "Location"));
388
392
  const recognizedFieldCount = fieldCount(block, "Severity") + fieldCount(block, "Rationale") +
389
- fieldCount(block, "Why") + fieldCount(block, "Location");
390
- const unconsumed = block.replace(/^\*\*(?:Severity|Rationale|Why|Location):\*\*.*$/gim, "").trim();
393
+ fieldCount(block, "Why") + fieldCount(block, "Confidence") + fieldCount(block, "Location");
394
+ const unconsumed = block.replace(/^\*\*(?:Severity|Rationale|Why|Confidence|Location):\*\*.*$/gim, "").trim();
391
395
  if (
392
- !explicitSeverity || !rationale?.trim() || recognizedFieldCount < 2 || unconsumed ||
396
+ !explicitSeverity || !rationale?.trim() || recognizedFieldCount < 2 || unconsumed || !validConfidence ||
393
397
  fieldCount(block, "Severity") !== 1 || fieldCount(block, "Rationale") + fieldCount(block, "Why") !== 1 ||
394
- fieldCount(block, "Location") > 1
398
+ fieldCount(block, "Confidence") !== 1 || fieldCount(block, "Location") > 1
395
399
  ) {
396
400
  complete = false;
397
401
  continue;
@@ -410,7 +414,7 @@ function parseFindings(text: string): { findings: ReviewFindingLike[]; count: nu
410
414
  severity,
411
415
  blocking: severity === "P0" || severity === "P1",
412
416
  body: rationale,
413
- confidence_score: 1,
417
+ confidence_score: confidence,
414
418
  code_location: location.location,
415
419
  });
416
420
  }
@@ -470,7 +474,6 @@ function syntheticReview(
470
474
  verdict,
471
475
  overall_correctness: verdict === "request_changes" ? "patch is incorrect" : "patch is correct",
472
476
  overall_explanation: overview.split(/\n\n|\n/)[0] ?? "Review completed.",
473
- overall_confidence_score: 1,
474
477
  };
475
478
  }
476
479
 
@@ -596,7 +596,7 @@ export interface CompletedReviewSessionIdentity {
596
596
  }
597
597
 
598
598
  export interface PersistedCompletedReview {
599
- schemaVersion: 2;
599
+ schemaVersion: 3;
600
600
  session: CompletedReviewSessionIdentity;
601
601
  invocation: ReviewInvocation;
602
602
  repository: RepositoryBinding;
@@ -819,7 +819,7 @@ export class CompletedReviewCache {
819
819
  const digest = reviewHash(record.review);
820
820
  const useReference = !!reviewEntryId && !!referencedReview && reviewHash(referencedReview) === digest;
821
821
  return {
822
- schemaVersion: 2,
822
+ schemaVersion: 3,
823
823
  session: { ...session },
824
824
  invocation: { ...record.invocation, autoPost: { ...record.invocation.autoPost } },
825
825
  repository: { ...record.repository },
@@ -849,7 +849,7 @@ export class CompletedReviewCache {
849
849
  ): boolean {
850
850
  if (
851
851
  !isObject(value) ||
852
- value.schemaVersion !== 2 ||
852
+ value.schemaVersion !== 3 ||
853
853
  !validSessionIdentity(value.session) ||
854
854
  !sameSessionIdentity(value.session, session) ||
855
855
  !validRepositoryBinding(value.repository)
@@ -866,7 +866,7 @@ export class CompletedReviewCache {
866
866
  const candidate = hasReference ? referencedReview : value.review;
867
867
  let parsed: PublishableReviewParseResult;
868
868
  try {
869
- parsed = parsePublishableReview(JSON.stringify(candidate));
869
+ parsed = canonicalReviewSnapshot(candidate as ReviewLike);
870
870
  } catch {
871
871
  return false;
872
872
  }
@@ -1083,7 +1083,10 @@ function publishableReviewEnvelope(text: string): { text: string; source: "json"
1083
1083
  * Publication accepts one complete JSON object. A single surrounding Markdown code
1084
1084
  * fence is tolerated for compatibility but remains distinguishable from exact JSON.
1085
1085
  */
1086
- export function parsePublishableReview(text: string): PublishableReviewParseResult {
1086
+ export function parsePublishableReview(
1087
+ text: string,
1088
+ options: { allowMissingOverallConfidence?: boolean } = {},
1089
+ ): PublishableReviewParseResult {
1087
1090
  const envelope = publishableReviewEnvelope(text);
1088
1091
  let value: unknown;
1089
1092
  try {
@@ -1140,7 +1143,11 @@ export function parsePublishableReview(text: string): PublishableReviewParseResu
1140
1143
  if (!new Set(["patch is correct", "patch is incorrect"]).has(String(value.overall_correctness))) {
1141
1144
  return { error: "overall_correctness is invalid" };
1142
1145
  }
1143
- if (!isConfidence(value.overall_confidence_score)) {
1146
+ if (
1147
+ value.overall_confidence_score !== undefined
1148
+ ? !isConfidence(value.overall_confidence_score)
1149
+ : !options.allowMissingOverallConfidence
1150
+ ) {
1144
1151
  return { error: "overall_confidence_score is invalid" };
1145
1152
  }
1146
1153
  return { review: value as unknown as ReviewLike, source: envelope.source };
@@ -1421,7 +1428,10 @@ export function canonicalReviewSnapshot(review: ReviewLike): PublishableReviewPa
1421
1428
  if (typeof serialized !== "string") {
1422
1429
  return { error: "review could not be serialized for publication" };
1423
1430
  }
1424
- return parsePublishableReview(serialized);
1431
+ // Host-synthesized canonical Markdown has no manufactured overall score,
1432
+ // but every parsed finding still requires its validated numeric confidence.
1433
+ // Assistant-authored strict JSON requires both at its public parse boundary.
1434
+ return parsePublishableReview(serialized, { allowMissingOverallConfidence: true });
1425
1435
  }
1426
1436
 
1427
1437
  export function validateInlineComments(
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-pr-review",
3
- "version": "1.17.2",
3
+ "version": "1.17.4",
4
4
  "description": "Parallel AI code review for GitHub pull requests in the Pi coding agent, with model-agnostic tiered subagents, structured findings, optional verification, and host-gated COMMENT or qualified APPROVE publishing.",
5
5
  "keywords": [
6
6
  "pi-package",
@@ -222,6 +222,7 @@ Return Markdown using these stable headings. Do not emit a GitHub API payload an
222
222
  ### [P1] Imperative finding title
223
223
  **Severity:** P1
224
224
  **Rationale:** Concise Markdown explaining the trigger, impact, and source-grounded evidence.
225
+ **Confidence:** 0.92
225
226
  **Location:** `path/to/file.ts:10-12 RIGHT`
226
227
 
227
228
  ## Lane completeness
@@ -231,6 +232,6 @@ Return Markdown using these stable headings. Do not emit a GitHub API payload an
231
232
  <useful strengths plus concise correctness, security, or performance notes>
232
233
  ```
233
234
 
234
- Repeat the `###` finding block for every confirmed finding. Omit `Location` when no safe diff location exists. Use `No findings.` under `## Findings` when empty. The extension retains the complete raw synthesis even when some or all finding blocks cannot be parsed. Canonical heading discovery ignores heading-like content in CommonMark fenced-code and HTML-block contexts. It validates safe P0–P3 anchors for inline placement, preserves ambiguous or unparsed substantive content in exactly one sanitized body-only `COMMENT`, and degrades malformed synthesis the same way; if synthesis is absent, it assembles a deterministic body-only review from retained lane artifacts. Optional formatting repair can never block this fallback.
235
+ Repeat the `###` finding block for every confirmed finding. `Confidence` is the independently validated numeric confidence for that finding from 0.0 through 1.0; do not use a fixed default. Omit `Location` when no safe diff location exists. Use `No findings.` under `## Findings` when empty. The extension retains the complete raw synthesis even when some or all finding blocks cannot be parsed. Canonical heading discovery ignores heading-like content in CommonMark fenced-code and HTML-block contexts. It validates safe P0–P3 anchors for inline placement, preserves ambiguous or unparsed substantive content in exactly one sanitized body-only `COMMENT`, and degrades malformed synthesis the same way; if synthesis is absent, it assembles a deterministic body-only review from retained lane artifacts. Optional formatting repair can never block this fallback.
235
236
 
236
237
  Host code captures repository, PR, reviewed head, lifecycle, posting authority, stale policy, and invocation identity before review execution, serializes the single GitHub request with `JSON.stringify`, sanitizes reserved markers, enforces payload limits, appends the canonical marker, and performs duplicate/stale/draft/lifecycle/final-head checks. The assistant's semantic `approve` is only one input to the host gates; assistant text cannot directly choose `APPROVE`, `REQUEST_CHANGES`, commit ID, API path, repository, or hostname. Parent candidate validation must gather all independent evidence in one wave and use at most one dependency-driven follow-up, as specified in Step 7.