pi-pr-review 1.17.2 → 1.17.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +14 -0
- package/README.md +3 -3
- package/lib/pr-review-deadlines.ts +1 -1
- package/lib/pr-review-markdown.ts +9 -6
- package/lib/pr-review-publish.ts +17 -7
- package/package.json +1 -1
- package/prompts/pr-review.md +2 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,19 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## [1.17.4](https://github.com/10ego/pi-pr-review/compare/v1.17.3...v1.17.4) (2026-09-01)
|
|
4
|
+
|
|
5
|
+
|
|
6
|
+
### Bug Fixes
|
|
7
|
+
|
|
8
|
+
* **review:** default heavy attempts to twelve minutes ([#130](https://github.com/10ego/pi-pr-review/issues/130)) ([7a3af3b](https://github.com/10ego/pi-pr-review/commit/7a3af3b3bf9f5beb0a8c309c3e7524f782a1d88f))
|
|
9
|
+
|
|
10
|
+
## [1.17.3](https://github.com/10ego/pi-pr-review/compare/v1.17.2...v1.17.3) (2026-08-31)
|
|
11
|
+
|
|
12
|
+
|
|
13
|
+
### Bug Fixes
|
|
14
|
+
|
|
15
|
+
* **review:** preserve validated confidence ratings ([#128](https://github.com/10ego/pi-pr-review/issues/128)) ([27843b7](https://github.com/10ego/pi-pr-review/commit/27843b7c6469827e65fc5d7c69511b20be3615ba))
|
|
16
|
+
|
|
3
17
|
## [1.17.2](https://github.com/10ego/pi-pr-review/compare/v1.17.1...v1.17.2) (2026-08-31)
|
|
4
18
|
|
|
5
19
|
|
package/README.md
CHANGED
|
@@ -175,7 +175,7 @@ Example:
|
|
|
175
175
|
},
|
|
176
176
|
"tools": ["read", "bash", "grep", "find", "ls"],
|
|
177
177
|
"deadlines": {
|
|
178
|
-
"attemptMs": { "light": 180000, "medium": 360000, "heavy":
|
|
178
|
+
"attemptMs": { "light": 180000, "medium": 360000, "heavy": 720000 },
|
|
179
179
|
"fallbackAttemptMs": 180000,
|
|
180
180
|
"batchMs": 720000,
|
|
181
181
|
"synthesisMs": 60000,
|
|
@@ -191,7 +191,7 @@ Example:
|
|
|
191
191
|
}
|
|
192
192
|
```
|
|
193
193
|
|
|
194
|
-
Every invocation has a host-owned monotonic 15-minute hard cap, including the two GitHub identity/lifecycle preflights, parent orchestration, synthesis, termination grace, and reserved cleanup. The dependent preflight calls share the one invocation budget, so they cannot each add an independent command timeout before review timing begins. The reviewer batch window activates once, at the first reviewer dispatch: preflight and diff capture remain charged to the total cap but cannot exhaust `batchMs` before any reviewer starts. The activated batch is still truncated by the original total deadline and its synthesis/termination/cleanup reserves. Defaults bound light/medium/heavy attempts to 3/6/
|
|
194
|
+
Every invocation has a host-owned monotonic 15-minute hard cap, including the two GitHub identity/lifecycle preflights, parent orchestration, synthesis, termination grace, and reserved cleanup. The dependent preflight calls share the one invocation budget, so they cannot each add an independent command timeout before review timing begins. The reviewer batch window activates once, at the first reviewer dispatch: preflight and diff capture remain charged to the total cap but cannot exhaust `batchMs` before any reviewer starts. The activated batch is still truncated by the original total deadline and its synthesis/termination/cleanup reserves. Defaults bound light/medium/heavy attempts to 3/6/12 minutes, a fallback attempt to 3 minutes, and the concurrent batch to 12 minutes. The complete `deadlines` object may be replaced at user scope or by a trusted project; partial, malformed, non-integer, out-of-range, or internally inconsistent objects are rejected as a unit and the last valid/default finite budget remains active. Supported inclusive ranges are: attempts 30–720 seconds, fallback 30–360 seconds, batch 60–840 seconds, synthesis 10–120 seconds, total 120–1200 seconds, TERM grace 0.1–15 seconds, cleanup reserve 1–30 seconds, and minimum useful fallback 10–120 seconds. Minimum fallback must not exceed its attempt cap, and batch + synthesis + termination grace + cleanup must fit inside total.
|
|
195
195
|
|
|
196
196
|
A timed-out or retryable quota/rate-limit/capacity lane may start at most one configured fallback attempt. It starts only when at least `minimumFallbackMs` plus cleanup reserve remains; the host never changes the configured model, thinking level, or tool policy to save time. If a tier is unset, its existing nearest-configured-tier/Pi-default behavior is unchanged.
|
|
197
197
|
|
|
@@ -293,7 +293,7 @@ Each finding includes:
|
|
|
293
293
|
|
|
294
294
|
- severity: `P0`, `P1`, `P2`, `P3`, or `nit`;
|
|
295
295
|
- whether it blocks the verdict;
|
|
296
|
-
- an explanation and confidence score;
|
|
296
|
+
- an explanation and independently validated confidence score (the host preserves the supplied 0.0–1.0 value instead of manufacturing `1.00`; scoreless or malformed current Markdown degrades fail-closed, and pre-fix cached completion records are invalidated);
|
|
297
297
|
- a diff-anchored file and line range when available.
|
|
298
298
|
|
|
299
299
|
| Severity | Meaning |
|
|
@@ -18,7 +18,7 @@ export interface ReviewDeadlineConfig {
|
|
|
18
18
|
|
|
19
19
|
/** Conservative initial caps; the 15 minute invocation cap makes 30 minute reviews impossible. */
|
|
20
20
|
export const DEFAULT_REVIEW_DEADLINES: Readonly<ReviewDeadlineConfig> = Object.freeze({
|
|
21
|
-
attemptMs: Object.freeze({ light: 180_000, medium: 360_000, heavy:
|
|
21
|
+
attemptMs: Object.freeze({ light: 180_000, medium: 360_000, heavy: 720_000 }),
|
|
22
22
|
fallbackAttemptMs: 180_000,
|
|
23
23
|
batchMs: 720_000,
|
|
24
24
|
synthesisMs: 60_000,
|
|
@@ -384,14 +384,18 @@ function parseFindings(text: string): { findings: ReviewFindingLike[]; count: nu
|
|
|
384
384
|
const severity = explicitSeverity?.toLowerCase() === "nit" ? "nit" : explicitSeverity?.toUpperCase();
|
|
385
385
|
const normalizedTag = tagged?.toLowerCase() === "nit" ? "nit" : tagged?.toUpperCase();
|
|
386
386
|
const rationale = field(block, "Rationale") ?? field(block, "Why");
|
|
387
|
+
const confidenceText = field(block, "Confidence");
|
|
388
|
+
const confidence = confidenceText === undefined ? undefined : Number(confidenceText.trim());
|
|
389
|
+
const validConfidence = confidenceText !== undefined &&
|
|
390
|
+
/^(?:0(?:\.\d+)?|1(?:\.0+)?)$/.test(confidenceText.trim()) && Number.isFinite(confidence);
|
|
387
391
|
const location = parseLocation(field(block, "Location"));
|
|
388
392
|
const recognizedFieldCount = fieldCount(block, "Severity") + fieldCount(block, "Rationale") +
|
|
389
|
-
fieldCount(block, "Why") + fieldCount(block, "Location");
|
|
390
|
-
const unconsumed = block.replace(/^\*\*(?:Severity|Rationale|Why|Location):\*\*.*$/gim, "").trim();
|
|
393
|
+
fieldCount(block, "Why") + fieldCount(block, "Confidence") + fieldCount(block, "Location");
|
|
394
|
+
const unconsumed = block.replace(/^\*\*(?:Severity|Rationale|Why|Confidence|Location):\*\*.*$/gim, "").trim();
|
|
391
395
|
if (
|
|
392
|
-
!explicitSeverity || !rationale?.trim() || recognizedFieldCount < 2 || unconsumed ||
|
|
396
|
+
!explicitSeverity || !rationale?.trim() || recognizedFieldCount < 2 || unconsumed || !validConfidence ||
|
|
393
397
|
fieldCount(block, "Severity") !== 1 || fieldCount(block, "Rationale") + fieldCount(block, "Why") !== 1 ||
|
|
394
|
-
fieldCount(block, "Location") > 1
|
|
398
|
+
fieldCount(block, "Confidence") !== 1 || fieldCount(block, "Location") > 1
|
|
395
399
|
) {
|
|
396
400
|
complete = false;
|
|
397
401
|
continue;
|
|
@@ -410,7 +414,7 @@ function parseFindings(text: string): { findings: ReviewFindingLike[]; count: nu
|
|
|
410
414
|
severity,
|
|
411
415
|
blocking: severity === "P0" || severity === "P1",
|
|
412
416
|
body: rationale,
|
|
413
|
-
confidence_score:
|
|
417
|
+
confidence_score: confidence,
|
|
414
418
|
code_location: location.location,
|
|
415
419
|
});
|
|
416
420
|
}
|
|
@@ -470,7 +474,6 @@ function syntheticReview(
|
|
|
470
474
|
verdict,
|
|
471
475
|
overall_correctness: verdict === "request_changes" ? "patch is incorrect" : "patch is correct",
|
|
472
476
|
overall_explanation: overview.split(/\n\n|\n/)[0] ?? "Review completed.",
|
|
473
|
-
overall_confidence_score: 1,
|
|
474
477
|
};
|
|
475
478
|
}
|
|
476
479
|
|
package/lib/pr-review-publish.ts
CHANGED
|
@@ -596,7 +596,7 @@ export interface CompletedReviewSessionIdentity {
|
|
|
596
596
|
}
|
|
597
597
|
|
|
598
598
|
export interface PersistedCompletedReview {
|
|
599
|
-
schemaVersion:
|
|
599
|
+
schemaVersion: 3;
|
|
600
600
|
session: CompletedReviewSessionIdentity;
|
|
601
601
|
invocation: ReviewInvocation;
|
|
602
602
|
repository: RepositoryBinding;
|
|
@@ -819,7 +819,7 @@ export class CompletedReviewCache {
|
|
|
819
819
|
const digest = reviewHash(record.review);
|
|
820
820
|
const useReference = !!reviewEntryId && !!referencedReview && reviewHash(referencedReview) === digest;
|
|
821
821
|
return {
|
|
822
|
-
schemaVersion:
|
|
822
|
+
schemaVersion: 3,
|
|
823
823
|
session: { ...session },
|
|
824
824
|
invocation: { ...record.invocation, autoPost: { ...record.invocation.autoPost } },
|
|
825
825
|
repository: { ...record.repository },
|
|
@@ -849,7 +849,7 @@ export class CompletedReviewCache {
|
|
|
849
849
|
): boolean {
|
|
850
850
|
if (
|
|
851
851
|
!isObject(value) ||
|
|
852
|
-
value.schemaVersion !==
|
|
852
|
+
value.schemaVersion !== 3 ||
|
|
853
853
|
!validSessionIdentity(value.session) ||
|
|
854
854
|
!sameSessionIdentity(value.session, session) ||
|
|
855
855
|
!validRepositoryBinding(value.repository)
|
|
@@ -866,7 +866,7 @@ export class CompletedReviewCache {
|
|
|
866
866
|
const candidate = hasReference ? referencedReview : value.review;
|
|
867
867
|
let parsed: PublishableReviewParseResult;
|
|
868
868
|
try {
|
|
869
|
-
parsed =
|
|
869
|
+
parsed = canonicalReviewSnapshot(candidate as ReviewLike);
|
|
870
870
|
} catch {
|
|
871
871
|
return false;
|
|
872
872
|
}
|
|
@@ -1083,7 +1083,10 @@ function publishableReviewEnvelope(text: string): { text: string; source: "json"
|
|
|
1083
1083
|
* Publication accepts one complete JSON object. A single surrounding Markdown code
|
|
1084
1084
|
* fence is tolerated for compatibility but remains distinguishable from exact JSON.
|
|
1085
1085
|
*/
|
|
1086
|
-
export function parsePublishableReview(
|
|
1086
|
+
export function parsePublishableReview(
|
|
1087
|
+
text: string,
|
|
1088
|
+
options: { allowMissingOverallConfidence?: boolean } = {},
|
|
1089
|
+
): PublishableReviewParseResult {
|
|
1087
1090
|
const envelope = publishableReviewEnvelope(text);
|
|
1088
1091
|
let value: unknown;
|
|
1089
1092
|
try {
|
|
@@ -1140,7 +1143,11 @@ export function parsePublishableReview(text: string): PublishableReviewParseResu
|
|
|
1140
1143
|
if (!new Set(["patch is correct", "patch is incorrect"]).has(String(value.overall_correctness))) {
|
|
1141
1144
|
return { error: "overall_correctness is invalid" };
|
|
1142
1145
|
}
|
|
1143
|
-
if (
|
|
1146
|
+
if (
|
|
1147
|
+
value.overall_confidence_score !== undefined
|
|
1148
|
+
? !isConfidence(value.overall_confidence_score)
|
|
1149
|
+
: !options.allowMissingOverallConfidence
|
|
1150
|
+
) {
|
|
1144
1151
|
return { error: "overall_confidence_score is invalid" };
|
|
1145
1152
|
}
|
|
1146
1153
|
return { review: value as unknown as ReviewLike, source: envelope.source };
|
|
@@ -1421,7 +1428,10 @@ export function canonicalReviewSnapshot(review: ReviewLike): PublishableReviewPa
|
|
|
1421
1428
|
if (typeof serialized !== "string") {
|
|
1422
1429
|
return { error: "review could not be serialized for publication" };
|
|
1423
1430
|
}
|
|
1424
|
-
|
|
1431
|
+
// Host-synthesized canonical Markdown has no manufactured overall score,
|
|
1432
|
+
// but every parsed finding still requires its validated numeric confidence.
|
|
1433
|
+
// Assistant-authored strict JSON requires both at its public parse boundary.
|
|
1434
|
+
return parsePublishableReview(serialized, { allowMissingOverallConfidence: true });
|
|
1425
1435
|
}
|
|
1426
1436
|
|
|
1427
1437
|
export function validateInlineComments(
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-pr-review",
|
|
3
|
-
"version": "1.17.
|
|
3
|
+
"version": "1.17.4",
|
|
4
4
|
"description": "Parallel AI code review for GitHub pull requests in the Pi coding agent, with model-agnostic tiered subagents, structured findings, optional verification, and host-gated COMMENT or qualified APPROVE publishing.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"pi-package",
|
package/prompts/pr-review.md
CHANGED
|
@@ -222,6 +222,7 @@ Return Markdown using these stable headings. Do not emit a GitHub API payload an
|
|
|
222
222
|
### [P1] Imperative finding title
|
|
223
223
|
**Severity:** P1
|
|
224
224
|
**Rationale:** Concise Markdown explaining the trigger, impact, and source-grounded evidence.
|
|
225
|
+
**Confidence:** 0.92
|
|
225
226
|
**Location:** `path/to/file.ts:10-12 RIGHT`
|
|
226
227
|
|
|
227
228
|
## Lane completeness
|
|
@@ -231,6 +232,6 @@ Return Markdown using these stable headings. Do not emit a GitHub API payload an
|
|
|
231
232
|
<useful strengths plus concise correctness, security, or performance notes>
|
|
232
233
|
```
|
|
233
234
|
|
|
234
|
-
Repeat the `###` finding block for every confirmed finding. Omit `Location` when no safe diff location exists. Use `No findings.` under `## Findings` when empty. The extension retains the complete raw synthesis even when some or all finding blocks cannot be parsed. Canonical heading discovery ignores heading-like content in CommonMark fenced-code and HTML-block contexts. It validates safe P0–P3 anchors for inline placement, preserves ambiguous or unparsed substantive content in exactly one sanitized body-only `COMMENT`, and degrades malformed synthesis the same way; if synthesis is absent, it assembles a deterministic body-only review from retained lane artifacts. Optional formatting repair can never block this fallback.
|
|
235
|
+
Repeat the `###` finding block for every confirmed finding. `Confidence` is the independently validated numeric confidence for that finding from 0.0 through 1.0; do not use a fixed default. Omit `Location` when no safe diff location exists. Use `No findings.` under `## Findings` when empty. The extension retains the complete raw synthesis even when some or all finding blocks cannot be parsed. Canonical heading discovery ignores heading-like content in CommonMark fenced-code and HTML-block contexts. It validates safe P0–P3 anchors for inline placement, preserves ambiguous or unparsed substantive content in exactly one sanitized body-only `COMMENT`, and degrades malformed synthesis the same way; if synthesis is absent, it assembles a deterministic body-only review from retained lane artifacts. Optional formatting repair can never block this fallback.
|
|
235
236
|
|
|
236
237
|
Host code captures repository, PR, reviewed head, lifecycle, posting authority, stale policy, and invocation identity before review execution, serializes the single GitHub request with `JSON.stringify`, sanitizes reserved markers, enforces payload limits, appends the canonical marker, and performs duplicate/stale/draft/lifecycle/final-head checks. The assistant's semantic `approve` is only one input to the host gates; assistant text cannot directly choose `APPROVE`, `REQUEST_CHANGES`, commit ID, API path, repository, or hostname. Parent candidate validation must gather all independent evidence in one wave and use at most one dependency-driven follow-up, as specified in Step 7.
|