jules-orchestrator-kit 0.69.0 → 0.70.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,23 @@ All notable changes to this project will be documented in this file.
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
 
8
+ ## [0.70.0] - 2026-09-05
9
+ *A session that has not finished is not a session that passed.*
10
+
11
+ An audit of the Jules session layer against the API it talks to. Twelve findings, each traced to a file and line in `docs/jules-quality-plan.md`; the three below are the ones that let the kit believe something about a session that was not true.
12
+
13
+ ### Fixed
14
+ - **The Retry Was Dispatched Without The Failure It Existed To Fix (`src/session-ops.mjs`)**: `retrySession` collected its diagnostics from `act.error`, `act.executionOutput`, `act.exitCode` and `act.status`. None of those four fields exists in the documented `Activity` type — what the API returns is `artifacts[].bashOutput.{command,output,exitCode}` and `sessionFailed.reason`. Measured against a response shaped exactly like the documentation, a `FAILED` session whose bash artifact carries `not ok 1 - invoice totals must round to cents … AssertionError: expected 10.01 to equal 10.00` with `exitCode: 1`, the retry went out with `[PREVIOUS_ATTEMPT_FAILURE_DIAGNOSTIC]` set to *"Previous session did not complete cleanly."* It was handed a sentence when it needed the assertion, the file and the line. `extractFailureDiagnostics` reads the documented fields, returns them highest-signal first so the 4000-character cut drops the least useful evidence rather than the one that failed, deduplicates repeats, and still reads the legacy spellings for provider shapes that have not been observed.
15
+ - **An Unfinished Session Was Reported As COMPLETED (`src/engine.mjs`)**: `pollSessionState` recognised two of the nine documented `SessionState` values and returned `String(session.status || "COMPLETED")` for every other exit — and `session.status` is never set on a fresh dispatch. Measured: a session in `AWAITING_USER_FEEDBACK` and one still `IN_PROGRESS` when the poll budget expired both came back `COMPLETED`, so `QUEUED`, `PLANNING`, `PAUSED` and a session waiting on a human were indistinguishable from success. Terminal, blocked and timed-out are now three distinct verdicts, a provider that stops answering is `unreachable` rather than a timeout describing the wrong thing, and nothing synthesises `COMPLETED` any more. The one path that still returns it is the dry-run simulation, and it is now flagged `simulated`.
16
+
17
+ ### Added
18
+ - **A Contract For The Poll (`test/session-poll.test.mjs`)**: 22 cases across all nine `SessionState` values, the timeout and wall-clock paths, the `approvePlan` side effect granted and refused, the unreachable paths, and the dry-run short-circuit. `pollSessionState` decides whether an agent session is believed to have finished writing — every later gate builds on that — and `grep -rn "pollSession" test/` returned nothing. That is the same class of hole `scripts/guard-reach-check.mjs` exists to close, in the one function that most needed it.
19
+ - **A Green Run Is Not Evidence (`test/session-ops.test.mjs`)**: the opposite failure of a diagnostic collector is one that flags everything, which buries the command that failed under a hundred that passed. `pytest` printing `13 passed`, node:test printing `# fail 0` and `go test` printing `ok` on exit 0 each yield no diagnostics.
20
+
21
+ ### Changed
22
+ - **`agentctl retry` Says When The Trace Is The Fallback (`bin/agentctl.mjs`, `src/session-ops.mjs`)**: `failureReason` alone cannot tell a real trace from the generic sentence — both are non-empty strings. `retrySession` now returns `diagnosticsFound` and `diagnosticSources`, and the CLI prints the count and where the evidence came from, or says plainly that the retry is going out with nothing but the generic line.
23
+ - **A Non-Terminal Session Is Announced Before The Gate Runs (`src/engine.mjs`)**: the repair loop polled the session and discarded the answer, then ran re-verification against a tree the agent might not have finished writing. The verdict is now read: a non-terminal session prints `[SESSION_NOT_TERMINAL]` naming what it is waiting on and appends a `session_not_terminal` telemetry event. The gate still runs either way — it is the authority on whether the change works — but it no longer runs silently on a half-applied patch.
24
+
8
25
  ## [0.69.0] - 2026-09-04
9
26
  *A denominator is not evidence if the things counted in it were never read.*
10
27
 
package/README.md CHANGED
@@ -208,7 +208,7 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
208
208
  * **Fail-Closed Security & Secret Redaction:** Evaluates explicit Deny rules before Allow rules against canonicalized, case-folded paths. Redacts high-entropy keys and base64-encoded credentials (such as Kubernetes `Secret` manifests).
209
209
  * **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, with a `node --check` syntax-verification gate that transparently escalates a FAST-tier result to the primary provider if it left broken JS on disk.
210
210
  * **Terminal UI & Diagnostic Matrix (`agentctl doctor`):** Interactive terminal dashboard, task sidecar manager, and automated transactional self-repair.
211
- * **Verified Test Suite:** Tested with **1223 unit tests across 173 suites**, green on every supported platform.
211
+ * **Verified Test Suite:** Tested with **1255 unit tests across 173 suites**, green on every supported platform.
212
212
 
213
213
  <br/>
214
214
 
package/ROADMAP_V1.md CHANGED
@@ -11,15 +11,21 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
11
11
  ## 📌 Release Milestones Overview
12
12
 
13
13
  ```
14
- v0.69.0 (Current Stable) ──► v0.70.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
15
- (Read, Not Just Counted) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
14
+ v0.70.0 (Current Stable) ──► v0.71.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
15
+ (Not Finished Is Not Passed) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
16
16
  ```
17
17
 
18
18
  ---
19
19
 
20
- ## ✅ Shipped Milestones (v0.20.0 – v0.69.0)
20
+ ## ✅ Shipped Milestones (v0.20.0 – v0.70.0)
21
21
 
22
22
 
23
+ ### v0.70.0: Not Finished Is Not Passed
24
+ - [x] **An Unfinished Session Is Not COMPLETED (`src/engine.mjs`)** — terminal, blocked and timed-out are three verdicts, not one.
25
+ - [x] **The Retry Carries The Failure (`src/session-ops.mjs`)** — it was reading four fields the API does not return.
26
+ - [x] **`agentctl retry` Says When The Trace Is The Fallback (`bin/agentctl.mjs`)** — a generic sentence must not look like evidence.
27
+ - [x] **A Contract For The Poll (`test/session-poll.test.mjs`)** — 22 cases over all nine documented session states.
28
+
23
29
  ### v0.69.0: Read, Not Just Counted
24
30
  - [x] **Expected-Value-First Assertions (`src/security.mjs`)** — JUnit and PHPUnit document the order the guard read as prose.
25
31
  - [x] **A Regex Is An Expected Value (`src/security.mjs`)** — a rewritten pattern was neither a change nor a loss.
package/bin/agentctl.mjs CHANGED
@@ -959,7 +959,13 @@ async function main() {
959
959
  console.log(`------------------------------------------------------------------`);
960
960
  console.log(` Original Session : ${res.originalSessionId}`);
961
961
  console.log(` New Session ID : ${res.newSession?.id || "N/A"}`);
962
- console.log(` Failure Trace Added : ${res.failureReason ? "YES" : "NO"}`);
962
+ console.log(
963
+ ` Failure Trace Added : ${
964
+ res.diagnosticsFound > 0
965
+ ? `YES (${res.diagnosticsFound} block${res.diagnosticsFound === 1 ? "" : "s"} from ${[...new Set(res.diagnosticSources || [])].join(", ")})`
966
+ : "NO — the session carried no readable diagnostics; the retry got the generic sentence"
967
+ }`
968
+ );
963
969
  console.log(`------------------------------------------------------------------\n`);
964
970
  }
965
971
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "jules-orchestrator-kit",
3
- "version": "0.69.0",
3
+ "version": "0.70.0",
4
4
  "description": "Zero-dependency safety gatekeeper, test oracle generator, and multi-agent coordination protocol for autonomous coding agents — Google Jules, Claude Code, Codex and Gemini CLI.",
5
5
  "repository": {
6
6
  "type": "git",
package/src/engine.mjs CHANGED
@@ -894,19 +894,49 @@ export async function repair(failure, opts = {}) {
894
894
  break;
895
895
  }
896
896
 
897
- attempts.push({ n, session, ok: true });
898
- appendTelemetry(root, "ooda_repair_attempt", { attempt: n, ok: true });
899
-
900
897
  // Poll async provider for terminal session state before executing re-verification gates
898
+ let pollVerdict = null;
901
899
  if (provider && session) {
902
- await pollSessionState(provider, session, {
900
+ pollVerdict = await pollSessionState(provider, session, {
903
901
  root,
904
902
  dryRun: opts.dryRun,
905
903
  pollIntervalMs: opts.pollIntervalMs,
906
904
  maxPollAttempts: opts.maxPollAttempts,
907
905
  });
906
+
907
+ // The gate below is the authority on whether the change works, so it runs
908
+ // either way. But a non-terminal session means it is about to judge a
909
+ // tree the agent may not have finished writing, and saying nothing here
910
+ // turns that into a confusing cascade of repair attempts against a
911
+ // half-applied patch.
912
+ if (pollVerdict && pollVerdict.terminal !== true) {
913
+ const why = pollVerdict.blockedOn
914
+ ? `waiting on an actor (${pollVerdict.blockedOn})`
915
+ : pollVerdict.unreachable
916
+ ? "the provider stopped answering"
917
+ : `still ${pollVerdict.status} when the poll budget ran out`;
918
+ console.warn(
919
+ `[SESSION_NOT_TERMINAL] Session ${session.id} is ${why}. Re-verification is running against a tree the agent may not have finished writing.`
920
+ );
921
+ appendTelemetry(root, "session_not_terminal", {
922
+ attempt: n,
923
+ sessionId: session.id,
924
+ status: pollVerdict.status,
925
+ blockedOn: pollVerdict.blockedOn || null,
926
+ timedOut: Boolean(pollVerdict.timedOut),
927
+ unreachable: Boolean(pollVerdict.unreachable),
928
+ });
929
+ }
908
930
  }
909
931
 
932
+ attempts.push({
933
+ n,
934
+ session,
935
+ ok: true,
936
+ poll: pollVerdict ? { status: pollVerdict.status, terminal: pollVerdict.terminal } : null,
937
+ });
938
+ appendTelemetry(root, "ooda_repair_attempt", { attempt: n, ok: true });
939
+
910
940
  // Re-verify after repair attempt
911
941
  const gateRes = await gate({ root, config, fix: false, progressBus, progressToken });
912
942
  if (gateRes.ok) {
@@ -1000,10 +1030,38 @@ export async function checkTaskPremise(task = {}, opts = {}) {
1000
1030
  }
1001
1031
 
1002
1032
  /**
1003
- * Polls the provider until the session terminates in COMPLETED, FAILED, or reaches timeout.
1033
+ * Session states the API documents as final. The full `SessionState` enum is
1034
+ * transcribed in `docs/jules-quality-plan.md`.
1035
+ */
1036
+ export const TERMINAL_SESSION_STATES = new Set(["COMPLETED", "FAILED"]);
1037
+
1038
+ /**
1039
+ * Session states that cannot advance without an actor — a human approving a
1040
+ * plan, a human answering a question, or whoever paused it resuming it.
1041
+ *
1042
+ * Polling these is not waiting, it is spending the budget: the state cannot
1043
+ * change on its own no matter how long the loop runs.
1044
+ */
1045
+ export const BLOCKING_SESSION_STATES = new Set(["AWAITING_PLAN_APPROVAL", "AWAITING_USER_FEEDBACK", "PAUSED"]);
1046
+
1047
+ /**
1048
+ * Polls the provider until the session reaches a terminal state, blocks on an
1049
+ * actor, or the poll budget runs out.
1050
+ *
1051
+ * The verdict is explicit about which of those three happened, because the
1052
+ * caller runs the verification gate on the assumption that the agent has
1053
+ * finished writing. A previous version returned `COMPLETED` for every
1054
+ * non-terminal exit — a session sitting in `AWAITING_USER_FEEDBACK` and one
1055
+ * still `IN_PROGRESS` when the budget expired were both reported as success,
1056
+ * which is the same shape this project exists to refuse.
1057
+ *
1058
+ * @returns {Promise<object>} Always carries `status` and `terminal`. Non-terminal
1059
+ * exits additionally carry one of `blockedOn`, `timedOut` or `unreachable`.
1060
+ * `status` is never synthesised: it is the last state the provider reported,
1061
+ * or `UNKNOWN` when it never answered.
1004
1062
  */
1005
1063
  export async function pollSessionState(provider, session, opts = {}) {
1006
- if (!session || !session.id) return { status: "COMPLETED" };
1064
+ if (!session || !session.id) return { status: "UNKNOWN", terminal: false, unpolled: true };
1007
1065
  const initialStatus = String(session.status || session.state || "").toUpperCase();
1008
1066
  if (
1009
1067
  initialStatus === "COMPLETED" ||
@@ -1013,13 +1071,20 @@ export async function pollSessionState(provider, session, opts = {}) {
1013
1071
  session.id.startsWith("mock-") ||
1014
1072
  session.id.startsWith("dry-run-")
1015
1073
  ) {
1016
- return { status: initialStatus || "COMPLETED" };
1074
+ // A dry run has no session to watch, so it reports the outcome a real
1075
+ // dispatch would have had to earn. Flagged as simulated so a caller can
1076
+ // tell the two apart; a terminal state already reported by the provider is
1077
+ // a fact, not a simulation, and is returned as itself.
1078
+ const status = initialStatus || "COMPLETED";
1079
+ return { status, terminal: TERMINAL_SESSION_STATES.has(status), simulated: true, polls: 0 };
1017
1080
  }
1018
1081
 
1019
1082
  const maxAttempts = opts.maxPollAttempts || 30;
1020
1083
  const pollIntervalMs = opts.pollIntervalMs || 1000;
1021
1084
  const timeoutMs = opts.pollTimeoutMs || 300000;
1022
1085
  const startTime = Date.now();
1086
+ let lastStatus = "";
1087
+ let polls = 0;
1023
1088
 
1024
1089
  for (let attempt = 0; attempt < maxAttempts; attempt++) {
1025
1090
  if (Date.now() - startTime > timeoutMs) break;
@@ -1035,30 +1100,59 @@ export async function pollSessionState(provider, session, opts = {}) {
1035
1100
  } catch (_) {}
1036
1101
  }
1037
1102
 
1038
- if (currentSession) {
1039
- const status = String(currentSession.status || currentSession.state || "").toUpperCase();
1040
- if (
1041
- (status === "AWAITING_PLAN_APPROVAL" || status === "PENDING_APPROVAL") &&
1042
- (opts.autoApprovePlan || opts.autoApprove || session.autoApprovePlan)
1043
- ) {
1044
- if (provider && typeof provider.approvePlan === "function") {
1045
- try {
1046
- await provider.approvePlan(session.id, opts);
1047
- } catch (_) {}
1103
+ if (!currentSession) {
1104
+ return {
1105
+ ...session,
1106
+ status: lastStatus || "UNKNOWN",
1107
+ terminal: false,
1108
+ unreachable: true,
1109
+ polls,
1110
+ };
1111
+ }
1112
+
1113
+ polls += 1;
1114
+ const status = String(currentSession.status || currentSession.state || "").toUpperCase();
1115
+ if (status) lastStatus = status;
1116
+
1117
+ if (TERMINAL_SESSION_STATES.has(status)) {
1118
+ return { ...currentSession, status, terminal: true, polls };
1119
+ }
1120
+
1121
+ if (BLOCKING_SESSION_STATES.has(status)) {
1122
+ // A pending plan approval is the one blocking state this loop is allowed
1123
+ // to resolve itself, and only when the caller said so.
1124
+ const isPlanApproval = status === "AWAITING_PLAN_APPROVAL";
1125
+ const wantsAutoApprove = Boolean(opts.autoApprovePlan || opts.autoApprove || session.autoApprovePlan);
1126
+ if (isPlanApproval && wantsAutoApprove && provider && typeof provider.approvePlan === "function") {
1127
+ try {
1128
+ await provider.approvePlan(session.id, opts);
1129
+ await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
1130
+ continue;
1131
+ } catch (err) {
1132
+ return {
1133
+ ...currentSession,
1134
+ status,
1135
+ terminal: false,
1136
+ blockedOn: status,
1137
+ approvePlanError: err && err.message ? err.message : String(err),
1138
+ polls,
1139
+ };
1048
1140
  }
1049
1141
  }
1050
1142
 
1051
- if (status === "COMPLETED" || status === "FAILED") {
1052
- return { ...currentSession, status };
1053
- }
1054
- } else {
1055
- break;
1143
+ return { ...currentSession, status, terminal: false, blockedOn: status, polls };
1056
1144
  }
1057
1145
 
1058
1146
  await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
1059
1147
  }
1060
1148
 
1061
- return { ...session, status: String(session.status || "COMPLETED").toUpperCase() };
1149
+ return {
1150
+ ...session,
1151
+ status: lastStatus || "UNKNOWN",
1152
+ terminal: false,
1153
+ timedOut: true,
1154
+ polls,
1155
+ };
1062
1156
  }
1063
1157
 
1064
1158
  function buildRepairPrompt(failure, attempt, _config, extraPromptDirective = null) {
@@ -31,6 +31,106 @@ export function parseAgeDuration(duration) {
31
31
  }
32
32
  }
33
33
 
34
+ /**
35
+ * Output shapes that mean "this command failed" even when the runner exited 0.
36
+ *
37
+ * Kept deliberately narrow: this list is only consulted for a `bashOutput`
38
+ * whose `exitCode` is `0` or absent, where the alternative is reporting
39
+ * nothing at all. A false positive here costs a few hundred characters of
40
+ * retry prompt; a false negative costs the retry its only evidence.
41
+ */
42
+ const BASH_FAILURE_HINTS = [
43
+ /(^|\n)\s*not ok\b/i,
44
+ /(^|\n)# fail [1-9]/i,
45
+ /\b[1-9]\d* failed\b/i,
46
+ /\b[1-9]\d* failing\b/i,
47
+ /\bAssertionError\b/,
48
+ /\bFAILED\b/,
49
+ /\bTraceback \(most recent call last\)/,
50
+ /\bpanic: /,
51
+ /\berror:/i,
52
+ ];
53
+
54
+ /**
55
+ * Collects the diagnostics a session actually carries.
56
+ *
57
+ * The documented Activity type — the Jules API types reference, transcribed
58
+ * in `docs/jules-quality-plan.md` — puts command output under
59
+ * `artifacts[].bashOutput.{command,output,exitCode}` and the failure reason
60
+ * under `sessionFailed.reason`. Neither `act.error` nor `act.executionOutput`
61
+ * — the only two fields this file used to read — exists in that schema, so a
62
+ * real failure came back as the generic fallback sentence and the retry session
63
+ * was dispatched without the assertion it existed to fix.
64
+ *
65
+ * Blocks are returned highest-signal first, because `retrySession` truncates
66
+ * the joined result to 4000 characters from the front: the ordering decides
67
+ * which evidence survives the cut.
68
+ *
69
+ * The legacy spellings are still read. They cost nothing, and an unrecognised
70
+ * provider shape that does carry an error is better served by it than by the
71
+ * fallback sentence.
72
+ *
73
+ * @param {Array<object>} activities - Activities as returned by `listActivities`.
74
+ * @returns {Array<{ source: string, text: string }>} Highest-signal first.
75
+ */
76
+ export function extractFailureDiagnostics(activities) {
77
+ const list = Array.isArray(activities) ? activities : [];
78
+ const failingBash = [];
79
+ const failureReasons = [];
80
+ const suspiciousBash = [];
81
+ const legacy = [];
82
+ const agentNotes = [];
83
+ const seen = new Set();
84
+
85
+ const push = (bucket, source, text) => {
86
+ const clean = String(text ?? "").trim();
87
+ if (!clean) return;
88
+ const key = `${source}\u0000${clean}`;
89
+ if (seen.has(key)) return;
90
+ seen.add(key);
91
+ bucket.push({ source, text: clean });
92
+ };
93
+
94
+ for (const act of list) {
95
+ if (!act || typeof act !== "object") continue;
96
+
97
+ const artifacts = Array.isArray(act.artifacts) ? act.artifacts : [];
98
+ for (const art of artifacts) {
99
+ const bash = art && typeof art === "object" ? art.bashOutput : null;
100
+ if (!bash || typeof bash !== "object") continue;
101
+ const command = typeof bash.command === "string" ? bash.command : "";
102
+ const output = typeof bash.output === "string" ? bash.output : "";
103
+ if (!command && !output) continue;
104
+ const rendered = `$ ${command || "(unknown command)"}\n${output || "(no output)"}`;
105
+ const exitCode = bash.exitCode;
106
+ if (typeof exitCode === "number" && exitCode !== 0) {
107
+ push(failingBash, "bashOutput", `${rendered}\n(exit code ${exitCode})`);
108
+ } else if (BASH_FAILURE_HINTS.some((re) => re.test(output))) {
109
+ const codeNote = typeof exitCode === "number" ? String(exitCode) : "not reported";
110
+ push(suspiciousBash, "bashOutput", `${rendered}\n(exit code ${codeNote})`);
111
+ }
112
+ }
113
+
114
+ const failed = act.sessionFailed && typeof act.sessionFailed === "object" ? act.sessionFailed : null;
115
+ if (failed && typeof failed.reason === "string") {
116
+ push(failureReasons, "sessionFailed", failed.reason);
117
+ }
118
+
119
+ if (act.error) {
120
+ push(legacy, "legacy.error", typeof act.error === "string" ? act.error : JSON.stringify(act.error));
121
+ }
122
+ if (act.executionOutput && (act.exitCode !== 0 || act.status === "FAILED")) {
123
+ push(legacy, "legacy.executionOutput", act.executionOutput);
124
+ }
125
+
126
+ if (act.originator === "agent" && typeof act.description === "string") {
127
+ push(agentNotes, "agentMessage", act.description);
128
+ }
129
+ }
130
+
131
+ return [...failingBash, ...failureReasons, ...suspiciousBash, ...legacy, ...agentNotes];
132
+ }
133
+
34
134
  /**
35
135
  * Extracts git diff patch, PR details, and modified files from a Jules session.
36
136
  *
@@ -213,7 +313,7 @@ export async function applySessionPatch(sessionId, opts = {}) {
213
313
  *
214
314
  * @param {string} sessionId
215
315
  * @param {object} opts
216
- * @returns {Promise<{ ok: boolean, originalSessionId: string, newSession: object, failureReason: string }>}
316
+ * @returns {Promise<{ ok: boolean, originalSessionId: string, newSession: object, failureReason: string, diagnosticsFound: number, diagnosticSources: string[] }>}
217
317
  */
218
318
  export async function retrySession(sessionId, opts = {}) {
219
319
  if (!sessionId || typeof sessionId !== "string") {
@@ -227,18 +327,19 @@ export async function retrySession(sessionId, opts = {}) {
227
327
  const activitiesRes = await provider.listActivities(sessionId, opts).catch(() => ({ activities: [] }));
228
328
  const activities = activitiesRes.activities || [];
229
329
 
230
- // Extract failure diagnostics
231
- const errors = [];
232
- for (const act of activities) {
233
- if (act.error) errors.push(typeof act.error === "string" ? act.error : JSON.stringify(act.error));
234
- if (act.executionOutput && (act.exitCode !== 0 || act.status === "FAILED")) {
235
- errors.push(act.executionOutput);
236
- }
237
- }
330
+ // Extract failure diagnostics. `extractFailureDiagnostics` reads the fields
331
+ // the API actually documents (artifacts[].bashOutput, sessionFailed.reason)
332
+ // and returns them highest-signal first, so the 4000-character cut below
333
+ // drops the least useful evidence rather than the assertion that failed.
334
+ const diagnostics = extractFailureDiagnostics(activities);
238
335
 
239
336
  const rawPrompt = session?.raw?.prompt || session?.prompt || "";
240
337
  const title = session?.raw?.title || `Retry of Session ${sessionId}`;
241
- const failureReason = errors.join("\n").slice(0, 4000) || "Previous session did not complete cleanly.";
338
+ const failureReason =
339
+ diagnostics
340
+ .map((d) => d.text)
341
+ .join("\n")
342
+ .slice(0, 4000) || "Previous session did not complete cleanly.";
242
343
 
243
344
  let synthesizedPrompt = rawPrompt;
244
345
  if (opts.withFailure !== false && failureReason) {
@@ -264,6 +365,12 @@ export async function retrySession(sessionId, opts = {}) {
264
365
  originalSessionId: sessionId,
265
366
  newSession: dispatchRes,
266
367
  failureReason,
368
+ // Non-zero only when the session carried evidence the retry could act on.
369
+ // A zero here with a non-empty `failureReason` means the fallback sentence
370
+ // was sent — the distinction the CLI and any telemetry need, because the
371
+ // two look identical in `failureReason` alone.
372
+ diagnosticsFound: diagnostics.length,
373
+ diagnosticSources: diagnostics.map((d) => d.source),
267
374
  };
268
375
  }
269
376