@try-works/dsh-recursive-mode 0.4.7 → 0.4.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -386,10 +386,17 @@ from `session/event`, with a cheap shape test first, because that event fires fo
386
386
  observer **waits** — bounded, polling, with a timeout — and reports "no settlement yet" only after actually
387
387
  waiting. The park remains as the fallback, so an unobserved round is never an approval.
388
388
 
389
- **Action records are written for every delegation**, and they carry `Execution Mode` and `Status`. After a
390
- failed delegation, they now also carry **`Failure:`** — because `Status: failed` on its own cannot distinguish a
391
- child that ran and failed from a child that never started. That distinction cost three rounds of investigation
392
- before the field existed.
389
+ **Action records are written for every delegation**, and they carry `Execution Mode` and a `Status` with
390
+ **three** values, not two: `accepted`, `failed`, and `parked (still running; no settlement yet)`. The third
391
+ state exists because a round whose child had not settled is **not** a failure — the delegation interface says
392
+ so, and `recursive_review` already reports it as `parked` — and a binary status forced it to read as one: a live
393
+ run's record said `Status: failed` with *"the child never reported, or never ran"*, the main agent concluded its
394
+ reviewer was dead and obtained the review another way, and the child replied eighteen minutes later. A `failed`
395
+ record carries **`Failure:`** with its cause; a parked record carries **`Parked:`** instead — stating only what
396
+ is known ("no settlement had landed when the wait ended … the child may still be working"), naming the
397
+ `childId`, and saying that resuming with that id is the next step. `operations/operations.jsonl` records the
398
+ same distinction (`parked` rather than `unaccepted`), because the operation log had the identical conflation.
399
+ That distinction cost three rounds of investigation before the third state existed.
393
400
 
394
401
  ---
395
402
 
@@ -336,9 +336,40 @@ export interface ActionRecordInput {
336
336
  * nothing else: a delegation that FAILED and a delegation that NEVER HAPPENED read identically, which is what
337
337
  * let me conclude for three rounds that the host was not scheduling children. The caller ALREADY passed
338
338
  * `stopReason`, and the live record said `n/a` — because there was no result to take a stop reason from.
339
+ *
340
+ * ⚠ AND IT IS EMITTED ONLY FOR A GENUINE FAILURE. A PARKED round is neither accepted nor failed, so it
341
+ * carries {@link ActionRecordInput.parked} instead: calling a live child a failure is the defect this
342
+ * field's own history is made of.
339
343
  */
340
344
  failure?: string;
345
+ /**
346
+ * ⚠ THE THIRD STATE, AND WHY `success` COULD NOT CARRY IT.
347
+ *
348
+ * A continuable round that has not settled is NOT a failure — `ContinuableDelegationLike.parked` says so in
349
+ * its own doc comment, and the tool result already surfaces it. The RECORD did not: a live run's child was
350
+ * parked, the record said `Status: failed` with "the child never reported, or never ran", the main agent
351
+ * read that as a dead child and obtained the review elsewhere — while the child was still working and
352
+ * replied eighteen minutes later. `success: false` cannot express "still in flight", because a genuinely
353
+ * dead child also produces `success: false`, so a second field is required rather than a cleverer boolean.
354
+ *
355
+ * When true, the record says `parked` and carries a `Parked:` line (never a `Failure:` line) whose text
356
+ * states only what is KNOWN, names the `childId`, and names the resume step. It takes precedence over
357
+ * `success`: an unobserved round is never an acceptance, whatever a caller passes alongside it.
358
+ */
359
+ parked?: boolean;
341
360
  }
361
+ /**
362
+ * The status a record states, in ONE place, because the three states are the fix and a second writer would
363
+ * drift from this one.
364
+ *
365
+ * The wording is chosen for two readers at once. The MODEL reads it to decide whether to resume or to give up
366
+ * and obtain the result another way — the exact decision the live defect got wrong — so the token must not
367
+ * read as a failure and must not need the rest of the document to be understood. A HUMAN reading the run tree
368
+ * months later needs to tell "died" from "still working" at a glance. Hence `parked (still running; no
369
+ * settlement yet)`: a third token rather than a renamed second, qualified with the two facts that separate it
370
+ * from `failed`, and short enough to sit in a status line.
371
+ */
372
+ export declare function actionRecordStatus(input: Pick<ActionRecordInput, 'success' | 'parked'>): string;
342
373
  /**
343
374
  * Write a durable action record under subagents/ in the shape this repo's own
344
375
  * linter accepts (ts-lint.ts lintSubagentActionRecordFile — every top-level .md
@@ -351,7 +382,10 @@ export interface ActionRecordInput {
351
382
  * `Code Refs` strictly inside ## Inputs Provided. The linter resolves each of
352
383
  * those through the heading body, so a field under another heading is not
353
384
  * found at all.
354
- * A success:false attempt is written with a failed status and is NOT accepted.
385
+ * A success:false attempt is written with a failed status and is NOT accepted. A round that PARKED is written
386
+ * with a status of its own (`parked (still running; no settlement yet)`, see {@link actionRecordStatus}) and a
387
+ * `Parked:` line instead of a `Failure:` one, because no settlement is not a death: it is the caller's signal to
388
+ * resume the same child. Only a genuine failure carries `Failure:`.
355
389
  */
356
390
  export declare function writeActionRecord(input: ActionRecordInput): string;
357
391
  /**
@@ -55,8 +55,14 @@ export declare const DEFAULT_ENFORCEMENT: EnforcementConfig;
55
55
  * every allow/deny/ask the guard hands back). `'none'` means no guard predicate
56
56
  * fired; `'transition'` marks a decision whose only dissenting gate was the
57
57
  * advisory transition gate (see consultTransitionGate).
58
+ *
59
+ * `'phase-order'` is the WRITE half of the ordering rule the owner states as *only
60
+ * one phase may be active at a time, and the phases should be sequential and the
61
+ * active phase must be locked before proceeding to next phase*: `'lock-order'` refuses
62
+ * locking ahead, `'phase-order'` refuses WRITING ahead. Two labels rather than one,
63
+ * because the guard log has to tell the owner which of the two the agent attempted.
58
64
  */
59
- export type GuardRule = 'lock-order' | 'tdd-evidence' | 'locked-write' | 'transition' | 'none';
65
+ export type GuardRule = 'lock-order' | 'phase-order' | 'tdd-evidence' | 'locked-write' | 'transition' | 'none';
60
66
  /** T15: the transition gate's verdict as attached to a decision (advisory only). */
61
67
  export interface GuardTransition {
62
68
  passed: boolean;
@@ -108,7 +114,7 @@ export interface ToolExecLike {
108
114
  * `recursive:policy` later — can read the effective rules instead of inferring
109
115
  * them from a code path.
110
116
  */
111
- export declare function resolveToolPolicyForGuard(worktreeRoot: string, runId: string): ToolPolicy;
117
+ export declare function resolveToolPolicyForGuard(worktreeRoot: string, runId: string, activePhaseArtifact?: string): ToolPolicy;
112
118
  /**
113
119
  * The artifact whose phase baseline applies: the HIGHEST-numbered phase artifact
114
120
  * present in the run (a run at phase 3 has `00`-`03` on disk). Read from the
package/lib/index.js CHANGED
@@ -1512,11 +1512,91 @@ function lockedWriteRule(target, worktreeRoot) {
1512
1512
  };
1513
1513
  }
1514
1514
  /**
1515
+ * The file name when `abs` is a DIRECT CHILD of `runDir`, and `null` otherwise.
1516
+ *
1517
+ * ⚠ THE DIRECT-CHILD TEST IS NOT TIDINESS — IT IS WHAT KEEPS THE RULE OFF THE SUPPORT
1518
+ * FILES. `phaseNumberForArtifact` reads the leading digits of a NAME, so
1519
+ * `<run>/evidence/01-as-is.md` or `<run>/subagents/child/03-brief.md` would look like
1520
+ * phase 1 and phase 3 artifacts if the name were all that was examined. A phase artifact
1521
+ * is a file the run tree holds DIRECTLY beside the others (`recursive_init` writes all
1522
+ * twelve into `<run>/` itself), so the parent directory is part of the definition.
1523
+ *
1524
+ * The comparison normalizes separators and case: the same run directory reached through
1525
+ * a Windows spelling that differs in case is the same directory, and the rule must not
1526
+ * abstain on one spelling and fire on the other.
1527
+ */
1528
+ function directChildName(abs, runDir) {
1529
+ const normalize = (path) => resolve(path).replace(/\\/g, "/").replace(/\/+$/, "").toLowerCase();
1530
+ if (normalize(dirname(abs)) !== normalize(runDir)) return null;
1531
+ return basename(abs);
1532
+ }
1533
+ /**
1534
+ * PHASE-ORDER rule: a denial when the target is a LATER phase's artifact than the phase
1535
+ * currently active, `null` when it is not.
1536
+ *
1537
+ * ⚠ THE HOLE THIS CLOSES. The monotonic rule was enforced on `recursive_lock` ONLY. An
1538
+ * agent could therefore write `08-memory-impact.md` while the run sat at phase 0 — and a
1539
+ * live run did exactly that: twelve artifacts, not one of them locked, written out of
1540
+ * order, with a single line in `operations/operations.jsonl`. Ordering that only binds
1541
+ * the lock tool is not ordering; the model's ordinary `write` is the path that mattered.
1542
+ *
1543
+ * The owner's rule, verbatim: *only one phase may be active at a time, and the phases
1544
+ * should be sequential and the active phase must be locked before proceeding to next
1545
+ * phase*. The ACTIVE phase is the one the selector names (`ctx.activePhaseArtifact`), so:
1546
+ *
1547
+ * - the target is the ACTIVE artifact, or shares its phase number (`00-requirements.md`
1548
+ * and `00-worktree.md` are both phase 0; `01-as-is.md` and `01.5-root-cause.md` are
1549
+ * both phase 1) -> ABSTAIN, the write is allowed;
1550
+ * - the target is an EARLIER phase -> ABSTAIN. Such an artifact is LOCKED by
1551
+ * construction (the active phase is the lowest UNLOCKED one), so the locked-artifact
1552
+ * rule above decides it, and its rule label is preserved;
1553
+ * - the target is a LATER phase -> DENY: working ahead.
1554
+ *
1555
+ * ⚠ THE ALLOW HALF IS LOAD-BEARING. An enforcement rule in this exact area was once the
1556
+ * bug: strict enforcement denied the run's OWN artifacts in every phase and made the
1557
+ * workflow unusable (see `resolveFrom` and `currentPhaseArtifact`). "The active artifact
1558
+ * stays writable at every phase" is therefore asserted by walking every phase, not by
1559
+ * one case — `tests/strict-run-tree.spec.ts` (d).
1560
+ *
1561
+ * WHAT IT ABSTAINS ON, deliberately:
1562
+ * - no `activePhaseArtifact` (a caller with no run context, or a run with no phase
1563
+ * artifacts yet) -> abstain, never guess a phase;
1564
+ * - a target outside the active run's own directory -> abstain. The rule is about THIS
1565
+ * run's sequence; another run's tree is a different question and denying it here
1566
+ * would be a false positive;
1567
+ * - a support file (anything not a direct child) -> abstain: `evidence/`, `scratch/`,
1568
+ * `addenda/`, `subagents/`, `operations/` and a plain `<run>/notes.md` are not phases.
1569
+ *
1570
+ * The predicate adds only the per-call particular — which artifact, which phase, which is
1571
+ * active; the rule keeps the static, auditable sentence.
1572
+ */
1573
+ function phaseOrderRule(target, ctx) {
1574
+ const active = ctx.activePhaseArtifact;
1575
+ if (!target || !ctx.runDir || !ctx.worktreeRoot) return null;
1576
+ if (typeof active !== "string" || active === "") return null;
1577
+ const activePhaseText = phaseNumberForArtifact(active);
1578
+ if (!activePhaseText) return null;
1579
+ const normalized = target.replace(/\\/g, "/");
1580
+ if (!normalized.endsWith(".md")) return null;
1581
+ const abs = resolveFrom(ctx.worktreeRoot, normalized);
1582
+ if (!abs) return null;
1583
+ const name = directChildName(abs, ctx.runDir);
1584
+ if (name === null) return null;
1585
+ const phaseText = phaseNumberForArtifact(name);
1586
+ if (!phaseText) return null;
1587
+ if (Number(phaseText) <= Number(activePhaseText)) return null;
1588
+ return {
1589
+ verdict: "deny",
1590
+ detail: name + " is phase " + phaseText + " but the ACTIVE phase is " + active + " (phase " + activePhaseText + ") - the active phase must be locked before writing a later phase"
1591
+ };
1592
+ }
1593
+ /**
1515
1594
  * The BUILT-IN default rule list — the pre-T16 guard behaviour expressed as
1516
1595
  * data:
1517
1596
  *
1518
1597
  * `recursive_lock*` -> the monotonic lock-order denial;
1519
1598
  * the write-tool ids -> the locked-artifact write denial;
1599
+ * the write-tool ids -> the phase-order (write-ahead) denial;
1520
1600
  * `*` -> allow, so an ordinary tool is not turned into an `ask`.
1521
1601
  *
1522
1602
  * Every `deny` precedes the `allow`, which is what makes "deny wins over allow" a
@@ -1543,6 +1623,13 @@ function builtInToolPolicyRules() {
1543
1623
  label: "locked-write",
1544
1624
  predicate: (id, args, ctx) => WRITE_TOOL_NAMES.has(id) ? lockedWriteRule(policyTargetPath(args), ctx.worktreeRoot) : null
1545
1625
  });
1626
+ for (const name of WRITE_TOOL_NAMES) rules.push({
1627
+ pattern: name,
1628
+ verdict: "deny",
1629
+ reason: "phase order: only one phase may be active at a time - the active phase must be locked before a later phase artifact is written",
1630
+ label: "phase-order",
1631
+ predicate: (id, args, ctx) => WRITE_TOOL_NAMES.has(id) ? phaseOrderRule(policyTargetPath(args), ctx) : null
1632
+ });
1546
1633
  rules.push({
1547
1634
  pattern: "*",
1548
1635
  verdict: "allow",
@@ -1581,6 +1668,10 @@ function attachPolicyPredicate(rule) {
1581
1668
  ...rule,
1582
1669
  predicate: (id, args, ctx) => LOCK_TOOL_NAMES.has(id) ? lockOrderRule(args.artifact, ctx.runDir) : null
1583
1670
  };
1671
+ if (rule.label === "phase-order") return {
1672
+ ...rule,
1673
+ predicate: (id, args, ctx) => WRITE_TOOL_NAMES.has(id) ? phaseOrderRule(policyTargetPath(args), ctx) : null
1674
+ };
1584
1675
  if (WRITE_TOOL_NAMES.has(rule.pattern)) return {
1585
1676
  ...rule,
1586
1677
  predicate: (id, args, ctx) => WRITE_TOOL_NAMES.has(id) ? lockedWriteRule(policyTargetPath(args), ctx.worktreeRoot) : null
@@ -7784,8 +7875,8 @@ const DEFAULT_ENFORCEMENT = {
7784
7875
  * `recursive:policy` later — can read the effective rules instead of inferring
7785
7876
  * them from a code path.
7786
7877
  */
7787
- function resolveToolPolicyForGuard(worktreeRoot, runId) {
7788
- return withPhaseBaseline(loadToolPolicyFile(worktreeRoot).policy, currentPhaseArtifact(worktreeRoot, runId));
7878
+ function resolveToolPolicyForGuard(worktreeRoot, runId, activePhaseArtifact) {
7879
+ return withPhaseBaseline(loadToolPolicyFile(worktreeRoot).policy, activePhaseArtifact ?? currentPhaseArtifact(worktreeRoot, runId));
7789
7880
  }
7790
7881
  /**
7791
7882
  * The artifact whose phase baseline applies: the HIGHEST-numbered phase artifact
@@ -7835,11 +7926,13 @@ function evaluateToolGuard(exec, worktreeRoot, activeRunId, mode = "advisory") {
7835
7926
  const runId = typeof activeRunId === "string" ? activeRunId.trim() : "";
7836
7927
  const runDir = join(worktreeRoot, ".recursive", "run", runId);
7837
7928
  const transition = consultTransitionGate(name, args, worktreeRoot, runId);
7838
- return advisory(verdictFor(mode, evaluateToolPolicy(resolveToolPolicyForGuard(worktreeRoot, runId), name, args, {
7929
+ const activePhaseArtifact = currentPhaseArtifact(worktreeRoot, runId);
7930
+ return advisory(verdictFor(mode, evaluateToolPolicy(resolveToolPolicyForGuard(worktreeRoot, runId, activePhaseArtifact), name, args, {
7839
7931
  args,
7840
7932
  runDir,
7841
7933
  runId,
7842
- worktreeRoot
7934
+ worktreeRoot,
7935
+ activePhaseArtifact
7843
7936
  })), transition);
7844
7937
  }
7845
7938
  /**
@@ -8086,6 +8179,7 @@ function renderStableContract(config = DEFAULT_ENFORCEMENT) {
8086
8179
  "- Gates in force: pre-step " + config.preStep + ", tool guards " + config.toolGuards + ", tamper detection " + config.tamper + ".",
8087
8180
  "- A transition that fails its gates is BLOCKED (strict) or warns (advisory); no rejected transition proceeds silently.",
8088
8181
  "- Writes to a Status: LOCKED phase doc are denied/asked; reopen explicitly to edit.",
8182
+ "- Phase order binds WRITES as well as locks: only the ACTIVE phase (the lowest-numbered artifact not yet LOCKED) may be written; a write to a LATER phase artifact is denied/asked. Run support files (evidence/, scratch/, addenda/, subagents/, operations/) are not phases.",
8089
8183
  "- Phase 3 lock requires TDD evidence (strict) or rationale (pragmatic); Phase 5 requires QA evidence.",
8090
8184
  "- The control-plane root is resolved STRICTLY from this session workspace (never scanned from another)."
8091
8185
  ].join("\n");
@@ -8779,7 +8873,7 @@ async function delegateContinuable(input) {
8779
8873
  const observed = await input.awaitRoundResult(childId, lastMessageId);
8780
8874
  if (observed === null) return {
8781
8875
  ok: false,
8782
- reason: "no settlement has landed for round " + (round + 1) + " yet (the child is still working)",
8876
+ reason: "no settlement has landed for round " + (round + 1) + " yet, so the round is PARKED, not failed: the child may still be working. Resume it on a later turn with childId " + String(childId) + " (the `resumeChild` argument) — do not start a second child.",
8783
8877
  childId,
8784
8878
  messageIds,
8785
8879
  rounds,
@@ -8993,6 +9087,21 @@ function validateReferences(root, references) {
8993
9087
  checked
8994
9088
  };
8995
9089
  }
9090
+ /**
9091
+ * The status a record states, in ONE place, because the three states are the fix and a second writer would
9092
+ * drift from this one.
9093
+ *
9094
+ * The wording is chosen for two readers at once. The MODEL reads it to decide whether to resume or to give up
9095
+ * and obtain the result another way — the exact decision the live defect got wrong — so the token must not
9096
+ * read as a failure and must not need the rest of the document to be understood. A HUMAN reading the run tree
9097
+ * months later needs to tell "died" from "still working" at a glance. Hence `parked (still running; no
9098
+ * settlement yet)`: a third token rather than a renamed second, qualified with the two facts that separate it
9099
+ * from `failed`, and short enough to sit in a status line.
9100
+ */
9101
+ function actionRecordStatus(input) {
9102
+ if (input.parked === true) return "parked (still running; no settlement yet)";
9103
+ return input.success ? "accepted" : "failed";
9104
+ }
8996
9105
  function slugify(value) {
8997
9106
  return value.toLowerCase().replace(/[^a-z0-9]+/g, "-").replace(/^-+|-+$/g, "") || "action";
8998
9107
  }
@@ -9008,7 +9117,10 @@ function slugify(value) {
9008
9117
  * `Code Refs` strictly inside ## Inputs Provided. The linter resolves each of
9009
9118
  * those through the heading body, so a field under another heading is not
9010
9119
  * found at all.
9011
- * A success:false attempt is written with a failed status and is NOT accepted.
9120
+ * A success:false attempt is written with a failed status and is NOT accepted. A round that PARKED is written
9121
+ * with a status of its own (`parked (still running; no settlement yet)`, see {@link actionRecordStatus}) and a
9122
+ * `Parked:` line instead of a `Failure:` one, because no settlement is not a death: it is the caller's signal to
9123
+ * resume the same child. Only a genuine failure carries `Failure:`.
9012
9124
  */
9013
9125
  function writeActionRecord(input) {
9014
9126
  const { root, runId } = input;
@@ -9046,8 +9158,8 @@ function writeActionRecord(input) {
9046
9158
  "- Phase: " + input.phase,
9047
9159
  "- Purpose: " + input.purpose,
9048
9160
  "- Execution Mode: " + input.executionMode,
9049
- "- Status: " + (input.success ? "accepted" : "failed"),
9050
- ...input.success === false && input.failure !== void 0 ? ["- Failure: " + input.failure] : [],
9161
+ "- Status: " + actionRecordStatus(input),
9162
+ ...input.success === false && input.failure !== void 0 ? [input.parked === true ? "- Parked: " + input.failure : "- Failure: " + input.failure] : [],
9051
9163
  "- Stop Reason: " + (input.stopReason ?? "n/a"),
9052
9164
  "- Timestamp: " + (/* @__PURE__ */ new Date()).toISOString(),
9053
9165
  "",
@@ -10827,7 +10939,10 @@ var RecursiveRuntime = class extends Service {
10827
10939
  } else if (continuable.ok && continuable.rounds.length > 0) {
10828
10940
  result = continuable.rounds[continuable.rounds.length - 1].result ?? null;
10829
10941
  if (!result) error = "continuable child produced no final result";
10830
- } else error = continuable.reason ?? "continuable delegation failed";
10942
+ } else {
10943
+ result = continuable.rounds[continuable.rounds.length - 1]?.result ?? null;
10944
+ if (result === null) error = continuable.reason ?? "continuable delegation failed";
10945
+ }
10831
10946
  } else try {
10832
10947
  result = await delegate({
10833
10948
  subagents,
@@ -10852,7 +10967,7 @@ var RecursiveRuntime = class extends Service {
10852
10967
  id: operation,
10853
10968
  act: "delegate-review",
10854
10969
  at: (/* @__PURE__ */ new Date()).toISOString().replace(/\.\d{3}Z$/, "Z"),
10855
- outcome: evaluation.accepted ? "accepted" : "unaccepted",
10970
+ outcome: parked ? "parked" : evaluation.accepted ? "accepted" : "unaccepted",
10856
10971
  phase: input.phase
10857
10972
  });
10858
10973
  const actionRecordPath = writeActionRecord({
@@ -10871,7 +10986,8 @@ var RecursiveRuntime = class extends Service {
10871
10986
  findings: evaluation.accepted && result?.structured ? [result.structured?.verdict ?? "accepted"] : void 0,
10872
10987
  success: evaluation.accepted,
10873
10988
  stopReason: result?.stopReason,
10874
- failure: evaluation.accepted ? void 0 : result == null ? input.mode !== "one-shot" ? "the continuable start was made and NO SETTLEMENT arrived within the wait (the child never reported, or never ran); tier " + decision.tier + ", provider " + (decision.provider ?? "none chosen") + ", names on offer [" + (this.lastProviderNames.join(", ") || "none") + "]; parent id " + (input.parent?.id ?? "none") + ", parent session keys [" + (input.parent === void 0 ? "no parent" : Object.keys(input.parent).join(", ")) + "]" : "no delegate result was produced by the one-shot path; tier " + decision.tier + ", provider " + (decision.provider ?? "none chosen") : "the delegation returned without acceptance; stop reason " + (result.stopReason ?? "none reported")
10989
+ ...parked ? { parked: true } : {},
10990
+ failure: evaluation.accepted ? void 0 : parked ? "no settlement had landed when the wait ended, so this round is PARKED, not failed: nothing was accepted and nothing was refused, and the child may still be working. The next step is to RESUME this round, not to re-dispatch it or replace the child: call `recursive_review` again on a later turn with childId " + String(continuable?.childId ?? input.childId) + " (the child the round was started for, which stays resumable). Diagnostics: tier " + decision.tier + ", provider " + (decision.provider ?? "none chosen") + ", names on offer [" + (this.lastProviderNames.join(", ") || "none") + "]; parent id " + (input.parent?.id ?? "none") + ", parent session keys [" + (input.parent === void 0 ? "no parent" : Object.keys(input.parent).join(", ")) + "]" : result == null ? input.mode !== "one-shot" ? "the continuable start was made and NO SETTLEMENT arrived within the wait (the child never reported, or never ran); tier " + decision.tier + ", provider " + (decision.provider ?? "none chosen") + ", names on offer [" + (this.lastProviderNames.join(", ") || "none") + "]; parent id " + (input.parent?.id ?? "none") + ", parent session keys [" + (input.parent === void 0 ? "no parent" : Object.keys(input.parent).join(", ")) + "]" : "no delegate result was produced by the one-shot path; tier " + decision.tier + ", provider " + (decision.provider ?? "none chosen") : "the delegation returned without acceptance; stop reason " + (result.stopReason ?? "none reported") + (result.success === false ? " (the child itself reported success:false)" : "")
10875
10991
  });
10876
10992
  const delegationMode = continuable !== null ? continuable.fellBackToOneShot ? "continuable-unavailable" : "continuable" : result !== null ? "one-shot" : "none";
10877
10993
  return {
@@ -14793,4 +14909,4 @@ function apply(ctx, config) {
14793
14909
  });
14794
14910
  }
14795
14911
  //#endregion
14796
- export { Config, DEFAULT_BUDGETS, DEFAULT_ENFORCEMENT, OPTIONAL_PHASES, PHASES, PHASE_POSITIONS, PHASE_SEQUENCE, RECURSIVE_API_PREFIX, RUN_ARTIFACT_SEQUENCE, RUN_STATES, RecursiveRuntime, apply, auditToPass, buildDelegationPrompt, buildReviewBundle, buildWorkSlice, builtInToolPolicy, capabilityProbe, childScratchPath, coerceAskToDecision, contentSha256, contractDigest, coupleGateBlockToGoal, createChildBrief, createHandoff, createRecursiveCloseoutTool, createRecursiveInitTool, createRecursiveLintTool, createRecursiveLockTool, createRecursivePhaseTool, createRecursiveScratchTool, createRecursiveStatusTool, createRecursiveWorktreeTool, currentPhaseArtifact, defaultReviewToolFilter, delegate, delegateContinuable, delegationDecisionBasis, delegationError, detectTamper, discoverRuns, drainContinuableChildren, drainContinuableDescendants, escapeRegExp, evaluateDelegationResult, evaluateToolGuard, foldDiagnostics, foldRun, foldRunCard, getAllStaleReceipts, getArtifactState, getGateStatus, getLatestRunDirectory, getLockStatus, getMdFieldValue, getNextLegalPhase, getPrerequisiteBlockers, getPrerequisites, getStaleDownstreamPhases, getTodoStats, getWorkflowProfile, inject, interruptContinuable, invalidateReceipt, isCoreArtifact, isTaskClaimedBy, loadRouterPolicy, lockHashFromContent, makeRecursiveRoutes, mountRecursiveRoutesOnce, name, normalizeForLockHash, parseReplyVerdict, pendingWork, phaseIndex, phasePosition, probeCapabilities, readReceipt, readRepairFromReply, readRepairFromStructured, readVerdictFromReply, readVerdictFromStructured, receiptPath, referencesFromResult, registerRecursiveSkill, remainingDepthFor, renderPhaseTail, renderRecursivePolicy, renderStableContract, renderTaskHistory, replyPath, resetFoldCache, resolveEnforcementConfig, resolveRole, resolveRunDir, resolveToolPolicyForGuard, reviewBundleDir, reviewOutputSchema, routerPolicyPath, snapshotWorkspace, tamperCandidatePath, trimMdValue, validateChain, validateReferences, validateTransition, writeActionRecord, writeReceipt };
14912
+ export { Config, DEFAULT_BUDGETS, DEFAULT_ENFORCEMENT, OPTIONAL_PHASES, PHASES, PHASE_POSITIONS, PHASE_SEQUENCE, RECURSIVE_API_PREFIX, RUN_ARTIFACT_SEQUENCE, RUN_STATES, RecursiveRuntime, actionRecordStatus, apply, auditToPass, buildDelegationPrompt, buildReviewBundle, buildWorkSlice, builtInToolPolicy, capabilityProbe, childScratchPath, coerceAskToDecision, contentSha256, contractDigest, coupleGateBlockToGoal, createChildBrief, createHandoff, createRecursiveCloseoutTool, createRecursiveInitTool, createRecursiveLintTool, createRecursiveLockTool, createRecursivePhaseTool, createRecursiveScratchTool, createRecursiveStatusTool, createRecursiveWorktreeTool, currentPhaseArtifact, defaultReviewToolFilter, delegate, delegateContinuable, delegationDecisionBasis, delegationError, detectTamper, discoverRuns, drainContinuableChildren, drainContinuableDescendants, escapeRegExp, evaluateDelegationResult, evaluateToolGuard, foldDiagnostics, foldRun, foldRunCard, getAllStaleReceipts, getArtifactState, getGateStatus, getLatestRunDirectory, getLockStatus, getMdFieldValue, getNextLegalPhase, getPrerequisiteBlockers, getPrerequisites, getStaleDownstreamPhases, getTodoStats, getWorkflowProfile, inject, interruptContinuable, invalidateReceipt, isCoreArtifact, isTaskClaimedBy, loadRouterPolicy, lockHashFromContent, makeRecursiveRoutes, mountRecursiveRoutesOnce, name, normalizeForLockHash, parseReplyVerdict, pendingWork, phaseIndex, phasePosition, probeCapabilities, readReceipt, readRepairFromReply, readRepairFromStructured, readVerdictFromReply, readVerdictFromStructured, receiptPath, referencesFromResult, registerRecursiveSkill, remainingDepthFor, renderPhaseTail, renderRecursivePolicy, renderStableContract, renderTaskHistory, replyPath, resetFoldCache, resolveEnforcementConfig, resolveRole, resolveRunDir, resolveToolPolicyForGuard, reviewBundleDir, reviewOutputSchema, routerPolicyPath, snapshotWorkspace, tamperCandidatePath, trimMdValue, validateChain, validateReferences, validateTransition, writeActionRecord, writeReceipt };
@@ -21,6 +21,18 @@ export interface ToolPolicyContext {
21
21
  runDir?: string;
22
22
  runId?: string;
23
23
  worktreeRoot?: string;
24
+ /**
25
+ * The ACTIVE phase artifact of the run — the LOWEST-numbered phase artifact that is not
26
+ * LOCKED, falling back to the highest when every one of them is locked.
27
+ *
28
+ * ⚠ THIS IS NOT A SECOND SELECTOR. It is the answer `currentPhaseArtifact` in
29
+ * `enforcement.ts` produced for this call, carried here so a rule in THIS module can
30
+ * compare against it. `policy-globs.ts` cannot import that function (enforcement.ts
31
+ * imports this module, and the repo has already paid once for a value-level import
32
+ * cycle), so the value is passed in rather than recomputed. A caller that omits it
33
+ * gets NO phase-order verdict — the rule abstains rather than guessing a phase.
34
+ */
35
+ activePhaseArtifact?: string;
24
36
  }
25
37
  /**
26
38
  * What a rule predicate returns when the condition it guards DOES apply:
@@ -152,6 +164,7 @@ export declare const LOCK_TOOL_NAMES: Set<string>;
152
164
  *
153
165
  * `recursive_lock*` -> the monotonic lock-order denial;
154
166
  * the write-tool ids -> the locked-artifact write denial;
167
+ * the write-tool ids -> the phase-order (write-ahead) denial;
155
168
  * `*` -> allow, so an ordinary tool is not turned into an `ask`.
156
169
  *
157
170
  * Every `deny` precedes the `allow`, which is what makes "deny wins over allow" a
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@try-works/dsh-recursive-mode",
3
3
  "description": "recursive-mode workflow as a DeepSeek Harness bundle: RecursiveRuntime service + 13 recursive_* tools (recursive_status, recursive_init, recursive_lock, recursive_lint, recursive_closeout, recursive_scratch, recursive_worktree, recursive_phase, recursive_audit_team, recursive_review, recursive_delegate, recursive_ask, recursive_preview)",
4
- "version": "0.4.7",
4
+ "version": "0.4.9",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",
7
7
  "types": "lib/index.d.ts",
package/src/delegation.ts CHANGED
@@ -559,9 +559,18 @@ export async function delegateContinuable(input: {
559
559
  // turn-shaped caller resumes on a later turn instead of treating the round
560
560
  // as lost, and so `accepted` stays false — an unobserved round is never an
561
561
  // approval.
562
+ //
563
+ // ⚠ THE SENTENCE NAMES THE CHILD, and that is not decoration. The advice this
564
+ // reason carries ("resume with the SAME child") is UNACTIONABLE without the id,
565
+ // so a caller that reads it and does not also read `childId` can only report the
566
+ // park, not act on it — and the record built from this reason is exactly where a
567
+ // live run read "no settlement" as "the child is dead" and went around its own
568
+ // reviewer. `resumeChild` is the field that consumes this id.
562
569
  return {
563
570
  ok: false,
564
- reason: 'no settlement has landed for round ' + (round + 1) + ' yet (the child is still working)',
571
+ reason: 'no settlement has landed for round ' + (round + 1) + ' yet, so the round is PARKED, not failed: '
572
+ + 'the child may still be working. Resume it on a later turn with childId ' + String(childId)
573
+ + ' (the `resumeChild` argument) — do not start a second child.',
565
574
  childId,
566
575
  messageIds,
567
576
  rounds,
@@ -774,8 +783,43 @@ export interface ActionRecordInput {
774
783
  * nothing else: a delegation that FAILED and a delegation that NEVER HAPPENED read identically, which is what
775
784
  * let me conclude for three rounds that the host was not scheduling children. The caller ALREADY passed
776
785
  * `stopReason`, and the live record said `n/a` — because there was no result to take a stop reason from.
786
+ *
787
+ * ⚠ AND IT IS EMITTED ONLY FOR A GENUINE FAILURE. A PARKED round is neither accepted nor failed, so it
788
+ * carries {@link ActionRecordInput.parked} instead: calling a live child a failure is the defect this
789
+ * field's own history is made of.
777
790
  */
778
791
  failure?: string
792
+ /**
793
+ * ⚠ THE THIRD STATE, AND WHY `success` COULD NOT CARRY IT.
794
+ *
795
+ * A continuable round that has not settled is NOT a failure — `ContinuableDelegationLike.parked` says so in
796
+ * its own doc comment, and the tool result already surfaces it. The RECORD did not: a live run's child was
797
+ * parked, the record said `Status: failed` with "the child never reported, or never ran", the main agent
798
+ * read that as a dead child and obtained the review elsewhere — while the child was still working and
799
+ * replied eighteen minutes later. `success: false` cannot express "still in flight", because a genuinely
800
+ * dead child also produces `success: false`, so a second field is required rather than a cleverer boolean.
801
+ *
802
+ * When true, the record says `parked` and carries a `Parked:` line (never a `Failure:` line) whose text
803
+ * states only what is KNOWN, names the `childId`, and names the resume step. It takes precedence over
804
+ * `success`: an unobserved round is never an acceptance, whatever a caller passes alongside it.
805
+ */
806
+ parked?: boolean
807
+ }
808
+
809
+ /**
810
+ * The status a record states, in ONE place, because the three states are the fix and a second writer would
811
+ * drift from this one.
812
+ *
813
+ * The wording is chosen for two readers at once. The MODEL reads it to decide whether to resume or to give up
814
+ * and obtain the result another way — the exact decision the live defect got wrong — so the token must not
815
+ * read as a failure and must not need the rest of the document to be understood. A HUMAN reading the run tree
816
+ * months later needs to tell "died" from "still working" at a glance. Hence `parked (still running; no
817
+ * settlement yet)`: a third token rather than a renamed second, qualified with the two facts that separate it
818
+ * from `failed`, and short enough to sit in a status line.
819
+ */
820
+ export function actionRecordStatus(input: Pick<ActionRecordInput, 'success' | 'parked'>): string {
821
+ if (input.parked === true) return 'parked (still running; no settlement yet)'
822
+ return input.success ? 'accepted' : 'failed'
779
823
  }
780
824
 
781
825
  function slugify(value: string): string {
@@ -794,7 +838,10 @@ function slugify(value: string): string {
794
838
  * `Code Refs` strictly inside ## Inputs Provided. The linter resolves each of
795
839
  * those through the heading body, so a field under another heading is not
796
840
  * found at all.
797
- * A success:false attempt is written with a failed status and is NOT accepted.
841
+ * A success:false attempt is written with a failed status and is NOT accepted. A round that PARKED is written
842
+ * with a status of its own (`parked (still running; no settlement yet)`, see {@link actionRecordStatus}) and a
843
+ * `Parked:` line instead of a `Failure:` one, because no settlement is not a death: it is the caller's signal to
844
+ * resume the same child. Only a genuine failure carries `Failure:`.
798
845
  */
799
846
  export function writeActionRecord(input: ActionRecordInput): string {
800
847
  const { root, runId } = input
@@ -844,11 +891,18 @@ export function writeActionRecord(input: ActionRecordInput): string {
844
891
  '- Phase: ' + input.phase,
845
892
  '- Purpose: ' + input.purpose,
846
893
  '- Execution Mode: ' + input.executionMode,
847
- '- Status: ' + (input.success ? 'accepted' : 'failed'),
894
+ '- Status: ' + actionRecordStatus(input),
848
895
  // ⚠ EMITTED ONLY WHEN A REASON IS GIVEN, so a caller that says nothing produces the record it always did.
849
896
  // Not politeness: this record's shape is asserted by specs, and a first attempt that always emitted the line
850
897
  // failed 13 tests across 5 files. A change to a shared surface should be additive where it can be.
851
- ...(input.success === false && input.failure !== undefined ? ['- Failure: ' + input.failure] : []),
898
+ //
899
+ // ⚠ AND `parked` IS THE ONE STATE THAT MUST NOT WEAR THE `Failure:` LABEL. This line used to be
900
+ // `input.success === false && input.failure !== undefined`, which is true of a PARKED round as well — so a
901
+ // live child that was still working was recorded, on the same line, as a failure. A parked round is not a
902
+ // failure, so it gets its own field and states only what is established.
903
+ ...(input.success === false && input.failure !== undefined
904
+ ? [input.parked === true ? '- Parked: ' + input.failure : '- Failure: ' + input.failure]
905
+ : []),
852
906
  '- Stop Reason: ' + (input.stopReason ?? 'n/a'),
853
907
  // Timestamp LAST in Metadata. This USED to be load-bearing: `getHeadingBody` ended
854
908
  // its capture with a `\Z` that JavaScript reads as a literal `Z`, so a body was
@@ -129,8 +129,14 @@ export const DEFAULT_ENFORCEMENT: EnforcementConfig = {
129
129
  * every allow/deny/ask the guard hands back). `'none'` means no guard predicate
130
130
  * fired; `'transition'` marks a decision whose only dissenting gate was the
131
131
  * advisory transition gate (see consultTransitionGate).
132
+ *
133
+ * `'phase-order'` is the WRITE half of the ordering rule the owner states as *only
134
+ * one phase may be active at a time, and the phases should be sequential and the
135
+ * active phase must be locked before proceeding to next phase*: `'lock-order'` refuses
136
+ * locking ahead, `'phase-order'` refuses WRITING ahead. Two labels rather than one,
137
+ * because the guard log has to tell the owner which of the two the agent attempted.
132
138
  */
133
- export type GuardRule = 'lock-order' | 'tdd-evidence' | 'locked-write' | 'transition' | 'none'
139
+ export type GuardRule = 'lock-order' | 'phase-order' | 'tdd-evidence' | 'locked-write' | 'transition' | 'none'
134
140
 
135
141
  /** T15: the transition gate's verdict as attached to a decision (advisory only). */
136
142
  export interface GuardTransition {
@@ -179,9 +185,12 @@ export interface ToolExecLike {
179
185
  * `recursive:policy` later — can read the effective rules instead of inferring
180
186
  * them from a code path.
181
187
  */
182
- export function resolveToolPolicyForGuard(worktreeRoot: string, runId: string): ToolPolicy {
188
+ export function resolveToolPolicyForGuard(worktreeRoot: string, runId: string, activePhaseArtifact?: string): ToolPolicy {
183
189
  const loaded = loadToolPolicyFile(worktreeRoot)
184
- return withPhaseBaseline(loaded.policy, currentPhaseArtifact(worktreeRoot, runId))
190
+ // A caller that already asked `currentPhaseArtifact` for this call passes the answer in,
191
+ // so one guard call reads the run tree's lock statuses ONCE rather than twice. The
192
+ // no-cache discipline is unchanged: an omitted argument still reads the filesystem here.
193
+ return withPhaseBaseline(loaded.policy, activePhaseArtifact ?? currentPhaseArtifact(worktreeRoot, runId))
185
194
  }
186
195
 
187
196
  /**
@@ -258,8 +267,14 @@ export function evaluateToolGuard(
258
267
  // T16: the verdict comes from the ordered policy, not from branches here. The
259
268
  // policy file is re-read per call on purpose: a policy a human just edited
260
269
  // must take effect on the next tool call, not after a restart.
261
- const policy = resolveToolPolicyForGuard(worktreeRoot, runId)
262
- const context: ToolPolicyContext = { args, runDir, runId, worktreeRoot }
270
+ //
271
+ // The ACTIVE phase is resolved ONCE, by the ONE selector, and is used twice: it
272
+ // selects the phase baseline (narrowing rules) and it is carried into the context so
273
+ // the phase-order rule can refuse a WRITE that is ahead of the active phase. Both
274
+ // halves of the ordering rule therefore read the same answer for the same call.
275
+ const activePhaseArtifact = currentPhaseArtifact(worktreeRoot, runId)
276
+ const policy = resolveToolPolicyForGuard(worktreeRoot, runId, activePhaseArtifact)
277
+ const context: ToolPolicyContext = { args, runDir, runId, worktreeRoot, activePhaseArtifact }
263
278
  const decision = evaluateToolPolicy(policy, name, args, context)
264
279
  return advisory(verdictFor(mode, decision), transition)
265
280
  }
@@ -40,10 +40,10 @@
40
40
  * just wrote and believes is in force.
41
41
  */
42
42
  import { existsSync, readFileSync } from 'node:fs'
43
- import { join } from 'node:path'
43
+ import { basename, dirname, join, resolve } from 'node:path'
44
44
  import { getLockStatus, getPrerequisiteBlockers } from './lock.ts'
45
45
  import { getMdFieldValue } from './status.ts'
46
- import { policyTargetPath, resolveFrom } from './phase-rules.ts'
46
+ import { phaseNumberForArtifact, policyTargetPath, resolveFrom } from './phase-rules.ts'
47
47
 
48
48
  /** The three verdicts a rule may carry. */
49
49
  export type Verdict = 'allow' | 'deny' | 'ask'
@@ -70,6 +70,18 @@ export interface ToolPolicyContext {
70
70
  runDir?: string
71
71
  runId?: string
72
72
  worktreeRoot?: string
73
+ /**
74
+ * The ACTIVE phase artifact of the run — the LOWEST-numbered phase artifact that is not
75
+ * LOCKED, falling back to the highest when every one of them is locked.
76
+ *
77
+ * ⚠ THIS IS NOT A SECOND SELECTOR. It is the answer `currentPhaseArtifact` in
78
+ * `enforcement.ts` produced for this call, carried here so a rule in THIS module can
79
+ * compare against it. `policy-globs.ts` cannot import that function (enforcement.ts
80
+ * imports this module, and the repo has already paid once for a value-level import
81
+ * cycle), so the value is passed in rather than recomputed. A caller that omits it
82
+ * gets NO phase-order verdict — the rule abstains rather than guessing a phase.
83
+ */
84
+ activePhaseArtifact?: string
73
85
  }
74
86
 
75
87
  /**
@@ -452,12 +464,95 @@ function lockedWriteRule(target: string | null, worktreeRoot: string | undefined
452
464
  return { verdict: 'deny', detail: normalized + ' carries Status: LOCKED (reopen explicitly to edit)' }
453
465
  }
454
466
 
467
+ /**
468
+ * The file name when `abs` is a DIRECT CHILD of `runDir`, and `null` otherwise.
469
+ *
470
+ * ⚠ THE DIRECT-CHILD TEST IS NOT TIDINESS — IT IS WHAT KEEPS THE RULE OFF THE SUPPORT
471
+ * FILES. `phaseNumberForArtifact` reads the leading digits of a NAME, so
472
+ * `<run>/evidence/01-as-is.md` or `<run>/subagents/child/03-brief.md` would look like
473
+ * phase 1 and phase 3 artifacts if the name were all that was examined. A phase artifact
474
+ * is a file the run tree holds DIRECTLY beside the others (`recursive_init` writes all
475
+ * twelve into `<run>/` itself), so the parent directory is part of the definition.
476
+ *
477
+ * The comparison normalizes separators and case: the same run directory reached through
478
+ * a Windows spelling that differs in case is the same directory, and the rule must not
479
+ * abstain on one spelling and fire on the other.
480
+ */
481
+ function directChildName(abs: string, runDir: string): string | null {
482
+ const normalize = (path: string) => resolve(path).replace(/\\/g, '/').replace(/\/+$/, '').toLowerCase()
483
+ if (normalize(dirname(abs)) !== normalize(runDir)) return null
484
+ return basename(abs)
485
+ }
486
+
487
+ /**
488
+ * PHASE-ORDER rule: a denial when the target is a LATER phase's artifact than the phase
489
+ * currently active, `null` when it is not.
490
+ *
491
+ * ⚠ THE HOLE THIS CLOSES. The monotonic rule was enforced on `recursive_lock` ONLY. An
492
+ * agent could therefore write `08-memory-impact.md` while the run sat at phase 0 — and a
493
+ * live run did exactly that: twelve artifacts, not one of them locked, written out of
494
+ * order, with a single line in `operations/operations.jsonl`. Ordering that only binds
495
+ * the lock tool is not ordering; the model's ordinary `write` is the path that mattered.
496
+ *
497
+ * The owner's rule, verbatim: *only one phase may be active at a time, and the phases
498
+ * should be sequential and the active phase must be locked before proceeding to next
499
+ * phase*. The ACTIVE phase is the one the selector names (`ctx.activePhaseArtifact`), so:
500
+ *
501
+ * - the target is the ACTIVE artifact, or shares its phase number (`00-requirements.md`
502
+ * and `00-worktree.md` are both phase 0; `01-as-is.md` and `01.5-root-cause.md` are
503
+ * both phase 1) -> ABSTAIN, the write is allowed;
504
+ * - the target is an EARLIER phase -> ABSTAIN. Such an artifact is LOCKED by
505
+ * construction (the active phase is the lowest UNLOCKED one), so the locked-artifact
506
+ * rule above decides it, and its rule label is preserved;
507
+ * - the target is a LATER phase -> DENY: working ahead.
508
+ *
509
+ * ⚠ THE ALLOW HALF IS LOAD-BEARING. An enforcement rule in this exact area was once the
510
+ * bug: strict enforcement denied the run's OWN artifacts in every phase and made the
511
+ * workflow unusable (see `resolveFrom` and `currentPhaseArtifact`). "The active artifact
512
+ * stays writable at every phase" is therefore asserted by walking every phase, not by
513
+ * one case — `tests/strict-run-tree.spec.ts` (d).
514
+ *
515
+ * WHAT IT ABSTAINS ON, deliberately:
516
+ * - no `activePhaseArtifact` (a caller with no run context, or a run with no phase
517
+ * artifacts yet) -> abstain, never guess a phase;
518
+ * - a target outside the active run's own directory -> abstain. The rule is about THIS
519
+ * run's sequence; another run's tree is a different question and denying it here
520
+ * would be a false positive;
521
+ * - a support file (anything not a direct child) -> abstain: `evidence/`, `scratch/`,
522
+ * `addenda/`, `subagents/`, `operations/` and a plain `<run>/notes.md` are not phases.
523
+ *
524
+ * The predicate adds only the per-call particular — which artifact, which phase, which is
525
+ * active; the rule keeps the static, auditable sentence.
526
+ */
527
+ function phaseOrderRule(target: string | null, ctx: ToolPolicyContext): ToolPolicyPredicateMatch | null {
528
+ const active = ctx.activePhaseArtifact
529
+ if (!target || !ctx.runDir || !ctx.worktreeRoot) return null
530
+ if (typeof active !== 'string' || active === '') return null
531
+ const activePhaseText = phaseNumberForArtifact(active)
532
+ if (!activePhaseText) return null
533
+ const normalized = target.replace(/\\/g, '/')
534
+ if (!normalized.endsWith('.md')) return null
535
+ const abs = resolveFrom(ctx.worktreeRoot, normalized)
536
+ if (!abs) return null
537
+ const name = directChildName(abs, ctx.runDir)
538
+ if (name === null) return null
539
+ const phaseText = phaseNumberForArtifact(name)
540
+ if (!phaseText) return null
541
+ if (Number(phaseText) <= Number(activePhaseText)) return null
542
+ return {
543
+ verdict: 'deny',
544
+ detail: name + ' is phase ' + phaseText + ' but the ACTIVE phase is ' + active
545
+ + ' (phase ' + activePhaseText + ') - the active phase must be locked before writing a later phase',
546
+ }
547
+ }
548
+
455
549
  /**
456
550
  * The BUILT-IN default rule list — the pre-T16 guard behaviour expressed as
457
551
  * data:
458
552
  *
459
553
  * `recursive_lock*` -> the monotonic lock-order denial;
460
554
  * the write-tool ids -> the locked-artifact write denial;
555
+ * the write-tool ids -> the phase-order (write-ahead) denial;
461
556
  * `*` -> allow, so an ordinary tool is not turned into an `ask`.
462
557
  *
463
558
  * Every `deny` precedes the `allow`, which is what makes "deny wins over allow" a
@@ -488,6 +583,20 @@ export function builtInToolPolicyRules(): ToolPolicyRule[] {
488
583
  predicate: (id, args, ctx) => (WRITE_TOOL_NAMES.has(id) ? lockedWriteRule(policyTargetPath(args), ctx.worktreeRoot) : null),
489
584
  })
490
585
  }
586
+ // The SAME write-tool ids get a SECOND conditional deny, the phase-order rule. Two rules
587
+ // share one pattern on purpose: the engine's predicates ABSTAIN (`null`) when their
588
+ // condition does not apply, so a clean write falls through both to the catch-all allow,
589
+ // and a target that is BOTH locked and a later phase is reported by the locked rule
590
+ // first (file order within a specificity tier), which is the pre-existing wording.
591
+ for (const name of WRITE_TOOL_NAMES) {
592
+ rules.push({
593
+ pattern: name,
594
+ verdict: 'deny',
595
+ reason: 'phase order: only one phase may be active at a time - the active phase must be locked before a later phase artifact is written',
596
+ label: 'phase-order',
597
+ predicate: (id, args, ctx) => (WRITE_TOOL_NAMES.has(id) ? phaseOrderRule(policyTargetPath(args), ctx) : null),
598
+ })
599
+ }
491
600
  rules.push({
492
601
  pattern: '*',
493
602
  verdict: 'allow',
@@ -523,6 +632,18 @@ export function attachPolicyPredicate(rule: ToolPolicyRule): ToolPolicyRule {
523
632
  if (rule.pattern === 'recursive_lock*') {
524
633
  return { ...rule, predicate: (id, args, ctx) => (LOCK_TOOL_NAMES.has(id) ? lockOrderRule(args.artifact, ctx.runDir) : null) }
525
634
  }
635
+ // ⚠ THE LABEL IS CONSULTED BEFORE THE PATTERN, because the phase-order rule and the
636
+ // locked-artifact rule share EVERY write-tool pattern (see `builtInToolPolicyRules`).
637
+ // Selecting by pattern alone would give BOTH rules the locked-artifact condition — the
638
+ // second would then deny a locked target with the wrong reason and the phase-order
639
+ // condition would never be attached at all, so the hole would stay open in every repo
640
+ // that ships a policy file. The `label` is already the rule's machine-readable identity
641
+ // (`firstPolicyDefect` admits it, `decide` surfaces it as the decision's `rule`), so the
642
+ // label is what names the condition. A rule with NO label keeps the pattern's condition,
643
+ // exactly as before.
644
+ if (rule.label === 'phase-order') {
645
+ return { ...rule, predicate: (id, args, ctx) => (WRITE_TOOL_NAMES.has(id) ? phaseOrderRule(policyTargetPath(args), ctx) : null) }
646
+ }
526
647
  if (WRITE_TOOL_NAMES.has(rule.pattern)) {
527
648
  return { ...rule, predicate: (id, args, ctx) => (WRITE_TOOL_NAMES.has(id) ? lockedWriteRule(policyTargetPath(args), ctx.worktreeRoot) : null) }
528
649
  }
package/src/policy.ts CHANGED
@@ -111,6 +111,7 @@ export function renderStableContract(config: EnforcementConfig = DEFAULT_ENFORCE
111
111
  + ', tamper detection ' + config.tamper + '.',
112
112
  '- A transition that fails its gates is BLOCKED (strict) or warns (advisory); no rejected transition proceeds silently.',
113
113
  '- Writes to a Status: LOCKED phase doc are denied/asked; reopen explicitly to edit.',
114
+ '- Phase order binds WRITES as well as locks: only the ACTIVE phase (the lowest-numbered artifact not yet LOCKED) may be written; a write to a LATER phase artifact is denied/asked. Run support files (evidence/, scratch/, addenda/, subagents/, operations/) are not phases.',
114
115
  '- Phase 3 lock requires TDD evidence (strict) or rationale (pragmatic); Phase 5 requires QA evidence.',
115
116
  '- The control-plane root is resolved STRICTLY from this session workspace (never scanned from another).',
116
117
  ].join('\n')
package/src/runtime.ts CHANGED
@@ -1099,7 +1099,16 @@ export class RecursiveRuntime extends Service {
1099
1099
  result = continuable.rounds[continuable.rounds.length - 1].result ?? null
1100
1100
  if (!result) error = 'continuable child produced no final result'
1101
1101
  } else {
1102
- error = continuable.reason ?? 'continuable delegation failed'
1102
+ // ⚠ A ROUND THAT SETTLED IS A RESULT, EVEN WHEN IT WAS NOT ACCEPTED — and this branch used to
1103
+ // discard it. The condition above requires `continuable.ok`, so a child that REPORTED and was
1104
+ // refused (`success: false`, or a non-completed stop reason) fell through to here: `result` stayed
1105
+ // null, the action record said "NO SETTLEMENT arrived within the wait", and the child's own stop
1106
+ // reason — the one fact that explains the refusal — was dropped on the floor. It is the same defect
1107
+ // as the parked one, one branch over: an absence asserted where the code had evidence. Keeping the
1108
+ // result is what lets the record say "the delegation returned without acceptance; stop reason error"
1109
+ // instead of blaming a wait that ended perfectly well.
1110
+ result = continuable.rounds[continuable.rounds.length - 1]?.result ?? null
1111
+ if (result === null) error = continuable.reason ?? 'continuable delegation failed'
1103
1112
  }
1104
1113
  } else {
1105
1114
  try {
@@ -1147,7 +1156,14 @@ export class RecursiveRuntime extends Service {
1147
1156
  id: operation,
1148
1157
  act: 'delegate-review',
1149
1158
  at: new Date().toISOString().replace(/\.\d{3}Z$/, 'Z'),
1150
- outcome: evaluation.accepted ? 'accepted' : 'unaccepted',
1159
+ // ⚠ A PARK IS NOT A REFUSAL, and `unaccepted` said it was. This line used to be
1160
+ // `accepted ? 'accepted' : 'unaccepted'`, so a round that had merely not settled yet was indexed
1161
+ // exactly like a delegation that was evaluated and refused — while the round it really was (still in
1162
+ // flight, resume the same child) was nowhere in the run's own operation log. The defect the action
1163
+ // record had, the log had too. The new value is honest on both readings that matter: it is not
1164
+ // `accepted`, so every retry gate still treats the operation as unfinished and retryable — which is
1165
+ // what a parked round is — and it no longer claims the delegation was judged and rejected.
1166
+ outcome: parked ? 'parked' : (evaluation.accepted ? 'accepted' : 'unaccepted'),
1151
1167
  phase: input.phase,
1152
1168
  })
1153
1169
  }
@@ -1158,8 +1174,8 @@ export class RecursiveRuntime extends Service {
1158
1174
  runId: input.runId,
1159
1175
  subagentId: input.childId,
1160
1176
  phase: input.phase,
1161
- // ⚠ FU-17 — the kind is stated in the record. `Status` says accepted or failed; nothing said whether the
1162
- // child PRODUCED the phase's work or JUDGED it, and a reader of a run could not tell the two apart.
1177
+ // ⚠ FU-17 — the kind is stated in the record. `Status` says accepted, failed or parked; nothing said whether
1178
+ // the child PRODUCED the phase's work or JUDGED it, and a reader of a run could not tell the two apart.
1163
1179
  purpose: input.role + (input.kind === 'work' ? ' (work)' : '') + ' for run ' + input.runId,
1164
1180
  executionMode: decision.tier + (input.mode !== 'one-shot' ? ' (continuable)' : ''),
1165
1181
  artifactPath: input.artifactPath,
@@ -1171,32 +1187,70 @@ export class RecursiveRuntime extends Service {
1171
1187
  findings: evaluation.accepted && result?.structured ? [(result.structured as { verdict?: string })?.verdict ?? 'accepted'] : undefined,
1172
1188
  success: evaluation.accepted,
1173
1189
  stopReason: result?.stopReason,
1190
+ // ⚠ A PARKED ROUND IS RECORDED AS PARKED — the whole defect in one field. `success: evaluation.accepted`
1191
+ // is false for a park (correct: nothing was accepted), and `writeActionRecord` reads this flag to state
1192
+ // the third state instead of collapsing it into `failed`.
1193
+ ...(parked ? { parked: true } : {}),
1174
1194
  // ⚠ FU-9 — AND SAY WHICH KIND OF FAILURE, because the record previously could not. `result == null` means
1175
1195
  // the provider never produced anything at all (never started, or returned nothing) — which is what the
1176
1196
  // live record's `Stop Reason: n/a` was quietly telling me — while a present result that failed to be
1177
1197
  // accepted means a child DID run and its work was refused. Different problems, identical artifacts.
1198
+ //
1199
+ // ⚠ AND `parked` IS BRANCHED FIRST. A parked round produced NO result at all — that IS what parking
1200
+ // means — so without this branch first it fell into the `result == null` text below: "the continuable
1201
+ // start was made and NO SETTLEMENT arrived within the wait (the child never reported, or never ran)".
1202
+ // That is the exact false conclusion this fix exists for, and it was reached whatever the record's
1203
+ // Status said. Order is therefore load-bearing here.
1178
1204
  failure: evaluation.accepted
1179
1205
  ? undefined
1180
- : result == null
1181
- // ⚠ THE TWO STATES ARE NOT THE SAME AND THE MESSAGE USED TO CONFLATE THEM. The one-shot path cannot
1182
- // resolve to nothing — the host's `start` returns a run or throws (assertCapabilities, expectProvider)
1183
- // — so a null result on the CONTINUABLE path means the opposite of what I first wrote: the start WAS
1184
- // made and NO SETTLEMENT ARRIVED within the wait. That distinction cost me two rounds of looking at
1185
- // provider names, so the record now states which path a run took and what it was waiting for.
1186
- ? (input.mode !== 'one-shot'
1187
- ? 'the continuable start was made and NO SETTLEMENT arrived within the wait (the child never reported,'
1188
- + ' or never ran); tier ' + decision.tier + ', provider ' + (decision.provider ?? 'none chosen')
1189
- + ', names on offer [' + (this.lastProviderNames.join(', ') || 'none') + ']'
1190
- // ⚠ FU-9 — THE PARENT IDENTITY, because the host refuses a prompt when it cannot resolve the
1191
- // parent session as a live Agent (`subagent/parent-unavailable`, index.ts L429-436), and the tool
1192
- // builds this handle with a CAST (`exec.agent as unknown as SubagentParentHandle`). A cast is not
1193
- // a contract: if the id here is not the one the host looks up, the refusal is real and the
1194
- // classifier's crash has been hiding it. Printing it here costs nothing and settles the question.
1195
- + '; parent id ' + ((input.parent as { id?: string } | undefined)?.id ?? 'none')
1196
- + ', parent session keys [' + (input.parent === undefined ? 'no parent' : Object.keys(input.parent as object).join(', ')) + ']'
1197
- : 'no delegate result was produced by the one-shot path; tier ' + decision.tier
1198
- + ', provider ' + (decision.provider ?? 'none chosen'))
1199
- : 'the delegation returned without acceptance; stop reason ' + (result.stopReason ?? 'none reported'),
1206
+ : parked
1207
+ // ⚠ WHAT IS KNOWN, AND ONLY WHAT IS KNOWN, WITH THE ID THE READER NEEDS TO ACT.
1208
+ //
1209
+ // The text this replaces said "the child never reported, or never ran" — a CONCLUSION drawn from an
1210
+ // ABSENCE, and it was false: a live child went on to complete three review rounds and reply eighteen
1211
+ // minutes later, while the main agent read `Status: failed`, concluded the child was dead, and
1212
+ // obtained its review by other means. So the parked message asserts nothing about the child's state
1213
+ // beyond "no settlement had landed when the wait ended", keeps "may still be working" as the
1214
+ // possibility it is, and NAMES the childId plus the exact next step, because advice to resume is
1215
+ // unactionable without the id. The identity diagnostics stay, because they are what makes a
1216
+ // misconfigured provider readable — but they are diagnostics, not the reason.
1217
+ ? 'no settlement had landed when the wait ended, so this round is PARKED, not failed: nothing was'
1218
+ + ' accepted and nothing was refused, and the child may still be working. The next step is to RESUME'
1219
+ + ' this round, not to re-dispatch it or replace the child: call `recursive_review` again on a later'
1220
+ + ' turn with childId ' + String(continuable?.childId ?? input.childId) + ' (the child the round was'
1221
+ + ' started for, which stays resumable). Diagnostics: tier ' + decision.tier
1222
+ + ', provider ' + (decision.provider ?? 'none chosen')
1223
+ + ', names on offer [' + (this.lastProviderNames.join(', ') || 'none') + ']'
1224
+ // ⚠ FU-9 — THE PARENT IDENTITY, because the host refuses a prompt when it cannot resolve the
1225
+ // parent session as a live Agent (`subagent/parent-unavailable`, index.ts L429-436), and the tool
1226
+ // builds this handle with a CAST (`exec.agent as unknown as SubagentParentHandle`). A cast is not
1227
+ // a contract: if the id here is not the one the host looks up, the refusal is real and the
1228
+ // classifier's crash has been hiding it. Printing it here costs nothing and settles the question.
1229
+ + '; parent id ' + ((input.parent as { id?: string } | undefined)?.id ?? 'none')
1230
+ + ', parent session keys [' + (input.parent === undefined ? 'no parent' : Object.keys(input.parent as object).join(', ')) + ']'
1231
+ : result == null
1232
+ // ⚠ THE TWO STATES ARE NOT THE SAME AND THE MESSAGE USED TO CONFLATE THEM. The one-shot path cannot
1233
+ // resolve to nothing — the host's `start` returns a run or throws (assertCapabilities, expectProvider)
1234
+ // — so a null result on the CONTINUABLE path means the opposite of what I first wrote: the start WAS
1235
+ // made and NO SETTLEMENT ARRIVED within the wait. That distinction cost me two rounds of looking at
1236
+ // provider names, so the record now states which path a run took and what it was waiting for.
1237
+ // (A park is handled above and never reaches this branch; this one is a continuable round that
1238
+ // produced neither a result nor the park signal, which IS a failure to report.)
1239
+ ? (input.mode !== 'one-shot'
1240
+ ? 'the continuable start was made and NO SETTLEMENT arrived within the wait (the child never reported,'
1241
+ + ' or never ran); tier ' + decision.tier + ', provider ' + (decision.provider ?? 'none chosen')
1242
+ + ', names on offer [' + (this.lastProviderNames.join(', ') || 'none') + ']'
1243
+ // ⚠ FU-9 — THE PARENT IDENTITY, because the host refuses a prompt when it cannot resolve the
1244
+ // parent session as a live Agent (`subagent/parent-unavailable`, index.ts L429-436), and the tool
1245
+ // builds this handle with a CAST (`exec.agent as unknown as SubagentParentHandle`). A cast is not
1246
+ // a contract: if the id here is not the one the host looks up, the refusal is real and the
1247
+ // classifier's crash has been hiding it. Printing it here costs nothing and settles the question.
1248
+ + '; parent id ' + ((input.parent as { id?: string } | undefined)?.id ?? 'none')
1249
+ + ', parent session keys [' + (input.parent === undefined ? 'no parent' : Object.keys(input.parent as object).join(', ')) + ']'
1250
+ : 'no delegate result was produced by the one-shot path; tier ' + decision.tier
1251
+ + ', provider ' + (decision.provider ?? 'none chosen'))
1252
+ : 'the delegation returned without acceptance; stop reason ' + (result.stopReason ?? 'none reported')
1253
+ + (result.success === false ? ' (the child itself reported success:false)' : ''),
1200
1254
  })
1201
1255
 
1202
1256
  // T35: report the mode that ACTUALLY ran, not the one that was asked for. A