devrites 4.0.5 → 4.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.md +12 -0
  2. package/README.md +4 -2
  3. package/docs/skills.md +2 -2
  4. package/pack/.claude/agents/devrites-plan-reviewer.md +7 -2
  5. package/pack/.claude/skills/devrites-debug-recovery/SKILL.md +19 -7
  6. package/pack/.claude/skills/devrites-debug-recovery/reference/cleanup-and-classify.md +6 -2
  7. package/pack/.claude/skills/devrites-lib/reference/reply-contract.md +5 -1
  8. package/pack/.claude/skills/devrites-lib/reference/standards/README.md +1 -0
  9. package/pack/.claude/skills/devrites-lib/reference/standards/afk-hitl.md +10 -0
  10. package/pack/.claude/skills/devrites-lib/reference/standards/core.md +4 -0
  11. package/pack/.claude/skills/devrites-lib/reference/standards/one-shot-actions.md +66 -0
  12. package/pack/.claude/skills/rite-autocomplete/SKILL.md +20 -3
  13. package/pack/.claude/skills/rite-autocomplete/reference/loop.md +11 -0
  14. package/pack/.claude/skills/rite-autocomplete/reference/stop-conditions.md +10 -0
  15. package/pack/.claude/skills/rite-prove/SKILL.md +20 -1
  16. package/pack/.claude/skills/rite-prove/reference/failure-triage.md +12 -3
  17. package/pack/.claude/skills/rite-spec/reference/state-workspace.md +7 -0
  18. package/pack/.claude/skills/rite-vet/SKILL.md +7 -1
  19. package/pack/.claude/skills/rite-vet/reference/artifacts.md +6 -0
  20. package/pack/generated/claude/agents/devrites-plan-reviewer.md +7 -2
  21. package/pack/generated/claude/skills/devrites-debug-recovery/SKILL.md +19 -7
  22. package/pack/generated/claude/skills/devrites-debug-recovery/reference/cleanup-and-classify.md +6 -2
  23. package/pack/generated/claude/skills/devrites-lib/reference/reply-contract.md +5 -1
  24. package/pack/generated/claude/skills/devrites-lib/reference/standards/README.md +1 -0
  25. package/pack/generated/claude/skills/devrites-lib/reference/standards/afk-hitl.md +10 -0
  26. package/pack/generated/claude/skills/devrites-lib/reference/standards/core.md +4 -0
  27. package/pack/generated/claude/skills/devrites-lib/reference/standards/one-shot-actions.md +66 -0
  28. package/pack/generated/claude/skills/rite-autocomplete/SKILL.md +20 -3
  29. package/pack/generated/claude/skills/rite-autocomplete/reference/loop.md +11 -0
  30. package/pack/generated/claude/skills/rite-autocomplete/reference/stop-conditions.md +10 -0
  31. package/pack/generated/claude/skills/rite-prove/SKILL.md +20 -1
  32. package/pack/generated/claude/skills/rite-prove/reference/failure-triage.md +12 -3
  33. package/pack/generated/claude/skills/rite-spec/reference/state-workspace.md +7 -0
  34. package/pack/generated/claude/skills/rite-vet/SKILL.md +7 -1
  35. package/pack/generated/claude/skills/rite-vet/reference/artifacts.md +6 -0
  36. package/pack/generated/codex/agents/devrites-plan-reviewer.toml +7 -2
  37. package/pack/generated/codex/skills/devrites-debug-recovery/SKILL.md +19 -7
  38. package/pack/generated/codex/skills/devrites-debug-recovery/reference/cleanup-and-classify.md +6 -2
  39. package/pack/generated/codex/skills/devrites-lib/reference/reply-contract.md +5 -1
  40. package/pack/generated/codex/skills/devrites-lib/reference/standards/README.md +1 -0
  41. package/pack/generated/codex/skills/devrites-lib/reference/standards/afk-hitl.md +10 -0
  42. package/pack/generated/codex/skills/devrites-lib/reference/standards/core.md +4 -0
  43. package/pack/generated/codex/skills/devrites-lib/reference/standards/one-shot-actions.md +66 -0
  44. package/pack/generated/codex/skills/rite-autocomplete/SKILL.md +20 -3
  45. package/pack/generated/codex/skills/rite-autocomplete/reference/loop.md +11 -0
  46. package/pack/generated/codex/skills/rite-autocomplete/reference/stop-conditions.md +10 -0
  47. package/pack/generated/codex/skills/rite-prove/SKILL.md +20 -1
  48. package/pack/generated/codex/skills/rite-prove/reference/failure-triage.md +12 -3
  49. package/pack/generated/codex/skills/rite-spec/reference/state-workspace.md +7 -0
  50. package/pack/generated/codex/skills/rite-vet/SKILL.md +7 -1
  51. package/pack/generated/codex/skills/rite-vet/reference/artifacts.md +6 -0
  52. package/package.json +1 -1
package/CHANGELOG.md CHANGED
@@ -2,6 +2,18 @@
2
2
 
3
3
  All notable changes to DevRites are documented here. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and DevRites adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). Releases are generated automatically by [semantic-release](https://semantic-release.gitbook.io/) from Conventional Commits on `main`.
4
4
 
5
+ ## [4.0.7](https://github.com/ViktorsBaikers/DevRites/compare/v4.0.6...v4.0.7) (2026-08-08)
6
+
7
+ ### Fixed
8
+
9
+ * **rite:** resume retained one-shot recovery ([4bd5897](https://github.com/ViktorsBaikers/DevRites/commit/4bd58971f1f3b2f6d8d9bfc7cb22b64066b604be))
10
+
11
+ ## [4.0.6](https://github.com/ViktorsBaikers/DevRites/compare/v4.0.5...v4.0.6) (2026-08-07)
12
+
13
+ ### Fixed
14
+
15
+ * **rite:** gate one-shot evidence retention ([621a549](https://github.com/ViktorsBaikers/DevRites/commit/621a549acdf9ec2a0d427b829b0df8852502380d))
16
+
5
17
  ## [4.0.5](https://github.com/ViktorsBaikers/DevRites/compare/v4.0.4...v4.0.5) (2026-08-07)
6
18
 
7
19
  ### Fixed
package/README.md CHANGED
@@ -28,7 +28,7 @@ project-conventional push, tag, or PR action, and archive the workspace.
28
28
  Unattended runs may create local WIP checkpoint commits along the way, but they
29
29
  remain local unless Ship's disclosed plan includes an approved remote action.
30
30
 
31
- **Status:** [`v4.0.5`](https://github.com/ViktorsBaikers/DevRites/releases/tag/v4.0.5): see [`CHANGELOG.md`](CHANGELOG.md) for release notes.
31
+ **Status:** [`v4.0.7`](https://github.com/ViktorsBaikers/DevRites/releases/tag/v4.0.7): see [`CHANGELOG.md`](CHANGELOG.md) for release notes.
32
32
 
33
33
  This is the latest published release; `main` may contain unreleased work.
34
34
 
@@ -99,7 +99,9 @@ Some work needs a different route:
99
99
  - [`/rite-autocomplete`](pack/.claude/skills/rite-autocomplete/SKILL.md) runs
100
100
  the reversible lifecycle unattended. With `--ship`, it continues through
101
101
  Ship preflight, discloses the exact Git plan, and waits for a fresh literal
102
- `GO` plus native approval; without that flag, it stops at Seal GO.
102
+ `GO` plus native approval; without that flag, it stops at Seal GO. A failed
103
+ consumptive proof action never retries blindly: retained evidence drives
104
+ offline repair and re-vetting, and only the next real attempt needs a new GO.
103
105
  - [`/rite-upgrade [slug]`](pack/.claude/skills/rite-upgrade/SKILL.md) is a
104
106
  compatibility route for an older active workspace that cannot resume. Age or
105
107
  cursor form alone never triggers repair; it is not a lifecycle phase.
package/docs/skills.md CHANGED
@@ -187,7 +187,7 @@ and `Shipped`. Utility commands keep the same compact labels and one-next-action
187
187
 
188
188
  | Skill | What It Does | Use When |
189
189
  |---|---|---|
190
- | [`rite-prove`](../pack/.claude/skills/rite-prove/SKILL.md) | Positive, discriminating tests + build/runtime/browser evidence. Requires the same candidate digest before/after commands and records the exact binding. | All slices built; ready for full verification. |
190
+ | [`rite-prove`](../pack/.claude/skills/rite-prove/SKILL.md) | Positive, discriminating tests + build/runtime/browser evidence. Requires the same candidate digest before/after commands and records the exact binding. Consumptive actions must retain bounded diagnostics through cleanup; a failed action continues offline recovery from that artifact, while any next real execution waits for fresh authorization. | All slices built; ready for full verification. |
191
191
  | [`devrites-browser-proof`](../pack/.claude/skills/devrites-browser-proof/SKILL.md) | Browser proof ladder: Playwright MCP → Chrome DevTools MCP → `/run`+`/verify` → project E2E → manual. Auto-emits the structured **Visual Verdict** (per-criterion PASS/FAIL vs `design-brief.md`) for UI slices: consumed by `devrites-frontend-reviewer` and gated at `/rite-seal`. | Scope touches UI. |
192
192
 
193
193
  ### Polish: normalize, then check the details
@@ -222,7 +222,7 @@ and `Shipped`. Utility commands keep the same compact labels and one-next-action
222
222
 
223
223
  | Skill | What It Does | Use When |
224
224
  |---|---|---|
225
- | [`rite-autocomplete`](../pack/.claude/skills/rite-autocomplete/SKILL.md) | Runs the whole lifecycle unattended (spec → clarify → … → seal → ship), choosing the recommended option at each soft gate and recording the rationale in `decisions.md`. A vague prompt triggers one up-front spec/clarify window; after decision coverage is CLEAR it runs without per-phase iteration, pausing only for genuine product/scope/policy decisions, irreversible risk, human-only access/actions, NO-GO, or budget exhaustion. Objective red checks use bounded technical recovery. Default stops at Seal GO; `--ship` (alias `--yolo`) reaches Ship preflight but still requires a fresh literal `GO` and native approval. | "Autocomplete", "do the whole thing", "run the full cycle", "one-shot this feature". |
225
+ | [`rite-autocomplete`](../pack/.claude/skills/rite-autocomplete/SKILL.md) | Runs the whole lifecycle unattended (spec → clarify → … → seal → ship), choosing the recommended option at each soft gate and recording the rationale in `decisions.md`. A vague prompt triggers one up-front spec/clarify window; after decision coverage is CLEAR it runs without per-phase iteration, pausing only for genuine product/scope/policy decisions, irreversible risk, human-only access/actions, NO-GO, or budget exhaustion. Objective red checks use bounded technical recovery, including cold-resume recovery from retained one-shot evidence; spending an action authorization blocks only another real execution, not offline repair. Default stops at Seal GO; `--ship` (alias `--yolo`) reaches Ship preflight but still requires a fresh literal `GO` and native approval. | "Autocomplete", "do the whole thing", "run the full cycle", "one-shot this feature". |
226
226
  | [`rite-zoom-out`](../pack/.claude/skills/rite-zoom-out/SKILL.md) | Map the modules, callers, callees, and decisions in an unfamiliar area using the project's domain glossary. | Explicit-only: `/rite-zoom-out` / `/rite zoom-out`. |
227
227
  | [`rite-prototype`](../pack/.claude/skills/rite-prototype/SKILL.md) | Throwaway code answering ONE design question: logic harness OR 2 to 4 UI variations on one route. | Explicit-only: `/rite-prototype` / `/rite prototype`. |
228
228
  | [`rite-handoff`](../pack/.claude/skills/rite-handoff/SKILL.md) | Compact chat session → handoff doc. References existing `.devrites/work/<slug>/` artifacts by path. | Explicit-only: `/rite-handoff` / `/rite handoff`. |
@@ -55,7 +55,11 @@ line, and then assign the band. Do not choose a score and justify it afterward:
55
55
  model changes conservatively. Every destructive step needs a rollback.
56
56
  7. **Failure-mode coverage:** for each new codepath, find a realistic failure such
57
57
  as a timeout, nil value, race, or stale state. A silent failure with **no test AND
58
- no error handling** is a critical gap.
58
+ no error handling** is a critical gap. For every consumptive action under
59
+ `one-shot-actions.md`, require durable bounded trust-safe evidence for every
60
+ terminal path plus unknown-well-formed, malformed/hostile, and cleanup-survival
61
+ fixtures. Missing evidence completeness is `broken`: the plan must not consume
62
+ the action to discover diagnostics that cleanup can erase.
59
63
 
60
64
  ## Confidence calibration + verification gate (mandatory)
61
65
  Give every finding a **confidence score from 1 to 10** and a quoted source:
@@ -73,7 +77,8 @@ Band each dimension `strong` / `adequate` / `thin` / `broken` (`broken` means
73
77
  Critical; `thin` means Important). For a borderline dimension, sample twice and
74
78
  take the **lower** band. The verdict uses the weakest dimension, not an average.
75
79
  Pass only when every dimension is at least `adequate`, no critical failure-mode gap
76
- remains, the shared-contract check passes, and the ID-and-meaning map contains no orphaned
80
+ remains, every consumptive action passes one-shot evidence completeness, the
81
+ shared-contract check passes, and the ID-and-meaning map contains no orphaned
77
82
  criterion, slice, or proof.
78
83
 
79
84
  ## Rules
@@ -18,10 +18,14 @@ has no clear next move.
18
18
  1. **Build the feedback loop:** create a fast, deterministic, agent-runnable pass/fail
19
19
  signal. Spend most of the investigation here.
20
20
  See [build-the-loop.md](reference/build-the-loop.md).
21
- 2. **Reproduce:** run the loop. Confirm the failure matches the user's report
22
- (not a nearby failure); capture the **exact error text**; confirm
23
- reproducibility (or a high enough repro rate for flaky bugs). Do not proceed
24
- without reproduction.
21
+ 2. **Reproduce:** run the loop for a repeatable action. Confirm the failure matches
22
+ the user's report (not a nearby failure); capture the **exact error text**;
23
+ confirm reproducibility (or a high enough repro rate for flaky bugs). For a
24
+ consumptive action under
25
+ [`one-shot-actions.md`](../devrites-lib/reference/standards/one-shot-actions.md),
26
+ the retained bounded artifact
27
+ is the reproduction input and the action MUST NOT be rerun during diagnosis.
28
+ Do not proceed without one of those reproduction inputs.
25
29
  3. **Ranked hypotheses (3-5, falsifiable):** generate the list before testing
26
30
  any of them. Each must state a prediction.
27
31
  **Completion:** 3-5 distinct hypotheses each state an observable prediction.
@@ -51,8 +55,14 @@ has no clear next move.
51
55
  - **Do NOT loosen / delete a failing assertion** to get green: check whether
52
56
  it's drift first (route via `/rite-plan repair`).
53
57
  - **Do NOT hide flakiness** with sleeps / retries: characterize it.
54
- - **Re-run the original loop after the fix.** The minimized regression test is not
55
- enough; prove the user-visible failure no longer reproduces.
58
+ - **Re-run the original loop after the fix when it is repeatable.** For a
59
+ consumptive action, first re-vet evidence completeness and obtain any required
60
+ fresh authorization; offline fixtures remain mandatory but cannot authorize the
61
+ real attempt.
62
+ - **A spent consumptive authorization is not a spent recovery budget.** When its
63
+ retained artifact supplies a new Critical/Important fingerprint, continue
64
+ offline diagnosis, correction, fixtures, and narrow Vet under that fingerprint's
65
+ no-progress budget. Stop for fresh authorization only before the next real action.
56
66
  - **Classify before routing** with
57
67
  [cleanup-and-classify.md](reference/cleanup-and-classify.md).
58
68
  - **Durably record class and rationale** in `decisions.md` and the applicable
@@ -62,7 +72,9 @@ has no clear next move.
62
72
  reproduction plus decisive signal rather than hashing symptom text.
63
73
  The caller and recovery attempts share one count: read the current context and
64
74
  recorded `## Dead ends` / `evidence.md`, then include every no-progress attempt
65
- with that fingerprint. Reclassify only on new causal evidence.
75
+ with that fingerprint. Reclassify only on new causal evidence. On cold resume,
76
+ a retained fingerprint with fewer than three such attempts remains runnable
77
+ even if the previous action wrote a terminal cursor.
66
78
  - **A maximum of three no-progress attempts per exact causal fingerprint stops the loop.**
67
79
  Count an attempt only when its recheck preserves the same decisive failure.
68
80
  Record attempt number, exact failure, hypothesis, probe, and failed idea after
@@ -14,7 +14,10 @@ record the dead end and classify anew. Failure alone never resets the budget.
14
14
  - `preexisting`: the same failure exists outside the candidate delta. Record the baseline and fix it only when it blocks acceptance.
15
15
  - `not_a_defect`: the observation matches current accepted authority. Record that authority and continue.
16
16
 
17
- Only human credentials/quotas/actions or irreversible work pause; never ask to retry.
17
+ Only human credentials/quotas/actions, irreversible work, or fresh authorization
18
+ required before the next consumptive execution pause. Fresh authorization never
19
+ pauses offline diagnosis or correction from retained evidence; never ask for a
20
+ blind retry.
18
21
 
19
22
  Record class/routing in `decisions.md` and each failed attempt in `evidence.md`
20
23
  or `## Dead ends`. Use one stable causal fingerprint shaped as `<affected
@@ -30,7 +33,8 @@ file or command.
30
33
 
31
34
  ## Cleanup checklist: required before declaring done
32
35
 
33
- - [ ] Original repro no longer reproduces (re-run the Phase 1 loop).
36
+ - [ ] Original repeatable repro no longer reproduces, or the consumptive action's
37
+ offline regression and evidence-completeness fixtures pass without a rerun.
34
38
  - [ ] Regression test passes (or absence of seam is documented).
35
39
  - [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix).
36
40
  - [ ] Throwaway harnesses deleted (or moved to a clearly marked debug location).
@@ -59,7 +59,11 @@ Claims such as proved, reviewed, sealed, shipped, or complete must point to real
59
59
  output or an artifact. Use exactly one recommended next action except for
60
60
  terminal agent-owned technical exhaustion, which has no runnable action.
61
61
 
62
- For that terminal case use:
62
+ Use that terminal case only after three recorded no-progress corrections of the
63
+ exact fingerprint, or when required evidence was irretrievably absent. A spent
64
+ consumptive-action authorization plus a retained new fingerprint is not terminal;
65
+ it continues offline recovery and waits for fresh authorization only after repair.
66
+ For a true terminal case use:
63
67
 
64
68
  ```text
65
69
  Stopped: Technical recovery exhausted
@@ -36,6 +36,7 @@ topic's owner.
36
36
  | `context-hygiene.md` | Choosing `/clear`, `/compact`, or a handoff. |
37
37
  | `anti-patterns.md` | A pack-wide rationalization or red flag appears. |
38
38
  | `afk-hitl.md` | A pause, question, resume, or AFK decision is possible. |
39
+ | `one-shot-actions.md` | A proof/action may be attempted once, needs fresh retry authorization, consumes external state/quota, or can delete its own failure evidence. |
39
40
  | `tooling.md` | Structural lookup, current external facts, or architecture memory is needed. |
40
41
  | `skill-authoring.md` | Creating, editing, routing, evaluating, or pruning a DevRites skill. |
41
42
  | `definition-of-done.md` | Prove, Seal, Ship, or Quick must decide whether work is finished. |
@@ -245,10 +245,20 @@ changes, not a request for permission to retry.
245
245
  new Critical or Important finding with a different failed invariant and exact
246
246
  evidence gets its own fingerprint and budget; Suggestion, Nit, FYI, or renamed
247
247
  evidence cannot open, reset, or extend recovery.
248
+ - **Separate consumptive authority from recovery.** Spending a one-shot action's
249
+ authorization forbids another real execution but does not consume the offline
250
+ recovery budget for a newly evidenced fingerprint. The retained artifact starts
251
+ caller-owned diagnosis and correction immediately; only the next consumptive
252
+ execution waits for fresh authorization.
248
253
  - **Persist accounting in existing records.** Record each fingerprint,
249
254
  reproduction, attempted correction, `progress: resolved|no-progress`, and
250
255
  decisive result in `drift.md` and `evidence.md`. Cold resume derives the count
251
256
  from those records. There is no recovery counter file or command.
257
+ - **Reconcile before honoring a terminal cursor.** A persisted distinct
258
+ Critical/Important fingerprint with retained evidence and fewer than three
259
+ recorded no-progress corrections still has recovery budget, even when an older
260
+ `state.md` says `Next step: none`. Resume it; do not treat session age or the
261
+ prior action's spent authorization as exhaustion.
252
262
  - **Classify exhaustion:** human-owned contract/risk/access gaps open their gate. Otherwise
253
263
  preserve reproduction/dead ends, set `Status: blocked` and `Next step: none — technical recovery exhausted for <causal fingerprint>; requires new evidence or changed failure conditions`.
254
264
  Do not emit `/rite-plan unblock`, another phase command, a question, or
@@ -54,6 +54,10 @@ re-reads `state.md`, follows the durable return cursor and intermediate
54
54
  `next_action`, and resumes its originating phase while no human-owned, safety,
55
55
  access, budget, or exhausted-recovery stop is active.
56
56
 
57
+ Derive `exhausted-recovery` from the exact fingerprint's recorded no-progress
58
+ attempts, not from a stale `state.md` label. A consumed authorization for one
59
+ real action does not exhaust offline recovery from its retained new evidence.
60
+
57
61
  An intermediate `Next step` is cold-resume metadata. Do not ask the human to
58
62
  copy routine `/rite-plan repair`, `/rite-vet`, `/rite-build`, or proof-rerun
59
63
  commands during the active recovery chain. Only the controlling caller emits
@@ -0,0 +1,66 @@
1
+ # One-shot evidence completeness
2
+
3
+ An action is **consumptive** when a failed attempt is not safely equivalent to a
4
+ normal rerun. This includes commands limited to one attempt, commands whose retry
5
+ needs fresh human authorization, actions that spend external quota or mutate
6
+ privileged/external state so a rerun is not equivalent, and actions whose cleanup
7
+ can destroy the failure state needed for diagnosis. Successful cleanup does not
8
+ make a consumptive action repeatable.
9
+
10
+ ## Pre-attempt gate
11
+
12
+ Before Vet can emit READY, and again immediately before Prove executes the action,
13
+ the approved `test-plan.md` must bind all of the following:
14
+
15
+ 1. **Durable retention:** an operator-controlled evidence artifact outside the
16
+ disposable runtime/cleanup tree, created before the first side effect, written
17
+ durably before cleanup, least-privilege, and bounded by schema, size, and
18
+ cardinality.
19
+ 2. **Trust-safe diagnostics:** known semantic values use the normal validator;
20
+ unknown but lexically well-formed non-secret values survive in bounded sanitized
21
+ fields; malformed, hostile, or secret-bearing values become fixed reason codes
22
+ rather than retained raw input.
23
+ 3. **Terminal completeness:** every success, nonzero exit, rejection, timeout,
24
+ signal, and cleanup failure either names the retained artifact or proves that
25
+ no diagnostic state exists. Failure retention preserves the original safe
26
+ failure family and cause through clean convergence.
27
+ 4. **Discriminating proof:** fixtures cover success, a known failure, an unknown
28
+ well-formed failure, malformed/hostile input, and cleanup after failure. They
29
+ prove cleanup cannot delete or overwrite the retained failure evidence.
30
+ 5. **Recovery sufficiency:** the retained bounded evidence is enough to choose an
31
+ offline correction or a truthful terminal classification without consuming
32
+ another attempt.
33
+
34
+ Missing or stale evidence is an agent-owned technical plan gap: Vet returns
35
+ `NEEDS REPLAN`, and Prove returns to Vet inline without executing the action. Never
36
+ weaken the trust validator or spend the attempt merely to discover what the
37
+ retention design should have preserved.
38
+
39
+ ## Failure handling
40
+
41
+ After a consumptive action fails, its retained artifact is the reproduction input.
42
+ Do not rerun the action during triage.
43
+
44
+ Keep two budgets separate:
45
+
46
+ - **Action authorization:** the failed execution consumes only the authorization
47
+ for that consumptive execution. Zero remaining action attempts prohibits another
48
+ real execution; it does not exhaust offline diagnosis or correction.
49
+ - **Causal-fingerprint recovery:** when the retained artifact supplies a new
50
+ Critical/Important failed invariant, the controlling caller immediately runs
51
+ bounded offline triage, repair, fixtures, and narrow Vet in the same invocation.
52
+ Count only no-progress corrections of that exact fingerprint under
53
+ `afk-hitl.md`; do not stop merely because the action authorization was consumed.
54
+
55
+ Cold resume does not make that fingerprint old or exhausted. Derive its offline
56
+ no-progress count from `drift.md` and `evidence.md`; while the count is below the
57
+ cap, resume recovery even if a prior writer stored `blocked` / `Next step: none`.
58
+ That terminal cursor is valid only for missing retention, a human/safety gate, or
59
+ an actually exhausted fingerprint.
60
+
61
+ A new real attempt is admissible only after the affected plan and fixtures are
62
+ re-vetted, the failure condition is shown changed, and any required fresh
63
+ authorization is obtained. Stop at that authorization boundary; never infer it
64
+ from successful offline repair. If the attempt ran without complete retention,
65
+ stop with the observed technical blocker because a blind retry cannot manufacture
66
+ the destroyed evidence.
@@ -15,6 +15,9 @@ default and **Full** for high-risk scope or explicit `--full`; profiles are defi
15
15
 
16
16
  ## Rules consulted (read on demand from `.claude/skills/devrites-lib/reference/standards/`)
17
17
  **Step 0:** Read `.claude/skills/devrites-lib/reference/standards/core.md` and `.claude/skills/devrites-lib/reference/standards/afk-hitl.md` first.
18
+ When the cursor or approved plan contains a consumptive action, also read
19
+ `.claude/skills/devrites-lib/reference/standards/one-shot-actions.md` before any
20
+ stop/continue or execution decision.
18
21
 
19
22
  ## Operating rules
20
23
  - **Use one initial human window.** Run spec and topology-first clarify; arm AFK
@@ -29,6 +32,11 @@ default and **Full** for high-risk scope or explicit `--full`; profiles are defi
29
32
  through repair, Vet, bounded implementation correction, and re-proof inside
30
33
  the active run. A nested phase `STOP` or intermediate `Next step` is not a
31
34
  user handoff; pause only on the shared real stop conditions.
35
+ - **Do not confuse an action budget with recovery exhaustion.** After a failed
36
+ consumptive action, zero remaining real attempts blocks only another execution.
37
+ New retained Critical/Important evidence starts offline repair and narrow Vet
38
+ inside this Autocomplete invocation. Stop for a fresh GO only after that repair
39
+ changes the failure condition and Vet is READY.
32
40
  - **Budget from the post-vet slice count.** Vet may split/add slices.
33
41
  `--max-slices N` may lower the cap for a partial run; otherwise build all.
34
42
  - **Parse flags only from this invocation.** `--ship`, `--yolo`, `--max-slices`,
@@ -55,7 +63,12 @@ default and **Full** for high-risk scope or explicit `--full`; profiles are defi
55
63
 
56
64
  ## Workflow
57
65
  1. **Orient + parse args.** Resolve the explicit or active slug, require its
58
- `state.md`, and read the cursor directly.
66
+ `state.md`, and read the cursor directly. Before honoring `blocked` with a
67
+ terminal `next_action`, reconcile retained consumptive-action artifacts and
68
+ per-fingerprint attempts from `drift.md` / `evidence.md`. A distinct retained
69
+ fingerprint below its no-progress cap reopens caller-owned offline recovery;
70
+ reconstruct any missing return cursor only from the current phase plus the
71
+ exact approved action in `test-plan.md` / evidence, never from chat or guesswork.
59
72
  The idea + flags: `--ship` / `--yolo` (continue through ship preflight, then stop
60
73
  for literal-GO and native approval), `--max-slices N` (optional lower safety cap for a partial run; default =
61
74
  the plan's slice count, i.e. run all planned slices). Parse only the current
@@ -90,7 +103,9 @@ default and **Full** for high-risk scope or explicit `--full`; profiles are defi
90
103
  before any later phase runs; the cursor names the last completed phase and the
91
104
  effective remaining budget.
92
105
  5. **Apply stop conditions at every gate** ([reference/stop-conditions.md](reference/stop-conditions.md)):
93
- on hard-risk / blocking / escalating / NO-GO / budget-exhausted / still-low-confidence
106
+ first route an agent-owned red technical result through the loop's bounded
107
+ recovery; its `blocked` label alone is not a stop condition. After that, on
108
+ hard-risk / human-owned blocking / escalating / NO-GO / budget-exhausted / still-low-confidence
94
109
  → write `state.md` (`Status`, `Next step`), surface *why*, and **STOP**.
95
110
  Exhausted agent-owned technical recovery uses the terminal `Next step: none`
96
111
  marker and never hands the user `/rite-plan unblock` or another routine
@@ -120,7 +135,9 @@ Do not write a narrative recap.
120
135
  ## Clean baseline and checkpoint mode
121
136
  - Before autonomy, require a clean or accepted baseline; refuse unrelated dirty work and record plan artifacts.
122
137
  - Arm `.devrites/CHECKPOINT`; `/rite-ship` collapses its local-only proven-slice checkpoints.
123
- - Stop on risky steps, red gates, NO-GO, stale evidence, or exhausted budget.
138
+ - Red gates block forward advancement and enter caller-owned bounded recovery.
139
+ Stop only when the shared stop contract classifies risk/HITL or the exact
140
+ fingerprint's recovery budget is exhausted.
124
141
  - One autocomplete pass spans the lifecycle. Each wright returns after one slice;
125
142
  explicit `.devrites/AFK` lets the Build root chain green slices under cap/pause rules;
126
143
  otherwise HITL stops.
@@ -57,6 +57,13 @@ Read each phase's `SKILL.md` and execute that workflow. Workspace files such as
57
57
  When a later phase finds an agent-owned technical gap in an earlier phase, the
58
58
  Autocomplete root remains the caller:
59
59
 
60
+ On cold resume, reconcile the terminal cursor against durable fingerprint
61
+ accounting first. A retained distinct fingerprint with fewer than three
62
+ no-progress corrections is unfinished recovery, not an unchanged terminal stop.
63
+ Restore `return_phase` from the current phase and `return_next_action` only from
64
+ the exact approved action recorded in `test-plan.md` / evidence; ambiguity returns
65
+ to the applicable Vet contract and never licenses execution.
66
+
60
67
  1. Save the originating phase/action in the native return cursor unless a valid
61
68
  one already exists.
62
69
  2. Invoke the required repair, Vet, remediation, and proof skills inline. After
@@ -71,6 +78,10 @@ Autocomplete root remains the caller:
71
78
  a genuinely new Critical/Important fingerprint. Charge only a no-progress
72
79
  outcome against the same fingerprint; lower-severity novelty cannot prolong
73
80
  the chain.
81
+ For a failed consumptive action, its spent authorization blocks only another
82
+ real execution. A retained artifact that identifies a new fingerprint is the
83
+ offline reproduction input: diagnose, repair, and narrow-Vet it now, then pause
84
+ for fresh authorization before any next consumptive execution.
74
85
  5. When the prerequisite chain is green, restore and consume the return cursor,
75
86
  resume the originating phase, and continue the forward table.
76
87
 
@@ -9,6 +9,12 @@ Closure of a prior fingerprint is progress, not exhaustion. A separately evidenc
9
9
  Critical/Important failed invariant starts its own bounded fingerprint; it never
10
10
  resets or extends the budget of the one just closed.
11
11
 
12
+ An exhausted consumptive-action authorization is not technical-recovery
13
+ exhaustion. It blocks another real action, but retained evidence of a new
14
+ Critical/Important fingerprint must enter offline caller-owned recovery while its
15
+ own no-progress budget remains. After affected Vet is READY, pause for fresh action
16
+ authorization; never execute from the old GO.
17
+
12
18
  ## Always stop (irreversible-risk list: from `afk-hitl.md`)
13
19
 
14
20
  Regardless of `allow_gates` or `--ship`:
@@ -35,6 +41,10 @@ the remaining choice is a real human/safety/access gate.
35
41
  On technical exhaustion, preserve the fingerprint, reproduction, attempts, and
36
42
  dead ends, then stop without `/rite-plan unblock` or another phase command.
37
43
  Reinvocation with unchanged evidence remains blocked and does not reset the cap.
44
+ Here `unchanged` means the same fingerprint already has three recorded
45
+ no-progress corrections. A retained fingerprint with remaining offline budget is
46
+ not terminal merely because `state.md` was written by the failed action or a prior
47
+ session ended.
38
48
 
39
49
  ## Stop on gate severity
40
50
 
@@ -38,6 +38,8 @@ Pull these via `Read` when relevant:
38
38
  webhook / config / error messages / getting-started): **measure** the DX scorecard (run the flow,
39
39
  measure time-to-hello-world, and capture verbatim error text), rather than asserting it.
40
40
  - `definition-of-done.md`: acceptance, proof, gates, scope, rollback/docs.
41
+ - `one-shot-actions.md`: pre-attempt evidence completeness and no-rerun handling for
42
+ consumptive proof actions.
41
43
 
42
44
 
43
45
  ## Operating rules
@@ -84,7 +86,10 @@ failure is a blocker; this Upgrade admission never authorizes source or test cha
84
86
 
85
87
  ## Workflow
86
88
  0. Read core, load relevant rules above, then resolve the active slug,
87
- require its `state.md`, and read the cursor directly.
89
+ require its `state.md`, and read the cursor directly. On a blocked Prove cold
90
+ resume, reconcile any retained consumptive-action artifact against recorded
91
+ no-progress corrections before accepting `next_action: none`; a distinct
92
+ fingerprint below its cap resumes offline triage without another real action.
88
93
  1. **Confirm all slices built.** Read `spec.md`, `tasks.md`, `state.md`,
89
94
  `test-plan.md`, and the full diff.
90
95
  A missing `test-plan.md` enters caller-owned Vet backtracking; invoke Vet
@@ -107,6 +112,16 @@ failure is a blocker; this Upgrade admission never authorizes source or test cha
107
112
  substituted commands, malformed manifests, zero-test/skipped/filtered results used as
108
113
  behavioral proof, success inferred only from exit status, and any source drift. Static
109
114
  gates prove only their named static criterion.
115
+ Immediately before any consumptive action, apply `one-shot-actions.md` to the
116
+ live candidate: require the vetted retained-artifact identity, bounds,
117
+ sanitization, terminal-path fixtures, and cleanup-survival proof. A missing,
118
+ stale, or disposable-only evidence surface returns to Vet inline and consumes
119
+ no attempt. Record the admitted artifact identity before execution. After a
120
+ failed consumptive action, triage from that artifact; never reproduce it by
121
+ rerunning the action. Consuming its authorization blocks only another real
122
+ execution. If the artifact identifies a new Critical/Important fingerprint,
123
+ continue offline recovery inside this Prove invocation; do not label the new
124
+ fingerprint exhausted because the action budget is zero.
110
125
  4. **UI feature?** The root applies the browser proof ladder with
111
126
  `design-brief.md`, `references.md`, the requested routes, browser harness, and
112
127
  allowed scratch path:
@@ -136,6 +151,10 @@ failure is a blocker; this Upgrade admission never authorizes source or test cha
136
151
  Vet inside this invocation, and resume the exact failed proof rung. Ask only
137
152
  for a human-owned decision; causal-fingerprint exhaustion stops once with a
138
153
  technical blocker rather than another phase command.
154
+ For a consumptive-action failure, use fixtures or another non-consumptive
155
+ reproduction to fix and re-vet the retained fingerprint. Once the failure
156
+ condition is demonstrably changed, stop at the fresh-authorization boundary;
157
+ only a new human GO may admit the next real attempt.
139
158
  8. The root updates `evidence.md`, `browser-evidence.md` (when present),
140
159
  `traceability.md`, and `state.md`. Record exactly one binding for the observed
141
160
  digest in evidence and browser evidence. New proof goes to canonical
@@ -4,17 +4,26 @@ When a test/build/runtime/browser check fails, triage before fixing. Delegates t
4
4
  loop to `devrites-debug-recovery`.
5
5
 
6
6
  ## Triage ladder
7
- 1. **Reproduce:** run the failing command again; capture the exact error (quote it).
7
+ 1. **Reproduce:** for a repeatable command, run it again and capture the exact
8
+ error (quote it). For a consumptive action under
9
+ [`one-shot-actions.md`](../../devrites-lib/reference/standards/one-shot-actions.md),
10
+ never rerun it: use the retained bounded artifact as the reproduction input.
8
11
  2. **Classify** the failure:
9
12
  - test is right, code is wrong → fix the code (in scope).
10
13
  - test is wrong/outdated → that may be **spec drift** (acceptance criteria wrong).
11
- - environment/setup (missing dep, wrong node/ruby) → fix setup or record blocker.
14
+ - environment/setup (missing dep, wrong node/ruby) → fix it when agent-owned;
15
+ record a blocker only for human/access ownership, scope overflow, or exhausted
16
+ fingerprint recovery.
12
17
  - flaky (passes on re-run, timing/order) → note it; don't paper over with retries.
13
18
  - external dependency down → record blocker; don't fake the result.
14
19
  3. **Isolate:** minimize to the smallest failing case; bisect the diff if needed.
15
20
  4. **Fix within scope:** the current slice/feature only. If the fix reaches outside
16
21
  scope, stop and record a blocker in `state.md` (then `/rite-plan unblock`).
17
- 5. **Re-run** the same command; confirm green; record both attempts in `evidence.md`.
22
+ 5. **Re-run** the same command when it is repeatable; confirm green and record both
23
+ attempts in `evidence.md`. A consumptive action needs re-vetted evidence
24
+ completeness and any fresh authorization before a new attempt. Its spent
25
+ action authorization does not stop offline recovery: a retained new fingerprint
26
+ enters caller-owned diagnosis, correction, fixtures, and narrow Vet immediately.
18
27
 
19
28
  ## Rules
20
29
  - Quote the real error text; don't paraphrase it away.
@@ -131,6 +131,13 @@ resume. Never overwrite an existing valid return cursor, and preserve every
131
131
  unrelated Markdown byte. `/rite-clarify` applies its stricter native cursor
132
132
  protocol below.
133
133
 
134
+ On cold resume, a terminal `next_action` is a claim to verify, not a counter.
135
+ Reconcile the exact fingerprint and its attempts in `drift.md` / `evidence.md`.
136
+ If a consumptive action retained a distinct fingerprint below the no-progress
137
+ cap, rewrite the cursor into caller-owned recovery and preserve or reconstruct
138
+ the return rows only from the current phase and approved recorded action. The
139
+ spent action authorization blocks another execution, not that rewrite.
140
+
134
141
  `afk_slices_remaining` is mutable runtime state, not `.devrites/AFK`
135
142
  configuration. Only the controlling root writes it under the shared
136
143
  [`afk-hitl.md`](../../devrites-lib/reference/standards/afk-hitl.md#the-sentinel-devritesafk)
@@ -20,7 +20,8 @@ but every profile keeps the exact plan-reviewer gate.
20
20
  Pull the standard named by the active axis: `principles.md`, `patterns.md`,
21
21
  `coding-style.md`, `testing.md`, `spec-grammar.md`, `performance.md`,
22
22
  `error-handling.md`, `development-workflow.md`, `afk-hitl.md`,
23
- `developer-experience.md`, `elicitation.md`, and `definition-of-done.md`.
23
+ `one-shot-actions.md`, `developer-experience.md`, `elicitation.md`, and
24
+ `definition-of-done.md`.
24
25
 
25
26
 
26
27
  ## Operating rules
@@ -97,6 +98,11 @@ Pull the standard named by the active axis: `principles.md`, `patterns.md`,
97
98
  unmeasurable conflict = gap. Record complete SHA-256 provenance inputs. Require
98
99
  every behavioral mapping to name a positive,
99
100
  discriminating assertion and decisive signal, not merely a command or expected exit zero.
101
+ Identify every consumptive action under `one-shot-actions.md`. Before admitting
102
+ it, require the exact durable retention surface, trust-safe diagnostic schema,
103
+ cleanup ordering, terminal-path coverage, and discriminating fixtures in
104
+ `test-plan.md`. Missing or stale one-shot evidence completeness is a technical
105
+ preflight gap; do not spend the action to learn what cleanup would erase.
100
106
  Preflight observes; it need not make future behavior pass.
101
107
  2c. **Implementation-readiness audit.** Goal-backward map every REQ/AC/NFR, interaction,
102
108
  edge/prohibition, and decision-coverage row to a slice and executable proof. Verify
@@ -35,6 +35,7 @@ Inventory/currentness: <pass/gaps> · slice order/independence: <pass/gaps> ·
35
35
  UX/spec/architecture: <pass/n-a/gaps> · operations/rollout/rollback: <pass/n-a/gaps>
36
36
 
37
37
  Shared contract proof: <pass | gap: missing/one-sided/duplicated-contract/vague/non-consuming>
38
+ One-shot evidence completeness: <n/a | pass: action + retained artifact + bounds/sanitization + cleanup-survival fixtures | gap: exact missing proof>
38
39
 
39
40
  ## 3. Axis findings (floor-gated)
40
41
  | Axis | Floor band | Findings (sev · confidence) |
@@ -94,6 +95,11 @@ Commands in this durable artifact are portable repository commands: no RTK or lo
94
95
  aliases, user-specific absolute paths, or temporary proof trees. Evidence records
95
96
  the command actually executed.
96
97
 
98
+ ## Consumptive action gates
99
+ | Action | Why one-shot/consumptive | Retained artifact | Bounds + sanitization | Terminal-path fixtures | Cleanup survival | Retry authority |
100
+ |---|---|---|---|---|---|---|
101
+ | <exact approved action or n/a> | <one attempt/quota/state/evidence deletion> | <durable operator-controlled path/schema> | <size/cardinality + known/unknown/malformed handling> | <success/known/unknown/hostile> | <assertion> | <none/fresh human authorization> |
102
+
97
103
  Every behavioral row names a positive, discriminating assertion and the decisive output it
98
104
  produces. A command or expected exit zero alone is not a behavioral assertion; static gates
99
105
  prove only their named static criterion.
@@ -55,7 +55,11 @@ line, and then assign the band. Do not choose a score and justify it afterward:
55
55
  model changes conservatively. Every destructive step needs a rollback.
56
56
  7. **Failure-mode coverage:** for each new codepath, find a realistic failure such
57
57
  as a timeout, nil value, race, or stale state. A silent failure with **no test AND
58
- no error handling** is a critical gap.
58
+ no error handling** is a critical gap. For every consumptive action under
59
+ `one-shot-actions.md`, require durable bounded trust-safe evidence for every
60
+ terminal path plus unknown-well-formed, malformed/hostile, and cleanup-survival
61
+ fixtures. Missing evidence completeness is `broken`: the plan must not consume
62
+ the action to discover diagnostics that cleanup can erase.
59
63
 
60
64
  ## Confidence calibration + verification gate (mandatory)
61
65
  Give every finding a **confidence score from 1 to 10** and a quoted source:
@@ -73,7 +77,8 @@ Band each dimension `strong` / `adequate` / `thin` / `broken` (`broken` means
73
77
  Critical; `thin` means Important). For a borderline dimension, sample twice and
74
78
  take the **lower** band. The verdict uses the weakest dimension, not an average.
75
79
  Pass only when every dimension is at least `adequate`, no critical failure-mode gap
76
- remains, the shared-contract check passes, and the ID-and-meaning map contains no orphaned
80
+ remains, every consumptive action passes one-shot evidence completeness, the
81
+ shared-contract check passes, and the ID-and-meaning map contains no orphaned
77
82
  criterion, slice, or proof.
78
83
 
79
84
  ## Rules
@@ -18,10 +18,14 @@ has no clear next move.
18
18
  1. **Build the feedback loop:** create a fast, deterministic, agent-runnable pass/fail
19
19
  signal. Spend most of the investigation here.
20
20
  See [build-the-loop.md](reference/build-the-loop.md).
21
- 2. **Reproduce:** run the loop. Confirm the failure matches the user's report
22
- (not a nearby failure); capture the **exact error text**; confirm
23
- reproducibility (or a high enough repro rate for flaky bugs). Do not proceed
24
- without reproduction.
21
+ 2. **Reproduce:** run the loop for a repeatable action. Confirm the failure matches
22
+ the user's report (not a nearby failure); capture the **exact error text**;
23
+ confirm reproducibility (or a high enough repro rate for flaky bugs). For a
24
+ consumptive action under
25
+ [`one-shot-actions.md`](../devrites-lib/reference/standards/one-shot-actions.md),
26
+ the retained bounded artifact
27
+ is the reproduction input and the action MUST NOT be rerun during diagnosis.
28
+ Do not proceed without one of those reproduction inputs.
25
29
  3. **Ranked hypotheses (3-5, falsifiable):** generate the list before testing
26
30
  any of them. Each must state a prediction.
27
31
  **Completion:** 3-5 distinct hypotheses each state an observable prediction.
@@ -51,8 +55,14 @@ has no clear next move.
51
55
  - **Do NOT loosen / delete a failing assertion** to get green: check whether
52
56
  it's drift first (route via `/rite-plan repair`).
53
57
  - **Do NOT hide flakiness** with sleeps / retries: characterize it.
54
- - **Re-run the original loop after the fix.** The minimized regression test is not
55
- enough; prove the user-visible failure no longer reproduces.
58
+ - **Re-run the original loop after the fix when it is repeatable.** For a
59
+ consumptive action, first re-vet evidence completeness and obtain any required
60
+ fresh authorization; offline fixtures remain mandatory but cannot authorize the
61
+ real attempt.
62
+ - **A spent consumptive authorization is not a spent recovery budget.** When its
63
+ retained artifact supplies a new Critical/Important fingerprint, continue
64
+ offline diagnosis, correction, fixtures, and narrow Vet under that fingerprint's
65
+ no-progress budget. Stop for fresh authorization only before the next real action.
56
66
  - **Classify before routing** with
57
67
  [cleanup-and-classify.md](reference/cleanup-and-classify.md).
58
68
  - **Durably record class and rationale** in `decisions.md` and the applicable
@@ -62,7 +72,9 @@ has no clear next move.
62
72
  reproduction plus decisive signal rather than hashing symptom text.
63
73
  The caller and recovery attempts share one count: read the current context and
64
74
  recorded `## Dead ends` / `evidence.md`, then include every no-progress attempt
65
- with that fingerprint. Reclassify only on new causal evidence.
75
+ with that fingerprint. Reclassify only on new causal evidence. On cold resume,
76
+ a retained fingerprint with fewer than three such attempts remains runnable
77
+ even if the previous action wrote a terminal cursor.
66
78
  - **A maximum of three no-progress attempts per exact causal fingerprint stops the loop.**
67
79
  Count an attempt only when its recheck preserves the same decisive failure.
68
80
  Record attempt number, exact failure, hypothesis, probe, and failed idea after
@@ -14,7 +14,10 @@ record the dead end and classify anew. Failure alone never resets the budget.
14
14
  - `preexisting`: the same failure exists outside the candidate delta. Record the baseline and fix it only when it blocks acceptance.
15
15
  - `not_a_defect`: the observation matches current accepted authority. Record that authority and continue.
16
16
 
17
- Only human credentials/quotas/actions or irreversible work pause; never ask to retry.
17
+ Only human credentials/quotas/actions, irreversible work, or fresh authorization
18
+ required before the next consumptive execution pause. Fresh authorization never
19
+ pauses offline diagnosis or correction from retained evidence; never ask for a
20
+ blind retry.
18
21
 
19
22
  Record class/routing in `decisions.md` and each failed attempt in `evidence.md`
20
23
  or `## Dead ends`. Use one stable causal fingerprint shaped as `<affected
@@ -30,7 +33,8 @@ file or command.
30
33
 
31
34
  ## Cleanup checklist: required before declaring done
32
35
 
33
- - [ ] Original repro no longer reproduces (re-run the Phase 1 loop).
36
+ - [ ] Original repeatable repro no longer reproduces, or the consumptive action's
37
+ offline regression and evidence-completeness fixtures pass without a rerun.
34
38
  - [ ] Regression test passes (or absence of seam is documented).
35
39
  - [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix).
36
40
  - [ ] Throwaway harnesses deleted (or moved to a clearly marked debug location).