task-pipeline-skill 1.79.1 → 1.81.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. package/CHANGELOG.md +224 -0
  2. package/CONTRIBUTING.md +15 -0
  3. package/README.md +4 -3
  4. package/SKILL-CARD.md +1 -1
  5. package/cursor/rules/task-pipeline.mdc +3 -1
  6. package/evals/RESULTS.md +212 -9
  7. package/evals/evidence-docs.evals.json +109 -0
  8. package/evals/project-audit.evals.json +108 -0
  9. package/evals/run.py +47 -19
  10. package/package.json +4 -2
  11. package/plugins/task-pipeline/.claude-plugin/plugin.json +3 -2
  12. package/plugins/task-pipeline/commands/task-pipeline.md +2 -1
  13. package/plugins/task-pipeline/hooks/gate-observer.sh +19 -2
  14. package/plugins/task-pipeline/skills/evidence-docs/SKILL.md +1 -0
  15. package/plugins/task-pipeline/skills/project-audit/SKILL.md +7 -2
  16. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +49 -56
  17. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +1 -1
  18. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +1 -1
  19. package/plugins/task-pipeline/skills/task-pipeline/references/adoption.md +15 -4
  20. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +4 -2
  21. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +6 -1
  22. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +11 -1
  23. package/plugins/task-pipeline/skills/task-pipeline/references/certification.md +7 -0
  24. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +33 -2
  25. package/plugins/task-pipeline/skills/task-pipeline/references/documentation.md +2 -2
  26. package/plugins/task-pipeline/skills/task-pipeline/references/exposure.md +7 -3
  27. package/plugins/task-pipeline/skills/task-pipeline/references/gates.md +17 -185
  28. package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +13 -0
  29. package/plugins/task-pipeline/skills/task-pipeline/references/portability.md +1 -0
  30. package/plugins/task-pipeline/skills/task-pipeline/references/probing.md +202 -0
  31. package/plugins/task-pipeline/skills/task-pipeline/references/progress.md +8 -4
  32. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +60 -1
  33. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +32 -54
  34. package/plugins/task-pipeline/skills/task-pipeline/references/work-graph.md +1 -1
  35. package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +28 -5
  36. package/plugins/task-pipeline/skills/task-pipeline/templates/backlog.md +6 -2
  37. package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +2 -2
@@ -0,0 +1,109 @@
1
+ {
2
+ "_note": "Behavioural evaluations for the evidence-docs skill — the ten-canon navigator that ships in this plugin. Same format and dimensions as task-pipeline.evals.json; the negatives are NEAR MISSES on purpose (a commit message and a code comment are the excluded surfaces closest to the trigger ground), because a control that shares no words with the trigger proves nothing about the boundary.",
3
+ "skill": "evidence-docs",
4
+ "models": ["haiku", "sonnet", "opus"],
5
+ "evals": [
6
+ {
7
+ "id": "TRIG-01",
8
+ "category": "should_trigger",
9
+ "skills": ["evidence-docs"],
10
+ "query": "запиши решение: внутренние сервисы переходят с REST на gRPC, обратной дороги нет",
11
+ "expected_behavior": [
12
+ "Invokes the evidence-docs skill (or its routed documentation doctrine), not a freehand note",
13
+ "The decision lands in the project's one decision register, append-only, with an id",
14
+ "The record carries its evidence and its date rather than only the conclusion"
15
+ ],
16
+ "why": "«записать решение» is a declared Russian trigger; the failure it guards is a decision written into chat instead of the register."
17
+ },
18
+ {
19
+ "id": "TRIG-02",
20
+ "category": "should_trigger",
21
+ "skills": ["evidence-docs"],
22
+ "query": "Write the acceptance report for the export migration — and make sure every claim in it is actually backed",
23
+ "expected_behavior": [
24
+ "Invokes the evidence-docs skill for the canon layer (a claim carries its address; numbers computed, never restated)",
25
+ "Each verified claim names the check that was seen deciding it; unverified ones are marked as such rather than dropped"
26
+ ],
27
+ "why": "An acceptance report is the archetypal will-be-read-as-true document; the canons are its standard."
28
+ },
29
+ {
30
+ "id": "TRIG-03",
31
+ "category": "should_trigger",
32
+ "skills": ["evidence-docs"],
33
+ "query": "Is this verified? The README claims the importer handles 10k rows per second.",
34
+ "expected_behavior": [
35
+ "Invokes the evidence-docs skill — the question is whether a claim is documentation or an assertion",
36
+ "Answers with the claim's evidence condition: what observable signal would license the number, and whether one exists",
37
+ "Does not re-assert the number without an address"
38
+ ],
39
+ "why": "'is this verified' is a declared trigger; the failure is grading the claim by plausibility instead of by evidence."
40
+ },
41
+ {
42
+ "id": "NOTRIG-01",
43
+ "category": "should_not_trigger",
44
+ "skills": ["evidence-docs"],
45
+ "query": "Write a commit message for this diff — it renames the exporter and fixes the retry backoff",
46
+ "expected_behavior": [
47
+ "Writes the commit message directly",
48
+ "Does NOT invoke the evidence-docs skill or demand file:line receipts for a commit message"
49
+ ],
50
+ "why": "Near miss by design: a commit message describes a change truthfully but is an excluded surface — demanding canons here teaches routing around them."
51
+ },
52
+ {
53
+ "id": "NOTRIG-02",
54
+ "category": "should_not_trigger",
55
+ "skills": ["evidence-docs"],
56
+ "query": "добавь комментарий в код, почему в этом месте ретрай с экспоненциальной задержкой",
57
+ "expected_behavior": [
58
+ "Writes the code comment directly",
59
+ "Does NOT invoke the evidence-docs skill — a code comment is an excluded surface by the skill's own Not-for list"
60
+ ],
61
+ "why": "Near miss: 'почему' invites an evidence reflex, and the boundary says a code comment is not documentation-read-as-true."
62
+ },
63
+ {
64
+ "id": "NOTRIG-03",
65
+ "category": "should_not_trigger",
66
+ "skills": ["evidence-docs"],
67
+ "query": "Draft some rough notes on how we might approach caching — just thinking out loud, no docs",
68
+ "expected_behavior": [
69
+ "Produces the draft directly",
70
+ "Does NOT apply the canons — 'без доков'/'no docs' is the declared opt-out, and a draft is outside the boundary anyway"
71
+ ],
72
+ "why": "The refusal phrase plus the draft exclusion; the failure is refusing exploratory writing for want of receipts."
73
+ },
74
+ {
75
+ "id": "AMB-01",
76
+ "category": "ambiguous",
77
+ "skills": ["evidence-docs"],
78
+ "query": "Update the README",
79
+ "expected_behavior": [
80
+ "Names the route: a developer README section is direct work, while a claim a reader will take as true (a benchmark, a guarantee) passes the canons",
81
+ "Does not silently pick one reading of the request"
82
+ ],
83
+ "why": "A README holds both excluded and covered content; the observable is naming the split, not the choice itself."
84
+ },
85
+ {
86
+ "id": "COEX-01",
87
+ "category": "coexistence",
88
+ "skills": ["evidence-docs", "task-pipeline"],
89
+ "query": "прогони миграцию через конвейер и запиши архитектурное решение в реестр",
90
+ "expected_behavior": [
91
+ "task-pipeline carries the change; the decision record goes through the documentation doctrine evidence-docs routes",
92
+ "One register entry with an id — not a second decision home invented for the run"
93
+ ],
94
+ "why": "The two skills share a plugin and a boundary: the pipeline owns delivery, evidence-docs owns what is written as true."
95
+ },
96
+ {
97
+ "id": "INSTR-01",
98
+ "category": "instruction_following",
99
+ "skills": ["evidence-docs"],
100
+ "query": "Record the decision that we drop Python 3.8 support, and note it was discussed in Slack",
101
+ "expected_behavior": [
102
+ "The record is appended to the existing register, never a new file beside it",
103
+ "A correction or later reversal would be appended, not edited over",
104
+ "The Slack discussion is cited as the decision's source, not pasted as its evidence"
105
+ ],
106
+ "why": "Append-only and one-home are the canons most often broken by a helpful rewrite."
107
+ }
108
+ ]
109
+ }
@@ -0,0 +1,108 @@
1
+ {
2
+ "_note": "Behavioural evaluations for the project-audit skill. Same format and dimensions as task-pipeline.evals.json. The negatives are NEAR MISSES on purpose — «аудит модуля» against «аудит проекта» is one word apart and routes to a different skill (the pipeline's in-run ladder), which is exactly the boundary the descriptions draw.",
3
+ "skill": "project-audit",
4
+ "models": ["haiku", "sonnet", "opus"],
5
+ "evals": [
6
+ {
7
+ "id": "TRIG-01",
8
+ "category": "should_trigger",
9
+ "skills": ["project-audit"],
10
+ "query": "сделай аудит проекта — что реально готово, что наполовину, что сломано?",
11
+ "expected_behavior": [
12
+ "Invokes the project-audit skill, not the pipeline's in-run audit ladder",
13
+ "Starts with discovery (what the project IS) before choosing probes",
14
+ "Leaves the HTML report and the JSON sidecar; proposes board rows and commits nothing"
15
+ ],
16
+ "why": "«аудит проекта» is the declared trigger; the failure is running a fixed checklist or treating it as a change-task."
17
+ },
18
+ {
19
+ "id": "TRIG-02",
20
+ "category": "should_trigger",
21
+ "skills": ["project-audit"],
22
+ "query": "What is actually true of this project right now — what is finished, what is half-built, what has nobody looked at?",
23
+ "expected_behavior": [
24
+ "Invokes the project-audit skill — the subject is the whole project, not one change",
25
+ "Reads production evidence (published artefact vs source, CI history, telemetry presence), not only the working tree"
26
+ ],
27
+ "why": "The skill's own opening sentence as a user query; the failure is answering from the README."
28
+ },
29
+ {
30
+ "id": "TRIG-03",
31
+ "category": "should_trigger",
32
+ "skills": ["project-audit"],
33
+ "query": "Run a project health check on this repository and leave me the report",
34
+ "expected_behavior": [
35
+ "Invokes the project-audit skill",
36
+ "Blind probes are rendered as their own section with reasons — a probe that could not look is not a probe that found nothing"
37
+ ],
38
+ "why": "'project health check' is a declared trigger; the blind verdict is the design worth probing for."
39
+ },
40
+ {
41
+ "id": "NOTRIG-01",
42
+ "category": "should_not_trigger",
43
+ "skills": ["project-audit"],
44
+ "query": "сделай аудит модуля оплат",
45
+ "expected_behavior": [
46
+ "Routes to the pipeline's audit path (a finding that lands in the repository), not to project-audit",
47
+ "Does NOT start a cold whole-project discovery for a one-module question"
48
+ ],
49
+ "why": "Near miss by one word: «аудит проекта» is this skill, «аудит модуля» is one deliverable inside a run — the pipeline's own ladder."
50
+ },
51
+ {
52
+ "id": "NOTRIG-02",
53
+ "category": "should_not_trigger",
54
+ "skills": ["project-audit"],
55
+ "query": "Review PR #24 and tell me what is wrong with it",
56
+ "expected_behavior": [
57
+ "Routes to the pipeline's PR-review findings path or answers directly",
58
+ "Does NOT invoke project-audit — reviewing a diff is its declared Not-for"
59
+ ],
60
+ "why": "A diff has a change as its subject; project-audit's subject is the project."
61
+ },
62
+ {
63
+ "id": "NOTRIG-03",
64
+ "category": "should_not_trigger",
65
+ "skills": ["project-audit"],
66
+ "query": "проверь, соответствует ли этот скил стандарту Agent Skills",
67
+ "expected_behavior": [
68
+ "Routes to make-skill's audit (/skill-audit), not to project-audit",
69
+ "Does NOT run whole-project probes over a skill-construction question"
70
+ ],
71
+ "why": "The disambiguation table's own row: a skill's construction belongs to make-skill even when the word 'audit' appears."
72
+ },
73
+ {
74
+ "id": "AMB-01",
75
+ "category": "ambiguous",
76
+ "skills": ["project-audit"],
77
+ "query": "проверь проект",
78
+ "expected_behavior": [
79
+ "Names the route it is taking in one line — whole-project diagnosis (project-audit) versus a specific check the operator may mean",
80
+ "Does not silently start either the full audit or a random spot-check"
81
+ ],
82
+ "why": "Two words with no object; the observable is the named route, not the choice."
83
+ },
84
+ {
85
+ "id": "COEX-01",
86
+ "category": "coexistence",
87
+ "skills": ["project-audit", "task-pipeline"],
88
+ "query": "Audit the whole project, then fix the three worst things you find",
89
+ "expected_behavior": [
90
+ "project-audit produces the findings as proposed board rows, read-only",
91
+ "The fixes are carried by task-pipeline runs off those rows — the audit itself commits nothing"
92
+ ],
93
+ "why": "The seam the two skills share: diagnosis is read-only, delivery is the pipeline's; collapsing them makes the audit unrepeatable."
94
+ },
95
+ {
96
+ "id": "INSTR-01",
97
+ "category": "instruction_following",
98
+ "skills": ["project-audit"],
99
+ "query": "Audit this project",
100
+ "expected_behavior": [
101
+ "Discovery runs first and the probes are chosen from the profile, not from a fixed list",
102
+ "The three-verdict vocabulary is kept: clean, finding, blind — with blind reasons on the page",
103
+ "Findings leave priced with the board header's declared formula; effort never ranks"
104
+ ],
105
+ "why": "The procedure's own load-bearing steps, each of which a helpful shortcut would skip."
106
+ }
107
+ ]
108
+ }
package/evals/run.py CHANGED
@@ -17,12 +17,17 @@ What it does:
17
17
 
18
18
  Zero dependencies, same as the validator.
19
19
  """
20
+ import glob
20
21
  import json
21
22
  import os
22
23
  import re
23
24
  import sys
24
25
 
25
26
  ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
27
+ # EVERY suite, discovered — the directory held one suite per skill from
28
+ # 2026-08-31 (evidence-docs and project-audit joined task-pipeline), and a
29
+ # runner pinned to one filename would validate a third of what ships.
30
+ SUITES = sorted(glob.glob(os.path.join(ROOT, "evals", "*.evals.json")))
26
31
  SUITE = os.path.join(ROOT, "evals", "task-pipeline.evals.json")
27
32
  RESULTS = os.path.join(ROOT, "evals", "RESULTS.md")
28
33
 
@@ -34,14 +39,12 @@ REQUIRED = ("should_trigger", "should_not_trigger", "ambiguous",
34
39
  MIN_EVALS = 3 # Anthropic: "At least three evaluations created"
35
40
 
36
41
 
37
- def main(argv):
42
+ def validate_suite(path):
43
+ """Every gap in one suite, in a stable order. The rules are the same for
44
+ every skill's suite — a second rule set would drift."""
38
45
  errors = []
39
- if not os.path.isfile(SUITE):
40
- print(f"FAIL: no suite at {os.path.relpath(SUITE, ROOT)}")
41
- return 2
42
- suite = json.load(open(SUITE, encoding="utf-8"))
46
+ suite = json.load(open(path, encoding="utf-8"))
43
47
  evals = suite.get("evals") or []
44
-
45
48
  seen = set()
46
49
  for e in evals:
47
50
  where = e.get("id", "<no id>")
@@ -68,6 +71,18 @@ def main(argv):
68
71
  for cat in REQUIRED:
69
72
  if cat not in covered:
70
73
  errors.append(f"no eval covers {cat!r}")
74
+ return suite, evals, errors
75
+
76
+
77
+ def main(argv):
78
+ if not os.path.isfile(SUITE):
79
+ print(f"FAIL: no suite at {os.path.relpath(SUITE, ROOT)}")
80
+ return 2
81
+ parsed, errors = [], []
82
+ for _sp in SUITES:
83
+ _suite, _evals, _errs = validate_suite(_sp)
84
+ parsed.append((os.path.basename(_sp), _suite, _evals))
85
+ errors += [f"{os.path.basename(_sp)}: {e}" for e in _errs]
71
86
 
72
87
  if errors:
73
88
  print("FAIL: evaluation suite invalid")
@@ -75,31 +90,42 @@ def main(argv):
75
90
  print(" - " + e)
76
91
  return 1
77
92
 
93
+ suite = next(s for n, s, ev in parsed if n == "task-pipeline.evals.json")
94
+ evals = [e for _, _, ev in parsed for e in ev]
78
95
  by_cat = {}
79
96
  for e in evals:
80
97
  by_cat.setdefault(e["category"], []).append(e)
81
98
 
82
99
  if "--list" in argv:
83
- for cat in REQUIRED:
84
- for e in by_cat.get(cat, []):
85
- print(f" {e['id']:<10} {cat:<22} {e['query'][:60]}")
86
- print(f"\n{len(evals)} evals across {len(by_cat)} categories")
100
+ for name, _s, ev in parsed:
101
+ print(f"{name} — {_s.get('skill', '?')}")
102
+ for cat in REQUIRED:
103
+ for e in ev:
104
+ if e["category"] == cat:
105
+ print(f" {e['id']:<10} {cat:<22} {e['query'][:60]}")
106
+ print(f"\n{len(evals)} evals across {len(by_cat)} categories, "
107
+ f"{len(parsed)} suite(s)")
87
108
  return 0
88
109
 
89
110
  print("=" * 72)
90
- print("task-pipeline evaluation protocol")
111
+ print("task-pipeline plugin evaluation protocol (one section per suite)")
91
112
  print("=" * 72)
92
- print("Run each query in a FRESH session with the skill installed, once per")
113
+ print("Run each query in a FRESH session with the pack installed, once per")
93
114
  print("model in", suite.get("models", []), "— effectiveness varies by model.")
94
115
  print("Record every verdict in evals/RESULTS.md with the date and the model.")
95
116
  print("A query you did not run is not a pass; leave it blank and say so.\n")
96
- for cat in REQUIRED:
97
- print(f"\n--- {cat} ---")
98
- for e in by_cat.get(cat, []):
99
- print(f"\n[{e['id']}] {e['query']}")
100
- print(f" why: {e['why']}")
101
- for b in e["expected_behavior"]:
102
- print(f" [ ] {b}")
117
+ for name, _s, ev in parsed:
118
+ print(f"\n=== {name} — {_s.get('skill', '?')} ===")
119
+ for cat in REQUIRED:
120
+ group = [e for e in ev if e["category"] == cat]
121
+ if not group:
122
+ continue
123
+ print(f"\n--- {cat} ---")
124
+ for e in group:
125
+ print(f"\n[{e['id']}] {e['query']}")
126
+ print(f" why: {e['why']}")
127
+ for b in e["expected_behavior"]:
128
+ print(f" [ ] {b}")
103
129
 
104
130
  print("\n" + "=" * 72)
105
131
  if not os.path.isfile(RESULTS):
@@ -118,6 +144,8 @@ def main(argv):
118
144
  if not infence:
119
145
  outside.append(ln)
120
146
  runs = [l for l in outside if re.match(r"^## 20\d{2}-\d{2}-\d{2}\b", l)]
147
+ print("suites: " + " · ".join(f"{n.replace('.evals.json', '')} {len(ev)}"
148
+ for n, _s, ev in parsed))
121
149
  print(f"suite: {len(evals)} evals · recorded runs: {len(runs)}")
122
150
  if not runs:
123
151
  print("RESULTS.md carries no dated run — the suite is authored and unexecuted.")
package/package.json CHANGED
@@ -1,16 +1,18 @@
1
1
  {
2
2
  "name": "task-pipeline-skill",
3
- "version": "1.79.1",
3
+ "version": "1.81.1",
4
4
  "description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
5
5
  "bin": {
6
6
  "task-pipeline": "bin/task-pipeline.js"
7
7
  },
8
8
  "scripts": {
9
9
  "test": "python3 test/validate.py && python3 test/graph_test.py && python3 test/project_audit_test.py",
10
- "test:all": "python3 test/validate.py && python3 test/graph_test.py && python3 test/project_audit_test.py && python3 test/negatives.py && npm run test:certify && npm run test:exposure && npm run test:probe && npm run test:hooks && npm run test:artifacts && npm run test:docs",
10
+ "test:all": "python3 test/validate.py && python3 test/graph_test.py && python3 test/project_audit_test.py && python3 test/negatives.py && npm run test:certify && npm run test:exposure && npm run test:probe && npm run test:anchors && npm run test:runner && npm run test:hooks && npm run test:artifacts && npm run test:docs",
11
11
  "test:negatives": "python3 test/negatives.py",
12
12
  "test:exposure": "python3 test/exposure_test.py",
13
13
  "test:probe": "python3 test/probe.py --self-test",
14
+ "test:anchors": "python3 test/anchors_test.py && python3 test/anchors.py",
15
+ "test:runner": "python3 test/runner_test.py",
14
16
  "test:hooks": "python3 test/release_gate_test.py",
15
17
  "test:artifacts": "python3 test/artifact_root_test.py && python3 test/migrate_artifacts_test.py",
16
18
  "test:docs": "bash plugins/task-pipeline/skills/task-pipeline/templates/docgate.sh",
@@ -1,8 +1,9 @@
1
1
  {
2
+ "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
2
3
  "name": "task-pipeline",
3
4
  "displayName": "Task Pipeline",
4
- "description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
5
- "version": "1.79.1",
5
+ "description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/judgment/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
6
+ "version": "1.81.1",
6
7
  "author": {
7
8
  "name": "ssheleg",
8
9
  "url": "https://x.com/sshlg93"
@@ -100,7 +100,8 @@ Anything deferred enters the carry-over ledger the moment it is said.
100
100
  | 9 | Docs + wiki | **three artifacts, not two** — module docs, the wiki, **and the code graph** |
101
101
  | 10 | Acceptance | the ladder walk first, then the table, then the retrospective |
102
102
 
103
- **Honor every gate by its type**: `auto` — verify the check yourself; `manual` — wait
103
+ **Honor every gate by its type**: `auto` — verify the check yourself; `judgment` —
104
+ record the named judge's ruling as judgement, never as a measurement; `manual` — wait
104
105
  for an explicit go.
105
106
 
106
107
  ## Cross-cutting — the three that fire at any stage
@@ -73,11 +73,28 @@ if not matches:
73
73
  # PostToolUse fires on success; PostToolUseFailure carries the error. Both are
74
74
  # wired to this script, and `error` present means the command did not exit 0.
75
75
  failed = bool(data.get("error")) or data.get("hook_event_name") == "PostToolUseFailure"
76
- out = data.get("tool_output") or {}
77
- if isinstance(out, dict) and out.get("exit_code") is not None:
76
+ # The harness documents the result field as `tool_response`; `tool_output` is the
77
+ # name this script shipped reading, so it stays as a fallback rather than a
78
+ # breaking change. Reading only the wrong name left the exit-code branch dead.
79
+ # The fallback is BY FIELD, not by object: the real Bash `tool_response` is
80
+ # {stdout, stderr, interrupted} with no exit_code, and `resp or legacy` made a
81
+ # legacy exit_code unreachable behind it — found by the R-005 reader, measured,
82
+ # before this shipped. The same reading found `interrupted`: a gate cut short
83
+ # is not a gate that passed, whatever a stale exit_code says, so a failure
84
+ # event or an interruption is never recorded as exit 0.
85
+ out = {}
86
+ interrupted = False
87
+ for cand in (data.get("tool_response"), data.get("tool_output")):
88
+ if isinstance(cand, dict):
89
+ interrupted = interrupted or bool(cand.get("interrupted"))
90
+ if not out and cand.get("exit_code") is not None:
91
+ out = cand
92
+ if out:
78
93
  code = int(out["exit_code"])
79
94
  else:
80
95
  code = 1 if failed else 0
96
+ if (failed or interrupted) and code == 0:
97
+ code = 1
81
98
 
82
99
  stamp = datetime.datetime.now(datetime.timezone.utc).replace(microsecond=0).isoformat().replace("+00:00", "Z")
83
100
 
@@ -24,6 +24,7 @@ The full statement of each canon, its rationale and its enforcement live in
24
24
  7. **Silence is not a pass** — ask what a mechanism prints when it did not look.
25
25
  8. **An estimate is never announced as a measurement** — a rule states its evidence condition.
26
26
  9. **What was not checked is printed beside what was.**
27
+ - **9a. A measured zero and an unmeasured quantity may not print the same** — canon 9 says carry the absence; 9a says refuse the number when nothing measured it.
27
28
  10. **The document ships in the change that made it true** — and a correction is appended, never written over.
28
29
 
29
30
  They are **epistemic**: what makes a claim documentation. The operational layer — what to
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: project-audit
3
3
  description: "Use when someone asks what is actually true of a whole project right now — what is finished, what is half-built, what is broken, and what nobody has looked at. Walks a cold start: discover what the project is, run a registry of probes chosen from that, read production evidence (published artefact against source, CI history, telemetry present or absent), then leave a self-contained HTML report and a JSON sidecar so the next audit can say what moved. Read-only: it proposes board rows and commits nothing. Triggers - 'project audit', 'audit the project', 'codebase audit', 'state of the project', 'what is unfinished', 'project health check', 'аудит проекта', 'проаудируй проект', 'состояние проекта', 'что не доделано', 'аудит кодовой базы'. Not for: auditing one deliverable inside a run (that is the pipeline's own ladder), reviewing a diff, or checking a skill's construction — say 'без диагностики' to opt out."
4
+ compatibility: "The collector (scripts/audit.py) needs python3 and reads committed state, so it needs git. Probes needing gh, npm, network or a browser declare it and report blind when it is absent — degraded, never silent."
4
5
  ---
5
6
 
6
7
  # Project audit — what is true of this project right now
@@ -33,6 +34,7 @@ next audit reads.
33
34
  | this skill | the **procedure** — cold start, probes, production, the report | a whole project is the subject |
34
35
  | `/skill-audit` (make-skill) | a skill's construction against the standard | the thing audited is a skill or plugin |
35
36
  | `/ux-audit` (super-ux) | code against documented scenarios | the question is user-facing behaviour |
37
+ | `/seo-aeo-audit` (seo-aeo-audit) | a public surface's search and answer-engine visibility | the question is whether a machine will find it |
36
38
 
37
39
  **The method is not restated here.** Phase 4 below hands off to `audit.md` and
38
40
  comes back; a second copy of the ladder would be a second rule, and the two
@@ -96,8 +98,11 @@ the same object.
96
98
  ### 6. Propose — rows, not edits
97
99
 
98
100
  **This skill commits nothing.** Findings leave as board rows in the project's
99
- own vocabulary, priced with the project's own formula —
100
- `P = blast × (1 + age_runs) / effort` — and the operator accepts them. An audit
101
+ own vocabulary, priced with **the board header's declared formula** — the shipped
102
+ default is `Sev × Blast + age_bonus` (`references/backlog.md`, the pipeline's
103
+ board doctrine) — and the operator accepts them. Effort never ranks inside an
104
+ audit: what a fix costs is the fixer's decision, not the finder's
105
+ (`references/prioritisation.md`). An audit
101
106
  that edits while it reads cannot be re-run to check itself.
102
107
 
103
108
  ## Three verdicts, and why the third one exists