@tea-agent/loop-agent 0.13.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (270) hide show
  1. package/AGENTS.md +157 -157
  2. package/CHANGELOG.md +73 -301
  3. package/README.md +338 -334
  4. package/bin/agent-worker.js +22 -22
  5. package/bin/loop-agent.js +21 -21
  6. package/dist/commands/cursor-prompt.js +6 -6
  7. package/dist/commands/init.js +505 -505
  8. package/dist/commands/loop-benchmark.js +11 -11
  9. package/dist/commands/pi-reuse-benchmark.js +16 -16
  10. package/dist/executors/pi-event-serializer.js +33 -11
  11. package/dist/sidecars/cursor-prompt/executor.js +1 -1
  12. package/dist/task/runtime.js +27 -27
  13. package/dist/worker/observe/spec-evidence.js +19 -10
  14. package/dist/worker/observe/static/api.js +46 -46
  15. package/dist/worker/observe/static/app.js +151 -150
  16. package/dist/worker/observe/static/constants.js +156 -148
  17. package/dist/worker/observe/static/copy.js +67 -67
  18. package/dist/worker/observe/static/dag-helpers.js +201 -172
  19. package/dist/worker/observe/static/dag-layout.d.ts +31 -31
  20. package/dist/worker/observe/static/dag-layout.js +83 -83
  21. package/dist/worker/observe/static/dag-model.js +72 -72
  22. package/dist/worker/observe/static/dom.js +122 -122
  23. package/dist/worker/observe/static/format-pool.d.ts +71 -0
  24. package/dist/worker/observe/static/format-pool.js +134 -67
  25. package/dist/worker/observe/static/format.js +317 -292
  26. package/dist/worker/observe/static/index.html +350 -308
  27. package/dist/worker/observe/static/kpi.js +100 -94
  28. package/dist/worker/observe/static/markdown-render.js +124 -0
  29. package/dist/worker/observe/static/relations.js +133 -133
  30. package/dist/worker/observe/static/router.js +93 -93
  31. package/dist/worker/observe/static/run-processing.js +148 -148
  32. package/dist/worker/observe/static/shell-chrome.js +74 -68
  33. package/dist/worker/observe/static/state.js +273 -267
  34. package/dist/worker/observe/static/styles.css +2504 -1902
  35. package/dist/worker/observe/static/views/batch.js +227 -227
  36. package/dist/worker/observe/static/views/dag-graph.js +172 -172
  37. package/dist/worker/observe/static/views/dag-inspector.js +530 -627
  38. package/dist/worker/observe/static/views/dag.js +371 -371
  39. package/dist/worker/observe/static/views/dashboard.js +86 -100
  40. package/dist/worker/observe/static/views/failures.js +143 -143
  41. package/dist/worker/observe/static/views/feature.js +492 -492
  42. package/dist/worker/observe/static/views/pool.js +708 -350
  43. package/dist/worker/observe/static/views/run.js +453 -453
  44. package/dist/worker/observe/static/views/session-timeline.js +771 -219
  45. package/dist/worker/observe/static/views/shell.js +7 -7
  46. package/dist/worker/observe/static/views/task.js +314 -314
  47. package/dist/worker/observe/static/views/timeline.js +163 -163
  48. package/dist/workflows/dag/canvas-observer.js +275 -275
  49. package/docs/README.md +105 -104
  50. package/docs/agent-dag-recovery-playbook.md +195 -195
  51. package/docs/agent-dag-runner.md +67 -67
  52. package/docs/architecture/README.md +26 -26
  53. package/docs/architecture/dag-execution.md +140 -140
  54. package/docs/architecture/evolution.md +54 -54
  55. package/docs/architecture/facts-and-state.md +71 -71
  56. package/docs/architecture/runtime-boundaries.md +191 -191
  57. package/docs/architecture/system-overview.md +93 -93
  58. package/docs/architecture/worker-and-feature.md +85 -85
  59. package/docs/cursor-prompt-sidecar.md +36 -36
  60. package/docs/decisions/README.md +18 -18
  61. package/docs/design/README.md +167 -167
  62. package/docs/development-principles.md +73 -73
  63. package/docs/exec-plans/README.md +6 -6
  64. package/docs/exec-plans/active/README.md +2 -1
  65. package/docs/exec-plans/completed/README.md +105 -104
  66. package/docs/feature-workflow.md +414 -414
  67. package/docs/harness-methodology-debugging.md +153 -153
  68. package/docs/harness-methodology-tdd.md +130 -130
  69. package/docs/harness-methodology-verification.md +27 -27
  70. package/docs/init-surface.manifest.json +307 -307
  71. package/docs/loop-agent-harness.md +142 -142
  72. package/docs/production-readiness.md +96 -96
  73. package/docs/progress/README.md +59 -58
  74. package/docs/reports/README.md +123 -119
  75. package/docs/skills/README.md +7 -7
  76. package/docs/skills/vetted-skill-registry.md +29 -29
  77. package/docs/templates/adr.md +60 -60
  78. package/docs/templates/agent-dag-authority-surface-audit.prompt.md +94 -94
  79. package/docs/templates/agent-dag-decision-envelope.schema.json +213 -213
  80. package/docs/templates/agent-dag-decision-gate-dogfood-report.md +117 -117
  81. package/docs/templates/agent-dag-decision-gate.prompt.md +246 -246
  82. package/docs/templates/agent-dag-process-supervisor.prompt.md +98 -98
  83. package/docs/templates/agent-dag-report.schema.json +473 -473
  84. package/docs/templates/agent-dag-review-verdict.prompt.md +68 -68
  85. package/docs/templates/agent-dag.base.json +190 -190
  86. package/docs/templates/agent-dag.final-verification.json +185 -185
  87. package/docs/templates/agent-dag.schema.json +411 -411
  88. package/docs/templates/agent-dag.supervised-implementation.json +620 -620
  89. package/docs/templates/backend-test-analysis.schema.json +44 -44
  90. package/docs/templates/backend-test-case-manifest.schema.json +190 -190
  91. package/docs/templates/backend-test-dag.classify.prompt.md +75 -75
  92. package/docs/templates/backend-test-dag.generate-pytest.prompt.md +204 -204
  93. package/docs/templates/backend-test-dag.json +559 -559
  94. package/docs/templates/backend-test-dag.retrospect.prompt.md +139 -139
  95. package/docs/templates/backend-test-dag.review-cases.prompt.md +83 -83
  96. package/docs/templates/backend-test-execution.schema.json +133 -133
  97. package/docs/templates/backend-test-result.schema.json +99 -99
  98. package/docs/templates/exec-plan.md +64 -64
  99. package/docs/templates/feature-spec.md +53 -53
  100. package/docs/templates/frontend-design-contract.md +42 -42
  101. package/docs/templates/frontend-eval/fixtures/failures/01-type-build-error.md +17 -17
  102. package/docs/templates/frontend-eval/fixtures/failures/02-unit-component-test-fail.md +16 -16
  103. package/docs/templates/frontend-eval/fixtures/failures/03-fixture-schema-drift.md +16 -16
  104. package/docs/templates/frontend-eval/fixtures/failures/04-missing-loading-empty-error-state.md +16 -16
  105. package/docs/templates/frontend-eval/fixtures/failures/05-forbidden-write-writeset-expansion.md +16 -16
  106. package/docs/templates/frontend-eval/fixtures/failures/06-unapproved-dependency-add.md +16 -16
  107. package/docs/templates/frontend-eval/fixtures/failures/07-mock-production-on.md +21 -21
  108. package/docs/templates/frontend-eval/fixtures/functional/01-simple-component-style.md +29 -29
  109. package/docs/templates/frontend-eval/fixtures/functional/02-form-validation.md +28 -28
  110. package/docs/templates/frontend-eval/fixtures/functional/03-list-detail-page.md +28 -28
  111. package/docs/templates/frontend-eval/fixtures/functional/04-api-mock.md +29 -29
  112. package/docs/templates/frontend-eval/fixtures/functional/05-permission-auth-gated-ui.md +27 -27
  113. package/docs/templates/frontend-eval/fixtures/functional/06-ssr-server-client-boundary.md +28 -28
  114. package/docs/templates/frontend-eval/fixtures/functional/07-shared-public-component-api.md +28 -28
  115. package/docs/templates/frontend-eval/fixtures/functional/08-pure-local-no-remote.md +27 -27
  116. package/docs/templates/frontend-eval/metrics.md +138 -138
  117. package/docs/templates/frontend-eval/smoke-targets.md +53 -53
  118. package/docs/templates/frontend-implementation-contract.schema.json +27 -27
  119. package/docs/templates/frontend-task-constraints.md +35 -35
  120. package/docs/templates/frontend-task-requirement.md +70 -70
  121. package/docs/templates/frontend-test-dag.generate-cases.prompt.md +5 -5
  122. package/docs/templates/frontend-test-dag.json +23 -23
  123. package/docs/templates/frontend-test-dag.retrieve-context.prompt.md +3 -3
  124. package/docs/templates/frontend-test-dag.retrospect.prompt.md +3 -3
  125. package/docs/templates/frontend-test-dag.review-cases.prompt.md +3 -3
  126. package/docs/templates/frontend-test-dag.review-execution.prompt.md +3 -3
  127. package/docs/templates/harness.schema.json +221 -221
  128. package/docs/templates/hybrid-dag.json +188 -188
  129. package/docs/templates/init-evolution-review.md +35 -35
  130. package/docs/templates/interactive-ui-round2-experiment.md +66 -66
  131. package/docs/templates/knowledge-graph-bootstrap-dag.json +118 -118
  132. package/docs/templates/knowledge-sync-dag.json +178 -178
  133. package/docs/templates/knowledge-sync-draft.schema.json +71 -71
  134. package/docs/templates/product-line/AGENTS.md +8 -8
  135. package/docs/templates/product-line/README.md +9 -9
  136. package/docs/templates/product-line/acceptance.yaml +14 -14
  137. package/docs/templates/product-line/closeout.yaml +9 -9
  138. package/docs/templates/product-line/design.md +13 -13
  139. package/docs/templates/product-line/links.md +10 -10
  140. package/docs/templates/product-line/requirement.md +17 -17
  141. package/docs/templates/product-line/task-graph.yaml +15 -15
  142. package/docs/templates/product-line/task.yaml +64 -64
  143. package/docs/templates/product-line/test-plan.md +7 -7
  144. package/docs/templates/production-readiness-checklist.md +57 -57
  145. package/docs/templates/progress-log.md +17 -17
  146. package/docs/templates/project-start-checklist.md +9 -9
  147. package/docs/templates/qa-report.md +48 -48
  148. package/docs/templates/sprint-contract.md +29 -29
  149. package/docs/templates/worker-dogfood-evidence.md +80 -80
  150. package/docs/templates/worker-dogfood-setup.md +68 -68
  151. package/docs/verification-matrix.md +70 -70
  152. package/examples/decision-gate-agent-dag.json +173 -173
  153. package/examples/example-dag.json +46 -46
  154. package/examples/hybrid-loop-agent-dag.json +188 -188
  155. package/harness.json +66 -66
  156. package/package.json +88 -52
  157. package/scripts/check-product-line-docs.sh +29 -29
  158. package/scripts/check-task-pool-root.sh +32 -32
  159. package/scripts/kb-bootstrap-init-skeleton.sh +240 -240
  160. package/scripts/kb-graph-incremental-prepare.mjs +386 -386
  161. package/scripts/kb-graph-incremental-prepare.sh +5 -5
  162. package/scripts/kb-graph-materialize.mjs +105 -105
  163. package/scripts/kb-graph-materialize.sh +4 -4
  164. package/scripts/kb-graph-promote.mjs +164 -164
  165. package/scripts/kb-graph-promote.sh +4 -4
  166. package/scripts/kb-query.mjs +554 -554
  167. package/scripts/kb-query.sh +5 -5
  168. package/skills/agent-worker/SKILL.md +39 -39
  169. package/skills/agent-worker/references/agent-worker-operator.md +60 -60
  170. package/skills/ai-engineering-context/SKILL.md +48 -48
  171. package/skills/analyze-product-dependencies/SKILL.md +67 -67
  172. package/skills/analyze-product-dependencies/agents/openai.yaml +4 -4
  173. package/skills/analyze-product-dependencies/references/api-documentation-schema.md +30 -30
  174. package/skills/analyze-product-dependencies/references/dependency-analysis-schema.md +28 -28
  175. package/skills/analyze-product-dependencies/references/example.md +76 -76
  176. package/skills/analyze-product-dependencies/references/forward-test-cases.md +35 -35
  177. package/skills/analyze-product-dependencies/references/input-contract.md +11 -11
  178. package/skills/analyze-product-dependencies/references/scouting-rules.md +61 -61
  179. package/skills/analyze-product-dependencies/scripts/test-validators.mjs +267 -267
  180. package/skills/analyze-product-dependencies/scripts/validate-api-documentation.mjs +101 -101
  181. package/skills/analyze-product-dependencies/scripts/validate-dependency-analysis.mjs +142 -142
  182. package/skills/analyze-product-dependencies/scripts/validate-product-requirement-input.mjs +76 -76
  183. package/skills/analyze-product-dependencies/scripts/validation-helpers.mjs +146 -146
  184. package/skills/analyze-product-requirements/SKILL.md +90 -90
  185. package/skills/analyze-product-requirements/agents/openai.yaml +4 -4
  186. package/skills/analyze-product-requirements/references/acceptance-criteria.md +91 -91
  187. package/skills/analyze-product-requirements/references/clarification-and-knowledge.md +56 -56
  188. package/skills/analyze-product-requirements/references/example.md +86 -86
  189. package/skills/analyze-product-requirements/references/forward-test-cases.md +66 -66
  190. package/skills/analyze-product-requirements/references/product-analysis-schema.md +32 -32
  191. package/skills/analyze-product-requirements/references/product-requirement-schema.md +33 -33
  192. package/skills/analyze-product-requirements/references/requirement-clarification-schema.md +35 -35
  193. package/skills/analyze-product-requirements/scripts/test-validators.mjs +193 -193
  194. package/skills/analyze-product-requirements/scripts/validate-product-analysis.mjs +69 -69
  195. package/skills/analyze-product-requirements/scripts/validate-product-requirement.mjs +97 -97
  196. package/skills/analyze-product-requirements/scripts/validate-requirement-clarification.mjs +98 -98
  197. package/skills/analyze-product-requirements/scripts/validation-helpers.mjs +156 -156
  198. package/skills/browser-tools/SKILL.md +196 -196
  199. package/skills/browser-tools/browser-content.js +103 -103
  200. package/skills/browser-tools/browser-cookies.js +35 -35
  201. package/skills/browser-tools/browser-eval.js +53 -53
  202. package/skills/browser-tools/browser-hn-scraper.js +108 -108
  203. package/skills/browser-tools/browser-nav.js +44 -44
  204. package/skills/browser-tools/browser-pick.js +162 -162
  205. package/skills/browser-tools/browser-screenshot.js +34 -34
  206. package/skills/browser-tools/browser-start.js +86 -86
  207. package/skills/browser-tools/package-lock.json +2556 -2556
  208. package/skills/browser-tools/package.json +19 -19
  209. package/skills/code-review-core/SKILL.md +20 -20
  210. package/skills/codebase-scout/SKILL.md +19 -19
  211. package/skills/frontend-design-review/SKILL.md +66 -66
  212. package/skills/frontend-design-review/references/review-checklist.md +58 -58
  213. package/skills/frontend-implementation/SKILL.md +49 -49
  214. package/skills/frontend-implementation/references/code-standards.md +32 -32
  215. package/skills/frontend-implementation/references/design-spec.md +46 -46
  216. package/skills/frontend-implementation/references/node-contracts.md +27 -27
  217. package/skills/frontend-review/SKILL.md +59 -59
  218. package/skills/frontend-review/references/review-findings.md +47 -47
  219. package/skills/frontend-verification/SKILL.md +53 -53
  220. package/skills/frontend-verification/references/verification-checklist.md +68 -68
  221. package/skills/grill-me/SKILL.md +10 -10
  222. package/skills/grill-with-docs/SKILL.md +88 -88
  223. package/skills/grill-with-docs/adr-format.md +47 -47
  224. package/skills/grill-with-docs/context-format.md +60 -60
  225. package/skills/init-capability-evolution/SKILL.md +70 -70
  226. package/skills/loop-agent/SKILL.md +151 -151
  227. package/skills/loop-agent/references/README.md +67 -67
  228. package/skills/loop-agent/references/command-reference.md +527 -527
  229. package/skills/loop-agent/references/docs-converge.md +126 -126
  230. package/skills/loop-agent/references/harness-policy.md +263 -263
  231. package/skills/loop-agent/references/hybrid-dag.md +243 -243
  232. package/skills/loop-agent/references/learned/README.md +21 -21
  233. package/skills/loop-agent/references/long-running-loop.md +57 -57
  234. package/skills/loop-agent/references/model-routing.md +36 -36
  235. package/skills/loop-agent/references/multi-worktree.md +54 -54
  236. package/skills/loop-agent/references/one-shot-runs.md +85 -85
  237. package/skills/loop-agent/references/orchestrator-and-interventions.md +169 -169
  238. package/skills/loop-agent/references/pi-prompt.md +23 -23
  239. package/skills/loop-agent/references/pi-subagent-assisted-mode.md +84 -84
  240. package/skills/loop-agent/references/post-implementation-and-patterns.md +44 -44
  241. package/skills/loop-agent/references/task-workflow.md +89 -89
  242. package/skills/loop-agent/references/verification-and-failure-handling.md +141 -141
  243. package/skills/playwright-cli/SKILL.md +420 -420
  244. package/skills/playwright-cli/references/element-attributes.md +23 -23
  245. package/skills/playwright-cli/references/playwright-tests.md +39 -39
  246. package/skills/playwright-cli/references/request-mocking.md +87 -87
  247. package/skills/playwright-cli/references/running-code.md +241 -241
  248. package/skills/playwright-cli/references/session-management.md +225 -225
  249. package/skills/playwright-cli/references/storage-state.md +275 -275
  250. package/skills/playwright-cli/references/test-generation.md +433 -433
  251. package/skills/playwright-cli/references/tracing.md +139 -139
  252. package/skills/playwright-cli/references/video-recording.md +143 -143
  253. package/skills/playwright-cli-case-generator/SKILL.md +74 -74
  254. package/skills/requesting-code-review/SKILL.md +101 -101
  255. package/skills/requesting-code-review/code-reviewer.md +168 -168
  256. package/skills/systematic-debugging/CREATION-LOG.md +119 -119
  257. package/skills/systematic-debugging/SKILL.md +296 -296
  258. package/skills/systematic-debugging/condition-based-waiting-example.ts +158 -158
  259. package/skills/systematic-debugging/condition-based-waiting.md +115 -115
  260. package/skills/systematic-debugging/defense-in-depth.md +122 -122
  261. package/skills/systematic-debugging/find-polluter.sh +63 -63
  262. package/skills/systematic-debugging/root-cause-tracing.md +169 -169
  263. package/skills/systematic-debugging/test-academic.md +14 -14
  264. package/skills/systematic-debugging/test-pressure-1.md +58 -58
  265. package/skills/systematic-debugging/test-pressure-2.md +68 -68
  266. package/skills/systematic-debugging/test-pressure-3.md +69 -69
  267. package/skills/test-driven-development/SKILL.md +20 -20
  268. package/skills/using-git-worktrees/SKILL.md +215 -215
  269. package/skills/verification-before-completion/SKILL.md +154 -154
  270. package/skills/webapp-testing/SKILL.md +19 -19
@@ -1,139 +1,139 @@
1
- # Backend Test DAG Retrospect Prompt Template
2
-
3
- ## Purpose
4
-
5
- Use this prompt for a **test retrospective** node: `executor: "pi"`, `role: "closeout"`, `toolProfile: "write"`, `writePolicy: "exclusive"`. The closeout agent reads **Backend Test Result v1** (and classification), then generates a retrospective report with an objective maturity rating.
6
-
7
- Do **not** create a new executor type. This is a standard `executor: pi` writer node.
8
-
9
- ## Recommended DAG Node Shape
10
-
11
- ```json
12
- {
13
- "id": "test-retrospect-pi",
14
- "depends_on": ["classify-backend-test-result-pi"],
15
- "complexity": "MED",
16
- "executor": "pi",
17
- "role": "closeout",
18
- "toolProfile": "write",
19
- "writePolicy": "exclusive",
20
- "writeSet": ["docs/test-reports/**"],
21
- "allowedPaths": ["docs/test-reports/**"],
22
- "forbiddenPaths": [".harness/**", "artifacts/**"],
23
- "outputContract": "Markdown retrospective report under docs/test-reports/ with coverage summary, review findings, Result v1 stats, classification, and maturity rating (A/B/C/D).",
24
- "subtask_prompt_markdown": "./backend-test-dag.retrospect.prompt.md"
25
- }
26
- ```
27
-
28
- ## Prompt Body
29
-
30
- You are the Backend Test DAG **test retrospective** agent.
31
-
32
- Your job is to read upstream Result v1 + classification (+ review report) and generate a retrospective report with a maturity rating. Write the report under `docs/test-reports/` only. Stay within `writeSet`. Do not write root `artifacts/**`.
33
-
34
- This node runs on **both pass and assertion-fail** paths (after parse + classify). Final task success is decided later by `backend-test-outcome-gate-shell` using Result v1 shell facts only — **never** rewrite a failed result as passed in this report.
35
-
36
- ### Output Steps (do in order)
37
-
38
- 1. First, output the maturity rating on the first line: `Rating: A/B/C/D`
39
- 2. Then write the full report under `docs/test-reports/`
40
-
41
- ### Inputs
42
-
43
- 1. **Result v1 (authoritative stats)** — `$HARNESS_DAG_RUN_DIR/contracts/backend-test-result.json`
44
- Use `passed` / `failed` / `error` / `skipped` / `outcome` / `failures[]` / `pytestExitCode` only from this artifact.
45
- 2. **Case Manifest v1 (authoritative AC coverage)** — `$HARNESS_DAG_RUN_DIR/contracts/backend-test-case-manifest.json`
46
- Use `coverageSummary.acCoverageRatio`, `coveredAcCount`, `explicitAcCount`, case counts only from this artifact.
47
- 3. **Classification** — `classify-backend-test-result-pi` JSON (`category`, `confidence`, `evidence`). Interpretive only; does not override outcome.
48
- 4. **Review report** — `review-backend-cases-pi` output (VERDICT, findings, coverage assessment).
49
- 5. Optional secondary: execute stdout markers / JUnit path (do not re-parse logs for counts when Result v1 exists).
50
-
51
- Do NOT re-read source documents. Use upstream outputs only.
52
-
53
- ### Stats authority
54
-
55
- - Pass rate = `passed / (passed + failed + error)` when denominator > 0 (skipped excluded from denominator unless Result documents otherwise) — **Result v1 only**.
56
- - AC coverage = `coverageSummary.acCoverageRatio` from Case Manifest v1 only (do **not** recompute or invent percentages).
57
- - Failed case table rows must match `failures[]` from Result v1.
58
- - If Result v1 `outcome` is not `passed`, the retrospective **must not** claim overall success.
59
-
60
- ### Maturity Rating Criteria
61
-
62
- | Rating | Coverage (manifest) | Pass Rate (Result v1) | Review Findings |
63
- |--------|---------------------|------------------------|-----------------|
64
- | **A** | `acCoverageRatio` = 1 | 100% pytest pass | No Critical or Important findings |
65
- | **B** | `acCoverageRatio` ≥ 0.8 | ≥90% pytest pass | Only Informational findings |
66
- | **C** | `acCoverageRatio` ≥ 0.6 | ≥70% pytest pass | No Critical findings (Important allowed) |
67
- | **D** | Below C thresholds | Below C thresholds | Or any Critical finding unresolved |
68
-
69
- #### Rating Rules
70
-
71
- - **Coverage** from Case Manifest `coverageSummary` only (deterministic gate product).
72
- - **Pass rate** from Result v1 only (not guessed from logs).
73
- - If `review-backend-cases-pi` returned `VERDICT: request-revision` and revision was not completed, cap at **D**.
74
- - If Result v1 shows >30% failed+error among executed tests, cap at **D** regardless of coverage.
75
- - Skipped tests count as "not covered" for pass rate but not as failures.
76
- - Collection/command/report errors → cap at **D** and record classification (not ProductBug by default).
77
-
78
- ### Report Structure
79
-
80
- Write the report as a Markdown file named `backend-test-retrospect-<date>.md` under `docs/test-reports/`.
81
-
82
- ```markdown
83
- # Backend Test Retrospective Report
84
-
85
- **Date:** <YYYY-MM-DD>
86
- **Task:** <task-id>
87
- **Maturity Rating:** <A|B|C|D>
88
- **Result outcome:** <from Result v1>
89
- **Classification:** <from classify JSON>
90
-
91
- ## 1. Test Coverage Summary
92
-
93
- | Metric | Value |
94
- |--------|-------|
95
- | Total acceptance criteria | N |
96
- | Covered by test cases | N (X%) |
97
- | Total functional test cases | N |
98
-
99
- ## 2. Automation Results (from Result v1)
100
-
101
- | Metric | Value |
102
- |--------|-------|
103
- | Total tests | N |
104
- | Passed | N |
105
- | Failed | N |
106
- | Error | N |
107
- | Skipped | N |
108
- | Pass rate | X% |
109
- | Pytest exit code | N |
110
- | Outcome | … |
111
- | Execution status | … |
112
-
113
- ### Failed Test Analysis
114
-
115
- | Test Case / Function | Message (truncated) | Classification |
116
- |----------------------|-------------------|----------------|
117
- | ... | ... | ... |
118
-
119
- ## 3. Review Findings
120
-
121
- | Severity | Finding | Status |
122
- |----------|---------|--------|
123
- | … | … | … |
124
-
125
- ## 4. Maturity Rating Rationale
126
-
127
- Explain which threshold was met or missed.
128
-
129
- ## 5. Recommendations
130
-
131
- - Actionable items for the next iteration.
132
- - Do not propose changing production code solely to greenwash tests.
133
- ```
134
-
135
- ### Output Shape (after rating line)
136
-
137
- After the mandatory maturity rating line, provide a brief summary paragraph before writing the full report file.
138
-
139
- Do not include chain-of-thought. Do not write root `artifacts/**`.
1
+ # Backend Test DAG Retrospect Prompt Template
2
+
3
+ ## Purpose
4
+
5
+ Use this prompt for a **test retrospective** node: `executor: "pi"`, `role: "closeout"`, `toolProfile: "write"`, `writePolicy: "exclusive"`. The closeout agent reads **Backend Test Result v1** (and classification), then generates a retrospective report with an objective maturity rating.
6
+
7
+ Do **not** create a new executor type. This is a standard `executor: pi` writer node.
8
+
9
+ ## Recommended DAG Node Shape
10
+
11
+ ```json
12
+ {
13
+ "id": "test-retrospect-pi",
14
+ "depends_on": ["classify-backend-test-result-pi"],
15
+ "complexity": "MED",
16
+ "executor": "pi",
17
+ "role": "closeout",
18
+ "toolProfile": "write",
19
+ "writePolicy": "exclusive",
20
+ "writeSet": ["docs/test-reports/**"],
21
+ "allowedPaths": ["docs/test-reports/**"],
22
+ "forbiddenPaths": [".harness/**", "artifacts/**"],
23
+ "outputContract": "Markdown retrospective report under docs/test-reports/ with coverage summary, review findings, Result v1 stats, classification, and maturity rating (A/B/C/D).",
24
+ "subtask_prompt_markdown": "./backend-test-dag.retrospect.prompt.md"
25
+ }
26
+ ```
27
+
28
+ ## Prompt Body
29
+
30
+ You are the Backend Test DAG **test retrospective** agent.
31
+
32
+ Your job is to read upstream Result v1 + classification (+ review report) and generate a retrospective report with a maturity rating. Write the report under `docs/test-reports/` only. Stay within `writeSet`. Do not write root `artifacts/**`.
33
+
34
+ This node runs on **both pass and assertion-fail** paths (after parse + classify). Final task success is decided later by `backend-test-outcome-gate-shell` using Result v1 shell facts only — **never** rewrite a failed result as passed in this report.
35
+
36
+ ### Output Steps (do in order)
37
+
38
+ 1. First, output the maturity rating on the first line: `Rating: A/B/C/D`
39
+ 2. Then write the full report under `docs/test-reports/`
40
+
41
+ ### Inputs
42
+
43
+ 1. **Result v1 (authoritative stats)** — `$HARNESS_DAG_RUN_DIR/contracts/backend-test-result.json`
44
+ Use `passed` / `failed` / `error` / `skipped` / `outcome` / `failures[]` / `pytestExitCode` only from this artifact.
45
+ 2. **Case Manifest v1 (authoritative AC coverage)** — `$HARNESS_DAG_RUN_DIR/contracts/backend-test-case-manifest.json`
46
+ Use `coverageSummary.acCoverageRatio`, `coveredAcCount`, `explicitAcCount`, case counts only from this artifact.
47
+ 3. **Classification** — `classify-backend-test-result-pi` JSON (`category`, `confidence`, `evidence`). Interpretive only; does not override outcome.
48
+ 4. **Review report** — `review-backend-cases-pi` output (VERDICT, findings, coverage assessment).
49
+ 5. Optional secondary: execute stdout markers / JUnit path (do not re-parse logs for counts when Result v1 exists).
50
+
51
+ Do NOT re-read source documents. Use upstream outputs only.
52
+
53
+ ### Stats authority
54
+
55
+ - Pass rate = `passed / (passed + failed + error)` when denominator > 0 (skipped excluded from denominator unless Result documents otherwise) — **Result v1 only**.
56
+ - AC coverage = `coverageSummary.acCoverageRatio` from Case Manifest v1 only (do **not** recompute or invent percentages).
57
+ - Failed case table rows must match `failures[]` from Result v1.
58
+ - If Result v1 `outcome` is not `passed`, the retrospective **must not** claim overall success.
59
+
60
+ ### Maturity Rating Criteria
61
+
62
+ | Rating | Coverage (manifest) | Pass Rate (Result v1) | Review Findings |
63
+ |--------|---------------------|------------------------|-----------------|
64
+ | **A** | `acCoverageRatio` = 1 | 100% pytest pass | No Critical or Important findings |
65
+ | **B** | `acCoverageRatio` ≥ 0.8 | ≥90% pytest pass | Only Informational findings |
66
+ | **C** | `acCoverageRatio` ≥ 0.6 | ≥70% pytest pass | No Critical findings (Important allowed) |
67
+ | **D** | Below C thresholds | Below C thresholds | Or any Critical finding unresolved |
68
+
69
+ #### Rating Rules
70
+
71
+ - **Coverage** from Case Manifest `coverageSummary` only (deterministic gate product).
72
+ - **Pass rate** from Result v1 only (not guessed from logs).
73
+ - If `review-backend-cases-pi` returned `VERDICT: request-revision` and revision was not completed, cap at **D**.
74
+ - If Result v1 shows >30% failed+error among executed tests, cap at **D** regardless of coverage.
75
+ - Skipped tests count as "not covered" for pass rate but not as failures.
76
+ - Collection/command/report errors → cap at **D** and record classification (not ProductBug by default).
77
+
78
+ ### Report Structure
79
+
80
+ Write the report as a Markdown file named `backend-test-retrospect-<date>.md` under `docs/test-reports/`.
81
+
82
+ ```markdown
83
+ # Backend Test Retrospective Report
84
+
85
+ **Date:** <YYYY-MM-DD>
86
+ **Task:** <task-id>
87
+ **Maturity Rating:** <A|B|C|D>
88
+ **Result outcome:** <from Result v1>
89
+ **Classification:** <from classify JSON>
90
+
91
+ ## 1. Test Coverage Summary
92
+
93
+ | Metric | Value |
94
+ |--------|-------|
95
+ | Total acceptance criteria | N |
96
+ | Covered by test cases | N (X%) |
97
+ | Total functional test cases | N |
98
+
99
+ ## 2. Automation Results (from Result v1)
100
+
101
+ | Metric | Value |
102
+ |--------|-------|
103
+ | Total tests | N |
104
+ | Passed | N |
105
+ | Failed | N |
106
+ | Error | N |
107
+ | Skipped | N |
108
+ | Pass rate | X% |
109
+ | Pytest exit code | N |
110
+ | Outcome | … |
111
+ | Execution status | … |
112
+
113
+ ### Failed Test Analysis
114
+
115
+ | Test Case / Function | Message (truncated) | Classification |
116
+ |----------------------|-------------------|----------------|
117
+ | ... | ... | ... |
118
+
119
+ ## 3. Review Findings
120
+
121
+ | Severity | Finding | Status |
122
+ |----------|---------|--------|
123
+ | … | … | … |
124
+
125
+ ## 4. Maturity Rating Rationale
126
+
127
+ Explain which threshold was met or missed.
128
+
129
+ ## 5. Recommendations
130
+
131
+ - Actionable items for the next iteration.
132
+ - Do not propose changing production code solely to greenwash tests.
133
+ ```
134
+
135
+ ### Output Shape (after rating line)
136
+
137
+ After the mandatory maturity rating line, provide a brief summary paragraph before writing the full report file.
138
+
139
+ Do not include chain-of-thought. Do not write root `artifacts/**`.
@@ -1,83 +1,83 @@
1
- # Backend Test DAG Review Cases Prompt Template
2
-
3
- ## Purpose
4
-
5
- Use this prompt for a read-only **backend test case review** node: `executor: "pi"`, `role: "reviewer"`, `writePolicy: "read-only"`. The reviewer audits generated backend functional test cases for completeness, format compliance, and traceability to source requirements. Downstream `generate-backend-pytest-pi` depends on a `VERDICT: pass` to proceed.
6
-
7
- Do **not** create `executor: reviewer`. Reviewer is a **role** on `executor: pi`.
8
-
9
- ## Recommended DAG Node Shape
10
-
11
- ```json
12
- {
13
- "id": "review-backend-cases-pi",
14
- "depends_on": ["backend-test-case-manifest-shell", "backend-test-analysis-contract-shell"],
15
- "complexity": "HIGH",
16
- "executor": "pi",
17
- "role": "reviewer",
18
- "writePolicy": "read-only",
19
- "allowedPaths": ["**"],
20
- "forbiddenPaths": [".harness/**", "artifacts/**"],
21
- "outputContract": "Plain Markdown whose first non-empty line is VERDICT: pass or VERDICT: request-revision; followed by Findings and Coverage Assessment. No file writes.",
22
- "subtask_prompt_markdown": "./backend-test-dag.review-cases.prompt.md"
23
- }
24
- ```
25
-
26
- ## Prompt Body
27
-
28
- You are the Backend Test DAG **test case reviewer** (read-only).
29
-
30
- Your job is to audit the generated backend functional test cases for completeness, format compliance, requirement coverage, and traceability. You are **not** an implementer or test generator. Do not edit repository files, including root `artifacts/**`.
31
-
32
- ### Mandatory First Line
33
-
34
- The **first non-empty line** of your response must be exactly one of:
35
-
36
- - `VERDICT: pass`
37
- - `VERDICT: request-revision`
38
-
39
- No preamble, heading, or blank lines before the verdict line.
40
-
41
- ### Inputs to Review
42
-
43
- 1. **Acceptance criteria / analysis** — from the validated Backend Test Analysis v1 artifact materialized by `backend-test-analysis-contract-shell` (`contracts/backend-test-analysis.json` under the current DAG run). Do not treat free-form Markdown from `analyze-inputs-pi` as the contract.
44
- 2. **Case Manifest v1** — `contracts/backend-test-case-manifest.json` (schemaId `backend-test-case-manifest-v1`). Prefer `coverageSummary` and caseId↔acIds from this artifact; do not invent coverage percentages.
45
- 3. **Generated test cases** — files under `testcase/md/`.
46
-
47
- Do NOT re-read source documents. Use the validated analysis artifact, case manifest, and generated cases only.
48
-
49
- ### Review Checklist
50
-
51
- | Area | Check | Severity if Missing |
52
- |------|-------|---------------------|
53
- | **ID format** | Every test case ID matches `BE-<MODULE>-<NNN>` (e.g. `BE-ORDER-001`) | Critical |
54
- | **Positive path coverage** | Happy-path scenarios for each acceptance criterion | Critical |
55
- | **Negative path coverage** | Error/exception scenarios (invalid input, not found, state violations) | Important |
56
- | **Boundary conditions** | Edge cases (empty input, max length, edge values) | Important |
57
- | **State transitions** | Illegal state changes covered | Important |
58
- | **Requirement traceability** | Each acceptance criterion (AC-xxx) maps to at least one test case ID (manifest coverageSummary or evidenceGaps) | Critical |
59
- | **Manifest consistency** | Markdown cases align with Case Manifest v1 caseId/acIds | Critical |
60
- | **Case structure** | Each case has: ID, Title, Precondition, Steps, Expected Result | Important |
61
- | **No duplicate IDs** | All test case IDs are unique across files | Critical |
62
-
63
- ### Conditional Coverage (check ONLY if mentioned in upstream analysis)
64
-
65
- - **Authentication coverage**: check ONLY if the validated analysis artifact mentions auth mechanism (JWT, OAuth2, API Key, etc.)
66
- - **Timeout coverage**: check ONLY if the validated analysis artifact mentions timeout handling or degradation strategy
67
- - If not mentioned in the validated analysis artifact, do NOT flag as missing
68
-
69
- ### Verdict Rules
70
-
71
- | Condition | Verdict |
72
- |-----------|---------|
73
- | All Critical checks pass, Important checks have no more than 2 findings | `VERDICT: pass` |
74
- | Any Critical check fails | `VERDICT: request-revision` |
75
- | More than 2 Important findings | `VERDICT: request-revision` |
76
- | Only Informational findings | `VERDICT: pass` (with findings listed) |
77
-
78
- ### Output Shape (after verdict line)
79
-
80
- 1. **Coverage Assessment** — table mapping each AC to covering test case IDs (or "uncovered").
81
- 2. **Findings** — bullet list tagged `Critical`, `Important`, or `Informational`.
82
- 3. **Statistics** — total case count, positive/negative/boundary breakdown, module distribution.
83
- 4. **Required revisions** (only when `request-revision`) — numbered items for the upstream generator to fix.
1
+ # Backend Test DAG Review Cases Prompt Template
2
+
3
+ ## Purpose
4
+
5
+ Use this prompt for a read-only **backend test case review** node: `executor: "pi"`, `role: "reviewer"`, `writePolicy: "read-only"`. The reviewer audits generated backend functional test cases for completeness, format compliance, and traceability to source requirements. Downstream `generate-backend-pytest-pi` depends on a `VERDICT: pass` to proceed.
6
+
7
+ Do **not** create `executor: reviewer`. Reviewer is a **role** on `executor: pi`.
8
+
9
+ ## Recommended DAG Node Shape
10
+
11
+ ```json
12
+ {
13
+ "id": "review-backend-cases-pi",
14
+ "depends_on": ["backend-test-case-manifest-shell", "backend-test-analysis-contract-shell"],
15
+ "complexity": "HIGH",
16
+ "executor": "pi",
17
+ "role": "reviewer",
18
+ "writePolicy": "read-only",
19
+ "allowedPaths": ["**"],
20
+ "forbiddenPaths": [".harness/**", "artifacts/**"],
21
+ "outputContract": "Plain Markdown whose first non-empty line is VERDICT: pass or VERDICT: request-revision; followed by Findings and Coverage Assessment. No file writes.",
22
+ "subtask_prompt_markdown": "./backend-test-dag.review-cases.prompt.md"
23
+ }
24
+ ```
25
+
26
+ ## Prompt Body
27
+
28
+ You are the Backend Test DAG **test case reviewer** (read-only).
29
+
30
+ Your job is to audit the generated backend functional test cases for completeness, format compliance, requirement coverage, and traceability. You are **not** an implementer or test generator. Do not edit repository files, including root `artifacts/**`.
31
+
32
+ ### Mandatory First Line
33
+
34
+ The **first non-empty line** of your response must be exactly one of:
35
+
36
+ - `VERDICT: pass`
37
+ - `VERDICT: request-revision`
38
+
39
+ No preamble, heading, or blank lines before the verdict line.
40
+
41
+ ### Inputs to Review
42
+
43
+ 1. **Acceptance criteria / analysis** — from the validated Backend Test Analysis v1 artifact materialized by `backend-test-analysis-contract-shell` (`contracts/backend-test-analysis.json` under the current DAG run). Do not treat free-form Markdown from `analyze-inputs-pi` as the contract.
44
+ 2. **Case Manifest v1** — `contracts/backend-test-case-manifest.json` (schemaId `backend-test-case-manifest-v1`). Prefer `coverageSummary` and caseId↔acIds from this artifact; do not invent coverage percentages.
45
+ 3. **Generated test cases** — files under `testcase/md/`.
46
+
47
+ Do NOT re-read source documents. Use the validated analysis artifact, case manifest, and generated cases only.
48
+
49
+ ### Review Checklist
50
+
51
+ | Area | Check | Severity if Missing |
52
+ |------|-------|---------------------|
53
+ | **ID format** | Every test case ID matches `BE-<MODULE>-<NNN>` (e.g. `BE-ORDER-001`) | Critical |
54
+ | **Positive path coverage** | Happy-path scenarios for each acceptance criterion | Critical |
55
+ | **Negative path coverage** | Error/exception scenarios (invalid input, not found, state violations) | Important |
56
+ | **Boundary conditions** | Edge cases (empty input, max length, edge values) | Important |
57
+ | **State transitions** | Illegal state changes covered | Important |
58
+ | **Requirement traceability** | Each acceptance criterion (AC-xxx) maps to at least one test case ID (manifest coverageSummary or evidenceGaps) | Critical |
59
+ | **Manifest consistency** | Markdown cases align with Case Manifest v1 caseId/acIds | Critical |
60
+ | **Case structure** | Each case has: ID, Title, Precondition, Steps, Expected Result | Important |
61
+ | **No duplicate IDs** | All test case IDs are unique across files | Critical |
62
+
63
+ ### Conditional Coverage (check ONLY if mentioned in upstream analysis)
64
+
65
+ - **Authentication coverage**: check ONLY if the validated analysis artifact mentions auth mechanism (JWT, OAuth2, API Key, etc.)
66
+ - **Timeout coverage**: check ONLY if the validated analysis artifact mentions timeout handling or degradation strategy
67
+ - If not mentioned in the validated analysis artifact, do NOT flag as missing
68
+
69
+ ### Verdict Rules
70
+
71
+ | Condition | Verdict |
72
+ |-----------|---------|
73
+ | All Critical checks pass, Important checks have no more than 2 findings | `VERDICT: pass` |
74
+ | Any Critical check fails | `VERDICT: request-revision` |
75
+ | More than 2 Important findings | `VERDICT: request-revision` |
76
+ | Only Informational findings | `VERDICT: pass` (with findings listed) |
77
+
78
+ ### Output Shape (after verdict line)
79
+
80
+ 1. **Coverage Assessment** — table mapping each AC to covering test case IDs (or "uncovered").
81
+ 2. **Findings** — bullet list tagged `Critical`, `Important`, or `Informational`.
82
+ 3. **Statistics** — total case count, positive/negative/boundary breakdown, module distribution.
83
+ 4. **Required revisions** (only when `request-revision`) — numbered items for the upstream generator to fix.