@tea-agent/loop-agent 0.12.0 → 0.13.0-beta.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (284) hide show
  1. package/AGENTS.md +155 -153
  2. package/CHANGELOG.md +338 -265
  3. package/README.md +345 -298
  4. package/bin/agent-worker.js +22 -22
  5. package/bin/loop-agent.js +21 -21
  6. package/dist/application/dag/generate-task-dag.js +28 -28
  7. package/dist/application/evaluation/candidate-hash.js +75 -0
  8. package/dist/application/evaluation/candidate.js +52 -0
  9. package/dist/application/evaluation/replay.js +289 -0
  10. package/dist/application/evaluation/types.js +130 -0
  11. package/dist/cli/command-definitions.js +27 -7
  12. package/dist/cli/program.js +8 -4
  13. package/dist/commands/cursor-prompt.js +6 -6
  14. package/dist/commands/eval.js +235 -0
  15. package/dist/commands/init.js +544 -506
  16. package/dist/commands/knowledge.js +129 -31
  17. package/dist/commands/loop-benchmark.js +11 -11
  18. package/dist/commands/pi-reuse-benchmark.js +16 -16
  19. package/dist/executors/pi-sdk-executor.js +38 -24
  20. package/dist/executors/shell-executor.js +34 -2
  21. package/dist/executors/shell-presets.js +20 -0
  22. package/dist/executors/shell-verification.js +7 -0
  23. package/dist/governance/manifest-types.js +4 -0
  24. package/dist/infrastructure/evaluation/candidate-store.js +435 -0
  25. package/dist/infrastructure/evaluation/store.js +40 -0
  26. package/dist/sidecars/cursor-prompt/executor.js +1 -1
  27. package/dist/task/config-types.js +28 -1
  28. package/dist/task/runtime.js +27 -27
  29. package/dist/worker/cli.js +96 -1
  30. package/dist/worker/delivery/package.js +3 -3
  31. package/dist/worker/feature/decision-loader.js +37 -6
  32. package/dist/worker/feature/next-action.js +10 -2
  33. package/dist/worker/feature/ready-plan-projection.js +81 -0
  34. package/dist/worker/feature/reducer.js +2 -1
  35. package/dist/worker/feature/review.js +19 -2
  36. package/dist/worker/feature/run.js +27 -2
  37. package/dist/worker/follow-up/approve.js +5 -2
  38. package/dist/worker/follow-up/factory.js +1 -1
  39. package/dist/worker/observability/read-model.js +246 -41
  40. package/dist/worker/observe/routes.js +173 -15
  41. package/dist/worker/observe/spec-evidence.js +281 -0
  42. package/dist/worker/observe/static/api.js +46 -27
  43. package/dist/worker/observe/static/app.js +150 -150
  44. package/dist/worker/observe/static/constants.js +148 -148
  45. package/dist/worker/observe/static/copy.js +67 -67
  46. package/dist/worker/observe/static/dag-helpers.js +172 -172
  47. package/dist/worker/observe/static/dag-layout.d.ts +31 -31
  48. package/dist/worker/observe/static/dag-layout.js +83 -83
  49. package/dist/worker/observe/static/dag-model.js +72 -72
  50. package/dist/worker/observe/static/dom.js +61 -61
  51. package/dist/worker/observe/static/format-pool.js +67 -67
  52. package/dist/worker/observe/static/format.js +292 -292
  53. package/dist/worker/observe/static/index.html +308 -308
  54. package/dist/worker/observe/static/kpi.js +94 -94
  55. package/dist/worker/observe/static/relations.js +133 -128
  56. package/dist/worker/observe/static/router.js +93 -85
  57. package/dist/worker/observe/static/run-processing.js +148 -148
  58. package/dist/worker/observe/static/shell-chrome.js +68 -68
  59. package/dist/worker/observe/static/state.js +253 -253
  60. package/dist/worker/observe/static/styles.css +1902 -1890
  61. package/dist/worker/observe/static/views/batch.js +227 -226
  62. package/dist/worker/observe/static/views/dag-graph.js +172 -172
  63. package/dist/worker/observe/static/views/dag-inspector.js +607 -477
  64. package/dist/worker/observe/static/views/dag.js +362 -362
  65. package/dist/worker/observe/static/views/dashboard.js +445 -442
  66. package/dist/worker/observe/static/views/failures.js +143 -143
  67. package/dist/worker/observe/static/views/feature.js +492 -453
  68. package/dist/worker/observe/static/views/pool.js +350 -347
  69. package/dist/worker/observe/static/views/run.js +453 -453
  70. package/dist/worker/observe/static/views/session-timeline.js +205 -205
  71. package/dist/worker/observe/static/views/shell.js +7 -7
  72. package/dist/worker/observe/static/views/task.js +314 -260
  73. package/dist/worker/observe/static/views/timeline.js +163 -163
  74. package/dist/worker/pool/doctor.js +165 -0
  75. package/dist/worker/pool/migrate-state.js +303 -0
  76. package/dist/worker/pool/run-store.js +205 -17
  77. package/dist/worker/pool/types.js +17 -1
  78. package/dist/worker/pool/validation.js +100 -15
  79. package/dist/worker/report/morning-report.js +12 -2
  80. package/dist/worker/runner/run-ready.js +41 -26
  81. package/dist/worker/task-graph/ready-planner.js +136 -0
  82. package/dist/workflows/dag/backend-test-analysis-contract.js +120 -0
  83. package/dist/workflows/dag/canvas-observer.js +275 -275
  84. package/dist/workflows/dag/convergence/controller.js +16 -8
  85. package/dist/workflows/dag/dynamic-runtime/map.js +90 -2
  86. package/dist/workflows/dag/failure-routing.js +12 -1
  87. package/dist/workflows/dag/init-hybrid.js +2404 -360
  88. package/dist/workflows/dag/node-execution.js +9 -0
  89. package/dist/workflows/dag/prompt.js +9 -0
  90. package/dist/workflows/dag/report.js +35 -1
  91. package/dist/workflows/dag/runner.js +28 -2
  92. package/dist/workflows/dag/task-demand-routing.js +383 -0
  93. package/dist/workflows/dag/types.js +51 -13
  94. package/dist/workflows/dag/upstream-artifacts.js +1 -0
  95. package/dist/workflows/dag/validate.js +59 -1
  96. package/docs/README.md +106 -104
  97. package/docs/agent-dag-recovery-playbook.md +195 -184
  98. package/docs/agent-dag-runner.md +67 -67
  99. package/docs/architecture/README.md +26 -26
  100. package/docs/architecture/dag-execution.md +140 -140
  101. package/docs/architecture/evolution.md +54 -53
  102. package/docs/architecture/facts-and-state.md +71 -58
  103. package/docs/architecture/runtime-boundaries.md +191 -191
  104. package/docs/architecture/system-overview.md +93 -93
  105. package/docs/architecture/worker-and-feature.md +85 -81
  106. package/docs/cursor-prompt-sidecar.md +36 -36
  107. package/docs/decisions/README.md +18 -15
  108. package/docs/design/README.md +167 -77
  109. package/docs/development-principles.md +73 -73
  110. package/docs/exec-plans/README.md +6 -6
  111. package/docs/exec-plans/active/README.md +15 -9
  112. package/docs/exec-plans/completed/README.md +85 -73
  113. package/docs/feature-workflow.md +389 -261
  114. package/docs/harness-methodology-debugging.md +153 -153
  115. package/docs/harness-methodology-tdd.md +130 -130
  116. package/docs/harness-methodology-verification.md +27 -27
  117. package/docs/init-surface.manifest.json +289 -280
  118. package/docs/loop-agent-harness.md +142 -130
  119. package/docs/production-readiness.md +96 -96
  120. package/docs/progress/README.md +64 -54
  121. package/docs/reports/README.md +117 -94
  122. package/docs/skills/README.md +7 -7
  123. package/docs/skills/vetted-skill-registry.md +29 -27
  124. package/docs/templates/adr.md +60 -60
  125. package/docs/templates/agent-dag-authority-surface-audit.prompt.md +94 -94
  126. package/docs/templates/agent-dag-decision-envelope.schema.json +213 -213
  127. package/docs/templates/agent-dag-decision-gate-dogfood-report.md +117 -117
  128. package/docs/templates/agent-dag-decision-gate.prompt.md +246 -246
  129. package/docs/templates/agent-dag-process-supervisor.prompt.md +98 -98
  130. package/docs/templates/agent-dag-report.schema.json +473 -473
  131. package/docs/templates/agent-dag-review-verdict.prompt.md +68 -68
  132. package/docs/templates/agent-dag.base.json +190 -190
  133. package/docs/templates/agent-dag.final-verification.json +185 -185
  134. package/docs/templates/agent-dag.schema.json +411 -383
  135. package/docs/templates/agent-dag.supervised-implementation.json +501 -501
  136. package/docs/templates/backend-test-analysis.schema.json +44 -0
  137. package/docs/templates/backend-test-dag.generate-pytest.prompt.md +202 -139
  138. package/docs/templates/backend-test-dag.json +311 -276
  139. package/docs/templates/backend-test-dag.retrospect.prompt.md +125 -125
  140. package/docs/templates/backend-test-dag.review-cases.prompt.md +81 -81
  141. package/docs/templates/exec-plan.md +64 -64
  142. package/docs/templates/feature-spec.md +53 -53
  143. package/docs/templates/frontend-design-contract.md +42 -33
  144. package/docs/templates/frontend-task-constraints.md +35 -25
  145. package/docs/templates/frontend-task-requirement.md +70 -61
  146. package/docs/templates/frontend-test-dag.generate-cases.prompt.md +5 -0
  147. package/docs/templates/frontend-test-dag.json +23 -0
  148. package/docs/templates/frontend-test-dag.retrieve-context.prompt.md +3 -0
  149. package/docs/templates/frontend-test-dag.retrospect.prompt.md +3 -0
  150. package/docs/templates/frontend-test-dag.review-cases.prompt.md +3 -0
  151. package/docs/templates/frontend-test-dag.review-execution.prompt.md +3 -0
  152. package/docs/templates/harness.schema.json +221 -221
  153. package/docs/templates/hybrid-dag.json +188 -188
  154. package/docs/templates/init-evolution-review.md +35 -35
  155. package/docs/templates/interactive-ui-round2-experiment.md +66 -66
  156. package/docs/templates/knowledge-graph-bootstrap-dag.json +118 -0
  157. package/docs/templates/knowledge-sync-dag.json +178 -0
  158. package/docs/templates/knowledge-sync-draft.schema.json +71 -0
  159. package/docs/templates/product-line/AGENTS.md +8 -8
  160. package/docs/templates/product-line/README.md +9 -9
  161. package/docs/templates/product-line/acceptance.yaml +14 -14
  162. package/docs/templates/product-line/closeout.yaml +9 -9
  163. package/docs/templates/product-line/design.md +13 -13
  164. package/docs/templates/product-line/links.md +10 -10
  165. package/docs/templates/product-line/requirement.md +17 -17
  166. package/docs/templates/product-line/task-graph.yaml +15 -15
  167. package/docs/templates/product-line/task.yaml +64 -64
  168. package/docs/templates/product-line/test-plan.md +7 -7
  169. package/docs/templates/production-readiness-checklist.md +57 -57
  170. package/docs/templates/progress-log.md +17 -17
  171. package/docs/templates/project-start-checklist.md +9 -9
  172. package/docs/templates/qa-report.md +48 -48
  173. package/docs/templates/sprint-contract.md +29 -29
  174. package/docs/templates/worker-dogfood-evidence.md +80 -80
  175. package/docs/templates/worker-dogfood-setup.md +68 -68
  176. package/docs/verification-matrix.md +70 -66
  177. package/examples/decision-gate-agent-dag.json +177 -177
  178. package/examples/example-dag.json +46 -46
  179. package/examples/hybrid-loop-agent-dag.json +189 -189
  180. package/harness.json +66 -66
  181. package/package.json +88 -46
  182. package/scripts/check-product-line-docs.sh +29 -29
  183. package/scripts/check-task-pool-root.sh +32 -32
  184. package/scripts/kb-bootstrap-init-skeleton.sh +240 -0
  185. package/scripts/kb-graph-incremental-prepare.mjs +386 -0
  186. package/scripts/kb-graph-incremental-prepare.sh +5 -0
  187. package/scripts/kb-graph-materialize.mjs +105 -0
  188. package/scripts/kb-graph-materialize.sh +4 -0
  189. package/scripts/kb-graph-promote.mjs +164 -0
  190. package/scripts/kb-graph-promote.sh +4 -0
  191. package/scripts/kb-query.mjs +554 -0
  192. package/scripts/kb-query.sh +5 -0
  193. package/skills/agent-worker/SKILL.md +39 -37
  194. package/skills/agent-worker/references/agent-worker-operator.md +60 -43
  195. package/skills/ai-engineering-context/SKILL.md +48 -48
  196. package/skills/analyze-product-dependencies/SKILL.md +67 -0
  197. package/skills/analyze-product-dependencies/agents/openai.yaml +4 -0
  198. package/skills/analyze-product-dependencies/references/api-documentation-schema.md +30 -0
  199. package/skills/analyze-product-dependencies/references/dependency-analysis-schema.md +28 -0
  200. package/skills/analyze-product-dependencies/references/example.md +76 -0
  201. package/skills/analyze-product-dependencies/references/forward-test-cases.md +35 -0
  202. package/skills/analyze-product-dependencies/references/input-contract.md +11 -0
  203. package/skills/analyze-product-dependencies/references/scouting-rules.md +61 -0
  204. package/skills/analyze-product-dependencies/scripts/test-validators.mjs +267 -0
  205. package/skills/analyze-product-dependencies/scripts/validate-api-documentation.mjs +101 -0
  206. package/skills/analyze-product-dependencies/scripts/validate-dependency-analysis.mjs +142 -0
  207. package/skills/analyze-product-dependencies/scripts/validate-product-requirement-input.mjs +76 -0
  208. package/skills/analyze-product-dependencies/scripts/validation-helpers.mjs +146 -0
  209. package/skills/analyze-product-requirements/SKILL.md +90 -0
  210. package/skills/analyze-product-requirements/agents/openai.yaml +4 -0
  211. package/skills/analyze-product-requirements/references/acceptance-criteria.md +91 -0
  212. package/skills/analyze-product-requirements/references/clarification-and-knowledge.md +56 -0
  213. package/skills/analyze-product-requirements/references/example.md +86 -0
  214. package/skills/analyze-product-requirements/references/forward-test-cases.md +66 -0
  215. package/skills/analyze-product-requirements/references/product-analysis-schema.md +32 -0
  216. package/skills/analyze-product-requirements/references/product-requirement-schema.md +33 -0
  217. package/skills/analyze-product-requirements/references/requirement-clarification-schema.md +35 -0
  218. package/skills/analyze-product-requirements/scripts/test-validators.mjs +193 -0
  219. package/skills/analyze-product-requirements/scripts/validate-product-analysis.mjs +69 -0
  220. package/skills/analyze-product-requirements/scripts/validate-product-requirement.mjs +97 -0
  221. package/skills/analyze-product-requirements/scripts/validate-requirement-clarification.mjs +98 -0
  222. package/skills/analyze-product-requirements/scripts/validation-helpers.mjs +156 -0
  223. package/skills/code-review-core/SKILL.md +20 -20
  224. package/skills/codebase-scout/SKILL.md +19 -19
  225. package/skills/frontend-design-review/SKILL.md +66 -59
  226. package/skills/frontend-design-review/references/review-checklist.md +58 -37
  227. package/skills/frontend-implementation/SKILL.md +47 -51
  228. package/skills/frontend-implementation/references/code-standards.md +32 -34
  229. package/skills/frontend-implementation/references/design-spec.md +46 -46
  230. package/skills/frontend-implementation/references/node-contracts.md +76 -32
  231. package/skills/frontend-review/SKILL.md +59 -53
  232. package/skills/frontend-review/references/review-findings.md +47 -42
  233. package/skills/frontend-verification/SKILL.md +53 -40
  234. package/skills/frontend-verification/references/verification-checklist.md +68 -56
  235. package/skills/grill-me/SKILL.md +10 -10
  236. package/skills/grill-with-docs/SKILL.md +88 -88
  237. package/skills/grill-with-docs/adr-format.md +47 -47
  238. package/skills/grill-with-docs/context-format.md +60 -60
  239. package/skills/init-capability-evolution/SKILL.md +70 -70
  240. package/skills/loop-agent/SKILL.md +151 -151
  241. package/skills/loop-agent/references/README.md +67 -67
  242. package/skills/loop-agent/references/command-reference.md +505 -452
  243. package/skills/loop-agent/references/docs-converge.md +126 -126
  244. package/skills/loop-agent/references/harness-policy.md +263 -263
  245. package/skills/loop-agent/references/hybrid-dag.md +238 -233
  246. package/skills/loop-agent/references/learned/README.md +21 -21
  247. package/skills/loop-agent/references/long-running-loop.md +57 -57
  248. package/skills/loop-agent/references/model-routing.md +36 -36
  249. package/skills/loop-agent/references/multi-worktree.md +54 -54
  250. package/skills/loop-agent/references/one-shot-runs.md +85 -85
  251. package/skills/loop-agent/references/orchestrator-and-interventions.md +169 -169
  252. package/skills/loop-agent/references/pi-prompt.md +23 -23
  253. package/skills/loop-agent/references/pi-subagent-assisted-mode.md +84 -84
  254. package/skills/loop-agent/references/post-implementation-and-patterns.md +44 -44
  255. package/skills/loop-agent/references/task-workflow.md +89 -89
  256. package/skills/loop-agent/references/verification-and-failure-handling.md +139 -139
  257. package/skills/playwright-cli/SKILL.md +420 -0
  258. package/skills/playwright-cli/references/element-attributes.md +23 -0
  259. package/skills/playwright-cli/references/playwright-tests.md +39 -0
  260. package/skills/playwright-cli/references/request-mocking.md +87 -0
  261. package/skills/playwright-cli/references/running-code.md +241 -0
  262. package/skills/playwright-cli/references/session-management.md +225 -0
  263. package/skills/playwright-cli/references/storage-state.md +275 -0
  264. package/skills/playwright-cli/references/test-generation.md +433 -0
  265. package/skills/playwright-cli/references/tracing.md +139 -0
  266. package/skills/playwright-cli/references/video-recording.md +143 -0
  267. package/skills/playwright-cli-case-generator/SKILL.md +74 -0
  268. package/skills/requesting-code-review/SKILL.md +101 -101
  269. package/skills/requesting-code-review/code-reviewer.md +168 -168
  270. package/skills/systematic-debugging/CREATION-LOG.md +119 -119
  271. package/skills/systematic-debugging/SKILL.md +296 -296
  272. package/skills/systematic-debugging/condition-based-waiting-example.ts +158 -158
  273. package/skills/systematic-debugging/condition-based-waiting.md +115 -115
  274. package/skills/systematic-debugging/defense-in-depth.md +122 -122
  275. package/skills/systematic-debugging/find-polluter.sh +63 -63
  276. package/skills/systematic-debugging/root-cause-tracing.md +169 -169
  277. package/skills/systematic-debugging/test-academic.md +14 -14
  278. package/skills/systematic-debugging/test-pressure-1.md +58 -58
  279. package/skills/systematic-debugging/test-pressure-2.md +68 -68
  280. package/skills/systematic-debugging/test-pressure-3.md +69 -69
  281. package/skills/test-driven-development/SKILL.md +20 -20
  282. package/skills/using-git-worktrees/SKILL.md +215 -215
  283. package/skills/verification-before-completion/SKILL.md +154 -154
  284. package/skills/webapp-testing/SKILL.md +19 -19
@@ -1,125 +1,125 @@
1
- # Backend Test DAG Retrospect Prompt Template
2
-
3
- ## Purpose
4
-
5
- Use this prompt for a **test retrospective** node: `executor: "pi"`, `role: "closeout"`, `toolProfile: "write"`, `writePolicy: "exclusive"`. The closeout agent reads upstream review reports and pytest execution results, then generates a retrospective report with an objective maturity rating.
6
-
7
- Do **not** create a new executor type. This is a standard `executor: pi` writer node.
8
-
9
- ## Recommended DAG Node Shape
10
-
11
- ```json
12
- {
13
- "id": "test-retrospect-pi",
14
- "depends_on": ["execute-backend-pytest-shell"],
15
- "complexity": "MED",
16
- "executor": "pi",
17
- "role": "closeout",
18
- "toolProfile": "write",
19
- "writePolicy": "exclusive",
20
- "writeSet": ["docs/test-reports/**"],
21
- "allowedPaths": ["docs/test-reports/**"],
22
- "forbiddenPaths": [".harness/**", "artifacts/**"],
23
- "outputContract": "Markdown retrospective report under docs/test-reports/ with coverage summary, review findings, pytest results, and maturity rating (A/B/C/D).",
24
- "subtask_prompt_markdown": "./backend-test-dag.retrospect.prompt.md"
25
- }
26
- ```
27
-
28
- ## Prompt Body
29
-
30
- You are the Backend Test DAG **test retrospective** agent.
31
-
32
- Your job is to read upstream outputs (review report + pytest results) and generate a retrospective report with a maturity rating. Write the report under `docs/test-reports/` only. Stay within `writeSet`. Do not write root `artifacts/**`.
33
-
34
- ### Output Steps (do in order)
35
-
36
- 1. First, output the maturity rating on the first line: `Rating: A/B/C/D`
37
- 2. Then write the full report under `docs/test-reports/`
38
-
39
- ### Inputs
40
-
41
- 1. **Review report** — `review-backend-cases-pi` output (VERDICT, findings, coverage assessment).
42
- 2. **Pytest output** — `execute-backend-pytest-shell` stdout/stderr and exit code.
43
- 3. **HTML report** — `reports/backend-test-report.html` (if generated).
44
-
45
- Do NOT re-read source documents. Use upstream outputs only.
46
-
47
- ### Maturity Rating Criteria
48
-
49
- | Rating | Coverage | Pass Rate | Review Findings |
50
- |--------|----------|-----------|-----------------|
51
- | **A** | 100% acceptance criteria covered | 100% pytest pass | No Critical or Important findings |
52
- | **B** | ≥80% acceptance criteria covered | ≥90% pytest pass | Only Informational findings |
53
- | **C** | ≥60% acceptance criteria covered | ≥70% pytest pass | No Critical findings (Important allowed) |
54
- | **D** | Below C thresholds | Below C thresholds | Or any Critical finding unresolved |
55
-
56
- #### Rating Rules
57
-
58
- - **Coverage** = (acceptance criteria with ≥1 covering test case) / (total acceptance criteria) × 100%
59
- - **Pass rate** = (passed pytest functions) / (total non-skipped pytest functions) × 100%
60
- - If `review-backend-cases-pi` returned `VERDICT: request-revision` and revision was not completed, cap at **D**.
61
- - If pytest exit code is non-zero and >30% tests failed, cap at **D** regardless of coverage.
62
- - Skipped tests (`@pytest.mark.skip`) count as "not covered" for pass rate but not as failures.
63
-
64
- ### Report Structure
65
-
66
- Write the report as a Markdown file named `backend-test-retrospect-<date>.md` under `docs/test-reports/`.
67
-
68
- ```markdown
69
- # Backend Test Retrospective Report
70
-
71
- **Date:** <YYYY-MM-DD>
72
- **Task:** <task-id>
73
- **Maturity Rating:** <A|B|C|D>
74
-
75
- ## 1. Test Coverage Summary
76
-
77
- | Metric | Value |
78
- |--------|-------|
79
- | Total acceptance criteria | N |
80
- | Covered by test cases | N (X%) |
81
- | Total functional test cases | N |
82
- | Positive path cases | N |
83
- | Negative path cases | N |
84
- | Boundary cases | N |
85
-
86
- ## 2. Automation Results
87
-
88
- | Metric | Value |
89
- |--------|-------|
90
- | Total pytest functions | N |
91
- | Passed | N |
92
- | Failed | N |
93
- | Skipped | N |
94
- | Pass rate | X% |
95
- | Pytest exit code | N |
96
-
97
- ### Failed Test Analysis
98
-
99
- | Test Case ID | Function | Failure Reason | Root Cause |
100
- |--------------|----------|----------------|------------|
101
- | ... | ... | ... | ... |
102
-
103
- ## 3. Review Findings
104
-
105
- | Severity | Finding | Status |
106
- |----------|---------|--------|
107
- | Critical | ... | Resolved / Unresolved |
108
- | Important | ... | Resolved / Unresolved |
109
- | Informational | ... | Resolved / Unresolved |
110
-
111
- ## 4. Maturity Rating Rationale
112
-
113
- Explain which threshold was met or missed, and why the specific rating was assigned.
114
-
115
- ## 5. Recommendations
116
-
117
- - Actionable items for improving the rating in the next iteration.
118
- - Specific gaps to close (uncovered criteria, flaky tests, missing negative paths).
119
- ```
120
-
121
- ### Output Shape (after rating line)
122
-
123
- After the mandatory maturity rating line, provide a brief summary paragraph before writing the full report file.
124
-
125
- Do not include chain-of-thought. Do not write root `artifacts/**`.
1
+ # Backend Test DAG Retrospect Prompt Template
2
+
3
+ ## Purpose
4
+
5
+ Use this prompt for a **test retrospective** node: `executor: "pi"`, `role: "closeout"`, `toolProfile: "write"`, `writePolicy: "exclusive"`. The closeout agent reads upstream review reports and pytest execution results, then generates a retrospective report with an objective maturity rating.
6
+
7
+ Do **not** create a new executor type. This is a standard `executor: pi` writer node.
8
+
9
+ ## Recommended DAG Node Shape
10
+
11
+ ```json
12
+ {
13
+ "id": "test-retrospect-pi",
14
+ "depends_on": ["execute-backend-pytest-shell"],
15
+ "complexity": "MED",
16
+ "executor": "pi",
17
+ "role": "closeout",
18
+ "toolProfile": "write",
19
+ "writePolicy": "exclusive",
20
+ "writeSet": ["docs/test-reports/**"],
21
+ "allowedPaths": ["docs/test-reports/**"],
22
+ "forbiddenPaths": [".harness/**", "artifacts/**"],
23
+ "outputContract": "Markdown retrospective report under docs/test-reports/ with coverage summary, review findings, pytest results, and maturity rating (A/B/C/D).",
24
+ "subtask_prompt_markdown": "./backend-test-dag.retrospect.prompt.md"
25
+ }
26
+ ```
27
+
28
+ ## Prompt Body
29
+
30
+ You are the Backend Test DAG **test retrospective** agent.
31
+
32
+ Your job is to read upstream outputs (review report + pytest results) and generate a retrospective report with a maturity rating. Write the report under `docs/test-reports/` only. Stay within `writeSet`. Do not write root `artifacts/**`.
33
+
34
+ ### Output Steps (do in order)
35
+
36
+ 1. First, output the maturity rating on the first line: `Rating: A/B/C/D`
37
+ 2. Then write the full report under `docs/test-reports/`
38
+
39
+ ### Inputs
40
+
41
+ 1. **Review report** — `review-backend-cases-pi` output (VERDICT, findings, coverage assessment).
42
+ 2. **Pytest output** — `execute-backend-pytest-shell` stdout/stderr and exit code.
43
+ 3. **Machine-readable report** — `$HARNESS_DAG_RUN_DIR/reports/backend-test-junit.xml` (runner-owned evidence; use the path reported by `execute-backend-pytest-shell`).
44
+
45
+ Do NOT re-read source documents. Use upstream outputs only.
46
+
47
+ ### Maturity Rating Criteria
48
+
49
+ | Rating | Coverage | Pass Rate | Review Findings |
50
+ |--------|----------|-----------|-----------------|
51
+ | **A** | 100% acceptance criteria covered | 100% pytest pass | No Critical or Important findings |
52
+ | **B** | ≥80% acceptance criteria covered | ≥90% pytest pass | Only Informational findings |
53
+ | **C** | ≥60% acceptance criteria covered | ≥70% pytest pass | No Critical findings (Important allowed) |
54
+ | **D** | Below C thresholds | Below C thresholds | Or any Critical finding unresolved |
55
+
56
+ #### Rating Rules
57
+
58
+ - **Coverage** = (acceptance criteria with ≥1 covering test case) / (total acceptance criteria) × 100%
59
+ - **Pass rate** = (passed pytest functions) / (total non-skipped pytest functions) × 100%
60
+ - If `review-backend-cases-pi` returned `VERDICT: request-revision` and revision was not completed, cap at **D**.
61
+ - If pytest exit code is non-zero and >30% tests failed, cap at **D** regardless of coverage.
62
+ - Skipped tests (`@pytest.mark.skip`) count as "not covered" for pass rate but not as failures.
63
+
64
+ ### Report Structure
65
+
66
+ Write the report as a Markdown file named `backend-test-retrospect-<date>.md` under `docs/test-reports/`.
67
+
68
+ ```markdown
69
+ # Backend Test Retrospective Report
70
+
71
+ **Date:** <YYYY-MM-DD>
72
+ **Task:** <task-id>
73
+ **Maturity Rating:** <A|B|C|D>
74
+
75
+ ## 1. Test Coverage Summary
76
+
77
+ | Metric | Value |
78
+ |--------|-------|
79
+ | Total acceptance criteria | N |
80
+ | Covered by test cases | N (X%) |
81
+ | Total functional test cases | N |
82
+ | Positive path cases | N |
83
+ | Negative path cases | N |
84
+ | Boundary cases | N |
85
+
86
+ ## 2. Automation Results
87
+
88
+ | Metric | Value |
89
+ |--------|-------|
90
+ | Total pytest functions | N |
91
+ | Passed | N |
92
+ | Failed | N |
93
+ | Skipped | N |
94
+ | Pass rate | X% |
95
+ | Pytest exit code | N |
96
+
97
+ ### Failed Test Analysis
98
+
99
+ | Test Case ID | Function | Failure Reason | Root Cause |
100
+ |--------------|----------|----------------|------------|
101
+ | ... | ... | ... | ... |
102
+
103
+ ## 3. Review Findings
104
+
105
+ | Severity | Finding | Status |
106
+ |----------|---------|--------|
107
+ | Critical | ... | Resolved / Unresolved |
108
+ | Important | ... | Resolved / Unresolved |
109
+ | Informational | ... | Resolved / Unresolved |
110
+
111
+ ## 4. Maturity Rating Rationale
112
+
113
+ Explain which threshold was met or missed, and why the specific rating was assigned.
114
+
115
+ ## 5. Recommendations
116
+
117
+ - Actionable items for improving the rating in the next iteration.
118
+ - Specific gaps to close (uncovered criteria, flaky tests, missing negative paths).
119
+ ```
120
+
121
+ ### Output Shape (after rating line)
122
+
123
+ After the mandatory maturity rating line, provide a brief summary paragraph before writing the full report file.
124
+
125
+ Do not include chain-of-thought. Do not write root `artifacts/**`.
@@ -1,81 +1,81 @@
1
- # Backend Test DAG Review Cases Prompt Template
2
-
3
- ## Purpose
4
-
5
- Use this prompt for a read-only **backend test case review** node: `executor: "pi"`, `role: "reviewer"`, `writePolicy: "read-only"`. The reviewer audits generated backend functional test cases for completeness, format compliance, and traceability to source requirements. Downstream `generate-backend-pytest-pi` depends on a `VERDICT: pass` to proceed.
6
-
7
- Do **not** create `executor: reviewer`. Reviewer is a **role** on `executor: pi`.
8
-
9
- ## Recommended DAG Node Shape
10
-
11
- ```json
12
- {
13
- "id": "review-backend-cases-pi",
14
- "depends_on": ["generate-backend-functional-cases-pi"],
15
- "complexity": "HIGH",
16
- "executor": "pi",
17
- "role": "reviewer",
18
- "writePolicy": "read-only",
19
- "allowedPaths": ["**"],
20
- "forbiddenPaths": [".harness/**", "artifacts/**"],
21
- "outputContract": "Plain Markdown whose first non-empty line is VERDICT: pass or VERDICT: request-revision; followed by Findings and Coverage Assessment. No file writes.",
22
- "subtask_prompt_markdown": "./backend-test-dag.review-cases.prompt.md"
23
- }
24
- ```
25
-
26
- ## Prompt Body
27
-
28
- You are the Backend Test DAG **test case reviewer** (read-only).
29
-
30
- Your job is to audit the generated backend functional test cases for completeness, format compliance, requirement coverage, and traceability. You are **not** an implementer or test generator. Do not edit repository files, including root `artifacts/**`.
31
-
32
- ### Mandatory First Line
33
-
34
- The **first non-empty line** of your response must be exactly one of:
35
-
36
- - `VERDICT: pass`
37
- - `VERDICT: request-revision`
38
-
39
- No preamble, heading, or blank lines before the verdict line.
40
-
41
- ### Inputs to Review
42
-
43
- 1. **Acceptance criteria** — from upstream `analyze-inputs-pi` output (AC-001, AC-002, ...).
44
- 2. **Generated test cases** — files under `testcase/md/`.
45
-
46
- Do NOT re-read source documents. Use upstream outputs only.
47
-
48
- ### Review Checklist
49
-
50
- | Area | Check | Severity if Missing |
51
- |------|-------|---------------------|
52
- | **ID format** | Every test case ID matches `BE-<MODULE>-<NNN>` (e.g. `BE-ORDER-001`) | Critical |
53
- | **Positive path coverage** | Happy-path scenarios for each acceptance criterion | Critical |
54
- | **Negative path coverage** | Error/exception scenarios (invalid input, not found, state violations) | Important |
55
- | **Boundary conditions** | Edge cases (empty input, max length, edge values) | Important |
56
- | **State transitions** | Illegal state changes covered | Important |
57
- | **Requirement traceability** | Each acceptance criterion (AC-xxx) maps to at least one test case ID | Critical |
58
- | **Case structure** | Each case has: ID, Title, Precondition, Steps, Expected Result | Important |
59
- | **No duplicate IDs** | All test case IDs are unique across files | Critical |
60
-
61
- ### Conditional Coverage (check ONLY if mentioned in upstream analysis)
62
-
63
- - **Authentication coverage**: check ONLY if `analyze-inputs-pi` mentions auth mechanism (JWT, OAuth2, API Key, etc.)
64
- - **Timeout coverage**: check ONLY if `analyze-inputs-pi` mentions timeout handling or degradation strategy
65
- - If not mentioned in upstream analysis, do NOT flag as missing
66
-
67
- ### Verdict Rules
68
-
69
- | Condition | Verdict |
70
- |-----------|---------|
71
- | All Critical checks pass, Important checks have no more than 2 findings | `VERDICT: pass` |
72
- | Any Critical check fails | `VERDICT: request-revision` |
73
- | More than 2 Important findings | `VERDICT: request-revision` |
74
- | Only Informational findings | `VERDICT: pass` (with findings listed) |
75
-
76
- ### Output Shape (after verdict line)
77
-
78
- 1. **Coverage Assessment** — table mapping each AC to covering test case IDs (or "uncovered").
79
- 2. **Findings** — bullet list tagged `Critical`, `Important`, or `Informational`.
80
- 3. **Statistics** — total case count, positive/negative/boundary breakdown, module distribution.
81
- 4. **Required revisions** (only when `request-revision`) — numbered items for the upstream generator to fix.
1
+ # Backend Test DAG Review Cases Prompt Template
2
+
3
+ ## Purpose
4
+
5
+ Use this prompt for a read-only **backend test case review** node: `executor: "pi"`, `role: "reviewer"`, `writePolicy: "read-only"`. The reviewer audits generated backend functional test cases for completeness, format compliance, and traceability to source requirements. Downstream `generate-backend-pytest-pi` depends on a `VERDICT: pass` to proceed.
6
+
7
+ Do **not** create `executor: reviewer`. Reviewer is a **role** on `executor: pi`.
8
+
9
+ ## Recommended DAG Node Shape
10
+
11
+ ```json
12
+ {
13
+ "id": "review-backend-cases-pi",
14
+ "depends_on": ["generate-backend-functional-cases-pi", "backend-test-analysis-contract-shell"],
15
+ "complexity": "HIGH",
16
+ "executor": "pi",
17
+ "role": "reviewer",
18
+ "writePolicy": "read-only",
19
+ "allowedPaths": ["**"],
20
+ "forbiddenPaths": [".harness/**", "artifacts/**"],
21
+ "outputContract": "Plain Markdown whose first non-empty line is VERDICT: pass or VERDICT: request-revision; followed by Findings and Coverage Assessment. No file writes.",
22
+ "subtask_prompt_markdown": "./backend-test-dag.review-cases.prompt.md"
23
+ }
24
+ ```
25
+
26
+ ## Prompt Body
27
+
28
+ You are the Backend Test DAG **test case reviewer** (read-only).
29
+
30
+ Your job is to audit the generated backend functional test cases for completeness, format compliance, requirement coverage, and traceability. You are **not** an implementer or test generator. Do not edit repository files, including root `artifacts/**`.
31
+
32
+ ### Mandatory First Line
33
+
34
+ The **first non-empty line** of your response must be exactly one of:
35
+
36
+ - `VERDICT: pass`
37
+ - `VERDICT: request-revision`
38
+
39
+ No preamble, heading, or blank lines before the verdict line.
40
+
41
+ ### Inputs to Review
42
+
43
+ 1. **Acceptance criteria / analysis** — from the validated Backend Test Analysis v1 artifact materialized by `backend-test-analysis-contract-shell` (`contracts/backend-test-analysis.json` under the current DAG run). Do not treat free-form Markdown from `analyze-inputs-pi` as the contract.
44
+ 2. **Generated test cases** — files under `testcase/md/`.
45
+
46
+ Do NOT re-read source documents. Use the validated analysis artifact and generated cases only.
47
+
48
+ ### Review Checklist
49
+
50
+ | Area | Check | Severity if Missing |
51
+ |------|-------|---------------------|
52
+ | **ID format** | Every test case ID matches `BE-<MODULE>-<NNN>` (e.g. `BE-ORDER-001`) | Critical |
53
+ | **Positive path coverage** | Happy-path scenarios for each acceptance criterion | Critical |
54
+ | **Negative path coverage** | Error/exception scenarios (invalid input, not found, state violations) | Important |
55
+ | **Boundary conditions** | Edge cases (empty input, max length, edge values) | Important |
56
+ | **State transitions** | Illegal state changes covered | Important |
57
+ | **Requirement traceability** | Each acceptance criterion (AC-xxx) maps to at least one test case ID | Critical |
58
+ | **Case structure** | Each case has: ID, Title, Precondition, Steps, Expected Result | Important |
59
+ | **No duplicate IDs** | All test case IDs are unique across files | Critical |
60
+
61
+ ### Conditional Coverage (check ONLY if mentioned in upstream analysis)
62
+
63
+ - **Authentication coverage**: check ONLY if the validated analysis artifact mentions auth mechanism (JWT, OAuth2, API Key, etc.)
64
+ - **Timeout coverage**: check ONLY if the validated analysis artifact mentions timeout handling or degradation strategy
65
+ - If not mentioned in the validated analysis artifact, do NOT flag as missing
66
+
67
+ ### Verdict Rules
68
+
69
+ | Condition | Verdict |
70
+ |-----------|---------|
71
+ | All Critical checks pass, Important checks have no more than 2 findings | `VERDICT: pass` |
72
+ | Any Critical check fails | `VERDICT: request-revision` |
73
+ | More than 2 Important findings | `VERDICT: request-revision` |
74
+ | Only Informational findings | `VERDICT: pass` (with findings listed) |
75
+
76
+ ### Output Shape (after verdict line)
77
+
78
+ 1. **Coverage Assessment** — table mapping each AC to covering test case IDs (or "uncovered").
79
+ 2. **Findings** — bullet list tagged `Critical`, `Important`, or `Informational`.
80
+ 3. **Statistics** — total case count, positive/negative/boundary breakdown, module distribution.
81
+ 4. **Required revisions** (only when `request-revision`) — numbered items for the upstream generator to fix.
@@ -1,64 +1,64 @@
1
- # 执行计划模板
2
-
3
- ## 标题
4
-
5
- ## 状态
6
-
7
- - draft / active / blocked / completed
8
-
9
- ## 背景
10
-
11
- - 当前问题或机会是什么:
12
- - 为什么值得现在做:
13
-
14
- ## 目标(Objective)
15
-
16
- - 本计划希望达成什么结果:
17
-
18
- ## 范围(Scope)
19
-
20
- - 包含:
21
- - 不包含:
22
-
23
- ## 约束(Constraints)
24
-
25
- - 架构/契约约束:
26
- - 时间/风险约束:
27
- - 验证约束:
28
-
29
- ## 里程碑(Milestones)
30
-
31
- | ID | Milestone | Status | Exit Criteria |
32
- |----|-----------|--------|---------------|
33
- | M1 | | todo | |
34
-
35
- ## 工作分解(Work Breakdown)
36
-
37
- | ID | Work Item | Status | Notes |
38
- |----|-----------|--------|-------|
39
- | W1 | | todo | |
40
-
41
- ## 验证关口(Verification Gates)
42
-
43
- - [ ] 文档/契约已同步
44
- - [ ] 最低必要测试已通过
45
- - [ ] 关键路径已验证
46
- - [ ] 风险与未覆盖项已记录
47
-
48
- ## 风险 / 阻塞项
49
-
50
- - 风险:
51
- - 阻塞:
52
-
53
- ## 回退 / 恢复(Rollback / Recovery)
54
-
55
- - 如何回退:
56
- - 回退后如何恢复基线:
57
-
58
- ## Open Questions
59
-
60
- -
61
-
62
- ## 推荐下一步
63
-
64
- -
1
+ # 执行计划模板
2
+
3
+ ## 标题
4
+
5
+ ## 状态
6
+
7
+ - draft / active / blocked / completed
8
+
9
+ ## 背景
10
+
11
+ - 当前问题或机会是什么:
12
+ - 为什么值得现在做:
13
+
14
+ ## 目标(Objective)
15
+
16
+ - 本计划希望达成什么结果:
17
+
18
+ ## 范围(Scope)
19
+
20
+ - 包含:
21
+ - 不包含:
22
+
23
+ ## 约束(Constraints)
24
+
25
+ - 架构/契约约束:
26
+ - 时间/风险约束:
27
+ - 验证约束:
28
+
29
+ ## 里程碑(Milestones)
30
+
31
+ | ID | Milestone | Status | Exit Criteria |
32
+ |----|-----------|--------|---------------|
33
+ | M1 | | todo | |
34
+
35
+ ## 工作分解(Work Breakdown)
36
+
37
+ | ID | Work Item | Status | Notes |
38
+ |----|-----------|--------|-------|
39
+ | W1 | | todo | |
40
+
41
+ ## 验证关口(Verification Gates)
42
+
43
+ - [ ] 文档/契约已同步
44
+ - [ ] 最低必要测试已通过
45
+ - [ ] 关键路径已验证
46
+ - [ ] 风险与未覆盖项已记录
47
+
48
+ ## 风险 / 阻塞项
49
+
50
+ - 风险:
51
+ - 阻塞:
52
+
53
+ ## 回退 / 恢复(Rollback / Recovery)
54
+
55
+ - 如何回退:
56
+ - 回退后如何恢复基线:
57
+
58
+ ## Open Questions
59
+
60
+ -
61
+
62
+ ## 推荐下一步
63
+
64
+ -
@@ -1,53 +1,53 @@
1
- # Feature Spec 模板
2
-
3
- ## 标题
4
-
5
- ## 状态
6
-
7
- - draft / active / completed
8
-
9
- ## 背景(Background)
10
-
11
- - 问题背景:
12
- - 用户/协作者痛点:
13
- - 与现有系统的关系:
14
-
15
- ## 目标(Goal)
16
-
17
- - 本 feature 要实现什么:
18
-
19
- ## 非目标(Non-goals)
20
-
21
- - 明确本轮不做什么:
22
-
23
- ## 关键场景 / 用户路径
24
-
25
- 1.
26
- 2.
27
- 3.
28
-
29
- ## Feature List
30
-
31
- | ID | Feature | Priority | Status | Acceptance | Notes |
32
- |----|---------|----------|--------|------------|-------|
33
- | F1 | | High | todo | | |
34
-
35
- ## 依赖(Dependencies)
36
-
37
- - 上游依赖:
38
- - 下游影响:
39
- - 文档/契约依赖:
40
-
41
- ## 风险(Risks)
42
-
43
- -
44
-
45
- ## 验证说明(Verification Notes)
46
-
47
- - 最低建议验证:
48
- - 加强验证:
49
- - 关键手工路径:
50
-
51
- ## Open Questions
52
-
53
- -
1
+ # Feature Spec 模板
2
+
3
+ ## 标题
4
+
5
+ ## 状态
6
+
7
+ - draft / active / completed
8
+
9
+ ## 背景(Background)
10
+
11
+ - 问题背景:
12
+ - 用户/协作者痛点:
13
+ - 与现有系统的关系:
14
+
15
+ ## 目标(Goal)
16
+
17
+ - 本 feature 要实现什么:
18
+
19
+ ## 非目标(Non-goals)
20
+
21
+ - 明确本轮不做什么:
22
+
23
+ ## 关键场景 / 用户路径
24
+
25
+ 1.
26
+ 2.
27
+ 3.
28
+
29
+ ## Feature List
30
+
31
+ | ID | Feature | Priority | Status | Acceptance | Notes |
32
+ |----|---------|----------|--------|------------|-------|
33
+ | F1 | | High | todo | | |
34
+
35
+ ## 依赖(Dependencies)
36
+
37
+ - 上游依赖:
38
+ - 下游影响:
39
+ - 文档/契约依赖:
40
+
41
+ ## 风险(Risks)
42
+
43
+ -
44
+
45
+ ## 验证说明(Verification Notes)
46
+
47
+ - 最低建议验证:
48
+ - 加强验证:
49
+ - 关键手工路径:
50
+
51
+ ## Open Questions
52
+
53
+ -