deviatdd 2.20.0__tar.gz → 2.20.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (400) hide show
  1. deviatdd-2.20.2/.worktrees/work/.env.example +8 -0
  2. {deviatdd-2.20.0 → deviatdd-2.20.2}/CHANGELOG.md +4 -0
  3. {deviatdd-2.20.0 → deviatdd-2.20.2}/PKG-INFO +1 -1
  4. {deviatdd-2.20.0 → deviatdd-2.20.2}/pyproject.toml +1 -1
  5. deviatdd-2.20.2/specs/005-acceptance-gates/002-task-acceptance-traceability/plan.md +150 -0
  6. deviatdd-2.20.2/specs/005-acceptance-gates/002-task-acceptance-traceability/tasks.jsonl +9 -0
  7. deviatdd-2.20.2/specs/005-acceptance-gates/002-task-acceptance-traceability/tasks.md +113 -0
  8. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/DeviaTDD-api.md +25 -20
  9. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/DeviaTDD-architecture.md +21 -10
  10. deviatdd-2.20.2/specs/adhoc/issues/016-single-source-prompt-templates.md +139 -0
  11. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/prd.md +16 -0
  12. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/issues.jsonl +3 -0
  13. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/meso.py +39 -7
  14. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/micro.py +52 -27
  15. deviatdd-2.20.2/src/deviate/core/tasks_ledger.py +125 -0
  16. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/validation.py +52 -0
  17. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/plan.md +6 -3
  18. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-adhoc.md +3 -2
  19. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-plan.md +5 -2
  20. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-tasks.md +3 -3
  21. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/state/ledger.py +36 -0
  22. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_meso_contracts.py +6 -5
  23. deviatdd-2.20.2/tests/test_core/test_tasks_ledger.py +404 -0
  24. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_validation.py +43 -0
  25. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_meso_resume.py +11 -10
  26. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_run.py +170 -5
  27. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_state/test_ledger.py +182 -0
  28. {deviatdd-2.20.0 → deviatdd-2.20.2}/uv.lock +1 -1
  29. deviatdd-2.20.0/src/deviate/core/tasks_ledger.py +0 -69
  30. deviatdd-2.20.0/tests/test_core/test_tasks_ledger.py +0 -145
  31. {deviatdd-2.20.0 → deviatdd-2.20.2}/.deviate/.gitignore +0 -0
  32. {deviatdd-2.20.0 → deviatdd-2.20.2}/.deviate/config.toml +0 -0
  33. {deviatdd-2.20.0 → deviatdd-2.20.2}/.deviate/evaluations/deviate-flows-skill-2026-06-26.json +0 -0
  34. {deviatdd-2.20.0 → deviatdd-2.20.2}/.env.example +0 -0
  35. {deviatdd-2.20.0 → deviatdd-2.20.2}/.gitattributes +0 -0
  36. {deviatdd-2.20.0 → deviatdd-2.20.2}/.githooks/pre-commit +0 -0
  37. {deviatdd-2.20.0 → deviatdd-2.20.2}/.githooks/pre-push +0 -0
  38. {deviatdd-2.20.0 → deviatdd-2.20.2}/.github/ISSUE_TEMPLATE/bug.md +0 -0
  39. {deviatdd-2.20.0 → deviatdd-2.20.2}/.github/ISSUE_TEMPLATE/feature.md +0 -0
  40. {deviatdd-2.20.0 → deviatdd-2.20.2}/.github/PULL_REQUEST_TEMPLATE.md +0 -0
  41. {deviatdd-2.20.0 → deviatdd-2.20.2}/.github/workflows/ci.yml +0 -0
  42. {deviatdd-2.20.0 → deviatdd-2.20.2}/.gitignore +0 -0
  43. {deviatdd-2.20.0 → deviatdd-2.20.2}/.opencode/.gitignore +0 -0
  44. {deviatdd-2.20.0 → deviatdd-2.20.2}/.opencode/opencode.json +0 -0
  45. {deviatdd-2.20.0/.worktrees/feat/005-acceptance-gates/002-task-acceptance-traceability → deviatdd-2.20.2/.worktrees/feat/005-acceptance-gates/003-micro-phase-gates-red-green}/.env.example +0 -0
  46. {deviatdd-2.20.0 → deviatdd-2.20.2}/.worktrees/feat/adhoc/012-rpc-streaming-tui-renderer/.env.example +0 -0
  47. {deviatdd-2.20.0 → deviatdd-2.20.2}/.worktrees/feat/adhoc/014-cwe-mapping-security-findings/.env.example +0 -0
  48. {deviatdd-2.20.0/.worktrees/work → deviatdd-2.20.2/.worktrees/feat/adhoc/016-single-source-prompt-templates}/.env.example +0 -0
  49. {deviatdd-2.20.0 → deviatdd-2.20.2}/AGENTS.md +0 -0
  50. {deviatdd-2.20.0 → deviatdd-2.20.2}/CLAUDE.md +0 -0
  51. {deviatdd-2.20.0 → deviatdd-2.20.2}/CODE_OF_CONDUCT.md +0 -0
  52. {deviatdd-2.20.0 → deviatdd-2.20.2}/CONTRIBUTING.md +0 -0
  53. {deviatdd-2.20.0 → deviatdd-2.20.2}/LICENSE +0 -0
  54. {deviatdd-2.20.0 → deviatdd-2.20.2}/README.md +0 -0
  55. {deviatdd-2.20.0 → deviatdd-2.20.2}/SECURITY.md +0 -0
  56. {deviatdd-2.20.0 → deviatdd-2.20.2}/deviatdd.png +0 -0
  57. {deviatdd-2.20.0 → deviatdd-2.20.2}/mise.toml +0 -0
  58. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat/001-deviate-cli-python/008-meso-macro-automated-orchestration.md +0 -0
  59. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat/adhoc/006-context-cli-integration.md +0 -0
  60. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-001-deviate-cli-python-001-cli-initialization-governance-provisioning-pr-11.md +0 -0
  61. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-001-deviate-cli-python-003-meso-layer-specification-task-decomposition.md +0 -0
  62. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-001-deviate-cli-python-004-micro-layer-tdd-sandbox-execution.md +0 -0
  63. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-001-deviate-cli-python-005-cli-architecture-realignment-skill-integration.md +0 -0
  64. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-001-deviate-cli-python-007-macro-meso-parity-backward-compatibility.md +0 -0
  65. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-002-deviatdd-gap-analysis-001-foundation-cli-infrastructure.md +0 -0
  66. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-002-deviatdd-gap-analysis-003-fast-path-commands.md +0 -0
  67. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-002-deviatdd-gap-analysis-004-governance-inspection.md +0 -0
  68. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-adhoc-001-streaming-pipeline-monitor.md +0 -0
  69. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat-adhoc-008-ast-phase-prioritization.md +0 -0
  70. {deviatdd-2.20.0 → deviatdd-2.20.2}/pr_descriptions/feat_002-deviatdd-gap-analysis_001-foundation-cli-infrastructure.md +0 -0
  71. {deviatdd-2.20.0 → deviatdd-2.20.2}/scripts/benchmark_lmstudio.py +0 -0
  72. {deviatdd-2.20.0 → deviatdd-2.20.2}/scripts/verify_install.py +0 -0
  73. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t001.json +0 -0
  74. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t002.json +0 -0
  75. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t004.json +0 -0
  76. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/spec.md +0 -0
  77. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/tasks.md +0 -0
  78. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/002-macro-layer-state-ledger-management/spec.md +0 -0
  79. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/002-macro-layer-state-ledger-management/tasks.md +0 -0
  80. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/003-meso-layer-specification-task-decomposition/spec.md +0 -0
  81. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/003-meso-layer-specification-task-decomposition/tasks.md +0 -0
  82. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/spec.md +0 -0
  83. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/tasks.jsonl +0 -0
  84. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/tasks.md +0 -0
  85. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/005-cli-architecture-realignment-skill-integration/spec.md +0 -0
  86. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/005-cli-architecture-realignment-skill-integration/tasks.md +0 -0
  87. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/007-macro-meso-parity-backward-compatibility/spec.md +0 -0
  88. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/007-macro-meso-parity-backward-compatibility/tasks.md +0 -0
  89. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/spec.md +0 -0
  90. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/tasks.jsonl +0 -0
  91. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/tasks.md +0 -0
  92. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/data-model.md +0 -0
  93. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/design.md +0 -0
  94. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/explore.md +0 -0
  95. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/001-cli-initialization-governance-provisioning.md +0 -0
  96. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/002-macro-layer-state-ledger-management.md +0 -0
  97. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/003-meso-layer-specification-task-decomposition.md +0 -0
  98. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/004-micro-layer-tdd-sandbox-execution.md +0 -0
  99. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/005-cli-architecture-realignment-skill-integration.md +0 -0
  100. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/006-state-persistence-concurrency-safety.md +0 -0
  101. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/007-macro-meso-parity-backward-compatibility.md +0 -0
  102. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/issues/008-meso-macro-automated-orchestration.md +0 -0
  103. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/001-deviate-cli-python/prd.md +0 -0
  104. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/spec.md +0 -0
  105. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/tasks.jsonl +0 -0
  106. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/tasks.md +0 -0
  107. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/spec.md +0 -0
  108. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/tasks.jsonl +0 -0
  109. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/tasks.md +0 -0
  110. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/004-governance-inspection/spec.md +0 -0
  111. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/004-governance-inspection/tasks.jsonl +0 -0
  112. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/004-governance-inspection/tasks.md +0 -0
  113. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/spec.md +0 -0
  114. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/tasks.jsonl +0 -0
  115. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/tasks.md +0 -0
  116. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/data-model.md +0 -0
  117. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/design.md +0 -0
  118. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/explore.md +0 -0
  119. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/issues/001-foundation-cli-infrastructure.md +0 -0
  120. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/issues/002-context-pipeline.md +0 -0
  121. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/issues/003-fast-path-commands.md +0 -0
  122. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/issues/004-governance-inspection.md +0 -0
  123. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/issues/005-micro-layer-integrity.md +0 -0
  124. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/plan-tdd-integration-gap.md +0 -0
  125. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/002-deviatdd-gap-analysis/prd.md +0 -0
  126. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/003-graphite-cli-integration/explore.md +0 -0
  127. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/004-per-task-security-profile/issues/001-security-profile-and-judge-checks.md +0 -0
  128. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/001-verification-mode-metadata/plan.md +0 -0
  129. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/001-verification-mode-metadata/tasks.jsonl +0 -0
  130. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/001-verification-mode-metadata/tasks.md +0 -0
  131. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/data-model.md +0 -0
  132. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/design.md +0 -0
  133. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/explore.md +0 -0
  134. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/issues/001-verification-mode-metadata.md +0 -0
  135. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/issues/002-task-acceptance-traceability.md +0 -0
  136. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/issues/003-micro-phase-gates-red-green.md +0 -0
  137. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/issues/004-refactor-regression-gate.md +0 -0
  138. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/issues/005-prompt-spec-alignment.md +0 -0
  139. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/005-acceptance-gates/prd.md +0 -0
  140. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/architecture.md +0 -0
  141. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/domain-model.md +0 -0
  142. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/flows/flows-product.md +0 -0
  143. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/flows/flows-streaming.md +0 -0
  144. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/flows/index.md +0 -0
  145. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/flows.jsonl +0 -0
  146. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/guildwright-current-system.md +0 -0
  147. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/guildwright-gap-register.md +0 -0
  148. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/guildwright-git-state-model.md +0 -0
  149. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/guildwright-rewrite.md +0 -0
  150. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/guildwright-rust-tui-requirements.md +0 -0
  151. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/_product/release-next.md +0 -0
  152. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/001-streaming-pipeline-monitor/spec.md +0 -0
  153. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/001-streaming-pipeline-monitor/tasks.jsonl +0 -0
  154. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/001-streaming-pipeline-monitor/tasks.md +0 -0
  155. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/003-meso-layer-restructuring/spec.md +0 -0
  156. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/003-meso-layer-restructuring/tasks.jsonl +0 -0
  157. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/003-meso-layer-restructuring/tasks.md +0 -0
  158. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/004-deviate-review-skill/spec.md +0 -0
  159. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/004-deviate-review-skill/tasks.jsonl +0 -0
  160. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/004-deviate-review-skill/tasks.md +0 -0
  161. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/005-per-phase-model-configuration/plan.md +0 -0
  162. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/005-per-phase-model-configuration/tasks.jsonl +0 -0
  163. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/005-per-phase-model-configuration/tasks.md +0 -0
  164. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/006-context-cli-integration/plan.md +0 -0
  165. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/006-context-cli-integration/tasks.jsonl +0 -0
  166. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/006-context-cli-integration/tasks.md +0 -0
  167. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/007-graphite-cli/plan.md +0 -0
  168. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/007-graphite-cli/tasks.jsonl +0 -0
  169. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/007-graphite-cli/tasks.md +0 -0
  170. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/008-ast-phase-prioritization/plan.md +0 -0
  171. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/008-ast-phase-prioritization/tasks.jsonl +0 -0
  172. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/008-ast-phase-prioritization/tasks.md +0 -0
  173. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/009-pi-agent-backend-integration/plan.md +0 -0
  174. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/009-pi-agent-backend-integration/tasks.jsonl +0 -0
  175. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/009-pi-agent-backend-integration/tasks.md +0 -0
  176. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/010-deviate-setup-product-layer/plan.md +0 -0
  177. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/010-deviate-setup-product-layer/tasks.jsonl +0 -0
  178. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/010-deviate-setup-product-layer/tasks.md +0 -0
  179. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/plan.md +0 -0
  180. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/tasks.jsonl +0 -0
  181. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/tasks.md +0 -0
  182. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/015-narrow-product-flow-scope/plan.md +0 -0
  183. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/015-narrow-product-flow-scope/tasks.md +0 -0
  184. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/001-streaming-pipeline-monitor.md +0 -0
  185. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/002-aider-agent-backend-integration.md +0 -0
  186. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/003-meso-layer-restructuring.md +0 -0
  187. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/004-deviate-review-skill.md +0 -0
  188. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/005-per-phase-model-configuration.md +0 -0
  189. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/006-context-cli-integration.md +0 -0
  190. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/007-graphite-cli.md +0 -0
  191. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/008-ast-phase-prioritization.md +0 -0
  192. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/009-pi-agent-backend-integration.md +0 -0
  193. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/010-deviate-setup-product-layer.md +0 -0
  194. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/012-rpc-streaming-tui-renderer.md +0 -0
  195. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/013-flow-ledger-canonical-source-of-truth.md +0 -0
  196. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/014-cwe-mapping-security-findings.md +0 -0
  197. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/adhoc/issues/015-narrow-product-flow-scope.md +0 -0
  198. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/constitution.md +0 -0
  199. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/ast-tree-sitter.md +0 -0
  200. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/flow-ledger.md +0 -0
  201. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/graphite-cli.md +0 -0
  202. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/pi-agent-backend.md +0 -0
  203. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/product-layer.md +0 -0
  204. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/rpc-streaming-tui.md +0 -0
  205. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/rpc-streaming.md +0 -0
  206. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/explore/security-hardening-cwe.md +0 -0
  207. {deviatdd-2.20.0 → deviatdd-2.20.2}/specs/implementation-gap.md +0 -0
  208. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/__init__.py +0 -0
  209. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/__init__.py +0 -0
  210. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/_common.py +0 -0
  211. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/_html.py +0 -0
  212. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/_safe_commands.py +0 -0
  213. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/adhoc.py +0 -0
  214. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/constitution.py +0 -0
  215. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/feature.py +0 -0
  216. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/flow_commands.py +0 -0
  217. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/init.py +0 -0
  218. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/inspect.py +0 -0
  219. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/macro.py +0 -0
  220. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/review.py +0 -0
  221. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/cli/walkthrough.py +0 -0
  222. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/__init__.py +0 -0
  223. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/_shared.py +0 -0
  224. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/agent.py +0 -0
  225. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/cache_discipline.py +0 -0
  226. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/commands.py +0 -0
  227. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/commit.py +0 -0
  228. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/complexity.py +0 -0
  229. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/constitution.py +0 -0
  230. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/contract.py +0 -0
  231. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/convention.py +0 -0
  232. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/epic.py +0 -0
  233. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/issues.py +0 -0
  234. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/prd.py +0 -0
  235. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/profile.py +0 -0
  236. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/repo.py +0 -0
  237. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/run_logger.py +0 -0
  238. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/core/worktree.py +0 -0
  239. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/html_templates/__init__.py +0 -0
  240. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/html_templates/architecture.html.tmpl +0 -0
  241. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/html_templates/domain-model.html.tmpl +0 -0
  242. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/html_templates/flows.html.tmpl +0 -0
  243. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/html_templates/plan.html.tmpl +0 -0
  244. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/html_templates/prd.html.tmpl +0 -0
  245. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/main.py +0 -0
  246. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/__init__.py +0 -0
  247. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/assembly.py +0 -0
  248. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/__init__.py +0 -0
  249. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/execute.md +0 -0
  250. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/explore.md +0 -0
  251. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/green.md +0 -0
  252. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/judge.md +0 -0
  253. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/prd.md +0 -0
  254. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/red.md +0 -0
  255. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/refactor.md +0 -0
  256. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/research.md +0 -0
  257. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/shard.md +0 -0
  258. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/specify.md +0 -0
  259. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/auto/tasks.md +0 -0
  260. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-architecture.md +0 -0
  261. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-constitution.md +0 -0
  262. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-e2e.md +0 -0
  263. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-execute.md +0 -0
  264. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-explore.md +0 -0
  265. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-flows.md +0 -0
  266. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-green.md +0 -0
  267. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-hotfix.md +0 -0
  268. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-html.md +0 -0
  269. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-init.md +0 -0
  270. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-judge.md +0 -0
  271. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-merge.md +0 -0
  272. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-pr.md +0 -0
  273. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-prd.md +0 -0
  274. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-prune.md +0 -0
  275. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-red.md +0 -0
  276. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-refactor.md +0 -0
  277. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-release.md +0 -0
  278. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-research.md +0 -0
  279. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-review.md +0 -0
  280. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-shard.md +0 -0
  281. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-triage.md +0 -0
  282. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/commands/deviate-walkthrough.md +0 -0
  283. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/constitution_seed.md +0 -0
  284. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/core.md +0 -0
  285. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/lifecycle-auto.md +0 -0
  286. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/lifecycle-manual.md +0 -0
  287. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/macro-shared.md +0 -0
  288. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/meso-shared.md +0 -0
  289. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/micro-shared.md +0 -0
  290. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/product-shared.md +0 -0
  291. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/core/style-ste.md +0 -0
  292. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/extras/deviate-pr-graphite-routing.md +0 -0
  293. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/governance/__init__.py +0 -0
  294. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/governance/agents_seed.md +0 -0
  295. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/governance/claudemd_seed.md +0 -0
  296. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/governance/graphite_seed.md +0 -0
  297. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/governance/libref_seed.md +0 -0
  298. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/prompts/skills/deviatdd/SKILL.md +0 -0
  299. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/state/__init__.py +0 -0
  300. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/state/config.py +0 -0
  301. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/ui/__init__.py +0 -0
  302. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/ui/monitor.py +0 -0
  303. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/ui/pipeline.py +0 -0
  304. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/ui/render.py +0 -0
  305. {deviatdd-2.20.0 → deviatdd-2.20.2}/src/deviate/visual/__init__.py +0 -0
  306. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/__init__.py +0 -0
  307. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/conftest.py +0 -0
  308. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/core/__init__.py +0 -0
  309. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/core/test_agent.py +0 -0
  310. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/core/test_smart_stall.py +0 -0
  311. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/e2e/test_macro_workflow.bats +0 -0
  312. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/__init__.py +0 -0
  313. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_adhoc.py +0 -0
  314. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_common.py +0 -0
  315. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_constitution.py +0 -0
  316. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_feature.py +0 -0
  317. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_flows_sync.py +0 -0
  318. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_help.py +0 -0
  319. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_html.py +0 -0
  320. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_init.py +0 -0
  321. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_inspect.py +0 -0
  322. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_macro_contracts.py +0 -0
  323. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_main_entrypoint.py +0 -0
  324. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_merge.py +0 -0
  325. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_meso.py +0 -0
  326. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_micro.py +0 -0
  327. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_real_descendant_kill.py +0 -0
  328. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_review.py +0 -0
  329. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_safe_commands.py +0 -0
  330. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_test_command_resolution.py +0 -0
  331. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_timeout_safe_command.py +0 -0
  332. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_cli/test_top_level_run.py +0 -0
  333. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_agent.py +0 -0
  334. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_cache_discipline.py +0 -0
  335. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_commands.py +0 -0
  336. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_commit.py +0 -0
  337. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_complexity.py +0 -0
  338. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_constitution.py +0 -0
  339. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_contract.py +0 -0
  340. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_convention.py +0 -0
  341. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_epic.py +0 -0
  342. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_flow_confirmation.py +0 -0
  343. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_issues.py +0 -0
  344. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_ledger.py +0 -0
  345. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_prd.py +0 -0
  346. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_profile.py +0 -0
  347. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_repo.py +0 -0
  348. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_run_logger.py +0 -0
  349. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_core/test_worktree.py +0 -0
  350. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_init.py +0 -0
  351. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/__init__.py +0 -0
  352. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/conftest.py +0 -0
  353. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_command_installation.py +0 -0
  354. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_init_export_cycle.py +0 -0
  355. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_macro_full_cycle.py +0 -0
  356. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_macro_layer.py +0 -0
  357. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_macro_orchestration.py +0 -0
  358. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_meso_layer.py +0 -0
  359. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_meso_orchestration.py +0 -0
  360. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_meso_task_ledger.py +0 -0
  361. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_parity.py +0 -0
  362. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_integration/test_skill_installation.py +0 -0
  363. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_macro/__init__.py +0 -0
  364. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_macro/test_explore.py +0 -0
  365. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_macro/test_macro_model_routing.py +0 -0
  366. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_macro/test_macro_orchestration.py +0 -0
  367. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_macro/test_prd.py +0 -0
  368. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_macro/test_research.py +0 -0
  369. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_macro/test_shard.py +0 -0
  370. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/__init__.py +0 -0
  371. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_auto_prompt_templates.py +0 -0
  372. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_meso_model_routing.py +0 -0
  373. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_meso_orchestration.py +0 -0
  374. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_plan_structure_injection.py +0 -0
  375. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_prompt_assembly.py +0 -0
  376. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_specify.py +0 -0
  377. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_meso/test_tasks.py +0 -0
  378. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/__init__.py +0 -0
  379. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/conftest.py +0 -0
  380. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_commit_failure.py +0 -0
  381. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_e2e.py +0 -0
  382. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_execute.py +0 -0
  383. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_green.py +0 -0
  384. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_hotfix.py +0 -0
  385. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_judge.py +0 -0
  386. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_orchestration.py +0 -0
  387. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_output_filter.py +0 -0
  388. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_red.py +0 -0
  389. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_refactor.py +0 -0
  390. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_rollback_safety.py +0 -0
  391. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_micro/test_task_label.py +0 -0
  392. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_state/__init__.py +0 -0
  393. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_state/test_config.py +0 -0
  394. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_state/test_security_profile.py +0 -0
  395. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_state/test_session.py +0 -0
  396. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_ui/__init__.py +0 -0
  397. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_ui/test_monitor.py +0 -0
  398. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_ui/test_pipeline.py +0 -0
  399. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_ui/test_render.py +0 -0
  400. {deviatdd-2.20.0 → deviatdd-2.20.2}/tests/test_visual_demo/test_tsk_001_01.py +0 -0
@@ -0,0 +1,8 @@
1
+ # PyPI API token for `mise run publish`.
2
+ #
3
+ # Generate at https://pypi.org/manage/account/token/
4
+ # - Scope the token to the `deviatdd` project only (not account-wide)
5
+ # - Copy this file to `.env` and paste your real token
6
+ # - `.env` is gitignored; never commit secrets
7
+
8
+ PYPI_API_TOKEN=
@@ -23,6 +23,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
23
23
  - **`deviate flows sync` — sole owner of `specs/_product/flows.jsonl` creation.** New top-level `flows` Typer group exposes a single `sync` subcommand that parses `specs/_product/flows/index.md` and appends one ``FlowRecord`` identity row plus ``FLOW_DISCOVERED`` and ``FLOW_DOCUMENTED`` events per flow. Re-running on a populated ledger is a no-op (compound-key idempotency on the underlying append helpers). Exits non-zero with ``FLOWS_INDEX_MISSING`` on stderr when the canonical index is absent, and ``FLOWS_INDEX_EMPTY`` when the index parses to zero rows — surfaces authoring defects instead of silently committing a half-baked catalog. New public service `seed_flow_ledger(flows_index, ledger_path)` in `src/deviate/state/ledger.py` plus `FlowIndexEmptyError` for the empty-index contract. `deviate explore post` no longer seeds identity or documentation events; it only renders the coverage report. ``flows.jsonl`` never receives ``FLOW_REFERENCED_BY_ISSUE`` events — the referenced-by relationship is derived read-only from ``specs/issues.jsonl::flow_refs`` at coverage time.
24
24
  - **`deviate html prd --bucket <slug>` targets a specific epic when more than one owns a `prd.md`.** Previously the command hard-failed with `HTML_AMBIGUOUS_PRD` whenever multiple numbered epics had a `prd.md` and offered no flag to disambiguate — the agent was told to "run from within the epic's worktree", which silently yielded `HTML_NO_PRD` because the resolver reads `specs/` from the repo root. The new option resolves `specs/<bucket>/prd.md` directly and bypasses ambiguity detection; the plain form keeps its existing behavior but its ambiguity banner now points at the flag. An unknown/absent bucket exits `PRD_NOT_FOUND`. The `/deviate-html` prompt (`src/deviate/prompts/commands/deviate-html.md`) and spec docs now reference `--bucket` instead of the broken cwd fallback. Pinned by `tests/test_cli/test_html.py::test_html_prd_bucket_targets_specific_epic`, `::test_html_prd_bucket_targets_unnumbered_dir`, and `::test_html_prd_bucket_missing_file_exits_cleanly`.
25
25
  ### Changed
26
+ - **Ad-hoc issue creation no longer marks the record COMPLETED at creation time.** `/deviate-adhoc` step 7 previously instructed the agent to “mark the record as completed” via `deviate adhoc post` — but that CLI reads a different ledger (`specs/adhoc.jsonl`) than the `specs/issues.jsonl` ledger the prompt populates, so agents instead appended a raw `COMPLETED` transition to `specs/issues.jsonl` immediately, before any plan/tasks/red-green work ran (visible as instant `BACKLOG → SPECIFIED → COMPLETED` sequences for `ISS-ADH-003/004/013`). Step 7 now commits artifacts with a plain `git commit`, leaves the record at `BACKLOG`, and forbids appending any `COMPLETED` transition at creation — completion is written only by later phase post-scripts. The re-installed mirrors (`.opencode/`, `.claude/`, `.omp/`, `.factory/`) carry the fix via `deviate setup`. Prompt-only change in `src/deviate/prompts/commands/deviate-adhoc.md`; pinned by the existing `TestConsumerRepositoryPromptBoundaries` prompt tests.
27
+ - **`/deviate-tasks` no longer waits for the removed Gate 2 approval.** The prompt still said the workflow “stops for joint human review” and that “TDD begins only after explicit artifact-bound Gate 2 approval” — but constitution 0.8.0 removed Gate 2, so the agent was being told to wait for an approval that no longer exists. The Meso Workflow Position line and the Tasks/TDD bullets now state that no human-approval gate sits between Tasks and Micro and that `deviate run` chains meso into micro end-to-end. Prompt-only change in `src/deviate/prompts/commands/deviate-tasks.md`, consistent with the constitution's Automated Advance invariant.
26
28
  - **RED phase now hard-fails when it produces no test files (`RED phase produced no test files`), instead of silently committing a vacuous 'failing test' commit.** Previously `_run_red_phase` skipped its test run entirely when `_find_test_files` returned empty, letting GREEN run against nothing and die in `TRAIN_EXHAUSTED` — a real stall. It now raises a `PhaseFailedError` naming `/deviate-execute` (DIRECT) or `/deviate-meso` re-sharding for tasks that legitimately need no code. It also treats a pytest-style exit code 5 (no tests collected) as a RED defect rather than a passing RED. The manual `deviate red post` `RedMustPassError` message now points the agent at the `failure_kind: already_satisfied` escape hatch instead of leading to an opaque crash. Pinned by the updated `_find_test_files`-mocked RED unit tests.
27
29
  - **Bare `deviate specify` now skips issues already claimed on a remote branch during auto-discovery.** Previously it auto-discovered the next unblocked BACKLOG issue via `select_next_unblocked_issue()` without checking the remote, so it could pick an issue whose `feat/{epic}/{issue}` branch already existed on `origin` (claimed elsewhere) and then hard-fail with `CLAIM_FAILED`. It now uses `_discover_claimable_issue()` — the same discovery `deviate meso run` uses — which skips issues whose branch is on the remote (or already COMPLETED) and reaches the next genuinely claimable one. When none remain it exits `NO_CLAIMABLE_ISSUES`. `deviate specify <id>` / `--local` explicit-claim behavior is unchanged. Pinned by `tests/test_meso/test_specify.py::TestSpecifySetup::test_specify_bare_arg_skips_remote_claimed_issue`.
28
30
  - **`deviate explore post` no longer writes ``FLOW_REFERENCED_BY_ISSUE`` events into ``specs/_product/flows.jsonl``.** The flows ledger never records issue-flow references — the referenced-by relationship is derived read-only at coverage time from ``specs/issues.jsonl::flow_refs`` (via ``load_flow_coverage``/``_latest_issue_reference``). `deviate explore post` (``_run_flow_ledger_cycle``) now renders the coverage report only; the macro-layer reverse-index helper ``_reverse_index_issue_flow_refs`` was removed. ``flows.jsonl`` still carries ``FLOW_DISCOVERED`` / ``FLOW_DOCUMENTED`` (from `deviate flows sync`), confirmation/evidence/release events (from merge/release), and deprecated events. Pinned by `tests/test_macro/test_explore.py::test_explore_post_never_writes_referenced_by_events`.
@@ -36,6 +38,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
36
38
  - **`/deviate-merge` no longer auto-pushes after the squash-merge commit; the push gate runs inline and the network push is opt-in.** The slash command previously ran `git push` as its final step. As of v2.4.0 the squash-merge commit lands on `main`, then a new `push_gate` step inlines the body of `.githooks/pre-push` (lint + format-check + testmon-driven affected tests with the warm-cache / full-suite fallback — bash 3.2 portable, `GIT_DIR` reset + trap preserved) so the safety net fires even though no `git push` happens yet. After the gate passes the skill asks the operator whether to `git push` (which fires the real `pre-push` hook and re-runs the same gate) or stop and push manually. The squash-merge commit and the ledger transition inside it are durable on `main` regardless of the push outcome — only the network push is deferred. New failure states: `Push_Gate_Failed` (inline gate non-zero), `Push_Failed` (`git push` non-zero, raw stderr surfaced), `Push_Deferred` (user chose "Stop — I'll push manually"). Inline gate body and `.githooks/pre-push` body must stay byte-equivalent; divergence is pinned by `tests/test_meso/test_auto_prompt_templates.py::TestMergePromptPushGate::test_hook_and_prompt_agree_on_gate_body` (which compares non-blank non-comment lines in both bodies and fails on drift in either direction) plus 3 supporting assertions on the upstream-first logic, the testmon fallback, and the prompt structure. Prompt: `src/deviate/prompts/commands/deviate-merge.md` (v2.3.0 → v2.4.0). Spec mirrors updated: `specs/DeviaTDD-architecture.md` (Merge bullet, new `**Push gate + opt-in push (v2.4.0)**` sub-bullet) and `specs/DeviaTDD-api.md` (new `**/deviate-merge push behavior (v2.4.0)**` entry under the `deviate merge` reference).
37
39
 
38
40
  ### Fixed
41
+ - **Language-agnostic test discovery: the RED gate no longer globs Python-only `tests/**/test_*.py`.** `_find_test_files` returned `[]` for non-Python projects, so `_run_red_phase` rejected a correctly authored Elixir test (`test/integration/tailwind_pipeline_test.exs`) with "RED phase produced no test files" before any command ran — blocking every task on an Elixir/Phoenix repo. The RED gate now trusts the resolved test command (`_test_command_candidates`: task `verification` → constitution `test_command` → `mise run test` → manifest table `mix.exs`→`mix test` / `Cargo.toml`→`cargo test` / `go.mod`→`go test ./...` / `package.json`→`npm test` / `pyproject.toml`→`pytest`), and a project with no resolvable command (returncode 127) routes to the same JUDGE adjudication as pytest's exit-5 no-tests case instead of dying. `deviate red post`, `deviate green post`, and `deviate refactor post` switched their `TEST_NOT_FOUND`/`NO_TESTS_TO_CHECK` gates from the Python glob to the same resolved-command check. `_find_test_files` remains only as the Python fallback signal. Pinned by `tests/test_micro/test_run.py::TestRunCommand::test_run_elixir_repo_red_accepts_exs_test` (fails on the old gate, passes after).
39
42
  - **Manifest schema recovery now drops out-of-enum fields instead of aborting the phase (GH-52, GH-55).** `AgentBackend.parse_output` re-constructed a `HandoverManifest` with the offending value intact after a `ValidationError`, so an out-of-enum `failure_kind` (RED) or `next_action` (JUDGE COMPLIANCE_PASS) re-raised the identical error unguarded and the phase died with a misleading `agent returned no manifest`. Recovery now deletes every top-level field that failed validation (the model default `None` applies); the JUDGE runner's existing `_coerce_judge_action` then routes an unknown `next_action` through the legacy pass path. Pinned by `test_recovery_drops_invalid_failure_kind` / `test_recovery_drops_invalid_next_action` in `tests/test_core/test_agent.py`.
40
43
  - **`deviate micro run` dedupes latest task records by `(issue_id, task_id)` so a same-numbered task in another issue no longer shadows the active issue's ledger.** `_collect_latest_task_records` previously keyed the latest record by task ID alone. Since every issue reuses the same `TSK-NNN-NN` numbering, a later-sorting ledger (e.g. `specs/adhoc/*` after `specs/005-*`) shadowed the active issue's records — bare `deviate micro run` reported `no ledger entry` against a committed ledger and explicit `<tid>` dispatch hit the wrong issue's task. Latest records are now keyed per issue, and `_find_task_record` prefers the record owned by the current worktree's branch issue. Pinned by `test_find_task_record_prefers_branch_issue` in `tests/test_micro/test_e2e.py`.
41
44
  - **`deviate micro run` re-keys a stale worktree session issue to the branch's issue (GH-54).** A freshly claimed worktree could carry the previous issue's id in `.deviate/session.json` while its branch pointed at the new issue; `_resolve_task_context` trusted the stale id, found no tasks board for it in the checkout, and emitted `NO_PENDING_TASKS` exit 0 for a queue that existed. Resolution now prefers the branch-derived issue whenever the session issue differs from it and has no tasks board in the checkout. Pinned by `test_stale_session_issue_rekeys_to_branch_issue` in `tests/test_cli/test_micro.py`.
@@ -56,6 +59,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
56
59
  - **`deviate shard post` now reliably commits the per-epic `issues/` directory alongside `specs/issues.jsonl`.** When the manifest omits `epic_slug` (or only passes it via `--epic`), the post-script now derives the epic from the first issue's `source_file` and uses a recursive `**/*.md` glob, so the slash command `/deviate-shard` (and its `shard`/`/shard` aliases) commits every shard issue file even if the LLM forgot to write `epic_slug` into the manifest. Missing `issues/` directories and unresolvable epics now emit a visible `SHARD_WARNING` instead of silently producing a ledger-only commit. Implementation lives in `shard_post` (`src/deviate/cli/macro.py`); the pre-script behavior is unchanged. Pinned by `tests/test_integration/test_macro_layer.py::TestShardPost::test_shard_post_commits_issues_directory`.
57
60
  - **`mise publish` no longer fails because a stray `.venv-tmp` directory leaks into the source distribution.** Hatchling's hard-coded sdist exclusion list covers `.venv` but not every venv variant, so a `.venv-tmp` left on disk (with its absolute `bin/python` symlink to `/usr/local/bin/python3.13`) was packaged into `dist/deviatdd-*.tar.gz` and `uv build` aborted with `Invalid tar file ... external symlinks are not allowed`. `pyproject.toml` now adds an explicit `[tool.hatch.build.targets.sdist]` exclude of `/.venv-tmp/`, making the build deterministic regardless of hatchling's built-in directory list. The wheel target already scoped packaging to `src/deviate`.
58
61
  - **JUDGE feedback commit now passes `--no-verify`.** `_commit_judge_feedback_and_advance` (`src/deviate/cli/micro.py`) issues the docs-only `docs(<tid>): add judge feedback for retry` commit with `--no-verify` so a repository pre-commit hook running the full test suite cannot deadlock the micro loop against a still-broken next-task deliverable. Mirrors the `no_verify=True` pattern already enforced at the RED / GREEN / REFACTOR micro-loop commit sites. Pinned by `tests/test_cli/test_micro.py::TestJudgeTrainRollback::test_commit_judge_feedback_and_advance_uses_no_verify`.
62
+ - **`deviate plan post` / `deviate meso tasks pre` / meso-run resume now auto-fill a missing `**Verification Mode**` line instead of rejecting the plan.** When a plan's acceptance contract fails *only* because scenarios lack the `**Verification Mode**:` line, `_validate_or_repair_plan` (`src/deviate/cli/meso.py`) injects the default `automated` value into each affected scenario body and persists the repaired `plan.md` with a `PLAN_MODE_REPAIR` banner on the console. An existing — even invalid or duplicated — mode literal is never touched, so `invalid Verification Mode` / `duplicate Verification Mode lines` still block with `PLAN_ACCEPTANCE_CONTRACT_INVALID` / `MESO_PLAN_INVALID`. The manual `/deviate-plan` prompt template (`src/deviate/prompts/commands/deviate-plan.md`) now also requires the mode line (per-scenario field 6, canonical example, forbidden patterns, and a pre-write self-check step), matching `src/deviate/prompts/auto/plan.md`. Pinned by `tests/test_core/test_validation.py::TestRepairMissingVerificationMode`, the rewritten `tests/test_meso/test_meso_resume.py::TestMesoIdempotentResume::test_modeless_contract_is_repaired_on_resume`, and `tests/test_cli/test_meso_contracts.py::TestMesoContracts::test_tasks_pre_repairs_missing_verification_mode`.
59
63
  ## [2.15.0] - 2026-07-29
60
64
 
61
65
  ### Added
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: deviatdd
3
- Version: 2.20.0
3
+ Version: 2.20.2
4
4
  Summary: DeviaTDD CLI — agent orchestration framework
5
5
  Project-URL: Homepage, https://github.com/wernerbisschoff/deviatdd
6
6
  Project-URL: Repository, https://github.com/wernerbisschoff/deviatdd
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "deviatdd"
3
- version = "2.20.0"
3
+ version = "2.20.2"
4
4
  description = "DeviaTDD CLI — agent orchestration framework"
5
5
  readme = "README.md"
6
6
  license = "MIT"
@@ -0,0 +1,150 @@
1
+ # Plan — Task Acceptance Traceability via acceptance_criteria Links
2
+
3
+ ## Plan Summary
4
+
5
+ - **Issue**: 005-002 — Task Acceptance Traceability via `acceptance_criteria` Links
6
+ - **Implementation Strategy**: This is a re-plan: prior `plan.md` was reviewed and refreshed against the completed issue `005-001`, whose `validate_acceptance_contract` (`src/deviate/core/validation.py:162`) now enforces one `**Verification Mode**` line per `AC-PLAN-NNN` scenario. Add a `CriterionLink` Pydantic model and an optional `acceptance_criteria: list[CriterionLink] | None` field to `TaskRecord` in `src/deviate/state/ledger.py`. Extend `generate_jsonl_from_md` and `_build_task_record` in `src/deviate/core/tasks_ledger.py` to parse per-task `**Acceptance Criteria**:` bullets from `tasks.md` and propagate the parsed links into each generated row. `CriterionLink` enforces the `AC-PLAN-\d{3}` id format, the three verification-mode literals, and the automated-link `test_ref` invariant, so malformed links fail generation with named errors. The field stays optional with default `None`, so legacy rows without the field parse under `model_config = {"extra": "forbid"}`.
7
+ - **Estimated Complexity**: Medium
8
+ - **Estimated Effort**: 2-4 hours
9
+
10
+ ## Product Layer Anchors
11
+
12
+ - **Flow References**: `[]`
13
+ - **Source**: `specs/005-acceptance-gates/issues/002-task-acceptance-traceability.md` (frontmatter field: `flow_refs`)
14
+ - **Release Context**: `specs/_product/release-next.md` Goal ships FLOW-04 (live-stream agent progress via subprocess RPC and a Rich TUI renderer capped at 10 lines). This issue is orthogonal; it hardens meso task-generation traceability, not the RPC/TUI transport.
15
+ - **Architecture Components Touched**: `C1` (the existing `deviate` CLI component; the `## Components` table C1 entry states it owns phase state, JSONL ledgers, and TOML config, which covers `src/deviate/state/ledger.py` and `src/deviate/core/tasks_ledger.py`)
16
+
17
+ ## Acceptance Contract
18
+
19
+ **Scenario AC-PLAN-001: Tasks.md criterion references propagate into generated task rows as acceptance_criteria links**
20
+ - **Source Outline**: `AO-002`
21
+ - **Upstream Traceability**: `US-005-03`, `FR-005-02`, `AC-005-02-01`
22
+ - **Current-Code Evidence**: `src/deviate/core/tasks_ledger.py:15`
23
+ - **Given**: a `tasks.md` whose task `TSK-005-01` carries an `**Acceptance Criteria**:` bullet that declares `AC-PLAN-001 (automated, tests/test_core/test_tasks_ledger.py)` and `AC-PLAN-002 (manual)`
24
+ - **When**: `generate_jsonl_from_md` runs on that `tasks.md`
25
+ - **Then**: the returned `TSK-005-01` `TaskRecord` carries an `acceptance_criteria` list whose two `CriterionLink` entries match the declared `criterion_id`, `verification_mode`, and `test_ref` values
26
+ - **Verification Mode**: automated
27
+
28
+ **Scenario AC-PLAN-002: A malformed criterion id or a missing test_ref on an automated link fails generation with a named error**
29
+ - **Source Outline**: `AO-002`
30
+ - **Upstream Traceability**: `US-005-03`, `FR-005-02`, `AC-005-02-01`
31
+ - **Current-Code Evidence**: `src/deviate/core/tasks_ledger.py:45`
32
+ - **Given**: a `tasks.md` whose task declares an `**Acceptance Criteria**:` bullet with a `criterion_id` outside the `AC-PLAN-\d{3}` pattern (for example `AC-PLAN-99`), or with an `automated` link whose `test_ref` is empty
33
+ - **When**: `generate_jsonl_from_md` runs on that `tasks.md`
34
+ - **Then**: generation raises a validation error whose message names the invalid `criterion_id` or the missing `test_ref`
35
+ - **Verification Mode**: automated
36
+
37
+ **Scenario AC-PLAN-003: A legacy task row without the acceptance_criteria field parses with the field absent under extra="forbid"**
38
+ - **Source Outline**: `AO-002`
39
+ - **Upstream Traceability**: `US-005-04`, `FR-005-02`, `AC-005-02-01`
40
+ - **Current-Code Evidence**: `src/deviate/state/ledger.py:81`
41
+ - **Given**: a `tasks.jsonl` row written by an older CLI version that carries `id`, `issue_id`, `description`, `status`, `execution_mode`, and `created_at` but no `acceptance_criteria` field
42
+ - **When**: `TaskRecord.model_validate` parses that row
43
+ - **Then**: the record parses with `acceptance_criteria` equal to `None`, and a row that carries a genuinely unknown field still fails validation
44
+ - **Verification Mode**: automated
45
+
46
+ **Scenario AC-PLAN-004: A task with no criterion references carries a null acceptance_criteria field, never an empty list**
47
+ - **Source Outline**: `AO-002`
48
+ - **Upstream Traceability**: `US-005-03`, `FR-005-02`, `AC-005-02-01`
49
+ - **Current-Code Evidence**: `src/deviate/state/ledger.py:97`
50
+ - **Given**: a `tasks.md` whose task `TSK-005-04` contains no `**Acceptance Criteria**:` bullet while a sibling task declares one
51
+ - **When**: `generate_jsonl_from_md` runs on that `tasks.md` and serializes the returned records
52
+ - **Then**: the record for `TSK-005-04` serializes `acceptance_criteria` as `null`, never as an empty list
53
+ - **Verification Mode**: automated
54
+
55
+ ## Workstation Mapping
56
+
57
+ - **`src/deviate/state/ledger.py:81-98`**: MODIFY — add the optional `acceptance_criteria: list[CriterionLink] | None = None` field to `TaskRecord`.
58
+ - **Current State**: `TaskRecord` (line 81) declares `id`, `issue_id`, `description`, `status` (seven-value `Literal`), `execution_mode`, `created_at`, and `security_profile`; `model_config = {"extra": "forbid"}` is at line 98. No `acceptance_criteria` field exists.
59
+ - **Changes Required**: Insert `acceptance_criteria: list[CriterionLink] | None = None` after the `security_profile` field, before `model_config`. Preserve `model_config = {"extra": "forbid"}` and the seven-value `status` `Literal` unchanged (PRD `RESOLVED-Q-004`).
60
+ - **Integration Surface**: `_build_task_record` (`src/deviate/core/tasks_ledger.py:45`) constructs the model; generated JSONL rows serialize via `append_task_record` (`src/deviate/state/ledger.py:181`); micro-phase runners parse rows via `TaskRecord.model_validate` (`src/deviate/cli/micro.py:1154,1259,1336,2524,3113,4437,4533,4806`).
61
+ - **`src/deviate/state/ledger.py`**: ADD — the `CriterionLink` Pydantic model.
62
+ - **Current State**: No `CriterionLink` type exists anywhere in the ledger module; `CriterionLink` appears only in the PRD/data-model schemas.
63
+ - **Changes Required**: Define `CriterionLink(BaseModel)` before `TaskRecord` with `criterion_id: str`, `verification_mode: Literal["automated", "manual", "deferred"]`, `test_ref: str | None = None`, and `model_config = {"extra": "forbid"}` (the ledger-family pattern; `SecurityProfile` at line 61-78 is the precedent). Add a `field_validator("criterion_id")` that rejects values that do not match `^AC-PLAN-\d{3}$`. Add a `model_validator(mode="after")` that raises when `verification_mode == "automated"` and `test_ref is None`. The data-model schema sketch (`specs/005-acceptance-gates/data-model.md:86-96`) shows the same field validator. `re` is already imported at line 4, so no new dependency is needed.
64
+ - **Integration Surface**: Embedded in `TaskRecord.acceptance_criteria`; validated per-link at model construction and at row parse.
65
+ - **`src/deviate/core/tasks_ledger.py:15`**: MODIFY — `generate_jsonl_from_md` propagates task-to-criterion and criterion-to-test links from `tasks.md` into each generated `TaskRecord`.
66
+ - **Current State**: The function scans task lines with `_TASK_LINE_PATTERN` (line 11), tracks `current_mode` from `**Mode**:` bullets, and appends `_build_task_record(current_id, issue_id, current_desc, current_mode)` per task. It has no awareness of criterion references.
67
+ - **Changes Required**: Add a `_CRITERIA_LINE_PATTERN` that matches `- **Acceptance Criteria**: <entries>` inside a task block and a `_LINK_PATTERN` that parses each comma-separated entry of the form `AC-PLAN-NNN (mode[, test_ref])`. Track `current_criteria` per task, reset on each new task line, and pass the collected entries to `_build_task_record`. The parse stays linear over the line list and adds no new dependencies (stdlib `re` only).
68
+ - **Integration Surface**: Consumed by `validate_tasks_jsonl` (`src/deviate/core/tasks_ledger.py:60`) and by the meso task-generation flow; the emitted rows serialize through `TaskRecord.model_dump_json`.
69
+ - **`src/deviate/core/tasks_ledger.py:45`**: MODIFY — `_build_task_record` carries the parsed links into the record.
70
+ - **Current State**: The helper builds `TaskRecord(id=..., issue_id=..., description=..., status="PENDING", execution_mode=...)` and returns it; no link handling exists.
71
+ - **Changes Required**: Accept the collected criteria entries (default empty), parse each entry with `_LINK_PATTERN`, and construct `CriterionLink` instances. An entry that fails `_LINK_PATTERN` raises a `ValueError` that names the offending text and the task id. A `criterion_id` outside `AC-PLAN-\d{3}`, a `verification_mode` outside the three literals, or an `automated` link with a null `test_ref` raises a `pydantic.ValidationError` from the `CriterionLink` validators, with the offending id or mode in the message. Pass `None` to `TaskRecord` when no links exist so the field serializes as `null`, never as `[]`.
72
+ - **Integration Surface**: `generate_jsonl_from_md` calls this helper once per task; `validate_tasks_jsonl` (line 60) stays the row-level validator and needs no structural change because `TaskRecord.model_validate` now owns the new field.
73
+ - **`src/deviate/core/tasks_ledger.py:60`**: REFERENCE — `validate_tasks_jsonl` stays the row-level validator.
74
+ - **Current State**: The function runs `TaskRecord.model_validate` per row and formats `Record {i}: {loc}: {msg}` errors; existing tests cover valid rows, invalid ids, missing fields, invalid status, and extra fields.
75
+ - **Changes Required**: None structurally. The new field flows through the existing validator because `TaskRecord` owns it; new tests assert that rows with valid links pass, rows with a malformed `criterion_id` or an `automated`-with-null-`test_ref` link produce errors whose `loc` names the link, and legacy rows pass unchanged.
76
+ - **Integration Surface**: Called by meso task-validation flows; error strings feed task-ledger inspection.
77
+ - **`specs/005-acceptance-gates/002-task-acceptance-traceability/tasks.md`**: TARGET — per-task criterion references (input artifact).
78
+ - **Current State**: The file does not exist yet; the tasks phase authors it. The sibling slice `specs/005-acceptance-gates/001-verification-mode-metadata/tasks.md` shows the bullet convention (per-task `**Mode**:`, `**Files**:`, `**Verification**:` bullets); no `tasks.md` anywhere in `specs/` carries an `**Acceptance Criteria**:` bullet today, so this plan pins the syntax.
79
+ - **Changes Required**: Each task that implements or verifies an `AC-PLAN-NNN` scenario declares one `- **Acceptance Criteria**:` bullet with comma-separated entries of the form `AC-PLAN-NNN (mode[, test_ref])`, where `mode` is `automated`, `manual`, or `deferred`. An `automated` entry includes the test path as `test_ref` (for example `AC-PLAN-001 (automated, tests/test_core/test_tasks_ledger.py)`). A task with no criterion references omits the bullet.
80
+ - **Integration Surface**: Parsed by `generate_jsonl_from_md`; the ids name `AC-PLAN-NNN` scenarios in this plan's `## Acceptance Contract`.
81
+ - **`deviate meso tasks pre`** (`src/deviate/cli/meso.py:939-1037`): GATE — emits the traced rows through the generation path and blocks on invalid links.
82
+ - **Current State**: `_tasks_pre` (line 939) validates the plan contract via `validate_acceptance_contract` at line 1004 (the 005-001 behavior, now at `src/deviate/core/validation.py:162`) and blocks on `PLAN_ACCEPTANCE_CONTRACT_INVALID` / `PLAN_ACCEPTANCE_CONTRACT_MISSING`; it does not itself rewrite rows. `_tasks_post` (line 1040) commits `tasks.md`. The meso flow consumes the generator in `src/deviate/core/tasks_ledger.py`.
83
+ - **Changes Required**: None in this issue. The defensive exclusion in the issue scope (`005-002` §Defensive Exclusions) forbids gate-behavior work here; gate behavior belongs to issues `005-003` and `005-004`. Invalid links fail inside `CriterionLink` construction, so any consumer of `generate_jsonl_from_md` raises before a row is appended.
84
+ - **Integration Surface**: The gate validates the contract whose ids `CriterionLink` references; the integration tier of this issue drives a fixture `tasks.md` through `generate_jsonl_from_md` and `validate_tasks_jsonl`, the exact code path `specs/005-acceptance-gates/data-model.md:190` documents (Flow: Acceptance Traceability, step 3).
85
+ - **`tests/test_core/test_tasks_ledger.py`**: TARGET — extend generation tests with link propagation and rejection.
86
+ - **Current State**: `TestGenerateJsonlFromMd` covers basic parsing (ids, issue ids, descriptions, modes, empty files); `TestValidateTasksJsonl` covers row-level rejections including the `extra="forbid"` case.
87
+ - **Changes Required**: Add propagation cases (single link, multiple links, mixed modes, `manual` without `test_ref`, `automated` with `test_ref`, a task without a bullet), rejection cases (malformed `criterion_id`, `automated` link with null `test_ref`, a `verification_mode` outside the literals, an unparseable entry), and the null-never-empty-list assertion. Extend `TestValidateTasksJsonl` with a valid-links row, a malformed-link row, and a legacy row without the field.
88
+ - **Integration Surface**: The functions under test are the same generator and row validator the meso flow consumes.
89
+ - **`tests/test_state/test_ledger.py`**: TARGET — extend `TaskRecord` parse tests; legacy rows without the field still parse.
90
+ - **Current State**: `TestTaskRecord` (line 150) covers creation, explicit status/mode, invalid status/mode, extra-field rejection, id format, empty description, and serialization round-trip; `TestAppendTaskRecord` covers append behavior.
91
+ - **Changes Required**: Add a round-trip test with a populated `acceptance_criteria` list, a legacy-row parse test (dict without the field yields `None`), a `CriterionLink` accept/reject set, and an `extra="forbid"` regression that a genuinely unknown field still fails. Keep append-only semantics intact (no rewrite of existing rows).
92
+ - **Integration Surface**: The model under test is the same `TaskRecord` the micro phase runners validate at `src/deviate/cli/micro.py:1154` and later.
93
+
94
+ ## Implementation Strategy
95
+
96
+ - **Phase 1**: `CriterionLink` model and the additive `TaskRecord` field
97
+ - **Files**: `src/deviate/state/ledger.py`, `tests/test_state/test_ledger.py`
98
+ - **Approach**: Add `CriterionLink(BaseModel)` before `TaskRecord` with `criterion_id`, `verification_mode: Literal["automated", "manual", "deferred"]`, `test_ref: str | None = None`, and `model_config = {"extra": "forbid"}`. Add a `field_validator("criterion_id")` matching `^AC-PLAN-\d{3}$` and a `model_validator(mode="after")` rejecting an `automated` link with a null `test_ref`. Add `acceptance_criteria: list[CriterionLink] | None = None` to `TaskRecord` after `security_profile`, before `model_config`. RED writes the round-trip, legacy-parse, extra-forbid, and `CriterionLink` accept/reject tests first.
99
+ - **Verification**: `uv run pytest tests/test_state/test_ledger.py -v`
100
+ - **Phase 2**: Link parsing and propagation in the generator
101
+ - **Files**: `src/deviate/core/tasks_ledger.py`, `tests/test_core/test_tasks_ledger.py`
102
+ - **Approach**: Add `_CRITERIA_LINE_PATTERN` for the `**Acceptance Criteria**:` bullet and `_LINK_PATTERN` for `AC-PLAN-NNN (mode[, test_ref])` entries. Track `current_criteria` inside `generate_jsonl_from_md`, reset per task, and pass it to `_build_task_record`. In `_build_task_record`, parse entries into `CriterionLink` instances; raise a named `ValueError` for unparseable entries and let the `CriterionLink` validators raise `pydantic.ValidationError` for malformed ids, illegal modes, and automated links without `test_ref`. Pass `None` when no links exist. RED writes the propagation and rejection tests first.
103
+ - **Verification**: `uv run pytest tests/test_core/test_tasks_ledger.py -v`
104
+ - **Phase 3**: Row-validator coverage and full check bundle
105
+ - **Files**: `tests/test_core/test_tasks_ledger.py`, `tests/test_state/test_ledger.py`
106
+ - **Approach**: Extend `TestValidateTasksJsonl` with valid-link, malformed-link, and legacy-row cases so `validate_tasks_jsonl` behavior is pinned without any structural change to it. Confirm the change touches only `ledger.py`, `tasks_ledger.py`, and their tests; run the full bundle. No CLI command in the new tests hits `_run_pytest`, so no `subprocess.CompletedProcess` mock is needed and the suite stays under 30 seconds.
107
+ - **Verification**: `uv run pytest tests/test_core/test_tasks_ledger.py tests/test_state/test_ledger.py -v`, then `mise run check`
108
+
109
+ ## Data Flow Analysis
110
+
111
+ 1. **Input**: `tasks.md` content is read by `generate_jsonl_from_md` (`src/deviate/core/tasks_ledger.py:16` via `Path.read_text`).
112
+ 2. **Extraction**: The scanner splits the content into lines; `_TASK_LINE_PATTERN` recognizes `TSK-\d{3}-\d{2}` task bullets; per-task `**Mode**:` and `**Acceptance Criteria**:` bullets accumulate on the current task.
113
+ 3. **Transformation**: `_build_task_record` converts each `**Acceptance Criteria**:` entry into a `CriterionLink` (`criterion_id`, `verification_mode`, optional `test_ref`) and constructs the `TaskRecord`; invalid entries raise `ValueError` or `pydantic.ValidationError` before any row is produced.
114
+ 4. **Output**: `TaskRecord.model_dump_json` serializes rows that carry `acceptance_criteria` when links exist and `"acceptance_criteria": null` when they do not; `append_task_record` (`src/deviate/state/ledger.py:181`) appends rows to `specs/**/tasks.jsonl` per the append-only protocol (`specs/constitution.md` §1).
115
+ 5. **Storage and consumption**: Micro-phase runners parse rows with `TaskRecord.model_validate` (`src/deviate/cli/micro.py:1154,1259,1336,2524,3113,4437,4533,4806`); rows without the field default to `None`, so mixed-version ledgers parse fully. `_read_ledger` (`src/deviate/state/ledger.py:41`) guards malformed JSONL lines with a warning instead of a crash.
116
+ 6. **Upstream constraint**: The `AC-PLAN-NNN` ids in `tasks.md` name scenarios in the owning slice's validated plan contract (`src/deviate/core/validation.py:162` `validate_acceptance_contract`, mode-enforced by `_validate_verification_mode` at `validation.py:139`); `deviate meso tasks pre` validates that contract at `src/deviate/cli/meso.py:1004` before generation, and this issue consumes the validated contract without re-checking (defensive exclusion).
117
+
118
+ ## Risk Assessment
119
+
120
+ | Risk | Impact | Likelihood | Mitigation |
121
+ |------|--------|------------|------------|
122
+ | `extra="forbid"` rejects older JSONL rows that lack the new field | High | Medium | Declare the field optional with default `None`; pin a legacy-row parse test; the design register mirrors this as `RSK-001` (`specs/005-acceptance-gates/design.md:63`). |
123
+ | Ambiguous or drifted `tasks.md` criterion syntax silently drops links | Medium | Medium | Pin the `**Acceptance Criteria**: AC-PLAN-NNN (mode[, test_ref])` syntax in this plan; the generator raises on unparseable entries instead of skipping them. |
124
+ | A `criterion_id` that matches `AC-PLAN-\d{3}` but names a criterion absent from the owning plan's contract escapes generation | Medium | Low | Documented boundary: `CriterionLink` validates the id pattern only; contract-level verification belongs to issue `005-001` and the micro/gate issues `005-003`/`005-004` (issue scope §Defensive Exclusions). No plan.md cross-check is added to the generator. |
125
+ | Generation-time error messages omit the offending link details | Low | Medium | Pydantic field/model validators pin named messages that include `criterion_id` and mode; tests assert the message substrings. |
126
+ | FLOW_CONTEXT_UNAVAILABLE — no existing flow mapping is available | Medium | Low | Preserve empty flow references and plan the application's requested behavior without creating flow or DeviaTDD setup work. |
127
+
128
+ ## Security Profile
129
+
130
+ Risk surfaces: file paths (reads `tasks.md` via `Path.read_text` — local read only), deserialization (`json.loads` on ledger rows in `_read_ledger`, already guarded by a try/except with a warning), and string fields (`criterion_id` and `test_ref` are plain strings; `test_ref` is never resolved or executed — existence checks are explicitly out of scope per the issue's Edge Cases).
131
+
132
+ Negative tests: a row with a genuinely unknown field still fails under `extra="forbid"`; a legacy row without the field parses with `acceptance_criteria` equal to `None`; an `automated` link without a `test_ref` fails generation; a `verification_mode` outside the three literals fails generation; an unparseable `**Acceptance Criteria**:` entry fails generation with a named error; a malformed JSONL line is skipped with a warning, never a crash.
133
+
134
+ Constraints: no new dependencies (stdlib `re` only); no hardcoded secrets; no changes to `src/deviate/core/validation.py`, `src/deviate/cli/micro.py`, `src/deviate/prompts/`, `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md`, or `CHANGELOG.md` (defensive exclusions); no subprocess in the new tests; the full suite stays under 30 seconds.
135
+
136
+ ## Integration Points
137
+
138
+ - **`src/deviate/cli/meso.py:1004`** (`_tasks_pre` contract gate): validates the plan `## Acceptance Contract` via `validate_acceptance_contract` (`src/deviate/core/validation.py:162`) before task generation; the `AC-PLAN-NNN` ids this issue propagates refer to that contract. No change here (gate behavior belongs to issues `005-003`/`005-004`).
139
+ - **`src/deviate/cli/micro.py` phase runners** (`TaskRecord.model_validate` at lines 1154, 1259, 1336, 2524, 3113, 4437, 4533, 4806): parse new and legacy task rows; the additive field must not break any runner. `src/deviate/prompts/auto/red.md:64` already instructs RED to trace `{TASK_ID}` to its `AC-PLAN-NNN` references in `tasks.md`.
140
+ - **`src/deviate/cli/inspect.py:244`**: reads `tasks.jsonl` for task inspection; rows carrying the new field render without a schema change.
141
+ - **`src/deviate/core/validation.py:166`** (`contract_pattern`, `AC-PLAN-\d{3}`) and **`:162`** (`validate_acceptance_contract`): provide and enforce the id vocabulary `CriterionLink` references; this issue consumes the validated contract and does not modify it (defensive exclusion).
142
+ - **`specs/005-acceptance-gates/data-model.md`** (`TaskRecord` and `CriterionLink` entity tables, lines 84-110): the authoritative schema this issue implements; `acceptance_criteria` is additive with default `None` and `CriterionLink` embeds per the relationship table (line 77).
143
+ - **`src/deviate/prompts/commands/deviate-tasks.md:47`** and **`src/deviate/prompts/auto/tasks.md:70`**: tasks-phase prompts require every task to cite its `AC-PLAN-NNN` scenarios; the new bullet syntax aligns with that requirement (prompt-template alignment itself belongs to issue `005-005`).
144
+
145
+ ## Constitutional Alignment
146
+
147
+ - **Architecture**: The change sits in the Meso layer (task generation carries criteria traceability into `specs/**/tasks.jsonl` rows) and feeds the Micro layer (phase runners parse the additive field). It implements the `acceptance_criteria` field and `CriterionLink` schema from `specs/005-acceptance-gates/data-model.md` and satisfies the append-only ledger protocol of `specs/constitution.md` §1 — the field is additive, and no existing JSONL line is ever modified.
148
+ - **Testing**: pytest unit tests in `tests/test_core/test_tasks_ledger.py` and `tests/test_state/test_ledger.py`; coverage target stays at or above 80%; the new tests are pure unit tests that make no CLI and no git calls, so the full suite stays under the 30-second contract (`AGENTS.md` test-performance pointer). `mise run check` (lint, format-check, types, full suite) is the final gate.
149
+ - **Git Isolation**: The orchestrator commits at phase boundaries; this plan's tests use only `tmp_path` fixtures and make zero branch-mutating or `git` calls, so the git-isolation invariants and the append-only ledger protocol remain intact.
150
+ - **Product Layer**: `flow_refs` is `[]` per the issue frontmatter, mirrored verbatim in `## Product Layer Anchors`; the release Goal (FLOW-04, subprocess RPC/TUI) is orthogonal and untouched. This plan adds no Product-layer, DeviaTDD setup, skill, flow-authoring, or workflow-ledger work; it only hardens the requested application behavior (task-to-criterion and criterion-to-test traceability) that the meso generation path delivers.
@@ -0,0 +1,9 @@
1
+ {"id":"TSK-005-01","issue_id":"005-002","description":"`CriterionLink` model plus the additive `TaskRecord.acceptance_criteria` field, with model-level accept/reject tests","status":"RED","execution_mode":"TDD","created_at":"2026-08-16T09:30:50.760884Z","security_profile":null}
2
+ {"id":"TSK-005-01","issue_id":"005-002","description":"`CriterionLink` model plus the additive `TaskRecord.acceptance_criteria` field, with model-level accept/reject tests","status":"GREEN","execution_mode":"TDD","created_at":"2026-08-16T09:41:57.012207Z","security_profile":null}
3
+ {"id":"TSK-005-01","issue_id":"005-002","description":"`CriterionLink` model plus the additive `TaskRecord.acceptance_criteria` field, with model-level accept/reject tests","status":"JUDGE","execution_mode":"TDD","created_at":"2026-08-16T09:44:26.864774Z","security_profile":null}
4
+ {"id":"TSK-005-01","issue_id":"005-002","description":"`CriterionLink` model plus the additive `TaskRecord.acceptance_criteria` field, with model-level accept/reject tests","status":"COMPLETED","execution_mode":"TDD","created_at":"2026-08-16T09:44:26.864774Z","security_profile":null}
5
+ {"id":"TSK-005-02","issue_id":"005-002","description":"Criterion-link parsing and propagation inside `generate_jsonl_from_md` and `_build_task_record`","status":"RED","execution_mode":"TDD","created_at":"2026-08-16T10:45:30.117531Z","security_profile":null}
6
+ {"id":"TSK-005-02","issue_id":"005-002","description":"Criterion-link parsing and propagation inside `generate_jsonl_from_md` and `_build_task_record`","status":"GREEN","execution_mode":"TDD","created_at":"2026-08-16T10:56:18.114404Z","security_profile":null}
7
+ {"id":"TSK-005-02","issue_id":"005-002","description":"Criterion-link parsing and propagation inside `generate_jsonl_from_md` and `_build_task_record`","status":"JUDGE","execution_mode":"TDD","created_at":"2026-08-16T11:01:16.110814Z","security_profile":null}
8
+ {"id":"TSK-005-02","issue_id":"005-002","description":"Criterion-link parsing and propagation inside `generate_jsonl_from_md` and `_build_task_record`","status":"COMPLETED","execution_mode":"TDD","created_at":"2026-08-16T11:13:02.162546Z","security_profile":null}
9
+ {"id":"TSK-005-03","issue_id":"005-002","description":"Row-validator accept/reject cases for link rows plus the mixed-version append regression","status":"COMPLETED","execution_mode":"IMMEDIATE","created_at":"2026-08-16T13:06:59.617574Z","security_profile":null}
@@ -0,0 +1,113 @@
1
+ # Implementation Tasks: feat/005-acceptance-gates/002-task-acceptance-traceability
2
+
3
+ ## Phase 1: CriterionLink Model and Additive TaskRecord Field
4
+ **Goal**: `CriterionLink` exists as a pydantic model with the `AC-PLAN-\d{3}` id format, the three verification-mode literals, and the automated-link `test_ref` invariant. `TaskRecord` gains the optional `acceptance_criteria: list[CriterionLink] | None = None` field while preserving `model_config = {"extra": "forbid"}` and the seven-value `status` Literal.
5
+
6
+ ### Tasks
7
+
8
+ - TSK-005-01: `CriterionLink` model plus the additive `TaskRecord.acceptance_criteria` field, with model-level accept/reject tests
9
+ - **Type**: Feature_Batch
10
+ - **Mode**: TDD
11
+ - **Test Strategy**: Sociable_Unit
12
+ - **Verification**: `uv run pytest tests/test_state/test_ledger.py -v`
13
+ - **Estimated Time**: 60 minutes
14
+ - **Flow References**: []
15
+ - **Acceptance Criteria**: AC-PLAN-002 (automated, tests/test_state/test_ledger.py), AC-PLAN-003 (automated, tests/test_state/test_ledger.py), AC-PLAN-004 (automated, tests/test_state/test_ledger.py)
16
+ - **Files**:
17
+ - `src/deviate/state/ledger.py`
18
+ - `tests/test_state/test_ledger.py`
19
+ - **Rationale**: US-005-04 requires legacy task rows without the new field to parse under `extra="forbid"`; US-005-03 requires each emitted row to carry valid links. This task owns the model layer: `CriterionLink` (new) and `TaskRecord.acceptance_criteria` (additive) in `src/deviate/state/ledger.py`, the single schema that `generate_jsonl_from_md` and the micro-phase runners (`src/deviate/cli/micro.py` `TaskRecord.model_validate` sites) share. `tests/test_state/test_ledger.py` pins AC-PLAN-003 (legacy row parses with `acceptance_criteria` equal to `None`), AC-PLAN-004 (serialization emits `null`, never `[]`), and the AC-PLAN-002 rejection half that lives at model construction (malformed `criterion_id`, `automated` without `test_ref`).
20
+ - **Details**:
21
+ - **Red**: Add `TestCriterionLink` to `tests/test_state/test_ledger.py` covering: accept a `manual` link with `test_ref` absent; accept a `deferred` link with `test_ref` absent; accept `manual` or `deferred` with `test_ref` present; reject `automated` without `test_ref` asserting the message names the missing `test_ref`; reject `criterion_id` `AC-PLAN-99` asserting the message names the id; reject `verification_mode` `soon`; reject an unknown field on the link under `extra="forbid"`. Extend `TestTaskRecord` with: the field defaults to `None` when omitted; a populated `acceptance_criteria` list round-trips through `model_dump_json` and `model_validate`; a legacy dict carrying `id`, `issue_id`, `description`, `status`, `execution_mode`, `created_at` and no `acceptance_criteria` parses with the field equal to `None`; `json.loads(record.model_dump_json())["acceptance_criteria"] is None` (never `[]`) when no links exist; a genuinely unknown field such as `unknown_field` still raises under `extra="forbid"`.
22
+ - **Green**: In `src/deviate/state/ledger.py`, define `CriterionLink(BaseModel)` before `TaskRecord` and after `SecurityProfile`, matching the ledger-family pattern, with `criterion_id: str`, `verification_mode: Literal["automated", "manual", "deferred"]`, `test_ref: str | None = None`, and `model_config = {"extra": "forbid"}`. Add a `field_validator("criterion_id")` that raises `ValueError(f"Invalid criterion ID format: {v}")` unless `re.match(r"^AC-PLAN-\d{3}$", v)`; `re` is already imported at ledger.py line 4. Add a `model_validator(mode="after")` that raises when `verification_mode == "automated"` and `test_ref is None`. Add `acceptance_criteria: list[CriterionLink] | None = None` to `TaskRecord` after the `security_profile` field and before `model_config`; preserve the seven-value `status` Literal and `model_config = {"extra": "forbid"}` unchanged (PRD `RESOLVED-Q-004`).
23
+ - **Refactor**: Follow the `SecurityProfile` precedent (lines 61-78) for the validator and config pattern. Keep the validator messages named so tests can assert message substrings. Preserve deterministic field order.
24
+ - **Edge Cases**: An `automated` link with `test_ref=""` fails — treat the empty string as missing. A `manual` or `deferred` link with `test_ref` present passes. Extra fields on the link or on the record fail under `extra="forbid"`.
25
+ - **Acceptance**: `uv run pytest tests/test_state/test_ledger.py -v` passes. Rows generated by later phases serialize `"acceptance_criteria": null` when a task declares no links, and a populated list otherwise.
26
+
27
+ ## Phase 2: Link Parsing and Propagation in the Generator
28
+ **Goal**: `generate_jsonl_from_md` reads per-task `- **Acceptance Criteria**:` bullets from `tasks.md`, parses each `AC-PLAN-NNN (mode[, test_ref])` entry, and emits every generated `TaskRecord` with its `acceptance_criteria` links. Malformed entries, illegal modes, and `automated` links without `test_ref` fail generation with named errors.
29
+
30
+ ### Tasks
31
+
32
+ - TSK-005-02: Criterion-link parsing and propagation inside `generate_jsonl_from_md` and `_build_task_record`
33
+ - **Type**: Feature_Batch
34
+ - **Mode**: TDD
35
+ - **Test Strategy**: Sociable_Unit
36
+ - **Verification**: `uv run pytest tests/test_core/test_tasks_ledger.py -v`
37
+ - **Estimated Time**: 90 minutes
38
+ - **Flow References**: []
39
+ - **Acceptance Criteria**: AC-PLAN-001 (automated, tests/test_core/test_tasks_ledger.py), AC-PLAN-002 (automated, tests/test_core/test_tasks_ledger.py), AC-PLAN-004 (automated, tests/test_core/test_tasks_ledger.py)
40
+ - **Dependency**: TSK-005-01
41
+ - **Files**:
42
+ - `src/deviate/core/tasks_ledger.py`
43
+ - `tests/test_core/test_tasks_ledger.py`
44
+ - **Rationale**: US-005-03 requires the generator to emit each task row with its `acceptance_criteria` links so the task points at the criteria it satisfies and the tests that verify them. AC-PLAN-001 pins the propagation (its Given names task `TSK-005-01` carrying two links, which is this slice's own first task), AC-PLAN-002 pins generation-time rejection of malformed `criterion_id` values and of `automated` links with a missing `test_ref`, and AC-PLAN-004 pins the null-never-empty-list behavior at the generation boundary. `src/deviate/core/tasks_ledger.py` is the generator that `deviate meso tasks pre` and `validate_tasks_jsonl` consume; `tests/test_core/test_tasks_ledger.py` is its unit-level contract.
45
+ - **Details**:
46
+ - **Red**: In `tests/test_core/test_tasks_ledger.py`, extend `TestGenerateJsonlFromMd` with `tmp_path` fixtures. Propagation: a task line followed by `- **Acceptance Criteria**: AC-PLAN-001 (automated, tests/test_core/test_tasks_ledger.py), AC-PLAN-002 (manual)` yields a record whose `acceptance_criteria` list holds two `CriterionLink` entries matching the declared `criterion_id`, `verification_mode`, and `test_ref` values, with `test_ref` equal to `None` for the `manual` entry. Null-never-empty: a sibling task with no bullet carries `acceptance_criteria is None`. Rejection: `AC-PLAN-99 (automated, tests/test_core/test_tasks_ledger.py)` raises with the invalid id named; `AC-PLAN-001 (automated)` without `test_ref` raises with the missing `test_ref` named; `AC-PLAN-001 (soon)` raises; an unparseable entry such as `AC-PLAN-001 automated` raises a `ValueError` that names the offending text and the task id. Assert the propagated ids reference `AC-PLAN-NNN` scenarios from this plan's `## Acceptance Contract`.
47
+ - **Green**: In `src/deviate/core/tasks_ledger.py`, add `_CRITERIA_LINE_PATTERN` matching `- **Acceptance Criteria**: <entries>` within a task block and `_LINK_PATTERN` matching `AC-PLAN-\d{3} \((automated|manual|deferred)(?:,\s*(\S+))?\)`. Track `current_criteria: list[str]` in `generate_jsonl_from_md`, reset it to `[]` on each new `_TASK_LINE_PATTERN` match, accumulate the entries text on a criteria-line match, and pass the collected entries to `_build_task_record` at both append sites (line 27 and line 39). In `_build_task_record`, accept the criteria entries (default empty), split them on top-level commas — commas inside the parenthesized entry must not split — and construct `CriterionLink` instances per entry; a fragment that fails `_LINK_PATTERN` raises `ValueError` naming the fragment and the task id; malformed ids, illegal modes, and `automated` links with a null `test_ref` raise `pydantic.ValidationError` from the `CriterionLink` validators. Pass `None` to `TaskRecord` when no links exist so the field serializes as `null`, never as `[]`. Add no new dependencies (stdlib `re` only).
48
+ - **Refactor**: Keep the scan linear over the line list exactly as today. Reuse the existing `current_id` / `current_desc` / `current_mode` accumulation pattern for `current_criteria`.
49
+ - **Edge Cases**: A criteria bullet that appears before the task's `**Mode**:` bullet still attaches to the current task. An `automated` entry with an empty `test_ref` (`AC-PLAN-001 (automated, )`) fails. `manual` and `deferred` entries with `test_ref` present pass. Whitespace around ids and modes is tolerated by the pattern. A `criterion_id` that matches `AC-PLAN-\d{3}` but names a criterion absent from the owning plan's contract is a documented boundary — no plan.md cross-check is added here (defensive exclusion).
50
+ - **Acceptance**: `uv run pytest tests/test_core/test_tasks_ledger.py -v` passes. Generation error messages name the invalid `criterion_id` or the missing `test_ref`. A row with links serializes `"acceptance_criteria": [{"criterion_id": ..., "verification_mode": ..., "test_ref": ...}]`.
51
+
52
+ ## Phase 3: Row-Validator Coverage and Full Check Bundle
53
+ **Goal**: `validate_tasks_jsonl` behavior with link rows is pinned with no structural change to it, mixed-version ledgers append and parse intact, and the full check bundle exits green.
54
+
55
+ ### Tasks
56
+
57
+ - TSK-005-03: Row-validator accept/reject cases for link rows plus the mixed-version append regression
58
+ - **Type**: Verification_Batch
59
+ - **Mode**: IMMEDIATE
60
+ - **Test Strategy**: Sociable_Unit
61
+ - **Verification**: `uv run pytest tests/test_core/test_tasks_ledger.py tests/test_state/test_ledger.py -v` then `mise run check`
62
+ - **Estimated Time**: 60 minutes
63
+ - **Flow References**: []
64
+ - **Acceptance Criteria**: AC-PLAN-001 (automated, tests/test_core/test_tasks_ledger.py), AC-PLAN-002 (automated, tests/test_core/test_tasks_ledger.py), AC-PLAN-003 (automated, tests/test_state/test_ledger.py)
65
+ - **Dependency**: TSK-005-02
66
+ - **Files**:
67
+ - `tests/test_core/test_tasks_ledger.py`
68
+ - `tests/test_state/test_ledger.py`
69
+ - **Rationale**: US-005-03 and US-005-04 both terminate in the row-level validator and ledger consumption. `validate_tasks_jsonl` (`src/deviate/core/tasks_ledger.py:60`) stays structurally unchanged — `TaskRecord.model_validate` owns the new field — but its behavior with link rows must be pinned so mixed-version ledgers never break the micro-phase runners. AC-PLAN-001 (a row with valid links passes), AC-PLAN-002 (a row whose link is malformed produces an error whose `loc` names the link), AC-PLAN-003 (a legacy row without the field passes). `tests/test_state/test_ledger.py` receives the mixed-version append case proving append-only semantics with the new field.
70
+ - **Details**:
71
+ - **Red**: Extend `TestValidateTasksJsonl` in `tests/test_core/test_tasks_ledger.py`: a row with a valid `acceptance_criteria` list passes with `errors == []`; a row whose link carries `criterion_id: "AC-PLAN-99"` returns an error whose `loc` names `acceptance_criteria`; a row whose `automated` link lacks `test_ref` returns an error naming `acceptance_criteria`; a legacy row without the field passes; a row that carries `acceptance_criteria` plus a genuinely unknown field still fails. In `tests/test_state/test_ledger.py`, extend `TestAppendTaskRecord`: append a legacy task row, then append a task record carrying links, then read the file back — both rows are present and unchanged, each line is appended (append-only protocol, constitution §1), and each parsed row carries the field respectively as `None` and as the populated list.
72
+ - **Green**: No production change. `CriterionLink` (TSK-005-01) and the generator (TSK-005-02) already produce the rows these tests exercise; `validate_tasks_jsonl` and `append_task_record` pass them through. Run the new tests against the committed implementation and confirm `git status` shows no modification to `src/deviate/cli/meso.py` or `src/deviate/core/validation.py`.
73
+ - **Refactor**: Reuse the dict-based record style of the existing `TestValidateTasksJsonl` cases. Keep fixture records minimal. Assert on `loc` substrings rather than full message text.
74
+ - **Edge Cases**: A malformed JSONL line in the append-read path is skipped with a warning, never a crash (existing `_read_ledger` guard). Mixed-version rows parse fully with the field present or absent. The append stays idempotent on the `(id, status)` compound key.
75
+ - **Acceptance**: `uv run pytest tests/test_core/test_tasks_ledger.py tests/test_state/test_ledger.py -v` passes, then `mise run check` (lint, format-check, types, full suite) exits 0. The complete diff touches only `src/deviate/state/ledger.py`, `src/deviate/core/tasks_ledger.py`, `tests/test_core/test_tasks_ledger.py`, and `tests/test_state/test_ledger.py`.
76
+
77
+ ---
78
+
79
+ ## Implementation Strategy
80
+ **Execution Order**:
81
+ 1. Phase 1 -> Phase 2 -> Phase 3 (logical dependency order)
82
+
83
+ **Critical Dependency Chains**:
84
+ - TSK-005-01 must precede TSK-005-02 (`_build_task_record` constructs `CriterionLink` instances and sets `TaskRecord.acceptance_criteria`, both defined by TSK-005-01 in `src/deviate/state/ledger.py`)
85
+ - TSK-005-02 must precede TSK-005-03 (TSK-005-03 pins row-validator behavior that exists only after the model field and the generator propagation land)
86
+
87
+ **Risk Hotspots**:
88
+ - `extra="forbid"` rejects older JSONL rows that lack the new field. The field is optional with default `None`, and a legacy-row parse test pins the behavior; the design register mirrors this as `RSK-001` (`specs/005-acceptance-gates/design.md:63`).
89
+ - Ambiguous or drifted `tasks.md` criterion syntax silently drops links. The `**Acceptance Criteria**: AC-PLAN-NNN (mode[, test_ref])` syntax is pinned in this plan, and the generator raises on unparseable entries instead of skipping them.
90
+ - `_LINK_PATTERN` entry splitting must respect parentheses: the `test_ref` comma sits inside `(...)`, so splitting the bullet on bare commas would break entries.
91
+ - A `criterion_id` that matches `AC-PLAN-\d{3}` but names a criterion absent from the owning plan's contract escapes generation by design; contract-level verification belongs to issues `005-001`, `005-003`, and `005-004` (issue scope Defensive Exclusions).
92
+
93
+ **Merge Conflict Boundaries**:
94
+ - `tests/test_core/test_tasks_ledger.py` is touched by TSK-005-02 (extends `TestGenerateJsonlFromMd`) and TSK-005-03 (extends `TestValidateTasksJsonl`). The edits target different classes; the tasks run sequentially.
95
+ - `tests/test_state/test_ledger.py` is touched by TSK-005-01 (adds `TestCriterionLink`, extends `TestTaskRecord`) and TSK-005-03 (extends `TestAppendTaskRecord`). The edits do not overlap.
96
+ - `src/deviate/state/ledger.py` is touched only by TSK-005-01.
97
+
98
+ **Product-Layer Anchors** (mirrored from plan.md):
99
+ - **Flow References**: `[]`
100
+ - **Source**: `specs/005-acceptance-gates/002-task-acceptance-traceability/plan.md`
101
+ - Downstream micro phases inherit this list per-task. Empty references mean no matching existing flow, not permission for enabling, setup, tooling, skill, release, or workflow-ledger tasks.
102
+
103
+ ---
104
+
105
+ ## Universal Test Constraints (ALL TASKS)
106
+
107
+ - **Git Isolation Mandatory**: Any test that invokes git operations MUST operate on a temporary directory initialized as a fresh git repo. Tests MUST NOT run git commands within the real repository's working tree.
108
+ - **Implementation Pattern**: Use a shared `tmp_git_repo` fixture from `tests/conftest.py`. Pass `repo=tmp_git_repo` to all git-interacting functions. Never reference `Path.cwd()` or the real repo root.
109
+ - **Rationale**: Prevent accidental commits, branch creation, or state mutation in the actual project repo during test execution.
110
+
111
+ ## Universal API Design Constraint (ALL CORE MODULES)
112
+
113
+ Every git-interacting function in core modules MUST accept an optional `repo_path: Path | None = None` parameter. When `None`, default to `Path.cwd()`.
@@ -425,13 +425,13 @@ accepts `--json` (emit JSON contract to stdout) and `--quiet` (suppress output).
425
425
  5. **Given / When / Then** — exactly three bold-labelled clauses in this order: `**Given**:`, `**When**:`, `**Then**:`. Each clause is a single imperative sentence and MUST NOT embed additional `**Given**` / `**When**` / `**Then**` markers. The `**Then**` clause MUST state a verifiable observable outcome.
426
426
  * **Required sections in canonical order**: `## Plan Summary` → `## Product Layer Anchors` → `## Acceptance Contract` → `## Workstation Mapping` → `## Implementation Strategy` → `## Data Flow Analysis` → `## Risk Assessment` → `## Security Profile` → `## Integration Points` → `## Constitutional Alignment`.
427
427
  * **Acceptance Coverage Invariant:** Every AO from the issue's `## Acceptance Outline` MUST appear as the Source Outline of at least one AC-PLAN scenario. Behavioural coverage that does not map cleanly to a single AO (e.g. an HMAC failure, an RLS isolation invariant, a defensive boundary) belongs under an existing AO's Error Category or Boundary Category. If no existing AO fits, the issue's outline is incomplete — halt with `INCOMPLETE_ISSUE_OUTLINE` and request that shard/adhoc regenerate the issue.
428
- * **Forbidden patterns** (any one triggers `PLAN_ACCEPTANCE_CONTRACT_INVALID` from `deviate plan post`): Source Outline labelled `Edge Cases`, `Boundary`, `Constitutional §…`, `RLS`, `Tenant Isolation`, `Hardening`, `Security`, or any non-AO string; missing `**Source Outline**` / `**Upstream Traceability**` / `**Current-Code Evidence**` / any of `**Given**` / `**When**` / `**Then**`; missing, repeated, or illegal `**Verification Mode**: <automated|manual|deferred>` (a scenario MUST carry exactly one legal mode line; an empty or non-alphabetic value is treated as missing); an issue AO not used by any AC-PLAN scenario; duplicate or non-sequential `AC-PLAN-NNN` identifiers; wrapping the plan body in any XML tag / code fence / preamble. The validator lives at `src/deviate/core/validation.py::validate_acceptance_contract`.
428
+ * **Forbidden patterns** (any one triggers `PLAN_ACCEPTANCE_CONTRACT_INVALID` from `deviate plan post`): Source Outline labelled `Edge Cases`, `Boundary`, `Constitutional §…`, `RLS`, `Tenant Isolation`, `Hardening`, `Security`, or any non-AO string; missing `**Source Outline**` / `**Upstream Traceability**` / `**Current-Code Evidence**` / any of `**Given**` / `**When**` / `**Then**`; a repeated or illegal `**Verification Mode**: <automated|manual|deferred>` literal (a scenario MUST carry exactly one legal mode line; an empty or non-alphabetic value is treated as missing); an issue AO not used by any AC-PLAN scenario; duplicate or non-sequential `AC-PLAN-NNN` identifiers; wrapping the plan body in any XML tag / code fence / preamble. A *missing* mode line is not a stopping error: the meso gates auto-fill the default `automated` value into the scenario body (see `deviate plan post`). The validator lives at `src/deviate/core/validation.py::validate_acceptance_contract`; the repair helper is `repair_missing_verification_mode`.
429
429
  * **Input Parameters:** `--issue`, `--force`, `--dry-run`; common `--json` / `--quiet` wrappers apply.
430
430
  * **Session:** force-transitions to PLAN with `active_issue_id` set.
431
431
 
432
432
  #### `deviate plan post [--force] [--issue-id]`
433
433
 
434
- Validates plan.md exists, is non-empty, and contains a valid Acceptance Contract; auto-renders HTML when changed, commits with convention-aware messaging, and transitions to TASKS. Missing/malformed contracts fail as `PLAN_ACCEPTANCE_CONTRACT_MISSING` / invalid contract diagnostics.
434
+ Validates plan.md exists, is non-empty, and contains a valid Acceptance Contract; auto-renders HTML when changed, commits with convention-aware messaging, and transitions to TASKS. Missing/malformed contracts fail as `PLAN_ACCEPTANCE_CONTRACT_MISSING` / invalid contract diagnostics. When the contract fails *only* because scenarios lack the `**Verification Mode**:` line, the gate auto-fills `automated` into each affected scenario body, persists the repaired `plan.md` (`PLAN_MODE_REPAIR` banner), and proceeds; an existing invalid or duplicated mode literal still blocks with `PLAN_ACCEPTANCE_CONTRACT_INVALID`.
435
435
 
436
436
  #### `deviate tasks pre [--force] [--dry-run]`
437
437
 
@@ -439,7 +439,7 @@ Validates plan.md exists, is non-empty, and contains a valid Acceptance Contract
439
439
  * **Two-source input:** `spec_path` supplies macro intent; `plan_path` supplies strategy and authoritative scenarios. Plan wins over legacy issue/spec Gherkin.
440
440
  * **Contract:** detects worktree/branch, resolves constitution commands, and emits `spec_path`, `plan_path`, `tasks_target`, worktree metadata, status, and flags. The issue is resolved from `session.active_issue_id`, falling back to a branch-derived lookup via the `feat/{epic}/{issue}` regex against `specs/issues.jsonl`.
441
441
  * **Plan digest:** TASKS receives a bounded 16 KiB UTF-8 `plan_digest` plus `plan_path`; truncation inserts `PLAN_DIGEST_TRUNCATED`, requiring a full read.
442
- * **Validation:** reports PLAN_NOT_FOUND, PLAN_ACCEPTANCE_CONTRACT_MISSING, or PLAN_ACCEPTANCE_CONTRACT_INVALID; no Gherkin fallback.
442
+ * **Validation:** reports PLAN_NOT_FOUND, PLAN_ACCEPTANCE_CONTRACT_MISSING, or PLAN_ACCEPTANCE_CONTRACT_INVALID; no Gherkin fallback. A contract that fails only for a missing `**Verification Mode**:` line is auto-repaired in place (default `automated`) before the status is computed.
443
443
  * **Common Flags:** `--json`, `--quiet`.
444
444
 
445
445
  #### `deviate tasks post [--force] [--issue-id]`
@@ -538,17 +538,21 @@ accepts `--json` and `--quiet`. `pre` emits a JSON contract describing the envir
538
538
  #### `deviate red post`
539
539
 
540
540
  * **Source:** `src/deviate/cli/micro.py`
541
- * **Description:** Runs `pytest -v` on all test files. Validates the test fails explicitly
542
- (ASSERTION_FAILURE, not PASS or SYNTAX_ERROR), runs the test command, and reports whether the
543
- test failed as expected. `deviate micro run`'s internal RED phase (`_run_red_phase`) applies the
544
- same contract: when the test command exits 0 (all tests passed) or collects no tests (pytest
545
- exit 5), it does NOT die — it routes the decision to JUDGE (``failure_kind: no_failing_test``),
546
- which either rules the behavior already exists (task COMPLETED, the uncommitted passing test
547
- discarded) or rules the test wrong (``revert_before``, RED re-authors a genuinely failing test).
548
- When RED produces no test files at all, `deviate micro run` raises "RED phase produced no test
549
- files" naming `/deviate-execute` (DIRECT) or `/deviate-meso` re-sharding. On a genuine failing
550
- test it appends the RED status transition to the task ledger, forces session to RED, and commits
551
- with `test({scope}): RED phase - failing test`.
541
+ * **Description:** Runs the project's resolved test command (language-agnostic: `mix test`,
542
+ `cargo test`, `npm test`, `go test ./...`, or `pytest` chosen via the `_test_command_candidates`
543
+ resolution order — task `verification`, constitution `test_command`, `mise run test`, manifest
544
+ table, Python fallback). Validates the test fails explicitly (ASSERTION_FAILURE, not PASS or
545
+ SYNTAX_ERROR), runs the test command, and reports whether the test failed as expected.
546
+ `deviate micro run`'s internal RED phase (`_run_red_phase`) applies the same contract: when the
547
+ test command exits 0 (all tests passed), collects no tests (pytest exit 5), or resolves to no
548
+ command at all (returncode 127), it does NOT die — it routes the decision to JUDGE
549
+ (``failure_kind: no_failing_test``), which either rules the behavior already exists (task
550
+ COMPLETED, the uncommitted passing test discarded) or rules the test wrong (``revert_before``,
551
+ RED re-authors a genuinely failing test). The RED gate does not require a Python
552
+ ``tests/**/test_*.py`` glob — test discovery follows the project's own convention (e.g.
553
+ ``test/**/*_test.exs`` for Elixir). On a genuine failing test it appends the RED status
554
+ transition to the task ledger, forces session to RED, and commits with
555
+ `test({scope}): RED phase - failing test`.
552
556
  Commit messages are convention-aware: when the project declares an emoji convention in
553
557
  ``CONTRIBUTING.md`` / ``.commit-convention.md``, the appropriate gitmoji is prepended
554
558
  automatically. RED phase `test:` commits are prefixed with 🚨 to flag the failing test (see
@@ -564,8 +568,9 @@ accepts `--json` and `--quiet`. `pre` emits a JSON contract describing the envir
564
568
  #### `deviate green post`
565
569
 
566
570
  * **Source:** `src/deviate/cli/micro.py`
567
- * **Description:** Verifies a RED transition exists for the active issue. Runs `pytest -v`,
568
- requires returncode 0. Appends GREEN transition to ledger, forces session to GREEN,
571
+ * **Description:** Verifies a RED transition exists for the active issue. Runs the project's
572
+ resolved test command (language-agnostic, e.g. `mix test` / `cargo test` / `pytest`), requires
573
+ returncode 0. Appends GREEN transition to ledger, forces session to GREEN,
569
574
  commits with `feat({scope}): GREEN phase - implementation passes tests`.
570
575
 
571
576
  #### `deviate judge pre`
@@ -584,9 +589,9 @@ accepts `--json` and `--quiet`. `pre` emits a JSON contract describing the envir
584
589
 
585
590
  * **Source:** `src/deviate/cli/micro.py`
586
591
  * **Description:** Verifies a GREEN transition exists. Appends REFACTOR transition, runs
587
- AST-based return type mismatch check, runs pytest before/after to detect regression.
588
- On regression, restores via `git restore .` and halts. Commits with
589
- `refactor({scope}): REFACTOR phase - code cleanup`.
592
+ the AST-based return type mismatch check (Python only), runs the resolved test command
593
+ before/after to detect regression. On regression, restores via `git restore .` and halts.
594
+ Commits with `refactor({scope}): REFACTOR phase - code cleanup`.
590
595
 
591
596
  #### `deviate execute pre [--task <id>]`
592
597
 
@@ -944,7 +949,7 @@ accepts `--json` and `--quiet`. `pre` emits a JSON contract describing the envir
944
949
  * Valid `plan.md` and no `tasks.md`: emit `MESO_RESUME`, skip Plan, and run Tasks.
945
950
  * Valid `plan.md` and non-empty `tasks.md`: emit `MESO_ALREADY_COMPLETE`, skip both agents,
946
951
  preserve ledger progress, and return the current worktree path.
947
- * Existing invalid `plan.md`: emit `MESO_PLAN_INVALID` and stop without overwrite.
952
+ * Existing `plan.md` valid only after repair: when the contract fails solely for a missing `**Verification Mode**:` line, it is auto-filled (`PLAN_MODE_REPAIR`) and treated as valid; a genuinely invalid `plan.md` (missing clauses, bad AO traceability, illegal/duplicated mode) emits `MESO_PLAN_INVALID` and stops without overwrite.
948
953
  * Existing empty `tasks.md`: emit `MESO_TASKS_INVALID` and stop without overwrite.
949
954
  A fresh claim does not use inherited main-branch artifacts as resume evidence. It runs Plan
950
955
  and Tasks in the new worktree.
@@ -274,7 +274,7 @@ to invoke, but model selection is delegated to the calling environment.
274
274
  * **Layer discipline:** GREEN's only invariant is "make the RED test pass via the library/API surface declared in scope." It does NOT make scope, spec-drift, or HITL-routing judgments — those belong to JUDGE. When a RED test cannot be satisfied within GREEN's mechanical scope, GREEN emits `status: FAILURE` with a concrete `rationale:` naming the test path and why; `status: "ERROR"` is reserved strictly for tool/orchestration failure. The runner's `_is_hitl_escalation` is a narrow defensive fallback that ONLY promotes structured `contract_drift` / `escalates_to` / `hitl_options` dict keys to `HITL_PENDING` — loose-string `error_kind` discriminators and free-form scope-conflict text do NOT trigger HITL escalation.
275
275
  * **Mechanical Failure → JUDGE Routing:** When GREEN emits `status: FAILURE` with a concrete `rationale:` (the mechanical scope-boundary case above), the runner routes control to JUDGE instead of raising `PhaseFailedError`. `_run_green_phase` sets `session.train_feedback = rationale` + `session.failure_kind = "mechanical"` and returns the session; `_run_judge_phase` injects a `<failure_kind>mechanical</failure_kind>` discriminator block into the JUDGE prompt that instructs the agent to emit `verdict: COMPLIANCE_PASS` + `next_action: proceed_to_refactor_no_diff` (when the slice is intrinsically RED-only and REFACTOR's no-op commit + COMPLETED transition is the right termination) OR `verdict: COMPLIANCE_VIOLATION` + one of three `next_action` values (`revert_before` / `revert_to_red` / `skip_refactor`) instead of attempting to satisfy the test itself. This closes the loop where mechanical FAILURE (e.g. slice-scope conflict, CLI-surface-out-of-sco…
276
276
  * **Test-Defect Failure → JUDGE Routing:** A second routable failure class, parallel to mechanical but pre-decided. When GREEN observes that the RED test itself is wrong (it asserts behavior the spec does not require, exercises the wrong abstraction, or encodes an assumption that contradicts spec/data-model), GREEN emits `status: FAILURE` with a concrete `rationale:` citing the FR/AC the test contradicts, plus `failure_kind: test_defect` on the manifest (`HandoverManifest.failure_kind: Literal["mechanical", "test_defect", "already_satisfied"] | None` in `src/deviate/core/agent.py` — session mirrors the discriminator as `SessionState.failure_kind: Literal["", "mechanical", "test_defect", "no_failing_test"]` in `src/deviate/state/config.py`). `_run_green_phase` reads the manifest's discriminator, sets `session.failure_kind = "test_defect"`, and routes to JUDGE. `_run_judge_phase` injects a `<failure_kind>test_defect</failure_kind>` discriminator block that pre-decides the routing — `verdict: COMPLIANCE_VIOLATION` + `next_action: revert_before` (re-run RED with the GREEN rationale as feedback) — because test defect has only one sensible outcome: the test itself must be re-authored. The discriminator intentionally narrows the JUDGE routing vocabulary compared to mechanical (`revert_to_red` / `skip_refactor` / `proceed_to_refactor_no_diff` are NOT options in the test_defect block). Default `manifest.failure_kind = None` falls through to `"mechanical"` in the runner to preserve prior behavior.
277
- * **RED No-Failing-Test → JUDGE Adjudication (RED → JUDGE direct route):** When RED completes but its test command exits 0 (all tests passed) or collects no tests (pytest exit 5), `_run_red_phase` does NOT raise a raw `PhaseFailedError` and does NOT let GREEN run against a vacuous test. It calls `_adjudicate_red_no_failing_test` (`src/deviate/cli/micro.py`), which sets `session.failure_kind = "no_failing_test"`, injects a `<failure_kind>no_failing_test</failure_kind>` discriminator block into the JUDGE prompt, and dispatches `_run_judge_phase(...)`. The judge diff spans the uncommitted RED test (the `red_baseline` parameter makes `_run_judge_phase` skip the `RED→HEAD` committed diff and surface the agent's uncommitted test through the dirty-parts scan) so JUDGE reviews what the agent actually wrote. JUDGE decides between two outcomes: `verdict: COMPLIANCE_PASS` + `next_action: skip_refactor` (or a bare PASS verdict) when the required behavior already exists — the runner discards the uncommitted passing test via `_restore_worktree_to_baseline` and marks the task COMPLETED without landing it; or `verdict: COMPLIANCE_VIOLATION` + `next_action: revert_before` (forced by the `_coerce_judge_action` runner-level override for `failure_kind` `no_failing_test`/`test_defect`) when the test is wrong — the runner resets to the RED baseline and re-dispatches RED so a fresh genuinely-failing test is authored. `--no-judge` makes the no-failing-test outcome a hard failure (adjudication disabled). When RED produces no test files at all, `_run_red_phase` raises `PhaseFailedError("RED phase produced no test files ...")` naming `/deviate-execute` (DIRECT) or `/deviate-meso` re-sharding — closing the prior stall where an empty `_find_test_files` let RED silently commit a vacuous 'failing test' and GREEN die in `TRAIN_EXHAUSTED`. The RED agent may steer the adjudication by declaring `failure_kind: already_satisfied` (behavior exists) or `failure_kind: test_defect` (test wrong) with a `rationale` on its handover manifest.
277
+ * **RED No-Failing-Test → JUDGE Adjudication (RED → JUDGE direct route):** When RED completes but its test command exits 0 (all tests passed), collects no tests (pytest exit 5), or resolves to no command at all (returncode 127 from `_run_test_cmd`), `_run_red_phase` does NOT raise a raw `PhaseFailedError` and does NOT let GREEN run against a vacuous test. It calls `_adjudicate_red_no_failing_test` (`src/deviate/cli/micro.py`), which sets `session.failure_kind = "no_failing_test"`, injects a `<failure_kind>no_failing_test</failure_kind>` discriminator block into the JUDGE prompt, and dispatches `_run_judge_phase(...)`. The judge diff spans the uncommitted RED test (the `red_baseline` parameter makes `_run_judge_phase` skip the `RED→HEAD` committed diff and surface the agent's uncommitted test through the dirty-parts scan) so JUDGE reviews what the agent actually wrote. JUDGE decides between two outcomes: `verdict: COMPLIANCE_PASS` + `next_action: skip_refactor` (or a bare PASS verdict) when the required behavior already exists — the runner discards the uncommitted passing test via `_restore_worktree_to_baseline` and marks the task COMPLETED without landing it; or `verdict: COMPLIANCE_VIOLATION` + `next_action: revert_before` (forced by the `_coerce_judge_action` runner-level override for `failure_kind` `no_failing_test`/`test_defect`) when the test is wrong — the runner resets to the RED baseline and re-dispatches RED so a fresh genuinely-failing test is authored. `--no-judge` makes the no-failing-test outcome a hard failure (adjudication disabled). Test discovery is language-agnostic: the RED gate no longer globs Python `tests/**/test_*.py` (`_find_test_files`); the project's own convention is honored through the resolved test command (e.g. `mix test` collecting `test/**/*_test.exs` for Elixir/Phoenix), so a correctly authored non-Python test is neither rejected up front nor silently conflated with "no test framework" — a project with no resolvable test command routes to the same JUDGE adjudication instead of failing. This closes the prior stall where an empty `_find_test_files` let RED silently commit a vacuous 'failing test' and GREEN die in `TRAIN_EXHAUSTED`. The RED agent may steer the adjudication by declaring `failure_kind: already_satisfied` (behavior exists) or `failure_kind: test_defect` (test wrong) with a `rationale` on its handover manifest.
278
278
  * **JUDGE / TRAIN (The Compliance Gate) — with Green → Judge → Green loop:**
279
279
  * **The Judge:** The CLI evaluates the committed RED-parent-to-HEAD diff against `spec.md` for invariant/security violations. If GREEN tests failed before the implementation commit, `_run_judge_phase` also appends staged/unstaged `git diff HEAD` output and per-file `git diff --no-index /dev/null <path>` output for untracked files, so JUDGE assesses the retained implementation rather than a false RED-only snapshot. This judge operates in a clean, zero-shared-history session to break recursive subjectivity. A `deviate-judge` skill (loaded from `_SKILL_NAMES["JUDGE"]`) guides the agent through supplementary compliance evaluation.
280
280
  * **The Train (Green → Judge → Green loop):** On `COMPLIANCE_VIOLATION` or test failure, the CLI safely resets without destroying task progress. The JUDGE phase honors `HandoverManifest.next_action` (see [specs/DeviaTDD-api.md](./DeviaTDD-api.md) for the routing table). The five routes:
@@ -621,17 +621,26 @@ framework's remaining HITL gates are Gate 1 and Gate 3 (Gate 2 was removed).
621
621
 
622
622
  ## 7. Multi-Framework Testing Abstraction
623
623
 
624
- DeviaTDD's current implementation (`src/deviate/cli/micro.py`) supports **pytest** as its
625
- test runner via `_run_pytest()`. The abstraction layer is designed to be extensible to other
626
- frameworks through the `_classify_pytest_outcome()` pattern, which parses stdout/stderr for
627
- syntax errors, assertion failures, and pass states. Currently, `_run_pytest()` collects all
628
- `tests/**/test_*.py` files and runs them with `python -m pytest -v`.
624
+ DeviaTDD's current implementation (`src/deviate/cli/micro.py`) runs tests through the
625
+ language-agnostic `_run_test_cmd()` → `_test_command_candidates()` resolution: the task's
626
+ `verification` value, the constitution `test_command`, a `mise run test` task, or the
627
+ `_MANIFEST_TEST_COMMANDS` manifest table (`mix.exs` → `mix test`, `Cargo.toml` → `cargo test`,
628
+ `go.mod` → `go test ./...`, `package.json` → `npm test`, `pyproject.toml` → `pytest`), with a
629
+ Python-only fallback (`_find_test_files` globbing `tests/**/test_*.py` → `pytest`) used only
630
+ when no other framework is detected. Test discovery follows each project's own convention
631
+ (e.g. `test/**/*_test.exs` for Elixir) — the RED gate does not require a Python-style
632
+ `tests/**/test_*.py` file. Outcome classification is shared: a red/green phase is decided by
633
+ the command's exit code plus `_is_no_tests_collected` (pytest exit 5) and `_is_no_test_command`
634
+ (returncode 127) sentinels. `_run_pytest()` remains the Python-specific subprocess runner
635
+ (`tests/**/test_*.py` + `python -m pytest -v`), used where a Python command was resolved.
629
636
 
630
637
  | Testing Framework | CLI Invocation Strategy | Success Validation | Error Parse Pattern | Scope Protection |
631
638
  | :--- | :--- | :--- | :--- | :--- |
632
- | **Python / pytest** | `python -m pytest tests/ -v` | `returncode == 0` | `_classify_pytest_outcome()`: checks `SYNTAX_ERROR` markers (SyntaxError, IndentationError, etc.), `ASSERTION_FAILURE`, `PASS`, `UNKNOWN_FAILURE`. | Reverts unauthorized test edits before running suite. |
633
- | **Node.js / Jest** | (Not implemented) | — | — | — |
634
- | **Go / testing** | (Not implemented) | — | — | — |
639
+ | **Python / pytest** | `python -m pytest tests/ -v` via `_run_pytest()` | `returncode == 0` | `_classify_pytest_outcome()`: checks `SYNTAX_ERROR` markers (SyntaxError, IndentationError, etc.), `ASSERTION_FAILURE`, `PASS`, `UNKNOWN_FAILURE`. | Reverts unauthorized test edits before running suite. |
640
+ | **Elixir / ExUnit** | `mix test` (via `_MANIFEST_TEST_COMMANDS`) | `returncode == 0` | exit code + `_is_no_tests_collected` / `_is_no_test_command` | Same. |
641
+ | **Rust / cargo** | `cargo test` (via `_MANIFEST_TEST_COMMANDS`) | `returncode == 0` | exit code + `_is_no_tests_collected` / `_is_no_test_command` | Same. |
642
+ | **Go / testing** | `go test ./...` (via `_MANIFEST_TEST_COMMANDS`) | `returncode == 0` | exit code + `_is_no_tests_collected` / `_is_no_test_command` | Same. |
643
+ | **Node.js / Jest** | `npm test` (via `_MANIFEST_TEST_COMMANDS`) | `returncode == 0` | exit code + `_is_no_tests_collected` / `_is_no_test_command` | Same. |
635
644
 
636
645
  ---
637
646
 
@@ -903,7 +912,9 @@ there is no `discover_skills()` abstraction.
903
912
  **Scope:** Unified Meso and Micro orchestration. The skill first invokes
904
913
  `deviate meso run`. In a linked feature worktree, Meso validates existing
905
914
  `plan.md` and `tasks.md`, skips completed phases, resumes at Tasks when only
906
- Plan is ready, and stops on invalid artifacts without overwrite. After Meso
915
+ Plan is ready, auto-repairs a plan whose acceptance scenarios lack the
916
+ `**Verification Mode**:` line (default `automated`, `PLAN_MODE_REPAIR`), and
917
+ stops on genuinely invalid artifacts without overwrite. After Meso
907
918
  succeeds, the skill invokes bare `deviate micro run` one task at a time.
908
919
  The existing failure triage and clean-slate safety flow remains. **v1.1.0 added a
909
920
  `## Troubleshooting failed runs` section** documenting the two