deviatdd 2.22.1__tar.gz → 2.23.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (474) hide show
  1. deviatdd-2.23.0/.deviate/config.toml +20 -0
  2. {deviatdd-2.22.1 → deviatdd-2.23.0}/CHANGELOG.md +7 -0
  3. {deviatdd-2.22.1 → deviatdd-2.23.0}/PKG-INFO +1 -1
  4. {deviatdd-2.22.1 → deviatdd-2.23.0}/pyproject.toml +1 -1
  5. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/DeviaTDD-api.md +21 -7
  6. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/DeviaTDD-architecture.md +13 -5
  7. deviatdd-2.23.0/specs/adhoc/031-judge-revert-boundary-no-failing-test/plan.md +149 -0
  8. deviatdd-2.23.0/specs/adhoc/031-judge-revert-boundary-no-failing-test/tasks.jsonl +5 -0
  9. deviatdd-2.23.0/specs/adhoc/031-judge-revert-boundary-no-failing-test/tasks.md +118 -0
  10. deviatdd-2.23.0/specs/adhoc/issues/029-ponytail-pruning-in-review.md +126 -0
  11. deviatdd-2.23.0/specs/adhoc/issues/030-config-rework.md +87 -0
  12. deviatdd-2.23.0/specs/adhoc/issues/031-judge-revert-boundary-no-failing-test.md +78 -0
  13. deviatdd-2.23.0/specs/adhoc/issues/032-judge-feedback-injection-fail-close.md +75 -0
  14. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/prd.md +55 -0
  15. deviatdd-2.23.0/specs/explore/config-rework.md +134 -0
  16. deviatdd-2.23.0/specs/explore/ponytail-pruning.md +136 -0
  17. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/issues.jsonl +6 -0
  18. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/meso.py +22 -7
  19. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/micro.py +46 -4
  20. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/agent.py +71 -2
  21. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/issues.py +6 -5
  22. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/worktree.py +4 -2
  23. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/judge.md +1 -1
  24. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/skills/deviatdd/SKILL.md +22 -5
  25. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/state/ledger.py +18 -3
  26. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/conftest.py +6 -1
  27. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_micro.py +129 -0
  28. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_agent.py +65 -0
  29. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_issues.py +29 -0
  30. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_worktree.py +45 -0
  31. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_meso_orchestration.py +51 -0
  32. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_judge.py +278 -0
  33. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_state/test_ledger.py +69 -0
  34. {deviatdd-2.22.1 → deviatdd-2.23.0}/uv.lock +1 -1
  35. deviatdd-2.22.1/.deviate/config.toml +0 -16
  36. {deviatdd-2.22.1 → deviatdd-2.23.0}/.deviate/.gitignore +0 -0
  37. {deviatdd-2.22.1 → deviatdd-2.23.0}/.deviate/evaluations/deviate-flows-skill-2026-06-26.json +0 -0
  38. {deviatdd-2.22.1 → deviatdd-2.23.0}/.env.example +0 -0
  39. {deviatdd-2.22.1 → deviatdd-2.23.0}/.gitattributes +0 -0
  40. {deviatdd-2.22.1 → deviatdd-2.23.0}/.githooks/pre-commit +0 -0
  41. {deviatdd-2.22.1 → deviatdd-2.23.0}/.githooks/pre-push +0 -0
  42. {deviatdd-2.22.1 → deviatdd-2.23.0}/.github/ISSUE_TEMPLATE/bug.md +0 -0
  43. {deviatdd-2.22.1 → deviatdd-2.23.0}/.github/ISSUE_TEMPLATE/feature.md +0 -0
  44. {deviatdd-2.22.1 → deviatdd-2.23.0}/.github/PULL_REQUEST_TEMPLATE.md +0 -0
  45. {deviatdd-2.22.1 → deviatdd-2.23.0}/.github/workflows/ci.yml +0 -0
  46. {deviatdd-2.22.1 → deviatdd-2.23.0}/.github/workflows/release.yml +0 -0
  47. {deviatdd-2.22.1 → deviatdd-2.23.0}/.gitignore +0 -0
  48. {deviatdd-2.22.1 → deviatdd-2.23.0}/.opencode/opencode.json +0 -0
  49. {deviatdd-2.22.1 → deviatdd-2.23.0}/AGENTS.md +0 -0
  50. {deviatdd-2.22.1 → deviatdd-2.23.0}/CLAUDE.md +0 -0
  51. {deviatdd-2.22.1 → deviatdd-2.23.0}/CODE_OF_CONDUCT.md +0 -0
  52. {deviatdd-2.22.1 → deviatdd-2.23.0}/CONTRIBUTING.md +0 -0
  53. {deviatdd-2.22.1 → deviatdd-2.23.0}/LICENSE +0 -0
  54. {deviatdd-2.22.1 → deviatdd-2.23.0}/README.md +0 -0
  55. {deviatdd-2.22.1 → deviatdd-2.23.0}/SECURITY.md +0 -0
  56. {deviatdd-2.22.1 → deviatdd-2.23.0}/deviatdd.png +0 -0
  57. {deviatdd-2.22.1 → deviatdd-2.23.0}/mise.toml +0 -0
  58. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat/001-deviate-cli-python/008-meso-macro-automated-orchestration.md +0 -0
  59. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat/adhoc/006-context-cli-integration.md +0 -0
  60. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat/adhoc/017-optional-push-as-lock.md +0 -0
  61. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-001-deviate-cli-python-001-cli-initialization-governance-provisioning-pr-11.md +0 -0
  62. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-001-deviate-cli-python-003-meso-layer-specification-task-decomposition.md +0 -0
  63. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-001-deviate-cli-python-004-micro-layer-tdd-sandbox-execution.md +0 -0
  64. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-001-deviate-cli-python-005-cli-architecture-realignment-skill-integration.md +0 -0
  65. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-001-deviate-cli-python-007-macro-meso-parity-backward-compatibility.md +0 -0
  66. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-002-deviatdd-gap-analysis-001-foundation-cli-infrastructure.md +0 -0
  67. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-002-deviatdd-gap-analysis-003-fast-path-commands.md +0 -0
  68. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-002-deviatdd-gap-analysis-004-governance-inspection.md +0 -0
  69. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-adhoc-001-streaming-pipeline-monitor.md +0 -0
  70. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-adhoc-008-ast-phase-prioritization.md +0 -0
  71. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat-adhoc-017-two-counter-tdd-retry.md +0 -0
  72. {deviatdd-2.22.1 → deviatdd-2.23.0}/pr_descriptions/feat_002-deviatdd-gap-analysis_001-foundation-cli-infrastructure.md +0 -0
  73. {deviatdd-2.22.1 → deviatdd-2.23.0}/scripts/benchmark_lmstudio.py +0 -0
  74. {deviatdd-2.22.1 → deviatdd-2.23.0}/scripts/next_version.py +0 -0
  75. {deviatdd-2.22.1 → deviatdd-2.23.0}/scripts/verify_install.py +0 -0
  76. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t001.json +0 -0
  77. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t002.json +0 -0
  78. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t004.json +0 -0
  79. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/spec.md +0 -0
  80. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/tasks.md +0 -0
  81. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/002-macro-layer-state-ledger-management/spec.md +0 -0
  82. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/002-macro-layer-state-ledger-management/tasks.md +0 -0
  83. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/003-meso-layer-specification-task-decomposition/spec.md +0 -0
  84. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/003-meso-layer-specification-task-decomposition/tasks.md +0 -0
  85. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/spec.md +0 -0
  86. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/tasks.jsonl +0 -0
  87. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/tasks.md +0 -0
  88. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/005-cli-architecture-realignment-skill-integration/spec.md +0 -0
  89. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/005-cli-architecture-realignment-skill-integration/tasks.md +0 -0
  90. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/007-macro-meso-parity-backward-compatibility/spec.md +0 -0
  91. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/007-macro-meso-parity-backward-compatibility/tasks.md +0 -0
  92. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/spec.md +0 -0
  93. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/tasks.jsonl +0 -0
  94. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/tasks.md +0 -0
  95. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/data-model.md +0 -0
  96. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/design.md +0 -0
  97. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/explore.md +0 -0
  98. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/001-cli-initialization-governance-provisioning.md +0 -0
  99. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/002-macro-layer-state-ledger-management.md +0 -0
  100. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/003-meso-layer-specification-task-decomposition.md +0 -0
  101. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/004-micro-layer-tdd-sandbox-execution.md +0 -0
  102. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/005-cli-architecture-realignment-skill-integration.md +0 -0
  103. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/006-state-persistence-concurrency-safety.md +0 -0
  104. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/007-macro-meso-parity-backward-compatibility.md +0 -0
  105. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/issues/008-meso-macro-automated-orchestration.md +0 -0
  106. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/001-deviate-cli-python/prd.md +0 -0
  107. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/spec.md +0 -0
  108. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/tasks.jsonl +0 -0
  109. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/tasks.md +0 -0
  110. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/spec.md +0 -0
  111. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/tasks.jsonl +0 -0
  112. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/tasks.md +0 -0
  113. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/004-governance-inspection/spec.md +0 -0
  114. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/004-governance-inspection/tasks.jsonl +0 -0
  115. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/004-governance-inspection/tasks.md +0 -0
  116. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/spec.md +0 -0
  117. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/tasks.jsonl +0 -0
  118. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/tasks.md +0 -0
  119. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/data-model.md +0 -0
  120. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/design.md +0 -0
  121. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/explore.md +0 -0
  122. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/issues/001-foundation-cli-infrastructure.md +0 -0
  123. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/issues/002-context-pipeline.md +0 -0
  124. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/issues/003-fast-path-commands.md +0 -0
  125. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/issues/004-governance-inspection.md +0 -0
  126. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/issues/005-micro-layer-integrity.md +0 -0
  127. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/plan-tdd-integration-gap.md +0 -0
  128. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/002-deviatdd-gap-analysis/prd.md +0 -0
  129. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/003-graphite-cli-integration/explore.md +0 -0
  130. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/004-per-task-security-profile/issues/001-security-profile-and-judge-checks.md +0 -0
  131. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/001-verification-mode-metadata/plan.md +0 -0
  132. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/001-verification-mode-metadata/tasks.jsonl +0 -0
  133. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/001-verification-mode-metadata/tasks.md +0 -0
  134. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/002-task-acceptance-traceability/plan.md +0 -0
  135. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/002-task-acceptance-traceability/tasks.jsonl +0 -0
  136. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/002-task-acceptance-traceability/tasks.md +0 -0
  137. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/data-model.md +0 -0
  138. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/design.md +0 -0
  139. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/explore.md +0 -0
  140. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/issues/001-verification-mode-metadata.md +0 -0
  141. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/issues/002-task-acceptance-traceability.md +0 -0
  142. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/issues/003-micro-phase-gates-red-green.md +0 -0
  143. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/issues/004-refactor-regression-gate.md +0 -0
  144. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/issues/005-prompt-spec-alignment.md +0 -0
  145. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/005-acceptance-gates/prd.md +0 -0
  146. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/architecture.md +0 -0
  147. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/domain-model.md +0 -0
  148. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/flows/flows-product.md +0 -0
  149. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/flows/flows-streaming.md +0 -0
  150. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/flows/index.md +0 -0
  151. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/flows.jsonl +0 -0
  152. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/guildwright-current-system.md +0 -0
  153. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/guildwright-gap-register.md +0 -0
  154. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/guildwright-git-state-model.md +0 -0
  155. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/guildwright-rewrite.md +0 -0
  156. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/guildwright-rust-tui-requirements.md +0 -0
  157. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/_product/release-next.md +0 -0
  158. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/001-streaming-pipeline-monitor/spec.md +0 -0
  159. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/001-streaming-pipeline-monitor/tasks.jsonl +0 -0
  160. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/001-streaming-pipeline-monitor/tasks.md +0 -0
  161. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/003-meso-layer-restructuring/spec.md +0 -0
  162. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/003-meso-layer-restructuring/tasks.jsonl +0 -0
  163. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/003-meso-layer-restructuring/tasks.md +0 -0
  164. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/004-deviate-review-skill/spec.md +0 -0
  165. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/004-deviate-review-skill/tasks.jsonl +0 -0
  166. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/004-deviate-review-skill/tasks.md +0 -0
  167. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/005-per-phase-model-configuration/plan.md +0 -0
  168. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/005-per-phase-model-configuration/tasks.jsonl +0 -0
  169. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/005-per-phase-model-configuration/tasks.md +0 -0
  170. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/006-context-cli-integration/plan.md +0 -0
  171. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/006-context-cli-integration/tasks.jsonl +0 -0
  172. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/006-context-cli-integration/tasks.md +0 -0
  173. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/007-graphite-cli/plan.md +0 -0
  174. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/007-graphite-cli/tasks.jsonl +0 -0
  175. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/007-graphite-cli/tasks.md +0 -0
  176. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/008-ast-phase-prioritization/plan.md +0 -0
  177. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/008-ast-phase-prioritization/tasks.jsonl +0 -0
  178. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/008-ast-phase-prioritization/tasks.md +0 -0
  179. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/009-pi-agent-backend-integration/plan.md +0 -0
  180. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/009-pi-agent-backend-integration/tasks.jsonl +0 -0
  181. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/009-pi-agent-backend-integration/tasks.md +0 -0
  182. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/010-deviate-setup-product-layer/plan.md +0 -0
  183. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/010-deviate-setup-product-layer/tasks.jsonl +0 -0
  184. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/010-deviate-setup-product-layer/tasks.md +0 -0
  185. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/plan.md +0 -0
  186. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/tasks.jsonl +0 -0
  187. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/tasks.md +0 -0
  188. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/015-narrow-product-flow-scope/plan.md +0 -0
  189. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/015-narrow-product-flow-scope/tasks.md +0 -0
  190. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/016-single-source-prompt-templates/plan.md +0 -0
  191. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/016-single-source-prompt-templates/tasks.jsonl +0 -0
  192. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/016-single-source-prompt-templates/tasks.md +0 -0
  193. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/017-optional-push-as-lock/plan.md +0 -0
  194. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/017-optional-push-as-lock/tasks.jsonl +0 -0
  195. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/017-optional-push-as-lock/tasks.md +0 -0
  196. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/017-two-counter-tdd-retry/plan.md +0 -0
  197. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/017-two-counter-tdd-retry/tasks.jsonl +0 -0
  198. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/017-two-counter-tdd-retry/tasks.md +0 -0
  199. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/018-one-behavior-rgr-granularity/plan.md +0 -0
  200. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/018-one-behavior-rgr-granularity/tasks.md +0 -0
  201. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/019-remote-aware-ordinal-allocation/plan.md +0 -0
  202. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/019-remote-aware-ordinal-allocation/tasks.jsonl +0 -0
  203. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/019-remote-aware-ordinal-allocation/tasks.md +0 -0
  204. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/020-judge-compliance-pass-evidence/plan.md +0 -0
  205. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/020-judge-compliance-pass-evidence/tasks.jsonl +0 -0
  206. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/020-judge-compliance-pass-evidence/tasks.md +0 -0
  207. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/021-no-failing-test-escalate-invokes-green/plan.md +0 -0
  208. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/021-no-failing-test-escalate-invokes-green/tasks.jsonl +0 -0
  209. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/021-no-failing-test-escalate-invokes-green/tasks.md +0 -0
  210. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/022-already-satisfied-red-requires-tests/plan.md +0 -0
  211. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/022-already-satisfied-red-requires-tests/tasks.jsonl +0 -0
  212. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/022-already-satisfied-red-requires-tests/tasks.md +0 -0
  213. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/023-pinned-micro-run-issue-scoped/plan.md +0 -0
  214. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/023-pinned-micro-run-issue-scoped/tasks.jsonl +0 -0
  215. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/023-pinned-micro-run-issue-scoped/tasks.md +0 -0
  216. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/024-worktree-session-stale-issue-id/plan.md +0 -0
  217. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/024-worktree-session-stale-issue-id/tasks.jsonl +0 -0
  218. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/024-worktree-session-stale-issue-id/tasks.md +0 -0
  219. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/025-green-stderr-noise-stall-detector/plan.md +0 -0
  220. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/025-green-stderr-noise-stall-detector/tasks.jsonl +0 -0
  221. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/025-green-stderr-noise-stall-detector/tasks.md +0 -0
  222. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/026-pi-spawn-lean-tool-schema/plan.md +0 -0
  223. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/026-pi-spawn-lean-tool-schema/tasks.jsonl +0 -0
  224. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/026-pi-spawn-lean-tool-schema/tasks.md +0 -0
  225. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/027-red-hang-timeout-rollback/plan.md +0 -0
  226. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/027-red-hang-timeout-rollback/tasks.jsonl +0 -0
  227. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/027-red-hang-timeout-rollback/tasks.md +0 -0
  228. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/028-task-scoped-judge-review-coverage/plan.md +0 -0
  229. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/028-task-scoped-judge-review-coverage/tasks.jsonl +0 -0
  230. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/028-task-scoped-judge-review-coverage/tasks.md +0 -0
  231. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/001-streaming-pipeline-monitor.md +0 -0
  232. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/002-aider-agent-backend-integration.md +0 -0
  233. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/003-meso-layer-restructuring.md +0 -0
  234. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/004-deviate-review-skill.md +0 -0
  235. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/005-per-phase-model-configuration.md +0 -0
  236. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/006-context-cli-integration.md +0 -0
  237. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/007-graphite-cli.md +0 -0
  238. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/008-ast-phase-prioritization.md +0 -0
  239. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/009-pi-agent-backend-integration.md +0 -0
  240. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/010-deviate-setup-product-layer.md +0 -0
  241. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/012-rpc-streaming-tui-renderer.md +0 -0
  242. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/013-flow-ledger-canonical-source-of-truth.md +0 -0
  243. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/014-cwe-mapping-security-findings.md +0 -0
  244. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/015-narrow-product-flow-scope.md +0 -0
  245. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/016-single-source-prompt-templates.md +0 -0
  246. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/017-optional-push-as-lock.md +0 -0
  247. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/017-two-counter-tdd-retry.md +0 -0
  248. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/018-one-behavior-rgr-granularity.md +0 -0
  249. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/019-remote-aware-ordinal-allocation.md +0 -0
  250. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/020-judge-compliance-pass-evidence.md +0 -0
  251. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/021-no-failing-test-escalate-invokes-green.md +0 -0
  252. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/022-already-satisfied-red-requires-tests.md +0 -0
  253. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/023-pinned-micro-run-issue-scoped.md +0 -0
  254. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/024-worktree-session-stale-issue-id.md +0 -0
  255. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/025-green-stderr-noise-stall-detector.md +0 -0
  256. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/026-pi-spawn-lean-tool-schema.md +0 -0
  257. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/027-red-hang-timeout-rollback.md +0 -0
  258. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc/issues/028-task-scoped-judge-review-coverage.md +0 -0
  259. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/adhoc.jsonl +0 -0
  260. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/constitution.md +0 -0
  261. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/ast-tree-sitter.md +0 -0
  262. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/flow-ledger.md +0 -0
  263. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/graphite-cli.md +0 -0
  264. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/pi-agent-backend.md +0 -0
  265. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/product-layer.md +0 -0
  266. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/rpc-streaming-tui.md +0 -0
  267. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/rpc-streaming.md +0 -0
  268. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/explore/security-hardening-cwe.md +0 -0
  269. {deviatdd-2.22.1 → deviatdd-2.23.0}/specs/implementation-gap.md +0 -0
  270. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/__init__.py +0 -0
  271. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/__init__.py +0 -0
  272. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/_common.py +0 -0
  273. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/_html.py +0 -0
  274. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/_safe_commands.py +0 -0
  275. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/adhoc.py +0 -0
  276. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/constitution.py +0 -0
  277. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/feature.py +0 -0
  278. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/flow_commands.py +0 -0
  279. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/init.py +0 -0
  280. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/inspect.py +0 -0
  281. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/macro.py +0 -0
  282. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/review.py +0 -0
  283. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/cli/walkthrough.py +0 -0
  284. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/__init__.py +0 -0
  285. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/_shared.py +0 -0
  286. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/cache_discipline.py +0 -0
  287. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/commands.py +0 -0
  288. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/commit.py +0 -0
  289. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/complexity.py +0 -0
  290. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/constitution.py +0 -0
  291. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/contract.py +0 -0
  292. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/convention.py +0 -0
  293. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/epic.py +0 -0
  294. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/judge_evidence.py +0 -0
  295. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/prd.py +0 -0
  296. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/profile.py +0 -0
  297. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/repo.py +0 -0
  298. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/review_coverage.py +0 -0
  299. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/run_logger.py +0 -0
  300. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/tasks_ledger.py +0 -0
  301. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/core/validation.py +0 -0
  302. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/html_templates/__init__.py +0 -0
  303. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/html_templates/architecture.html.tmpl +0 -0
  304. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/html_templates/domain-model.html.tmpl +0 -0
  305. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/html_templates/flows.html.tmpl +0 -0
  306. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/html_templates/plan.html.tmpl +0 -0
  307. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/html_templates/prd.html.tmpl +0 -0
  308. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/main.py +0 -0
  309. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/__init__.py +0 -0
  310. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/assembly.py +0 -0
  311. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/__init__.py +0 -0
  312. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/execute.md +0 -0
  313. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/explore.md +0 -0
  314. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/green.md +0 -0
  315. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/plan.md +0 -0
  316. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/prd.md +0 -0
  317. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/red.md +0 -0
  318. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/refactor.md +0 -0
  319. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/research.md +0 -0
  320. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/shard.md +0 -0
  321. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/specify.md +0 -0
  322. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/auto/tasks.md +0 -0
  323. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-adhoc.md +0 -0
  324. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-architecture.md +0 -0
  325. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-constitution.md +0 -0
  326. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-e2e.md +0 -0
  327. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-execute.md +0 -0
  328. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-explore.md +0 -0
  329. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-flows.md +0 -0
  330. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-green.md +0 -0
  331. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-hotfix.md +0 -0
  332. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-html.md +0 -0
  333. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-init.md +0 -0
  334. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-judge.md +0 -0
  335. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-merge.md +0 -0
  336. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-plan.md +0 -0
  337. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-pr.md +0 -0
  338. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-prd.md +0 -0
  339. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-prune.md +0 -0
  340. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-red.md +0 -0
  341. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-refactor.md +0 -0
  342. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-release.md +0 -0
  343. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-research.md +0 -0
  344. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-review.md +0 -0
  345. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-shard.md +0 -0
  346. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-tasks.md +0 -0
  347. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-triage.md +0 -0
  348. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/commands/deviate-walkthrough.md +0 -0
  349. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/constitution_seed.md +0 -0
  350. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/core.md +0 -0
  351. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/lifecycle-auto.md +0 -0
  352. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/lifecycle-manual.md +0 -0
  353. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/macro-shared.md +0 -0
  354. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/meso-shared.md +0 -0
  355. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/micro-shared.md +0 -0
  356. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/product-shared.md +0 -0
  357. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/core/style-ste.md +0 -0
  358. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/governance/__init__.py +0 -0
  359. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/governance/agents_seed.md +0 -0
  360. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/governance/claudemd_seed.md +0 -0
  361. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/prompts/governance/libref_seed.md +0 -0
  362. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/state/__init__.py +0 -0
  363. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/state/config.py +0 -0
  364. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/ui/__init__.py +0 -0
  365. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/ui/monitor.py +0 -0
  366. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/ui/pipeline.py +0 -0
  367. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/ui/render.py +0 -0
  368. {deviatdd-2.22.1 → deviatdd-2.23.0}/src/deviate/visual/__init__.py +0 -0
  369. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/__init__.py +0 -0
  370. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/core/__init__.py +0 -0
  371. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/core/test_agent.py +0 -0
  372. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/core/test_smart_stall.py +0 -0
  373. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_derived_command_install.bats +0 -0
  374. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_green_stderr_stall.bats +0 -0
  375. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_macro_workflow.bats +0 -0
  376. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_optional_push_as_lock.bats +0 -0
  377. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_pi_spawn_lean_tool_schema.bats +0 -0
  378. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_red_hang_timeout_rollback.bats +0 -0
  379. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_review_plan_ac_coverage.bats +0 -0
  380. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/e2e/test_worktree_session_stale_issue.bats +0 -0
  381. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/__init__.py +0 -0
  382. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_adhoc.py +0 -0
  383. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_common.py +0 -0
  384. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_constitution.py +0 -0
  385. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_feature.py +0 -0
  386. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_flows_sync.py +0 -0
  387. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_help.py +0 -0
  388. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_html.py +0 -0
  389. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_init.py +0 -0
  390. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_inspect.py +0 -0
  391. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_macro_contracts.py +0 -0
  392. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_main_entrypoint.py +0 -0
  393. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_merge.py +0 -0
  394. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_meso.py +0 -0
  395. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_meso_contracts.py +0 -0
  396. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_real_descendant_kill.py +0 -0
  397. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_review.py +0 -0
  398. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_safe_commands.py +0 -0
  399. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_test_command_resolution.py +0 -0
  400. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_timeout_safe_command.py +0 -0
  401. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_cli/test_top_level_run.py +0 -0
  402. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_cache_discipline.py +0 -0
  403. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_commands.py +0 -0
  404. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_commit.py +0 -0
  405. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_complexity.py +0 -0
  406. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_constitution.py +0 -0
  407. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_contract.py +0 -0
  408. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_convention.py +0 -0
  409. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_epic.py +0 -0
  410. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_flow_confirmation.py +0 -0
  411. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_judge_evidence.py +0 -0
  412. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_ledger.py +0 -0
  413. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_prd.py +0 -0
  414. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_profile.py +0 -0
  415. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_repo.py +0 -0
  416. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_run_logger.py +0 -0
  417. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_tasks_ledger.py +0 -0
  418. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_core/test_validation.py +0 -0
  419. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_init.py +0 -0
  420. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/__init__.py +0 -0
  421. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/conftest.py +0 -0
  422. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_command_installation.py +0 -0
  423. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_init_export_cycle.py +0 -0
  424. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_macro_full_cycle.py +0 -0
  425. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_macro_layer.py +0 -0
  426. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_macro_orchestration.py +0 -0
  427. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_meso_layer.py +0 -0
  428. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_meso_orchestration.py +0 -0
  429. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_meso_task_ledger.py +0 -0
  430. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_parity.py +0 -0
  431. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_integration/test_skill_installation.py +0 -0
  432. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_macro/__init__.py +0 -0
  433. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_macro/test_explore.py +0 -0
  434. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_macro/test_macro_model_routing.py +0 -0
  435. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_macro/test_macro_orchestration.py +0 -0
  436. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_macro/test_prd.py +0 -0
  437. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_macro/test_research.py +0 -0
  438. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_macro/test_shard.py +0 -0
  439. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/__init__.py +0 -0
  440. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_auto_prompt_templates.py +0 -0
  441. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_meso_model_routing.py +0 -0
  442. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_meso_resume.py +0 -0
  443. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_plan_structure_injection.py +0 -0
  444. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_pr_platform.py +0 -0
  445. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_prompt_assembly.py +0 -0
  446. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_specify.py +0 -0
  447. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_meso/test_tasks.py +0 -0
  448. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/__init__.py +0 -0
  449. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/conftest.py +0 -0
  450. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_commit_failure.py +0 -0
  451. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_completed_evidence.py +0 -0
  452. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_e2e.py +0 -0
  453. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_execute.py +0 -0
  454. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_green.py +0 -0
  455. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_hotfix.py +0 -0
  456. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_orchestration.py +0 -0
  457. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_output_filter.py +0 -0
  458. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_red.py +0 -0
  459. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_refactor.py +0 -0
  460. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_review_pause.py +0 -0
  461. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_rollback_safety.py +0 -0
  462. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_run.py +0 -0
  463. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_task_label.py +0 -0
  464. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_micro/test_two_counter_retry.py +0 -0
  465. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_release/test_next_version.py +0 -0
  466. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_state/__init__.py +0 -0
  467. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_state/test_config.py +0 -0
  468. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_state/test_security_profile.py +0 -0
  469. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_state/test_session.py +0 -0
  470. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_ui/__init__.py +0 -0
  471. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_ui/test_monitor.py +0 -0
  472. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_ui/test_pipeline.py +0 -0
  473. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_ui/test_render.py +0 -0
  474. {deviatdd-2.22.1 → deviatdd-2.23.0}/tests/test_visual_demo/test_tsk_001_01.py +0 -0
@@ -0,0 +1,20 @@
1
+ # Preset config group: "default", "full", "fast", or "secure"
2
+ profile = "default"
3
+ # CLI inactivity timeout in seconds (must be > 0)
4
+ timeout_seconds = 1800
5
+ # Agent export mode: "local" (project) or "global" (~/.claude/)
6
+ agent_export_mode = "local"
7
+ # Enable the libref CLI for offline documentation lookups
8
+ use_libref = true
9
+ # Trunk branch for worktrees, PR base, and review diffs
10
+ base_branch = "main"
11
+ # Push the claim branch as a distributed lock (default true)
12
+ claim_remote = true
13
+
14
+ # Agent backend configuration
15
+
16
+ [agent]
17
+ backend = "pi"
18
+ timeout = 600
19
+ pi_rpc = false
20
+ transport = "rpc"
@@ -31,6 +31,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
31
31
  - **`deviate html prd --bucket <slug>` targets a specific epic when more than one owns a `prd.md`.** Previously the command hard-failed with `HTML_AMBIGUOUS_PRD` whenever multiple numbered epics had a `prd.md` and offered no flag to disambiguate — the agent was told to "run from within the epic's worktree", which silently yielded `HTML_NO_PRD` because the resolver reads `specs/` from the repo root. The new option resolves `specs/<bucket>/prd.md` directly and bypasses ambiguity detection; the plain form keeps its existing behavior but its ambiguity banner now points at the flag. An unknown/absent bucket exits `PRD_NOT_FOUND`. The `/deviate-html` prompt (`src/deviate/prompts/commands/deviate-html.md`) and spec docs now reference `--bucket` instead of the broken cwd fallback. Pinned by `tests/test_cli/test_html.py::test_html_prd_bucket_targets_specific_epic`, `::test_html_prd_bucket_targets_unnumbered_dir`, and `::test_html_prd_bucket_missing_file_exits_cleanly`.
32
32
  - **Optional push-as-lock: `claim_remote` config plus `--local` on `meso run` and `run`.** Standing `.deviate/config.toml` key `claim_remote` defaults to `true` (absent file or absent key still push). `deviate setup --no-claim-remote` writes `claim_remote = false` without dropping `[models]`, `timeout_seconds`, or `[agent]`. Fresh setup without the flag writes `true`. Effective local mode is `--local` OR `claim_remote = false`; explicit `--local` always wins. Local mode still creates `.worktrees/feat/{epic}/{issue}/`, writes SPECIFIED, and commits the claim, and skips `branch_exists_on_remote` plus `git push`. `deviate specify --local` stays; omitted `--local` now honors config. `deviate meso run --local` and `deviate run --local` share that meaning. `--no-setup` remains a distinct skip of worktree plus claim. Local discovery does not treat an origin branch as claimed-elsewhere. Specs: `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md` (Atomic Concurrency Protocol: push-as-lock is the default, not mandatory). Pinned by `tests/test_state/test_config.py`, `tests/test_cli/test_meso.py`, `tests/test_meso/test_specify.py`, `tests/test_meso/test_meso_orchestration.py`, `tests/test_cli/test_init.py`, and `tests/test_cli/test_top_level_run.py`.
33
33
  ### Changed
34
+ - **The `deviatdd` skill checks for an existing OPEN GitHub issue before filing a harness-bug issue.** The "Filing deviatdd issues" section now runs `gh issue list --repo wernerbisschoff/deviatdd --state open --search` first; when an open issue already matches the same harness failure, it directs commenting the new evidence/task context onto that issue (`gh issue comment`) instead of `gh issue create`, and only creates a new issue when no match exists. Prompt-only change to the packaged skill (`src/deviate/prompts/skills/deviatdd/SKILL.md`), re-installed to all agent skill dirs.
34
35
  - **`deviate refactor pre` now scopes `files_to_refactor` to the RED+GREEN production set (GH-98).** The command no longer glob()s every `src/**/*.py` or discards `_resolve_task_context`. It lists production files from `HEAD~2..HEAD`, falling back to the task `Files:` list minus tests when that git range is empty or unavailable. Test files are never included. The JSON contract now emits the documented handover fields (`status`, `task_id`, `task_title`, `task_type`, `test_command`, `lint_command`, `spec_dir`, `verification`, `repo_root`, `git_branch`, `timestamp`) alongside `files_to_refactor`. Auto `_build_auto_prompt("refactor")` injects the same scoped list and keeps the `git log -2` / `git diff HEAD~2..HEAD` inspect step. Pinned by `tests/test_micro/test_refactor.py`.
35
36
  - **GREEN / REFACTOR / review prompts fold smallest-change into existing lines** (reuse stdlib or an already-installed dep; in-place refactor; Opportunities do not extract helpers). Prompt-only. Pinned by `tests/test_meso/test_auto_prompt_templates.py::TestSmallestChangeFoldedIntoExistingPrompts`.
36
37
  - **JUDGE evidence is task-scoped; `deviate review` fail-closes on an unclaimed plan AC.** TDD `_run_judge_phase` resolves required `AC-PLAN-NNN` tokens via `resolve_task_ac_tokens` (non-empty `acceptance_criteria` `criterion_id`s, else this task's `tasks.md` card, else none) and does not fall back to every token in `plan.md`. ISS-ADH-020 exact-substring, path-in-diff, and uniqueness-floor checks still apply to that set. EXECUTE, IMMEDIATE, and DIRECT stay ungated. `deviate review pre` / `post` run a runner-owned Gate 3 scan (`evaluate_review_coverage`) and exit non-zero with `COVERAGE_INCOMPLETE` when any this-issue plan token has no COMPLETED claim. Missing `plan.md` or missing plan tokens stay vacuously READY. PENDING, FAILED, and sibling-issue rows do not claim. Pinned by `tests/test_core/test_judge_evidence.py`, `tests/test_micro/test_judge.py`, and `tests/test_cli/test_review.py`.
@@ -54,6 +55,12 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
54
55
  - **`/deviate-merge` no longer auto-pushes after the squash-merge commit; the push gate runs inline and the network push is opt-in.** The slash command previously ran `git push` as its final step. As of v2.4.0 the squash-merge commit lands on `main`, then a new `push_gate` step inlines the body of `.githooks/pre-push` (lint + format-check + testmon-driven affected tests with the warm-cache / full-suite fallback — bash 3.2 portable, `GIT_DIR` reset + trap preserved) so the safety net fires even though no `git push` happens yet. After the gate passes the skill asks the operator whether to `git push` (which fires the real `pre-push` hook and re-runs the same gate) or stop and push manually. The squash-merge commit and the ledger transition inside it are durable on `main` regardless of the push outcome — only the network push is deferred. New failure states: `Push_Gate_Failed` (inline gate non-zero), `Push_Failed` (`git push` non-zero, raw stderr surfaced), `Push_Deferred` (user chose "Stop — I'll push manually"). Inline gate body and `.githooks/pre-push` body must stay byte-equivalent; divergence is pinned by `tests/test_meso/test_auto_prompt_templates.py::TestMergePromptPushGate::test_hook_and_prompt_agree_on_gate_body` (which compares non-blank non-comment lines in both bodies and fails on drift in either direction) plus 3 supporting assertions on the upstream-first logic, the testmon fallback, and the prompt structure. Prompt: `src/deviate/prompts/commands/deviate-merge.md` (v2.3.0 → v2.4.0). Spec mirrors updated: `specs/DeviaTDD-architecture.md` (Merge bullet, new `**Push gate + opt-in push (v2.4.0)**` sub-bullet) and `specs/DeviaTDD-api.md` (new `**/deviate-merge push behavior (v2.4.0)**` entry under the `deviate merge` reference).
55
56
 
56
57
  ### Fixed
58
+ - **JUDGE prompt injection now uses the Judge-Feedback-stripped task card (GH-118).** `_build_auto_prompt("judge")` still reads the raw card via `_task_card_text`, then applies the same `_strip_judge_feedback` pass that `resolve_task_ac_tokens` already used for token resolution (#89). Prior-round `**Judge Feedback**` bullets and their continuation lines (including false ownership claims such as "AC-PLAN-003 belongs to a later task") no longer appear in `<task_card>`. Token resolution, `_append_judge_feedback`, and the COMPLETED evidence gate are unchanged. Specs: `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md`. Pinned by `tests/test_micro/test_judge.py::TestJudgePromptStripsJudgeFeedback`.
59
+ - **Ledger appenders no longer concatenate a new JSONL record onto a last line that lacks a trailing newline (GH-117).** `_append_record` and `_append_with_compound_key` insert a leading `\n` when the file is non-empty and does not already end in a newline, so `_read_ledger` cannot skip a fused BACKLOG+SPECIFIED line. `claim_issue` now writes through `append_issue_transition` instead of a raw `"a"` append. Pinned by `tests/test_state/test_ledger.py` and `tests/test_core/test_issues.py`.
60
+ - **JUDGE handover parse now recovers unescaped `"` inside evidence quote fields (GH-116).** `AgentBackend.parse_output` retries `yaml.safe_load` after rewriting broken `quote` / `test_quote` / `impl_quote` double-quoted scalars as `|` block scalars, so citations such as `assert "YAGNI" in text` and `== ["AC-PLAN-002"]` no longer raise `MalformedHandoverManifestError`. Well-formed YAML is unchanged. The auto judge prompt also prefers `|` block scalars when a quote contains `"`. Pinned by `tests/test_core/test_agent.py`.
61
+ - **A `no_failing_test` already-exists JUDGE `COMPLIANCE_PASS` now completes via `skip_refactor` instead of hard-crashing with `ROLLBACK_BOUNDARY_MISSING`.** On the RED `no_failing_test` adjudication route (`session.failure_kind == "no_failing_test"`, `session.red_commit_sha == ""`), `_apply_judge_verdict` coerces any PASS `next_action` to `skip_refactor` through `_NO_FAILING_TEST_FORWARD_ROUTES`, skips the unmatched-PASS AC-token citation rewrite, and relaxes only the `COMPLETED_EVIDENCE_MISSING` token check for that route — the declared-regression-files gate (`_require_tdd_declared_regression_files`) stays fail-closed, a wrong test still routes to `revert_before` for RED re-author, and a genuine test-bearing RED with a real RED commit still rolls back via `revert_to_red`. Previously every retry crashed because the mechanical gate rewrote the pass to `revert_to_red` with no RED commit to roll back to. Specs updated: `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md`. Pinned by `tests/test_micro/test_judge.py` and `tests/test_cli/test_micro.py`.
62
+ - **Feature worktrees now base on the local trunk, so locally-authored (unpushed) issues plan without failing.** `resolve_start_point` (`src/deviate/core/worktree.py`) previously preferred `origin/<base>` over local `<base>`, so a `deviate meso run` that discovered an issue committed only to local `main` created its worktree from a stale `origin/main` that lacked the issue's `specs/issues.jsonl` row and source file — the plan agent wrote `plan.md`, then `deviate plan post` failed with `ISSUE_NOT_FOUND <IDS>` and exited 1. Meso reads the issue ledger from the local trunk, so the worktree base now comes from local `<base>` first (falling back to `origin/<base>`, then `HEAD`). Pinned by `tests/test_core/test_worktree.py::TestCreateWorktree::test_resolve_start_point_prefers_local_base_over_stale_origin`.
63
+ - **`deviate meso run` / `deviate specify` auto-discovery now skips issues already claimed locally.** A claim writes the issue's `BACKLOG → SPECIFIED` ledger row in the new worktree and commits it on the `feat/{epic}/{issue}` feature branch, so the main checkout's `specs/issues.jsonl` still shows the issue as `BACKLOG`. `_discover_claimable_issue` therefore re-returned the same first issue to a second terminal on the same checkout (parallel run) instead of claiming a different one; the origin-branch lock only caught issues whose claim branch had been `git push`-ed. Discovery now also skips a candidate whose `feat/{epic}/{issue}` branch already has a local worktree (treated as claimed here), in both local and default mode, mirroring the signal `_try_claim_issue` already uses for `ALREADY_CLAIMED_LOCAL`. Two terminals on one checkout can now claim two different BACKLOG issues in parallel. Pinned by `tests/test_meso/test_meso_orchestration.py::TestDiscoverClaimableIssue::test_skips_issues_with_local_worktree`.
57
64
  - **`_append_judge_feedback` writes under the rejected task card, and the evidence gate keeps judge `train_feedback` / `violations` (GH-102).** The appender locates the card with `_TASK_BULLET_HEAD_RE` exact-id match (same regex as `_read_judge_feedback_from_tasks_md`) and stops at the next task or a `##` / `###` phase heading, so a Phase 1 `TSK-004-01` rejection no longer drops `**Judge Feedback**` under the next phase `### Tasks` / `TSK-004-02`. When the evidence gate rewrites unmatched PASS, persist the judge's own `train_feedback` / `violations` (after the GH-103 citation strip) instead of replacing them with the generic `JUDGE evidence is missing...` string. Pinned by `tests/test_micro/test_judge.py::TestJudgeFeedbackLogging::test_append_judge_feedback_stays_under_rejected_card_across_phase_header` and `TestTddJudgeEvidenceGate::test_rewritten_pass_keeps_judge_train_feedback` / `::test_rewritten_pass_keeps_judge_violations`.
58
65
  - **JUDGE injected diff walks past docs-feedback `red_commit_sha` to the real RED commit (GH-88, GH-90).** `_maybe_advance_red_sha_past_feedback` still advances the session SHA onto `docs(...): add judge feedback for retry` so GREEN entry and `revert_to_red` keep a valid TRAIN boundary. `_assemble_judge_injected_diff` now resolves the diff base via `_resolve_judge_diff_base`, walking back through those feedback subjects to the RED-phase failing-test commit and running `git diff {red_sha}^..HEAD`. The evidence gate no longer false-rejects with `test_path is not in the injected diff` or missing AC tokens solely because the RED test dropped out of a docs-only range. Pinned by `tests/test_micro/test_judge.py::TestJudgeDiffBaseWalksPastFeedback`.
59
66
  - **`resolve_task_ac_tokens` no longer treats runner-appended `**Judge Feedback**` as required AC tokens (GH-89).** Card-text fallback still follows first-hit order (`acceptance_criteria` `criterion_id`s, else the task card, else none) but now drops `**Judge Feedback**` bullets and their continuation lines before scanning for `AC-PLAN-NNN`. A feedback line quoting `AC-PLAN-001` does not add that token when the card body does not name it. `_task_card_text` still returns the full card. Pinned by `tests/test_core/test_judge_evidence.py::TestResolveTaskAcTokens::test_judge_feedback_quote_does_not_add_token`.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: deviatdd
3
- Version: 2.22.1
3
+ Version: 2.23.0
4
4
  Summary: DeviaTDD CLI — agent orchestration framework
5
5
  Project-URL: Homepage, https://github.com/wernerbisschoff/deviatdd
6
6
  Project-URL: Repository, https://github.com/wernerbisschoff/deviatdd
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "deviatdd"
3
- version = "2.22.1"
3
+ version = "2.23.0"
4
4
  description = "DeviaTDD CLI — agent orchestration framework"
5
5
  readme = "README.md"
6
6
  license = "MIT"
@@ -436,7 +436,10 @@ accepts `--json` (emit JSON contract to stdout) and `--quiet` (suppress output).
436
436
  claims that specific issue. With **no argument**, auto-discovers the next claimable
437
437
  BACKLOG issue via `_discover_claimable_issue()` (the same discovery `deviate meso run`
438
438
  uses) and claims it. Default discovery skips issues whose `feat/{epic}/{issue}` branch
439
- already exists on remote (treated as claimed elsewhere). Local mode does not skip those
439
+ already exists on remote (treated as claimed elsewhere). In any mode it also skips
440
+ issues whose `feat/{epic}/{issue}` branch already has a local worktree (treated as claimed
441
+ here) — that local-worktree guard is what lets two parallel terminals claim two different
442
+ BACKLOG issues without re-claiming the same one. Local mode does not skip those
440
443
  origin branches. Stops after the worktree is created and the claim is committed — does
441
444
  NOT advance session state and does NOT run plan or tasks. To continue, run
442
445
  ``deviate plan pre`` or invoke the ``/deviate-plan`` slash command inside the new
@@ -594,9 +597,16 @@ accepts `--json` and `--quiet`. `pre` emits a JSON contract describing the envir
594
597
  in the injected ``<diff>`` or HEAD (constitution §3 Testing Protocols; §5 Definition of Done).
595
598
  Empty ``files`` / ``test_file`` is a RED defect (``PhaseFailedError``); the ledger writes no
596
599
  COMPLETED row. JUDGE ``skip_refactor`` / bare ``COMPLIANCE_PASS`` keeps those declared tests
597
- on disk via ``_restore_worktree_to_baseline(..., keep_paths=declared)``. A declared path
598
- missing from the snapshot rewrites PASS to ``revert_before`` / ``revert_to_red``. JUDGE still
599
- rules a wrong test as ``revert_before`` so RED re-authors a genuinely failing test. EXECUTE,
600
+ on disk via ``_restore_worktree_to_baseline(..., keep_paths=declared)``. On the already-exists
601
+ route (``session.red_commit_sha == ''``), a ``no_failing_test`` ``COMPLIANCE_PASS`` with any
602
+ ``next_action`` completes via ``skip_refactor`` even when the evidence cites only part of the
603
+ task's ``AC-PLAN-NNN`` tokens: ``_apply_judge_verdict`` skips the unmatched-PASS rewrite for
604
+ this route, and ``_require_tdd_completed_evidence`` relaxes the AC-token citation check while
605
+ keeping the declared regression-path presence gate. ``ROLLBACK_BOUNDARY_MISSING`` applies only
606
+ to a genuine TDD ``revert_to_red`` with an empty ``red_commit_sha`` — never on the
607
+ already-exists pass path. A declared path missing from the snapshot rewrites PASS to
608
+ ``revert_before`` / ``revert_to_red``. JUDGE still rules a wrong test as ``revert_before`` so
609
+ RED re-authors a genuinely failing test. EXECUTE,
600
610
  IMMEDIATE, and DIRECT stay ungated by this files rule. After ``revert_before`` / cycle
601
611
  ``no_failing_test_adjudicated``, the next ``INVOKE_AGENT`` is RED, or the loop raises
602
612
  ``TRAIN_EXHAUSTED`` / ``PhaseFailedError``. It never invokes GREEN while
@@ -946,7 +956,7 @@ accepts `--json` and `--quiet`. `pre` emits a JSON contract describing the envir
946
956
  populated `parse_errors` list and
947
957
  `HandoverManifest.is_success` returns `False` so existing
948
958
  `manifest.status.upper() in (...)` success gates keep rejecting them.
949
- (6) **Stricter mapping fallback** — the `_YAML_MAPPING_START_RE` fallback (`src/deviate/core/agent.py`) routes the candidate text through a `_looks_like_manifest` helper that requires `yaml.safe_load(candidate)` to return a `dict` with at least 2 keys before accepting. Single-key dicts (e.g. a stray `Status: complete` line in a JUDGE verdict with a verification matrix) look like prose, not manifests; the fallback now rejects them and the parser raises `MalformedHandoverManifestError` with the existing "No YAML handover manifest detected in agent output" hint. Multi-key partial dicts still flow through to schema recovery unchanged, so the existing `test_missing_phase_and_status_recover_as_unknown` contract is preserved.
959
+ (6) **Stricter mapping fallback** — the `_YAML_MAPPING_START_RE` fallback (`src/deviate/core/agent.py`) routes the candidate text through a `_looks_like_manifest` helper that requires `yaml.safe_load(candidate)` to return a `dict` with at least 2 keys before accepting. Single-key dicts (e.g. a stray `Status: complete` line in a JUDGE verdict with a verification matrix) look like prose, not manifests; the fallback now rejects them and the parser raises `MalformedHandoverManifestError` with the existing "No YAML handover manifest detected in agent output" hint. Multi-key partial dicts still flow through to schema recovery unchanged, so the existing `test_missing_phase_and_status_recover_as_unknown` contract is preserved. (6b) **Unescaped evidence-quote recovery (GH-116)** — when `yaml.safe_load` rejects a handover because an evidence `quote` / `test_quote` / `impl_quote` double-quoted scalar embeds raw `"`, `parse_output` rewrites those lines as `|` block scalars and reloads. Well-formed YAML is unchanged. Truly malformed YAML still raises `MalformedHandoverManifestError`. Verdict and evidence semantics are unchanged.
950
960
  (7) **Lean Pi spawn** — `AgentBackend.invoke` appends a lean tool policy after the existing Pi transport prefix. Print mode keeps `BACKEND_COMMANDS["pi"]` as `pi -p` (AC-009-07). RPC keeps `PI_RPC_COMMAND` as `pi --mode rpc --no-session` (AC-009-10). The helper `_pi_lean_flags` then adds `--no-extensions`, `--tools read,bash,edit,write`, and `--no-skills`. When `.pi/skills/deviatdd/SKILL.md` exists under the invoke `cwd` (or `Path.cwd()`), it also adds `--skill` to that relative path. A missing skill file keeps the four coding tools. The argv omits `--no-tools` and `--no-builtin-tools`. Non-Pi backends skip these flags.
951
961
  (8) **Schema-rejection fail-fast** — `_invoke_streaming`, `_invoke_blocking`, and `_invoke_rpc_blocking` scan each stderr and stdout line. The first line that contains `tool_count_limit` or `unsupported_tool_schema` kills the child. The helper raises `AgentSubprocessError` whose message carries those tokens. This path does not wait for `STREAM_STALL_TIMEOUT_SECONDS` (900s). It does not start the 30s timeout retry. It does not start the `EmptyOutputError` manifest retry. Schema tokens do not reset the stall clock. Stderr stays diagnostic for stall liveness (ISS-ADH-025). `_invoke_agent` logs `AGENT_ERROR` with the exception text. `_raise_schema_limit_phase_error` then raises `PhaseFailedError` so `deviate micro run` RED, GREEN, and REFACTOR include the tokens. The operator does not see only `agent returned no manifest`. EXECUTE stall stays 3600s (GH-53).
952
962
  * **GREEN Stub-PASS Guard (REMOVED):** An earlier revision of this spec
@@ -1483,7 +1493,11 @@ All state transitions are append-only. No existing line is ever modified or over
1483
1493
  - `append_issue_transition()`: Idempotent on `(issue_id, status)` compound key
1484
1494
  - `append_task_transition()`: Idempotent on `(id, status)` compound key
1485
1495
  - `_append_record()` / `_append_with_compound_key()`: Use `fcntl.flock` for file-level
1486
- locking on platforms that support it
1496
+ locking on platforms that support it. If the ledger is non-empty and the last
1497
+ line has no trailing newline, a leading `\n` is written before the new record
1498
+ so two JSON objects never share a line. Every successful write leaves a trailing
1499
+ newline. `claim_issue` writes through `append_issue_transition` (not a raw `"a"`
1500
+ append).
1487
1501
  - Canonical state: Issues derived bottom-up (latest entry per `issue_id`); tasks derived
1488
1502
  sequentially (latest entry per `(id, status)` compound key)
1489
1503
 
@@ -1540,7 +1554,7 @@ the action. EXECUTE `_run_execute_phase` and IMMEDIATE judge stay ungated.
1540
1554
 
1541
1555
  **Empty-diff sign-off:** `proceed_to_refactor_no_diff` (`src/deviate/cli/micro.py::_run_judge_phase`) is the forward-route escape for slices whose production-code scope is intrinsically nil — RED-only deliverable, fixture file, generated types, doc-only slice, or any task whose `failure_kind: mechanical` rationale asserts "no production code expected." The TDD evidence gate still requires a dirty-diff `test_quote` and omits `impl_quote`. The JUDGE-side responsibility is to emit the action on a `COMPLIANCE_PASS` verdict when the in-scope rationale is valid but the production diff cannot grow. The action lands the task at REFACTOR's no-op commit + COMPLETED transition in one step; unmatched empty-GREEN PASS does not COMPLETE.
1542
1556
 
1543
- **TDD mechanical evidence gate:** `HandoverManifest.evidence` is a first-class list of nested citations (`ac`, `test_path`, `test_quote`, `impl_path`, `impl_quote`) in `src/deviate/core/agent.py`. After `_coerce_judge_action`, TDD `_run_judge_phase` (`src/deviate/cli/micro.py`) resolves this task's required `AC-PLAN-NNN` tokens via `resolve_task_ac_tokens` (`src/deviate/core/judge_evidence.py`) and passes that list as `required_tokens` to `evaluate_judge_evidence`. First hit wins: non-empty `TaskRecord.acceptance_criteria` `criterion_id`s; else `AC-PLAN-NNN` tokens named in this task's `tasks.md` card after dropping `**Judge Feedback**` bullets and their continuation lines; else no AC tokens. The gate does not fall back to every token in `<authoritative_acceptance_contract source="plan.md">`. Omitting a later-shard plan token is legal at JUDGE. Auto and manual judge prompts require `evidence` only for the resolved task tokens. Quotes must copy from the already-built `<diff>` (`git diff <red>^..HEAD` where `<red>` is `_resolve_judge_diff_base(session.red_commit_sha)` — the RED-phase failing-test commit after walking back through `docs(...): add judge feedback for retry` subjects — plus dirty `git diff HEAD` and untracked `--no-index` hunks) or allowed HEAD files. ISS-ADH-020 quote checks still apply to that task set: missing this-task tokens, empty quotes, hallucinated paths, quotes below the uniqueness floor (≥ 12 non-whitespace characters, or the full added line if shorter), or quotes that are not exact substrings of the named file hunk rewrite the action to `revert_to_red` with runner-authored feedback in the `JUDGE_AGENT_NO_FEEDBACK` family. When the judge already emitted `train_feedback` or `violations`, persist that text (after the GH-103 citation strip) instead of replacing it with the generic missing-evidence string (GH-102). The task does not COMPLETE. `skip_refactor` on the already-exists path may quote HEAD file contents for this-task tokens; a named test file absent on disk fails. On a test-bearing TDD already-exists claim, every declared `files` / `test_file` path (and evidence `test_path`) must appear in `_assemble_judge_injected_diff` or `_evidence_head_contents`. The membership check runs even when the resolved set has no `AC-PLAN-*` tokens. Empty declared files remain a RED defect, not a COMPLETE. Tasks with no resolved `AC-PLAN-*` tokens may emit empty evidence quotes, but they still need named present test paths. `COMPLIANCE_VIOLATION` skips the quote gate. After the gate returns no feedback and the action is a completion path (`skip_refactor` / bare `COMPLIANCE_PASS` / post-REFACTOR complete / adjudicated already-exists), `_append_status_transition(..., "COMPLETED")` copies the validated `HandoverManifest.evidence` onto that COMPLETED `TaskRecord` as `evidence.items` and stamps `red` / `green` / `head` from `session.red_commit_sha` and `HEAD` (GH-84). TDD complete fail-closes when the injected plan contract has `AC-PLAN-NNN` tokens and the persisted bundle is missing or does not cover them — the same `evaluate_judge_evidence` matcher, with `use_head=True` so quotes resolve against HEAD at the COMPLETED write. Plans with no `AC-PLAN-*` tokens may complete with empty evidence. EXECUTE, IMMEDIATE, and DIRECT judge paths stay ungated.
1557
+ **TDD mechanical evidence gate:** `HandoverManifest.evidence` is a first-class list of nested citations (`ac`, `test_path`, `test_quote`, `impl_path`, `impl_quote`) in `src/deviate/core/agent.py`. After `_coerce_judge_action`, TDD `_run_judge_phase` (`src/deviate/cli/micro.py`) resolves this task's required `AC-PLAN-NNN` tokens via `resolve_task_ac_tokens` (`src/deviate/core/judge_evidence.py`) and passes that list as `required_tokens` to `evaluate_judge_evidence`. First hit wins: non-empty `TaskRecord.acceptance_criteria` `criterion_id`s; else `AC-PLAN-NNN` tokens named in this task's `tasks.md` card after dropping `**Judge Feedback**` bullets and their continuation lines; else no AC tokens. Auto `_build_auto_prompt("judge")` injects that same stripped card as `<task_card source="tasks.md">` (GH-118); `_task_card_text` still returns the raw card for token resolution, GREEN `<persisted_judge_feedback>`, and file-list parsing. The gate does not fall back to every token in `<authoritative_acceptance_contract source="plan.md">`. Omitting a later-shard plan token is legal at JUDGE. Auto and manual judge prompts require `evidence` only for the resolved task tokens. Quotes must copy from the already-built `<diff>` (`git diff <red>^..HEAD` where `<red>` is `_resolve_judge_diff_base(session.red_commit_sha)` — the RED-phase failing-test commit after walking back through `docs(...): add judge feedback for retry` subjects — plus dirty `git diff HEAD` and untracked `--no-index` hunks) or allowed HEAD files. ISS-ADH-020 quote checks still apply to that task set: missing this-task tokens, empty quotes, hallucinated paths, quotes below the uniqueness floor (≥ 12 non-whitespace characters, or the full added line if shorter), or quotes that are not exact substrings of the named file hunk rewrite the action to `revert_to_red` with runner-authored feedback in the `JUDGE_AGENT_NO_FEEDBACK` family. When the judge already emitted `train_feedback` or `violations`, persist that text (after the GH-103 citation strip) instead of replacing it with the generic missing-evidence string (GH-102). The task does not COMPLETE. `skip_refactor` on the already-exists path may quote HEAD file contents for this-task tokens; a named test file absent on disk fails. On a test-bearing TDD already-exists claim, every declared `files` / `test_file` path (and evidence `test_path`) must appear in `_assemble_judge_injected_diff` or `_evidence_head_contents`. The membership check runs even when the resolved set has no `AC-PLAN-*` tokens. Empty declared files remain a RED defect, not a COMPLETE. Tasks with no resolved `AC-PLAN-*` tokens may emit empty evidence quotes, but they still need named present test paths. `COMPLIANCE_VIOLATION` skips the quote gate. After the gate returns no feedback and the action is a completion path (`skip_refactor` / bare `COMPLIANCE_PASS` / post-REFACTOR complete / adjudicated already-exists), `_append_status_transition(..., "COMPLETED")` copies the validated `HandoverManifest.evidence` onto that COMPLETED `TaskRecord` as `evidence.items` and stamps `red` / `green` / `head` from `session.red_commit_sha` and `HEAD` (GH-84). TDD complete fail-closes when the injected plan contract has `AC-PLAN-NNN` tokens and the persisted bundle is missing or does not cover them — the same `evaluate_judge_evidence` matcher, with `use_head=True` so quotes resolve against HEAD at the COMPLETED write. Plans with no `AC-PLAN-*` tokens may complete with empty evidence. EXECUTE, IMMEDIATE, and DIRECT judge paths stay ungated.
1544
1558
 
1545
1559
  **Feedback-commit timeout:** The `revert_to_red` step's "append a feedback
1546
1560
  commit past RED" runs `_commit_judge_feedback_and_advance`
@@ -285,9 +285,9 @@ to invoke, but model selection is delegated to the calling environment.
285
285
  * **Layer discipline:** GREEN's only invariant is "make the RED test pass via the library/API surface declared in scope." It does NOT make scope, spec-drift, or HITL-routing judgments — those belong to JUDGE. When a RED test cannot be satisfied within GREEN's mechanical scope, GREEN emits `status: FAILURE` with a concrete `rationale:` naming the test path and why; `status: "ERROR"` is reserved strictly for tool/orchestration failure. The runner's `_is_hitl_escalation` is a narrow defensive fallback that ONLY promotes structured `contract_drift` / `escalates_to` / `hitl_options` dict keys to `HITL_PENDING` — loose-string `error_kind` discriminators and free-form scope-conflict text do NOT trigger HITL escalation.
286
286
  * **Mechanical Failure → JUDGE Routing:** When GREEN emits `status: FAILURE` with a concrete `rationale:` (the mechanical scope-boundary case above), the runner routes control to JUDGE instead of raising `PhaseFailedError`. `_run_green_phase` sets `session.train_feedback = rationale` + `session.failure_kind = "mechanical"` and returns the session; `_run_judge_phase` injects a `<failure_kind>mechanical</failure_kind>` discriminator block into the JUDGE prompt that instructs the agent to emit `verdict: COMPLIANCE_PASS` + `next_action: proceed_to_refactor_no_diff` (when the slice is intrinsically RED-only and REFACTOR's no-op commit + COMPLETED transition is the right termination) OR `verdict: COMPLIANCE_VIOLATION` + one of three `next_action` values (`revert_before` / `revert_to_red` / `skip_refactor`) instead of attempting to satisfy the test itself. This closes the loop where mechanical FAILURE (e.g. slice-scope conflict, CLI-surface-out-of-sco…
287
287
  * **Test-Defect Failure → JUDGE Routing:** A second routable failure class, parallel to mechanical but pre-decided. When GREEN observes that the RED test itself is wrong (it asserts behavior the spec does not require, exercises the wrong abstraction, or encodes an assumption that contradicts spec/data-model), GREEN emits `status: FAILURE` with a concrete `rationale:` citing the FR/AC the test contradicts, plus `failure_kind: test_defect` on the manifest (`HandoverManifest.failure_kind: Literal["mechanical", "test_defect", "already_satisfied"] | None` in `src/deviate/core/agent.py` — session mirrors the discriminator as `SessionState.failure_kind: Literal["", "mechanical", "test_defect", "no_failing_test"]` in `src/deviate/state/config.py`). `_run_green_phase` reads the manifest's discriminator, sets `session.failure_kind = "test_defect"`, and routes to JUDGE. `_run_judge_phase` injects a `<failure_kind>test_defect</failure_kind>` discriminator block that pre-decides the routing — `verdict: COMPLIANCE_VIOLATION` + `next_action: revert_before` (re-run RED with the GREEN rationale as feedback) — because test defect has only one sensible outcome: the test itself must be re-authored. The discriminator intentionally narrows the JUDGE routing vocabulary compared to mechanical (`revert_to_red` / `skip_refactor` / `proceed_to_refactor_no_diff` are NOT options in the test_defect block). Default `manifest.failure_kind = None` falls through to `"mechanical"` in the runner to preserve prior behavior.
288
- * **RED No-Failing-Test → JUDGE Adjudication (RED → JUDGE direct route):** When RED completes but its test command exits 0 (all tests passed), collects no tests (pytest exit 5), or resolves to no command at all (returncode 127 from `_run_test_cmd`), `_run_red_phase` does NOT raise a raw `PhaseFailedError` and does NOT let GREEN run against a vacuous test. It calls `_adjudicate_red_no_failing_test` (`src/deviate/cli/micro.py`), which sets `session.failure_kind = "no_failing_test"`, injects a `<failure_kind>no_failing_test</failure_kind>` discriminator block into the JUDGE prompt, and dispatches `_run_judge_phase(...)`. The judge diff spans the uncommitted RED test (the `red_baseline` parameter makes `_run_judge_phase` skip the `RED→HEAD` committed diff and surface the agent's uncommitted test through the dirty-parts scan) so JUDGE reviews what the agent actually wrote. On a test-bearing TDD task, `_require_tdd_declared_regression_files` requires a non-empty `files` set and/or `test_file` before this COMPLETE route (constitution §3 Testing Protocols; §5 Definition of Done). Empty declared files raise `PhaseFailedError`. JUDGE decides between two outcomes: `verdict: COMPLIANCE_PASS` + `next_action: skip_refactor` (or a bare PASS verdict) when the required behavior already exists and every declared regression path sits in the injected `<diff>` or `_evidence_head_contents` — the runner then calls `_restore_worktree_to_baseline(..., keep_paths=declared)` so those tests stay on disk, and marks the task COMPLETED; a declared path missing from that snapshot rewrites PASS to `revert_before` / `revert_to_red` with runner-authored feedback and writes no COMPLETED row; or `verdict: COMPLIANCE_VIOLATION` + `next_action: revert_before` (forced by the `_coerce_judge_action` runner-level override for `failure_kind` `no_failing_test`/`test_defect`) when the test is wrong — the runner resets to the RED baseline and re-dispatches RED so a fresh genuinely-failing test is authored. EXECUTE, IMMEDIATE, and DIRECT stay ungated by the files rule. The next `INVOKE_AGENT` is RED, or the loop raises `TRAIN_EXHAUSTED` / `PhaseFailedError`. It never invokes GREEN while `session.red_commit_sha` is empty. `--no-judge` makes the no-failing-test outcome a hard failure (adjudication disabled). Test discovery is language-agnostic: the RED gate no longer globs Python `tests/**/test_*.py` (`_find_test_files`); the project's own convention is honored through the resolved test command (e.g. `mix test` collecting `test/**/*_test.exs` for Elixir/Phoenix), so a correctly authored non-Python test is neither rejected up front nor silently conflated with "no test framework" — a project with no resolvable test command routes to the same JUDGE adjudication instead of failing. This closes the prior stall where an empty `_find_test_files` let RED silently commit a vacuous 'failing test' and GREEN die in `TRAIN_EXHAUSTED`. The RED agent may steer the adjudication by declaring `failure_kind: already_satisfied` (behavior exists) with a non-empty `files` set and/or `test_file`, or `failure_kind: test_defect` (test wrong) with a `rationale` on its handover manifest. A passing suite with no named test files is not a COMPLETE.
288
+ * **RED No-Failing-Test → JUDGE Adjudication (RED → JUDGE direct route):** When RED completes but its test command exits 0 (all tests passed), collects no tests (pytest exit 5), or resolves to no command at all (returncode 127 from `_run_test_cmd`), `_run_red_phase` does NOT raise a raw `PhaseFailedError` and does NOT let GREEN run against a vacuous test. It calls `_adjudicate_red_no_failing_test` (`src/deviate/cli/micro.py`), which sets `session.failure_kind = "no_failing_test"`, injects a `<failure_kind>no_failing_test</failure_kind>` discriminator block into the JUDGE prompt, and dispatches `_run_judge_phase(...)`. The judge diff spans the uncommitted RED test (the `red_baseline` parameter makes `_run_judge_phase` skip the `RED→HEAD` committed diff and surface the agent's uncommitted test through the dirty-parts scan) so JUDGE reviews what the agent actually wrote. On a test-bearing TDD task, `_require_tdd_declared_regression_files` requires a non-empty `files` set and/or `test_file` before this COMPLETE route (constitution §3 Testing Protocols; §5 Definition of Done). Empty declared files raise `PhaseFailedError`. JUDGE decides between two outcomes: `verdict: COMPLIANCE_PASS` + `next_action: skip_refactor` (or a bare PASS verdict, or any `next_action` on the already-exists route) when the required behavior already exists and every declared regression path sits in the injected `<diff>` or `_evidence_head_contents` — `_apply_judge_verdict` then coerces the action to `skip_refactor`, skips the `_rewrite_unmatched_tdd_pass` AC-token citation rewrite for this route, and `_require_tdd_completed_evidence` relaxes the AC-token citation check while retaining the declared-path presence gate, so partial evidence completes instead of rewriting to `revert_to_red`; the runner then calls `_restore_worktree_to_baseline(..., keep_paths=declared)` so those tests stay on disk, and marks the task COMPLETED. `_require_revert_to_red_boundary` is never invoked on this already-exists pass path, so `ROLLBACK_BOUNDARY_MISSING` applies only to a genuine TDD `revert_to_red` with an empty `red_commit_sha`. A declared path missing from that snapshot still fails closed via `_require_tdd_declared_regression_files` (rewriting PASS to `revert_before` / `revert_to_red` with runner-authored feedback and writing no COMPLETED row); or `verdict: COMPLIANCE_VIOLATION` + `next_action: revert_before` (forced by the `_coerce_judge_action` runner-level override for `failure_kind` `no_failing_test`/`test_defect`) when the test is wrong — the runner resets to the RED baseline and re-dispatches RED so a fresh genuinely-failing test is authored. EXECUTE, IMMEDIATE, and DIRECT stay ungated by the files rule. The next `INVOKE_AGENT` is RED, or the loop raises `TRAIN_EXHAUSTED` / `PhaseFailedError`. It never invokes GREEN while `session.red_commit_sha` is empty. `--no-judge` makes the no-failing-test outcome a hard failure (adjudication disabled). Test discovery is language-agnostic: the RED gate no longer globs Python `tests/**/test_*.py` (`_find_test_files`); the project's own convention is honored through the resolved test command (e.g. `mix test` collecting `test/**/*_test.exs` for Elixir/Phoenix), so a correctly authored non-Python test is neither rejected up front nor silently conflated with "no test framework" — a project with no resolvable test command routes to the same JUDGE adjudication instead of failing. This closes the prior stall where an empty `_find_test_files` let RED silently commit a vacuous 'failing test' and GREEN die in `TRAIN_EXHAUSTED`. The RED agent may steer the adjudication by declaring `failure_kind: already_satisfied` (behavior exists) with a non-empty `files` set and/or `test_file`, or `failure_kind: test_defect` (test wrong) with a `rationale` on its handover manifest. A passing suite with no named test files is not a COMPLETE.
289
289
  * **JUDGE / TRAIN (The Compliance Gate) — with Green → Judge → Green loop:**
290
- * **The Judge:** The CLI evaluates the committed RED-parent-to-HEAD diff against `spec.md` for invariant/security violations. When `session.red_commit_sha` has advanced onto a `docs(...): add judge feedback for retry` commit, `_assemble_judge_injected_diff` calls `_resolve_judge_diff_base` to walk back to the RED-phase failing-test commit before running `git diff {red_sha}^..HEAD`, so RED tests stay visible to the evidence gate (GH-88 / GH-90). If GREEN tests failed before the implementation commit, `_run_judge_phase` also appends staged/unstaged `git diff HEAD` output and per-file `git diff --no-index /dev/null <path>` output for untracked files, so JUDGE assesses the retained implementation rather than a false RED-only snapshot. This judge operates in a clean, zero-shared-history session to break recursive subjectivity. A `deviate-judge` skill (loaded from `_SKILL_NAMES["JUDGE"]`) guides the agent through supplementary compliance evaluation. TDD JUDGE evidence is task-scoped: `_run_judge_phase` resolves required `AC-PLAN-NNN` tokens via `resolve_task_ac_tokens` (non-empty `acceptance_criteria` `criterion_id`s, else this task's `tasks.md` card minus `**Judge Feedback**` bullets and their continuation lines, else none) and does not fall back to every token in `plan.md`. TDD `_run_judge_phase` then rejects `COMPLIANCE_PASS` unless `HandoverManifest.evidence` quotes for those task tokens match the injected diff (or HEAD file contents on the already-exists `skip_refactor` edge) and every declared `files` / `test_file` path exists in that snapshot. Unmatched PASS rewrites to `revert_to_red` / `revert_before`. On a forward PASS the runner stashes those citations on `SessionState.validated_evidence` (transient only) and writes them onto the COMPLETED `tasks.jsonl` row as `TaskRecord.evidence` (`items` plus `red` / `green` / `head` SHAs) so `deviate inspect tasks show` can display the proof after the session is gone (GH-84). TDD COMPLETE with `AC-PLAN-NNN` tokens refuse when that bundle is missing or does not cover the tokens (`evaluate_judge_evidence`). `.deviate/` session files are not the proof store. Earlier RED/GREEN/JUDGE rows stay lean. Persist the judge's own `train_feedback` / `violations` when present (GH-103 citation strip still applies); attach the generic missing-evidence string only when the judge emitted none (GH-102). `_append_judge_feedback` keys the card with `_TASK_BULLET_HEAD_RE` exact-id match and inserts under that card only (does not walk past a later phase heading). Gate 3 review (`deviate review`) owns plan-wide AC coverage. EXECUTE, IMMEDIATE, and DIRECT judge paths stay ungated.
290
+ * **The Judge:** The CLI evaluates the committed RED-parent-to-HEAD diff against `spec.md` for invariant/security violations. When `session.red_commit_sha` has advanced onto a `docs(...): add judge feedback for retry` commit, `_assemble_judge_injected_diff` calls `_resolve_judge_diff_base` to walk back to the RED-phase failing-test commit before running `git diff {red_sha}^..HEAD`, so RED tests stay visible to the evidence gate (GH-88 / GH-90). If GREEN tests failed before the implementation commit, `_run_judge_phase` also appends staged/unstaged `git diff HEAD` output and per-file `git diff --no-index /dev/null <path>` output for untracked files, so JUDGE assesses the retained implementation rather than a false RED-only snapshot. This judge operates in a clean, zero-shared-history session to break recursive subjectivity. A `deviate-judge` skill (loaded from `_SKILL_NAMES["JUDGE"]`) guides the agent through supplementary compliance evaluation. TDD JUDGE evidence is task-scoped: `_run_judge_phase` resolves required `AC-PLAN-NNN` tokens via `resolve_task_ac_tokens` (non-empty `acceptance_criteria` `criterion_id`s, else this task's `tasks.md` card minus `**Judge Feedback**` bullets and their continuation lines, else none) and does not fall back to every token in `plan.md`. Auto `_build_auto_prompt("judge")` injects that same stripped card as `<task_card source="tasks.md">` so prior-round Judge-Feedback prose cannot bias the next judge (GH-118); `_task_card_text` remains raw. TDD `_run_judge_phase` then rejects `COMPLIANCE_PASS` unless `HandoverManifest.evidence` quotes for those task tokens match the injected diff (or HEAD file contents on the already-exists `skip_refactor` edge) and every declared `files` / `test_file` path exists in that snapshot. Unmatched PASS rewrites to `revert_to_red` / `revert_before`. On a forward PASS the runner stashes those citations on `SessionState.validated_evidence` (transient only) and writes them onto the COMPLETED `tasks.jsonl` row as `TaskRecord.evidence` (`items` plus `red` / `green` / `head` SHAs) so `deviate inspect tasks show` can display the proof after the session is gone (GH-84). TDD COMPLETE with `AC-PLAN-NNN` tokens refuse when that bundle is missing or does not cover the tokens (`evaluate_judge_evidence`). `.deviate/` session files are not the proof store. Earlier RED/GREEN/JUDGE rows stay lean. Persist the judge's own `train_feedback` / `violations` when present (GH-103 citation strip still applies); attach the generic missing-evidence string only when the judge emitted none (GH-102). `_append_judge_feedback` keys the card with `_TASK_BULLET_HEAD_RE` exact-id match and inserts under that card only (does not walk past a later phase heading). Gate 3 review (`deviate review`) owns plan-wide AC coverage. EXECUTE, IMMEDIATE, and DIRECT judge paths stay ungated.
291
291
  * **The Train (Green → Judge → Green loop):** On `COMPLIANCE_VIOLATION` or test failure, the CLI safely resets without destroying task progress. The JUDGE phase honors `HandoverManifest.next_action` (see [specs/DeviaTDD-api.md](./DeviaTDD-api.md) for the routing table). The five routes:
292
292
  1. **`revert_before`** — discard this task's GREEN **and** its RED. `_resolve_pre_red_sha()` derives the SHA from `red_commit_sha^` (defended by a subject-match regex on the parent's commit message; logs `PRE_RED_AMBIGUOUS` when the parent isn't a RED-phase convention). The pre-RED anchor is threaded into `_execute_rollback(boundary_sha=<pre_red>, task_id=<tid>, attempt=<rollback_attempts>)`; the agent's pre-reset HEAD is captured on the per-attempt recovery ref `tmp/deviate-agent-work/<sanitized-task-id>/attempt-<rollback_attempts>` so a parent SIGTERM between `git reset` and `git clean` doesn't strand the discarded commit. `session.red_commit_sha` is cleared (the boundary was discarded) and `pending_judge_action = "revert_before"`. `force_transition_to("RED")` so the task retries from scratch. `_escalate_to_new_red` consumes `revert_before` only after a RED-phase `red_commit_sha` lands. While that SHA is empty, the next `INVOKE_AGENT` stays RED (or the loop raises `TRAIN_EXHAUSTED` / `PhaseFailedError`). **No implicit fallback**: if `_resolve_pre_red_sha` returns empty AND `session.red_commit_sha` is empty, the runner raises `PhaseFailedError("ROLLBACK_BOUNDARY_MISSING ...")` rather than fall back to `HEAD~1`. When `pre_red` is unresolvable but `session.red_commit_sha` is known, that cached SHA is used as the explicit boundary.
293
293
  **Runner-level override on `failure_kind=test_defect`:** when `session.failure_kind == "test_defect"` and JUDGE emits `COMPLIANCE_VIOLATION`, `_coerce_judge_action` (`src/deviate/cli/micro.py`) forces `next_action="revert_before"` regardless of what the JUDGE manifest declared or omitted. The override reflects a contract invariant: when the RED test itself is wrong, looping back to GREEN with the same test is futile. `_run_tdd_cycle` honours the resulting `pending_judge_action == "revert_before"` by escalating now. It resets `green_attempts` to 0, increments `red_attempts`, and persists both on `.deviate/session.json`. It then dispatches `_run_red_phase(task, ..., bypass_phase_done=True)` so a fresh RED record appends to the append-only ledger (never rewriting prior entries). The retry RED prompt receives a short `previous cycle failed because …` note, not the raw GREEN dump. `TRAIN_EXHAUSTED` prints only after three RED escalates. The override is silenced on `COMPLIANCE_PASS` verdicts; JUDGE's outcome remains authoritative when the implementation is sound. The override preserves the legacy resolution for `revert_to_red`, `continue_refactor`, `proceed_to_refactor_no_diff`, and `skip_refactor` — only the test_defect case is forced.
@@ -689,7 +689,7 @@ The orchestrator must maintain and enforce these structural constraints across a
689
689
 
690
690
  2. **The Scope Audit Law:** When entering or running the `GREEN` execution phase, the system checks for unauthorized changes to test, spec, and config directories. Protected files are reverted via `git restore <filepath>`. The JUDGE phase (`deviate judge pre`) additionally performs compliance verification by detecting changes to protected modules declared in `spec.md` `Module:` lines. Complements the GREEN stub-PASS guard in `_run_green_phase` (see `DeviaTDD-api.md` § GREEN Stub-PASS Guard): scope rejects writes the agent shouldn't have made; the stub-PASS guard rejects passes the agent shouldn't have emitted.
691
691
 
692
- 3. **Append-Only Ledger Protocol (issues.jsonl + tasks.jsonl):** All state transitions are append-only. The global `specs/issues.jsonl` serves as the authoritative issue registry. Issue-scoped micro-task ledgers live at `specs/{FEATURE_SLUG}/{ISSUE_ID}/tasks.jsonl` — the bucket directory (`{FEATURE_SLUG}`) is the epic scope (e.g. `001-…`, `002-…`, or `adhoc`); `{ISSUE_ID}` is the per-epic ordinal matching the issue markdown filename. The `source_file` recorded in `specs/issues.jsonl` follows `specs/{FEATURE_SLUG}/issues/{ISSUE_ID}.md` and the CLI strips the `issues/` segment when mapping to the tasks directory. Agents cannot edit any status fields directly — only the CLI may append events via `append_issue_transition()` and `append_task_transition()`. No existing line is ever modified or overwritten. Canonical state is derived by parsing each ledger using compound-key idempotency (bottom-up for `issues.jsonl`; `(id, status)` compound key for `tasks.jsonl`). For per-issue task ledgers, `COMPLETED` is terminal: once captured, no later non-`COMPLETED` transition may override it; among non-terminal entries, the last by file position wins. Ad-hoc issues bypass macro planning and route directly to isolated execution workspaces.
692
+ 3. **Append-Only Ledger Protocol (issues.jsonl + tasks.jsonl):** All state transitions are append-only. Append helpers (`_append_record`, `_append_with_compound_key`) insert a leading newline when a non-empty ledger does not already end in `\n`, so a missing trailing newline cannot fuse two records onto one line; every successful write leaves a trailing newline. `claim_issue` writes through `append_issue_transition`. The global `specs/issues.jsonl` serves as the authoritative issue registry. Issue-scoped micro-task ledgers live at `specs/{FEATURE_SLUG}/{ISSUE_ID}/tasks.jsonl` — the bucket directory (`{FEATURE_SLUG}`) is the epic scope (e.g. `001-…`, `002-…`, or `adhoc`); `{ISSUE_ID}` is the per-epic ordinal matching the issue markdown filename. The `source_file` recorded in `specs/issues.jsonl` follows `specs/{FEATURE_SLUG}/issues/{ISSUE_ID}.md` and the CLI strips the `issues/` segment when mapping to the tasks directory. Agents cannot edit any status fields directly — only the CLI may append events via `append_issue_transition()` and `append_task_transition()`. No existing line is ever modified or overwritten. Canonical state is derived by parsing each ledger using compound-key idempotency (bottom-up for `issues.jsonl`; `(id, status)` compound key for `tasks.jsonl`). For per-issue task ledgers, `COMPLETED` is terminal: once captured, no later non-`COMPLETED` transition may override it; among non-terminal entries, the last by file position wins. Ad-hoc issues bypass macro planning and route directly to isolated execution workspaces.
693
693
 
694
694
  4. **Deterministic Test Failure Check:** For a `RED` phase to be valid (`deviate red post`), `_classify_pytest_outcome()` must return `ASSERTION_FAILURE`. Return codes of `PASS` or `SYNTAX_ERROR` (SyntaxError, IndentationError, TabError, ImportError, ModuleNotFoundError) are rejected. Current implementation uses string-based parsing of `pytest -v` output; `pytest --json-report` migration is specified but not yet implemented.
695
695
 
@@ -704,7 +704,7 @@ The orchestrator must maintain and enforce these structural constraints across a
704
704
  composable overrides that take precedence over profile defaults. Execution
705
705
  profiles and agent backends are configured via `DeviateConfig.agent.backend`.
706
706
 
707
- 7. **Atomic Concurrency Protocol (Git Reference Locks):** To eliminate TOCTOU race conditions across distributed terminal instances, the issue claim workflow (formerly `deviate specify pre`, now part of the Plan phase orchestration) uses try-claim semantics. `select_unblocked_candidates()` returns all available BACKLOG issues. The worker iterates through them and attempts `claim_issue()` combined with `create_worktree()`. Next `NNN` is max(origin ledger, current ledger, remote `feat/<epic>/<NNN>-*` / `feat/adhoc/<NNN>-*`) + 1. Unmerged remote feat refs feed that max. Local-only unpushed feat branches do not reserve. The default serialization is push-as-lock: `git push -u <remote> <branch>`. The server serializes concurrent pushes. The first successful push wins. When `git push` of `feat/.../NNN-*` is rejected because the name exists, the claim path increments the ordinal and retries (cap 3). Collision retry does not set `--local`. Push-as-lock is the default, not a mandatory step. Local mode (`--local` on `deviate specify`, `deviate meso run`, or `deviate run`, or `claim_remote = false` in `.deviate/config.toml`) keeps the worktree and the ledger claim and skips the remote lock. Local mode is the explicit skip. It is not a collision winner. The `tasks.jsonl` ledger records the authoritative outcome.
707
+ 7. **Atomic Concurrency Protocol (Git Reference Locks):** To eliminate TOCTOU race conditions across distributed terminal instances, the issue claim workflow (formerly `deviate specify pre`, now part of the Plan phase orchestration) uses try-claim semantics. `select_unblocked_candidates()` returns all available BACKLOG issues. The worker iterates through them and attempts `claim_issue()` combined with `create_worktree()`. `_discover_claimable_issue` skips a candidate both when its `feat/<epic>/<slug>` branch already exists on origin (claimed elsewhere) and when that branch already has a local worktree (claimed here) — the local-worktree guard is what lets two parallel terminals on the same checkout claim two different BACKLOG issues even though each claim's SPECIFIED row lives on the feature branch and stays invisible to the main checkout's ledger. Next `NNN` is max(origin ledger, current ledger, remote `feat/<epic>/<NNN>-*` / `feat/adhoc/<NNN>-*`) + 1. Unmerged remote feat refs feed that max. Local-only unpushed feat branches do not reserve. The default serialization is push-as-lock: `git push -u <remote> <branch>`. The server serializes concurrent pushes. The first successful push wins. When `git push` of `feat/.../NNN-*` is rejected because the name exists, the claim path increments the ordinal and retries (cap 3). Collision retry does not set `--local`. Push-as-lock is the default, not a mandatory step. Local mode (`--local` on `deviate specify`, `deviate meso run`, or `deviate run`, or `claim_remote = false` in `.deviate/config.toml`) keeps the worktree and the ledger claim and skips the remote lock. Local mode is the explicit skip. It is not a collision winner. The `tasks.jsonl` ledger records the authoritative outcome.
708
708
 
709
709
  8. **The Session Continuity Principle:** Session state is persisted to `.deviate/session.json` after each CLI command. The `SessionState` class tracks `current_phase`, `active_issue_id`, and `last_command`. Macro and meso phases transition through `transition_to()` with validation from `_MACRO_TRANSITION_MAP`. Micro phases use `force_transition_to()`. The `_run_single()` function checks `session.current_phase` and supports resume from JUDGE/REFACTOR via optional `start_phase` parameter. Model continuity and KV cache management are delegated to the calling environment.
710
710
 
@@ -863,7 +863,15 @@ handing the manifest to the rest of the pipeline:
863
863
  pass a recovered manifest. `HandoverManifest` is imported by
864
864
  `scripts/verify_install.py` (the post-install smoke verifier)
865
865
  which checks the new constants and the recovery behaviour.
866
- 6. **Schema-rejection fail-fast** — the first stderr or stdout line
866
+ 6. **Unescaped evidence-quote recovery (GH-116)** — when
867
+ `yaml.safe_load` fails because an evidence `quote` /
868
+ `test_quote` / `impl_quote` double-quoted scalar embeds raw
869
+ `"`, `_safe_load_handover_yaml` rewrites those lines as `|`
870
+ block scalars and reloads. Well-formed YAML is unchanged.
871
+ Truly malformed YAML still raises
872
+ `MalformedHandoverManifestError`. Verdict and evidence
873
+ semantics are unchanged.
874
+ 7. **Schema-rejection fail-fast** — the first stderr or stdout line
867
875
  that contains `tool_count_limit` or `unsupported_tool_schema`
868
876
  kills the child. `invoke` raises `AgentSubprocessError` with those
869
877
  tokens. This path does not wait for the 900s stall clock. It does
@@ -0,0 +1,149 @@
1
+ # Plan — ISS-ADH-031
2
+
3
+ ## Plan Summary
4
+ - **Issue**: ISS-ADH-031 — Micro Judge already-exists COMPLIANCE_PASS must complete instead of hard-crashing with ROLLBACK_BOUNDARY_MISSING
5
+ - **Implementation Strategy**: Guard the mechanical evidence gate inside `_apply_judge_verdict` so a `failure_kind == "no_failing_test"` already-exists `COMPLIANCE_PASS` routes through `_NO_FAILING_TEST_FORWARD_ROUTES` to a graceful COMPLETED instead of being rewritten to `revert_to_red`, and relax the COMPLETED-write AC-token evidence check for that same route while keeping the declared-regression-files fail-closed gate.
6
+ - **Estimated Complexity**: Low
7
+ - **Estimated Effort**: 2-4 hours
8
+
9
+ ## Product Layer Anchors
10
+ - **Flow References**: `[]`
11
+ - **Source**: `specs/adhoc/issues/031-judge-revert-boundary-no-failing-test.md` (frontmatter field: `flow_refs`)
12
+ - **Release Context**: The next release (`specs/_product/release-next.md`, Goal anchor FLOW-04) ships RPC subprocess transport and a Rich TUI for live agent progress; it is unrelated to this micro judge routing bug, which carries no `flow_refs` anchor.
13
+ - **Architecture Components Touched**: C1 (`deviate` CLI — micro orchestrator in `src/deviate/cli/micro.py`). The fix does not touch C2-C6 (RPC transport, JSONL framing, command sender, event adapter, TUI renderer).
14
+
15
+ **Invariant**: This issue has an empty `flow_refs` list and the requested change is application-level micro-orchestration behavior only. It does not authorize flow-catalog, release, DeviaTDD-setup, skill, or workflow-ledger work.
16
+
17
+ ## Acceptance Contract
18
+
19
+ **Scenario AC-PLAN-001: A no_failing_test already-exists COMPLIANCE_PASS completes via skip_refactor without ROLLBACK_BOUNDARY_MISSING**
20
+ - **Source Outline**: `AO-031-01`
21
+ - **Upstream Traceability**: `US-031-01`, `FR-ADHOC-031`, `AC-ADHOC-031-01`
22
+ - **Current-Code Evidence**: `src/deviate/cli/micro.py:_adjudicate_red_no_failing_test` (line 1695), `src/deviate/cli/micro.py:_NO_FAILING_TEST_FORWARD_ROUTES` (line 1410)
23
+ - **Given**: `session.failure_kind == "no_failing_test"`, `session.red_commit_sha == ""`, and a RED no-failing-test task that declares a regression-pin `test_file` or `files` present in the worktree snapshot
24
+ - **When**: JUDGE emits `verdict: COMPLIANCE_PASS` with `next_action: skip_refactor` or a bare PASS verdict with no `next_action`
25
+ - **Then**: `_apply_judge_verdict` routes through `_NO_FAILING_TEST_FORWARD_ROUTES`, the task appends a COMPLETED transition, `pending_judge_action` is `skip_refactor`, the declared regression-pin test files remain on disk, and `ROLLBACK_BOUNDARY_MISSING` is never raised
26
+ - **Verification Mode**: automated
27
+
28
+ **Scenario AC-PLAN-002: A no_failing_test COMPLIANCE_PASS with partial evidence still completes instead of reverting to_red**
29
+ - **Source Outline**: `AO-031-01`
30
+ - **Upstream Traceability**: `US-031-01`, `FR-ADHOC-031`, `AC-ADHOC-031-01`
31
+ - **Current-Code Evidence**: `src/deviate/cli/micro.py:_apply_judge_verdict` (line 3226), `src/deviate/cli/micro.py:_rewrite_unmatched_tdd_pass` (line 3055), `src/deviate/core/judge_evidence.py:evaluate_judge_evidence` (line 56)
32
+ - **Given**: `session.failure_kind == "no_failing_test"`, `session.red_commit_sha == ""`, and a judge manifest with `verdict: COMPLIANCE_PASS`, `next_action: skip_refactor`, and evidence that cites all but one required `AC-PLAN-NNN` token
33
+ - **When**: `_apply_judge_verdict` runs the mechanical evidence gate
34
+ - **Then**: the gate does not rewrite the pass to `revert_to_red`, the task completes via `skip_refactor`, and the declared regression-pin tests remain on disk
35
+ - **Verification Mode**: automated
36
+
37
+ **Scenario AC-PLAN-003: A no_failing_test COMPLIANCE_PASS with no declared regression files fails closed**
38
+ - **Source Outline**: `AO-031-01`
39
+ - **Upstream Traceability**: `US-031-01`, `FR-ADHOC-031`, `AC-ADHOC-031-01`
40
+ - **Current-Code Evidence**: `src/deviate/cli/micro.py:_require_tdd_declared_regression_files` (line 1463), `src/deviate/cli/micro.py:_adjudicate_red_no_failing_test` (line 1801)
41
+ - **Given**: `session.failure_kind == "no_failing_test"` and a judge `COMPLIANCE_PASS` with `skip_refactor` on a test-bearing TDD task whose manifest declares an empty `files` set and no `test_file`
42
+ - **When**: `_adjudicate_red_no_failing_test` takes the forward-route COMPLETE branch
43
+ - **Then**: `_require_tdd_declared_regression_files` raises `PhaseFailedError`, no COMPLETED transition is appended, and the task fails closed with no empty test deliverable
44
+ - **Verification Mode**: automated
45
+
46
+ **Scenario AC-PLAN-004: A genuine test-bearing RED with a real RED commit still routes to revert_to_red**
47
+ - **Source Outline**: `AO-031-01`
48
+ - **Upstream Traceability**: `US-031-01`, `FR-ADHOC-031`, `AC-ADHOC-031-01`
49
+ - **Current-Code Evidence**: `src/deviate/cli/micro.py:_require_revert_to_red_boundary` (line 2492), `src/deviate/cli/micro.py:_rewrite_unmatched_tdd_pass` (line 3093)
50
+ - **Given**: a genuine test-bearing TDD task whose RED phase lands a failing test, records a non-empty `session.red_commit_sha`, and whose judge evidence omits a required `AC-PLAN-NNN` token
51
+ - **When**: `_apply_judge_verdict` runs the mechanical evidence gate
52
+ - **Then**: the unmatched pass rewrites to `revert_to_red`, `_require_revert_to_red_boundary` resolves the standing RED SHA, and the runner rolls back to RED without weakening the evidence gate
53
+ - **Verification Mode**: automated
54
+
55
+ **Scenario AC-PLAN-005: A no_failing_test COMPLIANCE_VIOLATION still routes to revert_before**
56
+ - **Source Outline**: `AO-031-02`
57
+ - **Upstream Traceability**: `US-031-02`, `FR-ADHOC-031`, `AC-ADHOC-031-02`
58
+ - **Current-Code Evidence**: `src/deviate/cli/micro.py:_coerce_judge_action` (line 2657), `src/deviate/cli/micro.py:_adjudicate_red_no_failing_test` (line 1783)
59
+ - **Given**: `session.failure_kind == "no_failing_test"`, `session.red_commit_sha == ""`, and a judge `verdict: COMPLIANCE_VIOLATION` or `next_action: revert_before`
60
+ - **When**: `_apply_judge_verdict` coerces the action
61
+ - **Then**: `_coerce_judge_action` forces `revert_before`, the runner resets to the RED baseline, `pending_judge_action` is `revert_before`, and the TDD loop re-dispatches RED to re-author a genuinely failing test
62
+ - **Verification Mode**: automated
63
+
64
+ **Scenario AC-PLAN-006: The already-exists pass path never raises ROLLBACK_BOUNDARY_MISSING**
65
+ - **Source Outline**: `AO-031-02`
66
+ - **Upstream Traceability**: `US-031-02`, `FR-ADHOC-031`, `AC-ADHOC-031-02`
67
+ - **Current-Code Evidence**: `src/deviate/cli/micro.py:_require_revert_to_red_boundary` (line 2497), `src/deviate/cli/micro.py:_apply_judge_verdict` (line 3383)
68
+ - **Given**: the already-exists `no_failing_test` `COMPLIANCE_PASS` pass path with `session.red_commit_sha == ""`
69
+ - **When**: the runner completes the task through `_NO_FAILING_TEST_FORWARD_ROUTES`
70
+ - **Then**: `_require_revert_to_red_boundary` is never invoked on that path, so `ROLLBACK_BOUNDARY_MISSING` is not raised, and the boundary helper remains reachable only for a genuine `revert_to_red` with a real RED commit
71
+ - **Verification Mode**: automated
72
+
73
+ **Scenario AC-PLAN-007: The shared judge-verdict helper does not regress the manual judge post path**
74
+ - **Source Outline**: `AO-031-01, AO-031-02`
75
+ - **Upstream Traceability**: `US-031-01`, `FR-ADHOC-031`, `AC-ADHOC-031-02`
76
+ - **Current-Code Evidence**: `src/deviate/cli/micro.py:_apply_judge_verdict` (line 3193), `src/deviate/cli/micro.py:_run_judge_phase` (line 3096)
77
+ - **Given**: `_apply_judge_verdict` is invoked from the manual `judge post` path on a no_failing_test task
78
+ - **When**: the manual path applies the same post-verdict side effects
79
+ - **Then**: a `skip_refactor` pass still completes, a `revert_to_red` with a standing RED SHA still rolls back, and the shared helper does not regress the non-auto path
80
+ - **Verification Mode**: automated
81
+
82
+ ## Workstation Mapping
83
+ - **src/deviate/cli/micro.py**: Primary fix for the micro judge routing crash.
84
+ - **Current State**: `_apply_judge_verdict` calls `_rewrite_unmatched_tdd_pass` unconditionally, which rewrites a `no_failing_test` `COMPLIANCE_PASS` to `revert_to_red` on partial evidence, then `_require_revert_to_red_boundary` raises `ROLLBACK_BOUNDARY_MISSING` because `session.red_commit_sha` is empty on the already-exists path.
85
+ - **Changes Required**: Skip `_rewrite_unmatched_tdd_pass` when `session.failure_kind == "no_failing_test"` and the verdict is a forward-route PASS. Ensure the COMPLETED-write AC-token evidence check (`_require_tdd_completed_evidence` via `_append_status_transition`) does not raise `COMPLETED_EVIDENCE_MISSING` on partial evidence for the same route, while still enforcing declared-regression-path presence. Keep `revert_before` and genuine `revert_to_red` routes unchanged.
86
+ - **Integration Surface**: `_coerce_judge_action`, `_rewrite_unmatched_tdd_pass`, `_NO_FAILING_TEST_FORWARD_ROUTES`, `_require_revert_to_red_boundary`, `_require_tdd_declared_regression_files`, `_append_status_transition`, and `SessionState.failure_kind` / `pending_judge_action` / `red_commit_sha`.
87
+ - **tests/test_micro/test_judge.py**: Unit sandbox for the fix.
88
+ - **Current State**: `test_already_exists_head_quotes_pass` (line 3315) drives `_run_tdd_judge` with `next_action=skip_refactor` on a RED-boundary task.
89
+ - **Changes Required**: Add a test that a `no_failing_test` `COMPLIANCE_PASS` with `red_commit_sha == ""` and partial evidence completes via `skip_refactor` instead of raising `ROLLBACK_BOUNDARY_MISSING`. Keep the genuine-test `revert_to_red` and `revert_before` assertions intact.
90
+ - **Integration Surface**: `_run_judge_phase`, `_apply_judge_verdict`, `_adjudicate_red_no_failing_test`, `HandoverManifest`, `SessionState`.
91
+ - **tests/test_cli/test_micro.py**: Integration sandbox reproducing the `TSK-029-02` crash.
92
+ - **Current State**: The `_run_pytest`-mocked CLI path covers judge routing and `ROLLBACK_BOUNDARY_MISSING` (lines 1760-1816).
93
+ - **Changes Required**: Add a `_run_pytest`-mocked CLI test that drives a `no_failing_test` already-exists task end to end and asserts it COMPLETES with no `ROLLBACK_BOUNDARY_MISSING` traceback.
94
+ - **Integration Surface**: `deviate micro run`, `_run_tdd_cycle`, `_run_red_phase`, `_run_judge_phase`.
95
+ - **specs/DeviaTDD-api.md**: Update the `no_failing_test` adjudication contract (line 595-604) to state that a `COMPLIANCE_PASS` already-exists pass completes via `skip_refactor` even with partial AC evidence, and that `ROLLBACK_BOUNDARY_MISSING` only applies to a genuine `revert_to_red` with a real RED commit.
96
+ - **specs/DeviaTDD-architecture.md**: Update §3 (line 288) to describe the guarded already-exists COMPLETE route and the retained fail-closed regression-files gate.
97
+ - **CHANGELOG.md**: Append a bullet under `[Unreleased]` `### Fixed` for the user-visible hard-crash fix.
98
+
99
+ ## Implementation Strategy
100
+ - **Phase 1**: Guard the evidence gate in `src/deviate/cli/micro.py` — deliverable: `no_failing_test` already-exists `COMPLIANCE_PASS` no longer rewrites to `revert_to_red`.
101
+ - **Files**: `src/deviate/cli/micro.py`
102
+ - **Approach**: In `_apply_judge_verdict`, run `_rewrite_unmatched_tdd_pass` only when `session.failure_kind != "no_failing_test"`. In the COMPLETED-write evidence check, skip the AC-token citation requirement for `failure_kind == "no_failing_test"` while retaining the declared-regression-path presence gate.
103
+ - **Verification**: `pytest tests/test_micro/test_judge.py -v` passes; the new partial-evidence no_failing_test test completes with no `ROLLBACK_BOUNDARY_MISSING`.
104
+ - **Phase 2**: Add regression tests — deliverable: failing-then-passing RED coverage for the fix.
105
+ - **Files**: `tests/test_micro/test_judge.py`, `tests/test_cli/test_micro.py`
106
+ - **Approach**: Add a unit test driving `_adjudicate_red_no_failing_test` / `_run_judge_phase` with `red_commit_sha == ""` and partial evidence, and a `_run_pytest`-mocked CLI test for the `TSK-029-02` shape. Assert COMPLETED + `skip_refactor`, regression-pin tests on disk, no crash.
107
+ - **Verification**: `pytest tests/ -v` (full suite under 30s with mocked `_run_pytest`); `ruff check .` clean.
108
+ - **Phase 3**: Align specs and changelog — deliverable: documentation reflects the final behavior.
109
+ - **Files**: `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md`, `CHANGELOG.md`
110
+ - **Approach**: Update the `no_failing_test` adjudication descriptions and add a `[Unreleased]` Fixed bullet.
111
+ - **Verification**: `mise run check` exits 0.
112
+
113
+ ## Data Flow Analysis
114
+ - RED phase calls `_run_red_phase`, which routes a no-failing-test outcome to `_adjudicate_red_no_failing_test`.
115
+ - `_adjudicate_red_no_failing_test` sets `session.failure_kind = "no_failing_test"` and dispatches `_run_judge_phase`.
116
+ - `_run_judge_phase` builds the JUDGE prompt with a `<failure_kind>no_failing_test</failure_kind>` block and the uncommitted RED test diff, then calls `_apply_judge_verdict`.
117
+ - `_apply_judge_verdict` coerces `next_action` via `_coerce_judge_action`, then applies the evidence gate. On the fixed route it keeps `skip_refactor`, stashes validated evidence, appends a COMPLETED transition, and parks the session at IDLE.
118
+ - `_adjudicate_red_no_failing_test` enforces `_require_tdd_declared_regression_files`, restores the worktree to the RED baseline while keeping declared regression paths, and clears the session for the next task.
119
+ - On `COMPLIANCE_VIOLATION`, `_coerce_judge_action` forces `revert_before`; the runner resets to the RED baseline and re-dispatches RED. Storage: append-only `specs/**/tasks.jsonl` COMPLETED row plus `.deviate/session.json` transient state.
120
+
121
+ ## Risk Assessment
122
+ | Risk | Impact | Likelihood | Mitigation |
123
+ |------|--------|------------|------------|
124
+ | Guarding on `failure_kind == "no_failing_test"` accidentally also skips the evidence gate for `test_defect` or `mechanical` routes | Medium | Low | Restrict the guard strictly to `no_failing_test`; keep the AC-token evidence gate for all other failure kinds and for the genuine test-bearing RED path. |
125
+ | Relaxing the COMPLETED-write evidence check weakens fail-closed on missing regression files | Medium | Low | Keep `_require_tdd_declared_regression_files` and the declared-path presence gate intact for the no_failing_test route; only the AC-token citation check is relaxed. |
126
+ | The shared `_apply_judge_verdict` regresses the manual `judge post` path | Medium | Low | The guard keys on `session.failure_kind`, which both auto and manual paths already set; add a judge-post regression scenario (AC-PLAN-007). |
127
+ | Double COMPLETED append on the no_failing_test forward route | Medium | Medium | The relaxed evidence check must tolerate the second `_append_status_transition`; verify the ledger keeps a single COMPLETED transition per task in tests. |
128
+ | FLOW_CONTEXT_UNAVAILABLE — no existing flow mapping is available | Medium | Low | Preserve the empty `flow_refs` and plan the application behavior without creating flow or DeviaTDD-setup work. |
129
+
130
+ ## Security Profile
131
+
132
+ Risk surfaces: subprocess (git reset / clean / restore during rollback), file paths (declared regression-test paths, `_evidence_head_contents`), ledger writes (append-only `tasks.jsonl`).
133
+
134
+ Negative tests: The `no_failing_test` already-exists COMPLIANCE_PASS never invokes `_require_revert_to_red_boundary`, so no git reset to an invented `HEAD~1` boundary occurs; a partial-evidence pass does not silently destroy the declared regression-pin tests; a `COMPLIANCE_VIOLATION` still fails closed to `revert_before` and never completes with an empty test deliverable.
135
+
136
+ Constraints: No new dependencies; no hardcoded secrets; no invented RED boundary or `HEAD~1` fallback; no changes to `flow_refs` or Product-layer flow artifacts; no file-path path traversal beyond the existing relative-path guard in `_evidence_head_contents`.
137
+
138
+ ## Integration Points
139
+ - **`_adjudicate_red_no_failing_test`**: Routes the RED no-failing-test outcome; owns the forward-route COMPLETE and the `revert_before` re-author contract.
140
+ - **`_apply_judge_verdict`**: Shared by auto `_run_judge_phase` and manual `judge post`; owns `next_action` coercion, the evidence gate, and the violation/forward side effects.
141
+ - **`_require_revert_to_red_boundary`**: Fatal `ROLLBACK_BOUNDARY_MISSING` for a genuine `revert_to_red` with no RED SHA; must stay reachable only for real revert routes.
142
+ - **`_require_tdd_declared_regression_files` / `_require_tdd_completed_evidence`**: The fail-closed files gate and the AC-token completed-evidence gate.
143
+ - **`SessionState` fields `failure_kind`, `red_commit_sha`, `pending_judge_action`, `last_judge_verdict`**: Discriminators that drive the guarded routing.
144
+
145
+ ## Constitutional Alignment
146
+ - **Architecture**: Aligns with the Micro layer (RED → GREEN → JUDGE → REFACTOR) and the Git Isolation Principle (§1) — the fix refuses to invent a RED boundary and completes the already-exists task without a destructive rollback.
147
+ - **Testing**: pytest under `tests/`; GREEN passes all tests and JUDGE verifies GREEN modified only allowed files; the full suite stays under 30s by mocking `deviate.cli.micro._run_pytest` with a `subprocess.CompletedProcess` fixture.
148
+ - **Git Isolation**: The fix runs on the dedicated issue branch/worktree; the already-exists pass writes no RED commit and does not mutate git history, and `_require_revert_to_red_boundary` stays reserved for genuine revert routes with a real RED commit.
149
+ - **Product Layer**: This issue carries an empty `flow_refs` and touches only application micro-orchestration behavior in C1; it preserves the existing user-visible flows and does not author or synchronize any Product-layer flow catalog.
@@ -0,0 +1,5 @@
1
+ {"id":"TSK-031-01","issue_id":"ISS-ADH-031","description":"Guard the already-exists pass route so a `no_failing_test` `COMPLIANCE_PASS` completes via `skip_refactor` without `ROLLBACK_BOUNDARY_MISSING`","status":"RED","execution_mode":"TDD","created_at":"2026-08-25T12:45:13.907479Z","security_profile":null,"acceptance_criteria":null}
2
+ {"id":"TSK-031-01","issue_id":"ISS-ADH-031","description":"Guard the already-exists pass route so a `no_failing_test` `COMPLIANCE_PASS` completes via `skip_refactor` without `ROLLBACK_BOUNDARY_MISSING`","status":"COMPLETED","execution_mode":"TDD","created_at":"2026-08-26T00:56:13.895039Z","security_profile":null,"acceptance_criteria":null,"evidence":{"items":[{"ac":"AC-PLAN-001","test_path":"tests/test_micro/test_judge.py","test_quote":"class TestNoFailingTestAlreadyExistsPass:","impl_path":"src/deviate/cli/micro.py","impl_quote":"no_failing_pass = ("},{"ac":"AC-PLAN-002","test_path":"tests/test_micro/test_judge.py","test_quote":"partial evidence must not rewrite to revert_to_red","impl_path":"src/deviate/cli/micro.py","impl_quote":"action = \"skip_refactor\""},{"ac":"AC-PLAN-003","test_path":"tests/test_micro/test_judge.py","test_quote":"guarded already-exists PASS route","impl_path":"src/deviate/cli/micro.py","impl_quote":"skip_token_citation=no_failing_pass,"},{"ac":"AC-PLAN-004","test_path":"tests/test_micro/test_judge.py","test_quote":"declared regression-pin test file must stay on disk","impl_path":"src/deviate/cli/micro.py","impl_quote":"skip_token_citation: bool = False,"},{"ac":"AC-PLAN-005","test_path":"tests/test_micro/test_judge.py","test_quote":"genuine COMPLIANCE_VIOLATION stays fail-closed","impl_path":"src/deviate/cli/micro.py","impl_quote":"if session.failure_kind == \"no_failing_test\":"},{"ac":"AC-PLAN-006","test_path":"tests/test_micro/test_judge.py","test_quote":"legacy bare COMPLIANCE_PASS coerces to skip_refactor","impl_path":"src/deviate/cli/micro.py","impl_quote":"extra_paths=declared,"},{"ac":"AC-PLAN-007","test_path":"tests/test_micro/test_judge.py","test_quote":"_run_no_failing_test_judge(","impl_path":"src/deviate/cli/micro.py","impl_quote":"_require_tdd_completed_evidence"}],"red":"b7cae3bc45798f3bc1cbae1485748143f3195a53","green":"285d80bd5d591cec2a56a44ec2657a4636ec289c","head":"285d80bd5d591cec2a56a44ec2657a4636ec289c"}}
3
+ {"id": "TSK-031-02", "issue_id": "ISS-ADH-031", "description": "Add a `_run_pytest`-mocked CLI test driving a `no_failing_test` already-exists task to COMPLETED", "status": "EXECUTE", "execution_mode": "IMMEDIATE", "created_at": "2026-08-26T01:08:19.991674+00:00", "security_profile": null, "acceptance_criteria": null}
4
+ {"id":"TSK-031-02","issue_id":"ISS-ADH-031","description":"Add a `_run_pytest`-mocked CLI test driving a `no_failing_test` already-exists task to COMPLETED","status":"COMPLETED","execution_mode":"IMMEDIATE","created_at":"2026-08-26T01:08:20.161218Z","security_profile":null,"acceptance_criteria":null,"evidence":{"items":[{"ac":"AC-PLAN-001","test_path":"tests/test_cli/test_micro.py","test_quote":"class TestNoFailingTestAlreadyExistsCliCompletes:","impl_path":"src/deviate/cli/micro.py","impl_quote":"if session.failure_kind == \"no_failing_test\":"},{"ac":"AC-PLAN-006","test_path":"tests/test_cli/test_micro.py","test_quote":"assert \"ROLLBACK_BOUNDARY_MISSING\" not in output","impl_path":"src/deviate/cli/micro.py","impl_quote":"no_failing_pass = ("}],"red":"285d80bd5d591cec2a56a44ec2657a4636ec289c","green":"7b5e70ee1f1fb92801974a24ad3ad1be4c2471a9","head":"7b5e70ee1f1fb92801974a24ad3ad1be4c2471a9"}}
5
+ {"id":"TSK-031-03","issue_id":"ISS-ADH-031","description":"Update the `no_failing_test` adjudication contract in the specs and add a changelog entry","status":"COMPLETED","execution_mode":"IMMEDIATE","created_at":"2026-08-26T01:11:08.855394Z","security_profile":null,"acceptance_criteria":null}
@@ -0,0 +1,118 @@
1
+ # Implementation Tasks: `feat/adhoc/031-judge-revert-boundary-no-failing-test`
2
+
3
+ ## Phase 1: Guard the already-exists `no_failing_test` pass route in `_apply_judge_verdict`
4
+ **Goal**: A `failure_kind == "no_failing_test"` already-exists `COMPLIANCE_PASS` completes via `skip_refactor` instead of being rewritten to `revert_to_red` and hard-crashing with `ROLLBACK_BOUNDARY_MISSING`.
5
+
6
+ ### Tasks
7
+
8
+ - TSK-031-01: Guard the already-exists pass route so a `no_failing_test` `COMPLIANCE_PASS` completes via `skip_refactor` without `ROLLBACK_BOUNDARY_MISSING`
9
+ - **Type**: Bugfix
10
+ - **Mode**: TDD
11
+ - **Test Strategy**: Sociable_Unit
12
+ - **Verification**: `pytest tests/test_micro/test_judge.py -v`
13
+ - **Estimated Time**: 60 minutes
14
+ - **Flow References**: `[]`
15
+ - **Files**:
16
+ - `src/deviate/cli/micro.py`
17
+ - `tests/test_micro/test_judge.py`
18
+ - **Rationale**: Fixes the micro-judge routing crash described by `US-031-01` and `US-031-02`. `_apply_judge_verdict` (line 3217) calls `_rewrite_unmatched_tdd_pass` unconditionally, which rewrites a `no_failing_test` `COMPLIANCE_PASS` with partial AC evidence to `revert_to_red`; the subsequent `_require_revert_to_red_boundary` (line 3383) then raises `ROLLBACK_BOUNDARY_MISSING` because `session.red_commit_sha` is empty on the already-exists path. Implements `AC-PLAN-001`, `AC-PLAN-002`, `AC-PLAN-004`, `AC-PLAN-005`, `AC-PLAN-006`, and `AC-PLAN-007`; guards the fail-closed behavior of `AC-PLAN-003`. `tests/test_micro/test_judge.py` is the unit sandbox pinned by the issue's Multi-Tiered Verification Targets.
19
+ - **Details**:
20
+ - **Red**: Add unit tests in `tests/test_micro/test_judge.py` driving `_adjudicate_red_no_failing_test` / `_run_judge_phase` with `session.red_commit_sha == ""` and `session.failure_kind == "no_failing_test"`. Assert: (1) `AC-PLAN-001`/`AC-PLAN-002` — a `COMPLIANCE_PASS` with `next_action: skip_refactor` (and a bare PASS with no `next_action`) and evidence omitting one required `AC-PLAN-NNN` token appends exactly one COMPLETED ledger row, sets `pending_judge_action == "skip_refactor"`, keeps the declared regression-pin test file on disk, and raises no `PhaseFailedError` / `ROLLBACK_BOUNDARY_MISSING`; (2) `AC-PLAN-003` — a `skip_refactor` pass with an empty `files` set and no `test_file` raises `PhaseFailedError`, appends no COMPLETED row, and fails closed; (3) `AC-PLAN-004` — a genuine test-bearing RED with a non-empty `red_commit_sha` and partial evidence still rewrites to `revert_to_red` and `_require_revert_to_red_boundary` resolves the standing RED SHA; (4) `AC-PLAN-005` — a `COMPLIANCE_VIOLATION` / `next_action: revert_before` still forces `revert_before` via `_coerce_judge_action`, resets to the RED baseline, and re-dispatches RED; (5) `AC-PLAN-007` — invoking `_apply_judge_verdict` directly (the manual `judge post` path) on a `no_failing_test` session still completes a `skip_refactor` pass and still rolls back a `revert_to_red` with a standing RED SHA.
21
+ - **Green**: In `_apply_judge_verdict` (`src/deviate/cli/micro.py`, line 3217), run `_rewrite_unmatched_tdd_pass` only when `session.failure_kind != "no_failing_test"`. Relax the COMPLETED-write AC-token citation check (`_require_tdd_completed_evidence`, invoked by `_append_status_transition`) for `failure_kind == "no_failing_test"` so partial AC evidence does not raise `COMPLETED_EVIDENCE_MISSING`, while retaining `_require_tdd_declared_regression_files` and the declared-path presence gate. Leave `_coerce_judge_action`, the genuine `revert_to_red` / `revert_before` routes, and the `_NO_FAILING_TEST_FORWARD_ROUTES` set unchanged.
22
+ - **Refactor**: Keep the guard strictly scoped to `no_failing_test`; do not touch `test_defect` or `mechanical` evidence gating. Confirm the forward-route `COMPLETED` write from `_adjudicate_red_no_failing_test` (line 1803) and the `skip_refactor` write from `_apply_judge_verdict` (line 3524) tolerate a single COMPLETED ledger row per task.
23
+ - **Edge Cases**: Handle the double-COMPLETED append risk by asserting the ledger keeps exactly one COMPLETED transition per task. Handle partial evidence without destroying the declared regression-pin tests. Handle `COMPLIANCE_VIOLATION` without completing with an empty test deliverable. Never invoke `_require_revert_to_red_boundary` on the already-exists pass path; keep it reachable only for a genuine `revert_to_red` with a real RED commit.
24
+ - **Acceptance**: The full `tests/test_micro/test_judge.py` suite passes; the existing `test_already_exists_head_quotes_pass` (line 3315) and `test_already_exists_missing_test_file_fails` (line 3332) still pass unchanged; the new partial-evidence `no_failing_test` test completes with no `ROLLBACK_BOUNDARY_MISSING`; `ruff check tests/test_micro/test_judge.py src/deviate/cli/micro.py` is clean.
25
+
26
+ ---
27
+
28
+ - **Judge Feedback**: Guarded already-exists route: no_failing_test COMPLIANCE_PASS coerces to skip_refactor and never reaches _require_revert_to_red_boundary; COMPLIANCE_VIOLATION still routes revert_before. All 3 new tests pass; full suite 1600 passed, 3 skipped; ruff clean.
29
+ - **Judge Feedback**: Guarded already-exists route verified: all 78 judge tests and the full suite pass; ruff clean.
30
+ ## Phase 2: Integration reproduction of the `TSK-029-02` crash
31
+ **Goal**: Reproduce the user-facing `deviate micro run` crash end to end through the `_run_pytest`-mocked CLI path and prove the task COMPLETES with no `ROLLBACK_BOUNDARY_MISSING` traceback.
32
+
33
+ ### Tasks
34
+
35
+ - TSK-031-02: Add a `_run_pytest`-mocked CLI test driving a `no_failing_test` already-exists task to COMPLETED
36
+ - **Type**: Verification_Batch
37
+ - **Mode**: IMMEDIATE
38
+ - **Test Strategy**: Integration
39
+ - **Verification**: `pytest tests/test_cli/test_micro.py -v`
40
+ - **Estimated Time**: 60 minutes
41
+ - **Flow References**: `[]`
42
+ - **Files**:
43
+ - `tests/test_cli/test_micro.py`
44
+ - `src/deviate/cli/micro.py`
45
+ - **Rationale**: The issue's demonstration path runs `deviate micro run --task TSK-029-02`; `tests/test_cli/test_micro.py` is the integration sandbox that reproduces the crash (`US-031-01`, `US-031-02`). The existing `_run_pytest`-mocked CLI path (lines 1760-1816) covers judge routing and `ROLLBACK_BOUNDARY_MISSING`; the new test drives the `no_failing_test` already-exists shape end to end. Implements `AC-PLAN-001` and `AC-PLAN-006`. If the CLI wiring exposes a gap not covered by `TSK-031-01`, fix it in `src/deviate/cli/micro.py`.
46
+ - **Details**:
47
+ - **Red**: Add a test in `tests/test_cli/test_micro.py` that drives the `deviate micro run` surface (`_run_tdd_cycle` → `_run_red_phase` → `_adjudicate_red_no_failing_test` → `_run_judge_phase`) on a `TSK-029-02`-style task whose RED phase routes to `failure_kind == "no_failing_test"` with `session.red_commit_sha == ""`. Mock `deviate.cli.micro._run_pytest` with a `subprocess.CompletedProcess(args=[], returncode=0, stdout="", stderr="")` fixture and mock `_invoke_agent` to return a `COMPLIANCE_PASS` / `skip_refactor` manifest that declares a regression-pin `test_file`. Assert the task reaches a COMPLETED status, `pending_judge_action == "skip_refactor"`, the declared regression-pin test remains on disk, and no `ROLLBACK_BOUNDARY_MISSING` traceback appears in the captured output.
48
+ - **Green**: Confirm the `_run_red_phase` → `_adjudicate_red_no_failing_test` → `_run_judge_phase` wiring drives the guarded `_apply_judge_verdict` forward route. If the integration path reveals a missing `declared_paths` or `red_baseline` thread, repair it in `src/deviate/cli/micro.py`; otherwise no production change is required beyond `TSK-031-01`.
49
+ - **Refactor**: Reuse the existing `_run_pytest`-mock fixture and `SessionState.load(session_path)` setup already present in `tests/test_cli/test_micro.py`; do not duplicate the git-env helper — use `_git_env()` and `tmp_git_repo` from `tests/conftest.py` so no git command runs against the real repo.
50
+ - **Edge Cases**: Assert the regression-pin test file survives the `_restore_worktree_to_baseline(..., keep_paths=declared)` restore. Assert the full suite stays under 30s because `_run_pytest` is mocked. Assert a `COMPLIANCE_VIOLATION` variant still routes to `revert_before` and never COMPLETES.
51
+ - **Acceptance**: `pytest tests/test_cli/test_micro.py -v` passes; `ruff check tests/test_cli/test_micro.py` is clean; `pytest tests/ -v` completes under 30s with the mocked `_run_pytest`.
52
+ - **Dependency**: TSK-031-01
53
+
54
+ ---
55
+
56
+ ## Phase 3: Align specs and changelog with the guarded routing
57
+ **Goal**: Reflect the guarded already-exists COMPLETE route and the retained fail-closed regression-files gate in the authoritative specs and the changelog.
58
+
59
+ ### Tasks
60
+
61
+ - TSK-031-03: Update the `no_failing_test` adjudication contract in the specs and add a changelog entry
62
+ - **Type**: Config
63
+ - **Mode**: IMMEDIATE
64
+ - **Verification**: `mise run check`
65
+ - **Estimated Time**: 30 minutes
66
+ - **Flow References**: `[]`
67
+ - **Files**:
68
+ - `specs/DeviaTDD-api.md`
69
+ - `specs/DeviaTDD-architecture.md`
70
+ - `CHANGELOG.md`
71
+ - **Rationale**: The constitution (§5 Definition of Done) and `AGENTS.md` require user-visible behavior changes to update the authoritative `specs/DeviaTDD-api.md` and `specs/DeviaTDD-architecture.md` in the same commit, and to append a `CHANGELOG.md` `[Unreleased]` bullet. The hard-crash fix is a user-visible behavior change (`US-031-01`). Documents `AC-PLAN-001`, `AC-PLAN-002`, `AC-PLAN-003`, and `AC-PLAN-006`.
72
+ - **Details**:
73
+ - **Red**: N/A (IMMEDIATE — docs-only).
74
+ - **Green**: N/A (IMMEDIATE — docs-only).
75
+ - **Implementation**: In `specs/DeviaTDD-api.md` (line 595-604), state that a `COMPLIANCE_PASS` already-exists pass on `failure_kind == "no_failing_test"` completes via `skip_refactor` even with partial AC evidence, and that `ROLLBACK_BOUNDARY_MISSING` applies only to a genuine `revert_to_red` with a real RED commit. In `specs/DeviaTDD-architecture.md` §3 (line 288), describe the guarded already-exists COMPLETE route, the relaxed AC-token citation check, and the retained `_require_tdd_declared_regression_files` fail-closed gate. Append a bullet under `CHANGELOG.md` `[Unreleased]` `### Fixed` describing the user-visible hard-crash fix.
76
+ - **Refactor**: Keep terminology consistent with the existing spec wording (`no_failing_test`, `skip_refactor`, `ROLLBACK_BOUNDARY_MISSING`, `_require_tdd_declared_regression_files`); do not rename existing tokens.
77
+ - **Edge Cases**: Verify the architecture §3 note still records that a genuine `revert_to_red` with an empty SHA is fatal, so the docs do not overstate the relaxation.
78
+ - **Acceptance**: `mise run check` exits 0; `grep` confirms the `no_failing_test` adjudication contract and the `### Fixed` changelog bullet are present.
79
+ - **Dependency**: TSK-031-01
80
+
81
+ ---
82
+
83
+ ## Implementation Strategy
84
+ **Execution Order**:
85
+ 1. Phase 1 (TSK-031-01) -> Phase 2 (TSK-031-02) -> Phase 3 (TSK-031-03)
86
+
87
+ **Critical Dependency Chains**:
88
+ - TSK-031-02 depends on TSK-031-01 (the CLI integration test requires the guarded routing)
89
+ - TSK-031-03 depends on TSK-031-01 (the specs and changelog document the final behavior)
90
+
91
+ **Risk Hotspots**:
92
+ - Guarding on `failure_kind == "no_failing_test"` must not skip the evidence gate for `test_defect` or `mechanical` routes — restrict the guard strictly to `no_failing_test`.
93
+ - Relaxing the COMPLETED-write evidence check must not weaken fail-closed on missing regression files — keep `_require_tdd_declared_regression_files` and the declared-path presence gate intact.
94
+ - The shared `_apply_judge_verdict` must not regress the manual `judge post` path — the guard keys on `session.failure_kind`, which both auto and manual paths set (AC-PLAN-007).
95
+ - Double COMPLETED append on the `no_failing_test` forward route — verify a single COMPLETED transition per task in the tests.
96
+ - No existing flow mapping is available — preserve the empty `flow_refs` and do not create Product-layer or DeviaTDD-setup work.
97
+
98
+ **Merge Conflict Boundaries**:
99
+ - Files touched by multiple phases: `src/deviate/cli/micro.py` (TSK-031-01, TSK-031-02).
100
+
101
+ **Product-Layer Anchors** (mirrored from plan.md):
102
+ - **Flow References**: `[]`
103
+ - **Source**: `specs/adhoc/031-judge-revert-boundary-no-failing-test/plan.md`
104
+ - Downstream micro phases inherit this list per-task. Empty references mean no matching existing flow, not permission for enabling, setup, tooling, skill, release, or workflow-ledger tasks.
105
+
106
+ **E2E Scope Note**: No `[E2E]` bats task is emitted. The plan scopes testing to the unit sandbox (`tests/test_micro/test_judge.py`) and the `_run_pytest`-mocked integration sandbox (`tests/test_cli/test_micro.py`), which reproduces the user-facing `deviate micro run` surface deterministically while keeping the full suite under 30s per constitution §3. A real-agent bats E2E is non-deterministic and would exceed the test-performance budget.
107
+
108
+ ---
109
+
110
+ ## Universal Test Constraints (ALL TASKS)
111
+
112
+ - **Git Isolation Mandatory**: Any test that invokes git operations MUST operate on a temporary directory initialized as a fresh git repo. Tests MUST NOT run git commands within the real repository's working tree.
113
+ - **Implementation Pattern**: Use a shared `tmp_git_repo` fixture from `tests/conftest.py`. Pass `repo=tmp_git_repo` to all git-interacting functions. Never reference `Path.cwd()` or the real repo root.
114
+ - **Rationale**: Prevent accidental commits, branch creation, or state mutation in the actual project repo during test execution.
115
+
116
+ ## Universal API Design Constraint (ALL CORE MODULES)
117
+
118
+ Every git-interacting function in core modules MUST accept an optional `repo_path: Path | None = None` parameter. When `None`, default to `Path.cwd()`.