deviatdd 2.23.0__tar.gz → 2.24.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.gitignore +7 -7
- {deviatdd-2.23.0 → deviatdd-2.24.0}/CHANGELOG.md +29 -1
- {deviatdd-2.23.0 → deviatdd-2.24.0}/CLAUDE.md +0 -4
- {deviatdd-2.23.0 → deviatdd-2.24.0}/PKG-INFO +90 -64
- {deviatdd-2.23.0 → deviatdd-2.24.0}/README.md +89 -63
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pyproject.toml +7 -1
- deviatdd-2.24.0/specs/006-setup-interactive-config/data-model.md +245 -0
- deviatdd-2.24.0/specs/006-setup-interactive-config/design.md +125 -0
- deviatdd-2.24.0/specs/006-setup-interactive-config/explore.md +143 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/DeviaTDD-api.md +239 -103
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/DeviaTDD-architecture.md +93 -52
- deviatdd-2.24.0/specs/adhoc/030-config-rework/plan.md +156 -0
- deviatdd-2.24.0/specs/adhoc/030-config-rework/tasks.jsonl +14 -0
- deviatdd-2.24.0/specs/adhoc/030-config-rework/tasks.md +196 -0
- deviatdd-2.24.0/specs/adhoc/034-setup-interactive-config/plan.md +159 -0
- deviatdd-2.24.0/specs/adhoc/034-setup-interactive-config/tasks.md +101 -0
- deviatdd-2.24.0/specs/adhoc/issues/033-prune-post-completed.md +83 -0
- deviatdd-2.24.0/specs/adhoc/issues/034-setup-interactive-config.md +106 -0
- deviatdd-2.24.0/specs/adhoc/issues/035-gate3-walkthrough-map-and-review-comments.md +98 -0
- deviatdd-2.24.0/specs/adhoc/issues/036-cli-help-readme-transparency.md +86 -0
- deviatdd-2.24.0/specs/adhoc/issues/037-setup-pack-tui-multiselect.md +81 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/prd.md +80 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc.jsonl +1 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/constitution.md +1 -1
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/issues.jsonl +21 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/__init__.py +784 -180
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/meso.py +24 -7
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/micro.py +37 -34
- deviatdd-2.24.0/src/deviate/cli/prune.py +67 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/review.py +54 -15
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/walkthrough.py +24 -7
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/agent.py +26 -6
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/commands.py +146 -15
- deviatdd-2.24.0/src/deviate/core/profile.py +50 -0
- deviatdd-2.24.0/src/deviate/core/prune.py +573 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/review_coverage.py +113 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/assembly.py +0 -1
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/red.md +3 -2
- deviatdd-2.24.0/src/deviate/prompts/commands/deviate-prune.md +165 -0
- deviatdd-2.24.0/src/deviate/prompts/commands/deviate-review.md +204 -0
- deviatdd-2.24.0/src/deviate/prompts/commands/deviate-walkthrough.md +165 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/core.md +2 -2
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/skills/deviatdd/SKILL.md +37 -8
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/state/config.py +143 -19
- deviatdd-2.24.0/src/deviate/ui/checkbox.py +242 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/core/test_agent.py +2 -2
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_derived_command_install.bats +18 -5
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_optional_push_as_lock.bats +6 -4
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_pi_spawn_lean_tool_schema.bats +5 -4
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_red_hang_timeout_rollback.bats +7 -6
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_review_plan_ac_coverage.bats +25 -15
- deviatdd-2.24.0/tests/e2e/test_setup_config_rework.bats +124 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_help.py +1 -1
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_init.py +417 -128
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_meso.py +94 -4
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_micro.py +61 -9
- deviatdd-2.24.0/tests/test_cli/test_prune.py +176 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_review.py +283 -36
- deviatdd-2.24.0/tests/test_cli/test_setup.py +936 -0
- deviatdd-2.24.0/tests/test_cli/test_walkthrough.py +129 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_agent.py +146 -25
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_commands.py +64 -0
- deviatdd-2.24.0/tests/test_core/test_profile.py +68 -0
- deviatdd-2.24.0/tests/test_core/test_prune.py +247 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_command_installation.py +25 -7
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_init_export_cycle.py +48 -28
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_skill_installation.py +25 -7
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_auto_prompt_templates.py +2 -9
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_meso_model_routing.py +7 -53
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_meso_orchestration.py +11 -297
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_prompt_assembly.py +1 -1
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_specify.py +100 -1
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_state/test_config.py +151 -7
- deviatdd-2.24.0/tests/test_ui/test_checkbox.py +90 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/uv.lock +1 -1
- deviatdd-2.23.0/.deviate/.gitignore +0 -6
- deviatdd-2.23.0/.deviate/config.toml +0 -20
- deviatdd-2.23.0/.deviate/evaluations/deviate-flows-skill-2026-06-26.json +0 -55
- deviatdd-2.23.0/.opencode/opencode.json +0 -4
- deviatdd-2.23.0/src/deviate/core/profile.py +0 -33
- deviatdd-2.23.0/src/deviate/prompts/auto/specify.md +0 -11
- deviatdd-2.23.0/src/deviate/prompts/commands/deviate-prune.md +0 -260
- deviatdd-2.23.0/src/deviate/prompts/commands/deviate-review.md +0 -374
- deviatdd-2.23.0/src/deviate/prompts/commands/deviate-walkthrough.md +0 -244
- deviatdd-2.23.0/tests/test_core/test_profile.py +0 -35
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.env.example +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.gitattributes +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.githooks/pre-commit +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.githooks/pre-push +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.github/ISSUE_TEMPLATE/bug.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.github/ISSUE_TEMPLATE/feature.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.github/PULL_REQUEST_TEMPLATE.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.github/workflows/ci.yml +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/.github/workflows/release.yml +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/AGENTS.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/CODE_OF_CONDUCT.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/CONTRIBUTING.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/LICENSE +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/SECURITY.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/deviatdd.png +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/mise.toml +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat/001-deviate-cli-python/008-meso-macro-automated-orchestration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat/adhoc/006-context-cli-integration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat/adhoc/017-optional-push-as-lock.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-001-deviate-cli-python-001-cli-initialization-governance-provisioning-pr-11.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-001-deviate-cli-python-003-meso-layer-specification-task-decomposition.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-001-deviate-cli-python-004-micro-layer-tdd-sandbox-execution.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-001-deviate-cli-python-005-cli-architecture-realignment-skill-integration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-001-deviate-cli-python-007-macro-meso-parity-backward-compatibility.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-002-deviatdd-gap-analysis-001-foundation-cli-infrastructure.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-002-deviatdd-gap-analysis-003-fast-path-commands.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-002-deviatdd-gap-analysis-004-governance-inspection.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-adhoc-001-streaming-pipeline-monitor.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-adhoc-008-ast-phase-prioritization.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat-adhoc-017-two-counter-tdd-retry.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/pr_descriptions/feat_002-deviatdd-gap-analysis_001-foundation-cli-infrastructure.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/scripts/benchmark_lmstudio.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/scripts/next_version.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/scripts/verify_install.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t001.json +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t002.json +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/manifest-t004.json +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/001-cli-initialization-governance-provisioning/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/002-macro-layer-state-ledger-management/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/002-macro-layer-state-ledger-management/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/003-meso-layer-specification-task-decomposition/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/003-meso-layer-specification-task-decomposition/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/004-micro-layer-tdd-sandbox-execution/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/005-cli-architecture-realignment-skill-integration/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/005-cli-architecture-realignment-skill-integration/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/007-macro-meso-parity-backward-compatibility/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/007-macro-meso-parity-backward-compatibility/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/008-meso-macro-automated-orchestration/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/data-model.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/design.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/explore.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/001-cli-initialization-governance-provisioning.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/002-macro-layer-state-ledger-management.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/003-meso-layer-specification-task-decomposition.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/004-micro-layer-tdd-sandbox-execution.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/005-cli-architecture-realignment-skill-integration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/006-state-persistence-concurrency-safety.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/007-macro-meso-parity-backward-compatibility.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/issues/008-meso-macro-automated-orchestration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/001-deviate-cli-python/prd.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/001-foundation-cli-infrastructure/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/003-fast-path-commands/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/004-governance-inspection/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/004-governance-inspection/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/004-governance-inspection/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/005-micro-layer-integrity/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/data-model.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/design.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/explore.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/issues/001-foundation-cli-infrastructure.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/issues/002-context-pipeline.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/issues/003-fast-path-commands.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/issues/004-governance-inspection.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/issues/005-micro-layer-integrity.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/plan-tdd-integration-gap.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/002-deviatdd-gap-analysis/prd.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/003-graphite-cli-integration/explore.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/004-per-task-security-profile/issues/001-security-profile-and-judge-checks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/001-verification-mode-metadata/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/001-verification-mode-metadata/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/001-verification-mode-metadata/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/002-task-acceptance-traceability/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/002-task-acceptance-traceability/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/002-task-acceptance-traceability/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/data-model.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/design.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/explore.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/issues/001-verification-mode-metadata.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/issues/002-task-acceptance-traceability.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/issues/003-micro-phase-gates-red-green.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/issues/004-refactor-regression-gate.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/issues/005-prompt-spec-alignment.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/005-acceptance-gates/prd.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/architecture.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/domain-model.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/flows/flows-product.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/flows/flows-streaming.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/flows/index.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/flows.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/guildwright-current-system.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/guildwright-gap-register.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/guildwright-git-state-model.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/guildwright-rewrite.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/guildwright-rust-tui-requirements.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/_product/release-next.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/001-streaming-pipeline-monitor/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/001-streaming-pipeline-monitor/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/001-streaming-pipeline-monitor/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/003-meso-layer-restructuring/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/003-meso-layer-restructuring/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/003-meso-layer-restructuring/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/004-deviate-review-skill/spec.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/004-deviate-review-skill/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/004-deviate-review-skill/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/005-per-phase-model-configuration/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/005-per-phase-model-configuration/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/005-per-phase-model-configuration/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/006-context-cli-integration/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/006-context-cli-integration/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/006-context-cli-integration/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/007-graphite-cli/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/007-graphite-cli/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/007-graphite-cli/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/008-ast-phase-prioritization/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/008-ast-phase-prioritization/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/008-ast-phase-prioritization/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/009-pi-agent-backend-integration/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/009-pi-agent-backend-integration/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/009-pi-agent-backend-integration/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/010-deviate-setup-product-layer/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/010-deviate-setup-product-layer/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/010-deviate-setup-product-layer/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/013-flow-ledger-canonical-source-of-truth/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/015-narrow-product-flow-scope/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/015-narrow-product-flow-scope/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/016-single-source-prompt-templates/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/016-single-source-prompt-templates/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/016-single-source-prompt-templates/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/017-optional-push-as-lock/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/017-optional-push-as-lock/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/017-optional-push-as-lock/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/017-two-counter-tdd-retry/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/017-two-counter-tdd-retry/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/017-two-counter-tdd-retry/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/018-one-behavior-rgr-granularity/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/018-one-behavior-rgr-granularity/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/019-remote-aware-ordinal-allocation/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/019-remote-aware-ordinal-allocation/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/019-remote-aware-ordinal-allocation/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/020-judge-compliance-pass-evidence/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/020-judge-compliance-pass-evidence/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/020-judge-compliance-pass-evidence/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/021-no-failing-test-escalate-invokes-green/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/021-no-failing-test-escalate-invokes-green/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/021-no-failing-test-escalate-invokes-green/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/022-already-satisfied-red-requires-tests/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/022-already-satisfied-red-requires-tests/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/022-already-satisfied-red-requires-tests/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/023-pinned-micro-run-issue-scoped/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/023-pinned-micro-run-issue-scoped/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/023-pinned-micro-run-issue-scoped/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/024-worktree-session-stale-issue-id/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/024-worktree-session-stale-issue-id/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/024-worktree-session-stale-issue-id/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/025-green-stderr-noise-stall-detector/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/025-green-stderr-noise-stall-detector/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/025-green-stderr-noise-stall-detector/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/026-pi-spawn-lean-tool-schema/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/026-pi-spawn-lean-tool-schema/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/026-pi-spawn-lean-tool-schema/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/027-red-hang-timeout-rollback/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/027-red-hang-timeout-rollback/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/027-red-hang-timeout-rollback/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/028-task-scoped-judge-review-coverage/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/028-task-scoped-judge-review-coverage/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/028-task-scoped-judge-review-coverage/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/031-judge-revert-boundary-no-failing-test/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/031-judge-revert-boundary-no-failing-test/tasks.jsonl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/031-judge-revert-boundary-no-failing-test/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/001-streaming-pipeline-monitor.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/002-aider-agent-backend-integration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/003-meso-layer-restructuring.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/004-deviate-review-skill.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/005-per-phase-model-configuration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/006-context-cli-integration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/007-graphite-cli.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/008-ast-phase-prioritization.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/009-pi-agent-backend-integration.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/010-deviate-setup-product-layer.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/012-rpc-streaming-tui-renderer.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/013-flow-ledger-canonical-source-of-truth.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/014-cwe-mapping-security-findings.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/015-narrow-product-flow-scope.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/016-single-source-prompt-templates.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/017-optional-push-as-lock.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/017-two-counter-tdd-retry.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/018-one-behavior-rgr-granularity.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/019-remote-aware-ordinal-allocation.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/020-judge-compliance-pass-evidence.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/021-no-failing-test-escalate-invokes-green.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/022-already-satisfied-red-requires-tests.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/023-pinned-micro-run-issue-scoped.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/024-worktree-session-stale-issue-id.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/025-green-stderr-noise-stall-detector.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/026-pi-spawn-lean-tool-schema.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/027-red-hang-timeout-rollback.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/028-task-scoped-judge-review-coverage.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/029-ponytail-pruning-in-review.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/030-config-rework.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/031-judge-revert-boundary-no-failing-test.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/adhoc/issues/032-judge-feedback-injection-fail-close.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/ast-tree-sitter.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/config-rework.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/flow-ledger.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/graphite-cli.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/pi-agent-backend.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/ponytail-pruning.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/product-layer.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/rpc-streaming-tui.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/rpc-streaming.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/explore/security-hardening-cwe.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/specs/implementation-gap.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/_common.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/_html.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/_safe_commands.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/adhoc.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/constitution.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/feature.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/flow_commands.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/init.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/inspect.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/cli/macro.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/_shared.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/cache_discipline.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/commit.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/complexity.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/constitution.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/contract.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/convention.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/epic.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/issues.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/judge_evidence.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/prd.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/repo.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/run_logger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/tasks_ledger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/validation.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/core/worktree.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/html_templates/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/html_templates/architecture.html.tmpl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/html_templates/domain-model.html.tmpl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/html_templates/flows.html.tmpl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/html_templates/plan.html.tmpl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/html_templates/prd.html.tmpl +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/main.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/execute.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/explore.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/green.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/judge.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/prd.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/refactor.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/research.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/shard.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/auto/tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-adhoc.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-architecture.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-constitution.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-e2e.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-execute.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-explore.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-flows.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-green.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-hotfix.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-html.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-init.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-judge.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-merge.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-plan.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-pr.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-prd.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-red.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-refactor.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-release.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-research.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-shard.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-tasks.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/commands/deviate-triage.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/constitution_seed.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/lifecycle-auto.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/lifecycle-manual.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/macro-shared.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/meso-shared.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/micro-shared.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/product-shared.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/core/style-ste.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/governance/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/governance/agents_seed.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/governance/claudemd_seed.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/prompts/governance/libref_seed.md +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/state/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/state/ledger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/ui/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/ui/monitor.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/ui/pipeline.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/ui/render.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/src/deviate/visual/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/conftest.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/core/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/core/test_smart_stall.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_green_stderr_stall.bats +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_macro_workflow.bats +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/e2e/test_worktree_session_stale_issue.bats +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_adhoc.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_common.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_constitution.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_feature.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_flows_sync.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_html.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_inspect.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_macro_contracts.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_main_entrypoint.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_merge.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_meso_contracts.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_real_descendant_kill.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_safe_commands.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_test_command_resolution.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_timeout_safe_command.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_cli/test_top_level_run.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_cache_discipline.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_commit.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_complexity.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_constitution.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_contract.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_convention.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_epic.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_flow_confirmation.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_issues.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_judge_evidence.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_ledger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_prd.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_repo.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_run_logger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_tasks_ledger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_validation.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_core/test_worktree.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_init.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/conftest.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_macro_full_cycle.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_macro_layer.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_macro_orchestration.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_meso_layer.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_meso_orchestration.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_meso_task_ledger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_integration/test_parity.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_macro/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_macro/test_explore.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_macro/test_macro_model_routing.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_macro/test_macro_orchestration.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_macro/test_prd.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_macro/test_research.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_macro/test_shard.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_meso_resume.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_plan_structure_injection.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_pr_platform.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_meso/test_tasks.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/conftest.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_commit_failure.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_completed_evidence.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_e2e.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_execute.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_green.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_hotfix.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_judge.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_orchestration.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_output_filter.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_red.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_refactor.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_review_pause.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_rollback_safety.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_run.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_task_label.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_micro/test_two_counter_retry.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_release/test_next_version.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_state/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_state/test_ledger.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_state/test_security_profile.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_state/test_session.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_ui/__init__.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_ui/test_monitor.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_ui/test_pipeline.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_ui/test_render.py +0 -0
- {deviatdd-2.23.0 → deviatdd-2.24.0}/tests/test_visual_demo/test_tsk_001_01.py +0 -0
|
@@ -10,15 +10,10 @@ __pycache__/
|
|
|
10
10
|
# Tooling
|
|
11
11
|
.rgr
|
|
12
12
|
.tdd-session.json
|
|
13
|
-
.deviate/
|
|
14
|
-
.opencode/
|
|
15
|
-
.codebase-index/
|
|
16
|
-
.opencode/codebase-index.json
|
|
13
|
+
.deviate/
|
|
14
|
+
.opencode/
|
|
17
15
|
node_modules/
|
|
18
16
|
|
|
19
|
-
.deviate/artifacts/
|
|
20
|
-
.deviate/logs/
|
|
21
|
-
.deviate/review/
|
|
22
17
|
*/commands/deviate-*.md
|
|
23
18
|
*/prompts/deviate-*.md
|
|
24
19
|
|
|
@@ -35,3 +30,8 @@ node_modules/
|
|
|
35
30
|
!.env.example
|
|
36
31
|
|
|
37
32
|
*/skills/deviatdd/
|
|
33
|
+
|
|
34
|
+
*/skills/deviate-*/
|
|
35
|
+
|
|
36
|
+
# Codebase index (regenerated by open-codebase-index)
|
|
37
|
+
.codebase-index/
|
|
@@ -8,6 +8,18 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
10
|
### Added
|
|
11
|
+
- **`deviate setup --agent pi --agent-export-mode global` now writes to pi's actual global dirs.** Prompt templates go to `~/.pi/agent/prompts/` and the deviatdd skill to `~/.pi/agent/skills/` — pi's `getPromptsDir()` / `getAgentDir()`-relative locations — instead of `~/.pi/prompts/` + `~/.pi/skills/`, which pi only treats as project paths. Local mode (`<workdir>/.pi/prompts`, `<workdir>/.pi/skills`) is unchanged. Pinned by `tests/test_cli/test_setup.py::TestSetupExportModeAndBaseBranch` and the fresh-global-install assertions.
|
|
12
|
+
- **Global installs no longer embed a project's constitution into the prompt.**
|
|
13
|
+
`install_command(..., export_mode=...)` / `compose_command_body(...,
|
|
14
|
+
include_constitution=...)` in `src/deviate/core/commands.py` skip tier 0
|
|
15
|
+
(constitution text) for `--agent-export-mode global`, because the global
|
|
16
|
+
prompt is project-agnostic and shared across every repo. Project-local
|
|
17
|
+
installs keep baking `<workdir>/specs/constitution.md` verbatim when the file
|
|
18
|
+
exists. Global prompts rely on the core block's Constitution Compliance
|
|
19
|
+
Mandate (invariant #10) to read `specs/constitution.md` at runtime.
|
|
20
|
+
- **Prune classifier is language-agnostic and drops fewer false positives.** The honeycomb classifier in `src/deviate/core/prune.py` now (a) ignores `self._helper(...)` calls — test-own helpers are no longer mistaken for private-state probes, (b) does not drop tests whose spy-assertions mock external boundaries (`subprocess`, `Popen`, agent spawn), (c) recognizes non-Python test declarations (Go `func TestXxx`, Rust `#[test] fn`, JS `describe`/`it`) via a regex fallback alongside the Python AST path, and (d) understands language-native markers (`#[behavioral]`, camelCase `TestBehavioral`) plus Go/JS assertion forms (`t.Fail`, `expect(...).toBe`). Test discovery now globs `*_test`, `*.spec.*`, `*_test.*` in addition to `test_*.py`. The `/deviate-prune` prompt and RED honeycomb-stamp instructions say "markers/annotations/tags" instead of pytest-only marks. Pinned by `tests/test_core/test_prune.py` (new Go/JS fallback test) and existing prune pins.
|
|
21
|
+
- **`deviate setup --agent codex` now pins Luna + high thinking for spawned micro runners.** Fresh Codex setup writes `[models].default = "gpt-5.6-luna"` and `[agent].reasoning_effort = "high"` alongside `backend = "codex"`. Re-running against an existing config upserts those keys only when they are missing or empty — a user-set `[models].default` or `[agent].reasoning_effort` is left alone. Non-Codex setup does not write Luna or a reasoning key. Spawned `codex exec` receives `-c model_reasoning_effort=<value>` from `[agent].reasoning_effort` (official values `minimal|low|medium|high|xhigh`); no repo-wide `.codex/config.toml` is written. Pinned by `tests/test_cli/test_setup.py::TestSetupCodex`, `tests/test_core/test_agent.py::TestCodexBackendRegistration`, and `tests/test_cli/test_micro.py::TestInvokeAgentCodexReasoning`.
|
|
22
|
+
- **ChatGPT Codex is a first-class setup + meso/micro backend.** `--agent codex` (and the interactive prompt) is a valid install target. Setup writes packaged skills to `.agents/skills/deviatdd/SKILL.md` plus one `.agents/skills/<command>/SKILL.md` per slash command (Codex CLI 0.117+ dropped `.codex/prompts`). `[agent].backend = "codex"` is persisted. Dispatch uses `codex exec --sandbox workspace-write --ask-for-approval never` with the existing stdin prompt path and `--model` injection. Transport stays `cli`. Pinned by `tests/test_cli/test_setup.py::TestSetupCodex` and `tests/test_core/test_agent.py::TestCodexBackendRegistration`.
|
|
11
23
|
- **On-demand GitHub Actions Release workflow (`workflow_dispatch`).** A maintainer runs Actions → Release → Run workflow. The job computes the next SemVer from conventional commits since the current `pyproject.toml` version (`feat` → minor, `fix` → patch, `BREAKING CHANGE` / `feat!` / `fix!` → major; chore/docs/test/ci/style/refactor still bump patch), writes `pyproject.toml` + the `deviatdd` version in `uv.lock`, commits `chore(release): version X.Y.Z` on the default branch, tags `vX.Y.Z`, and publishes with `uv build` / `uv publish --trusted-publishing always` (PyPI trusted publishing via GitHub OIDC; `contents: write` + `id-token: write`; no `PYPI_API_TOKEN` and no GitHub secret). Optional inputs: `bump` (`auto` / `patch` / `minor` / `major`) and `dry_run` (print the version; skip commit/tag/publish). Refuses a non-default branch unless `dry_run` is true. Does not rewrite `CHANGELOG.md` (stays `[Unreleased]`). Operator adds a pending PyPI trusted publisher for `wernerbisschoff/deviatdd` workflow `release.yml` (no environment). Local fallback remains `mise run publish`. Pinned by `tests/test_release/test_next_version.py`.
|
|
12
24
|
- **COMPLETED `tasks.jsonl` rows now persist runner-validated JUDGE evidence (GH-84).** After the #65 mechanical gate accepts a TDD completion path (`skip_refactor` / bare `COMPLIANCE_PASS` / post-REFACTOR complete / adjudicated already-exists), `_append_status_transition(..., "COMPLETED")` copies `HandoverManifest.evidence` (`ac`, `test_path`, `test_quote`, `impl_path`, `impl_quote`) onto an optional `TaskRecord.evidence` bundle and stamps `red` (`session.red_commit_sha`) plus `green`/`head` (`HEAD` at the write). TDD complete fail-closes when the plan contract has `AC-PLAN-NNN` tokens and the bundle is missing or does not cover them (`evaluate_judge_evidence`, no new matcher). EXECUTE / IMMEDIATE / DIRECT stay ungated; enabling/infra plans with no AC tokens may complete with empty evidence. `deviate inspect tasks show` prints the field when present. Session files are not the proof store. Pinned by `tests/test_micro/test_completed_evidence.py`, `tests/test_state/test_ledger.py`, and `tests/test_cli/test_inspect.py`.
|
|
13
25
|
- **`deviate micro run --review` pauses before each phase commit (GH-101).** After the agent finishes a phase and tests/lint have run, `--review` (and `--all --review`) halt before `_commit_phase` / `_commit_phase_with_recovery` (`git add -A`), print `REVIEW_PAUSE <phase> <task_id>`, leave the worktree dirty for herdr/hunk, and wait for TTY confirmation (`Enter` / yes). Then the runner commits as today and continues. Pause applies to RED, GREEN, REFACTOR, and EXECUTE — not JUDGE feedback commits. RED is still committed before GREEN so `session.red_commit_sha` remains available for JUDGE `revert_to_red`. Non-TTY / `--json` / missing stdin fail closed with `REVIEW_REQUIRES_TTY` and never auto-commit past the flag. Off by default; no config key. The pause is one helper (`_maybe_review_pause`) in front of `_commit_phase`, not a copy in each phase. Pinned by `tests/test_micro/test_review_pause.py`.
|
|
@@ -29,8 +41,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
29
41
|
- **`deviate specify --local` claims an issue without pushing to the remote.** When `--local` is set, `_try_claim_issue` skips the `branch_exists_on_remote` check and the `git push` step, so the worktree is created, the ledger row is written, and the claim commit is recorded locally only. If the local branch `feat/<epic>/<slug>` already exists, the command short-circuits with `ALREADY_CLAIMED_LOCAL` and reuses the existing worktree, treating branch presence as the claim signal without rewriting the ledger. NOTE: this semantic can false-positive on a manual `git checkout -b feat/<epic>/<slug>` that pre-dated any claim — there is no opt-in to disable it for the no-remote workflow, so prefer `deviate specify <id>` (the default, which checks remote first) when operating against a shared origin. The new option lives at `src/deviate/cli/meso.py::_try_claim_issue(..., local=...)` and the `specify` Typer command forwards it through `_specify_pre`. Pinned by `tests/test_meso/test_specify.py::TestSpecifyLocalFlag` (2 tests: branch-exists short-circuit, kwargs forwarded) and `tests/test_cli/test_meso.py::TestSpecifyLocalFlag` (4 tests: remote check skipped, push skipped, short-circuit dict shape, branch-exists short-circuit).
|
|
30
42
|
- **`deviate flows sync` — sole owner of `specs/_product/flows.jsonl` creation.** New top-level `flows` Typer group exposes a single `sync` subcommand that parses `specs/_product/flows/index.md` and appends one ``FlowRecord`` identity row plus ``FLOW_DISCOVERED`` and ``FLOW_DOCUMENTED`` events per flow. Re-running on a populated ledger is a no-op (compound-key idempotency on the underlying append helpers). Exits non-zero with ``FLOWS_INDEX_MISSING`` on stderr when the canonical index is absent, and ``FLOWS_INDEX_EMPTY`` when the index parses to zero rows — surfaces authoring defects instead of silently committing a half-baked catalog. New public service `seed_flow_ledger(flows_index, ledger_path)` in `src/deviate/state/ledger.py` plus `FlowIndexEmptyError` for the empty-index contract. `deviate explore post` no longer seeds identity or documentation events; it only renders the coverage report. ``flows.jsonl`` never receives ``FLOW_REFERENCED_BY_ISSUE`` events — the referenced-by relationship is derived read-only from ``specs/issues.jsonl::flow_refs`` at coverage time.
|
|
31
43
|
- **`deviate html prd --bucket <slug>` targets a specific epic when more than one owns a `prd.md`.** Previously the command hard-failed with `HTML_AMBIGUOUS_PRD` whenever multiple numbered epics had a `prd.md` and offered no flag to disambiguate — the agent was told to "run from within the epic's worktree", which silently yielded `HTML_NO_PRD` because the resolver reads `specs/` from the repo root. The new option resolves `specs/<bucket>/prd.md` directly and bypasses ambiguity detection; the plain form keeps its existing behavior but its ambiguity banner now points at the flag. An unknown/absent bucket exits `PRD_NOT_FOUND`. The `/deviate-html` prompt (`src/deviate/prompts/commands/deviate-html.md`) and spec docs now reference `--bucket` instead of the broken cwd fallback. Pinned by `tests/test_cli/test_html.py::test_html_prd_bucket_targets_specific_epic`, `::test_html_prd_bucket_targets_unnumbered_dir`, and `::test_html_prd_bucket_missing_file_exits_cleanly`.
|
|
32
|
-
- **Optional push-as-lock: `claim_remote` config plus `--local` on `meso run` and `run`.** Standing `.deviate/config.toml` key `claim_remote`
|
|
44
|
+
- **Optional push-as-lock: `claim_remote` config plus `--local` on `meso run` and `run`.** Standing `.deviate/config.toml` key `claim_remote` plus `--local` on `meso run` and `run`. `deviate setup --no-claim-remote` writes `claim_remote = false` without dropping `[models]`, `timeout_seconds`, or `[agent]`. Effective local mode is `--local` OR `claim_remote = false`; explicit `--local` always wins. Local mode still creates `.worktrees/feat/{epic}/{issue}/`, writes SPECIFIED, and commits the claim, and skips `branch_exists_on_remote` plus `git push`. `deviate specify --local` stays; omitted `--local` honors config. `deviate meso run --local` and `deviate run --local` share that meaning. `--no-setup` remains a distinct skip of worktree plus claim. Local discovery does not treat an origin branch as claimed-elsewhere. Specs: `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md`. Pinned by `tests/test_state/test_config.py`, `tests/test_cli/test_meso.py`, `tests/test_meso/test_specify.py`, `tests/test_meso/test_meso_orchestration.py`, `tests/test_cli/test_init.py`, and `tests/test_cli/test_top_level_run.py`.
|
|
33
45
|
### Changed
|
|
46
|
+
- **`deviate setup` no longer asks or writes `base_branch`.** Trunk is resolved at runtime by `resolve_base_branch`: hand-set `config.toml` `base_branch` if present, else `origin/HEAD`, else `main`. `--base-branch` is a script-only write-override. Pinned by `tests/test_state/test_config.py` and `tests/test_cli/test_setup.py::TestSetupConfigAllowlist`.
|
|
47
|
+
- **`deviate setup` TTY asks prompt/skill install `[l]ocal/[g]lobal`; generated `config.toml` is runner config only.** After the agent picker: export mode (default `l`, this run only — not persisted), then claim-remote, then the pack checkbox. Prompt copy uses ``rich.markup.escape`` so ``[y]es/[n]o`` and ``[l]ocal/[g]lobal`` render as literals. `--agent-export-mode` omitted on a TTY prompts (option default is `None`; non-TTY omitted installs local). `--base-branch` is a script-only write-override (omitted = do not write `base_branch`). `global` installs under `~/.{agent}/…` (Codex: `~/.agents/skills`). Fresh dump has inline comments and no `base_branch`, `agent_export_mode`, or `[agent].timeout` (`timeout_seconds` is the only timeout key). TTY/flag re-runs strip leftover `agent_export_mode` / `[agent].timeout` and leave a hand-set `base_branch` alone. Still no `~/.pi/agent/` writes. Pinned by `tests/test_cli/test_setup.py::TestSetupConfigAllowlist` / `::TestSetupExportModeAndBaseBranch` and `tests/test_cli/test_init.py::test_prompt_claim_remote_escapes_rich_brackets`.
|
|
48
|
+
- **`deviate setup` TTY optional-pack picker is a checkbox list (one pack per row).** Space toggles, Enter confirms, default nothing selected (macro+meso+micro only). Replaces the Rich `Prompt.ask` slash-separated list that wrapped mid-name (`pr` / `une`). `--packs` for scripts is unchanged. Non-TTY still default-only. Selection is not persisted in `config.toml`. Arrow keys stay in the list: CSI (`\\x1b[A`/`[B`) and SS3 (`\\x1bOA`/`OB`) are read with VMIN=0 VTIME=1 (no select race); lone ESC is ignored, not confirm. Pending stdin is flushed before the checklist so a leftover Enter from the agent/`claim_remote` `Prompt.ask` cannot auto-confirm empty. On a TTY re-run, the claim-remote `[y]es/[n]o` prompt always appears (default = current `claim_remote` in `config.toml`; accepts `y`/`n`/`yes`/`no`) and the answer is upserted; flags still skip it. Pinned by `tests/test_cli/test_setup.py::TestSetupPacks` (TTY helper invoked; mocked `product`+`pr` installs those two only), `tests/test_ui/test_checkbox.py`, and `tests/test_cli/test_init.py` claim-remote TTY re-run cases.
|
|
49
|
+
- **Gate 3 `/deviate-walkthrough` is a four-look map; `/deviate-review` is the PR review (comments by default, `--apply` CRITICAL-only).** Walkthrough maps THIS issue/PR: (a) brief path + this issue's plan AC lines if `plan.md` exists; (b) test hunks; (c) which production hunks claim which named check; (d) the command to run those checks. It must not reimplement, approve, hide hunks, tell the human to skip a look, or auto-edit. `deviate walkthrough pre` emits `issue_brief_path`, `plan_path` (null if absent), and classified `test_files` / `production_files`; it no longer sends constitution/prd as default inputs unless this brief names those paths. `/deviate-review` is comments-only by default (stdout and/or GitHub PR review event `COMMENT`); no edits, no `git add`/`git commit`, no `REQUEST_CHANGES`, not a merge gate. There is no always-on STEP 4. `deviate review --apply` (or `pre --apply`) may apply CRITICAL findings only (security / data loss / broken build / named-check fail with a concrete FIX) and commit only when such a fix landed; never auto-apply SUGGESTION or OPPORTUNITY. Same inputs → same comments, keyed by named-check tokens + test-weakening + this-issue cross-task drift. A brief with no named checks emits exactly `brief incomplete` (do not hunt Explore). Do not assume JUDGE already ran (`--profile fast` coworker path). `review_coverage.py` uncovered plan-AC tokens stay as comment input; `coverage_complete` is not an apply gate. Both remain optional packs (`walkthrough`, `review`); no `pr-review` pack. Pinned by CLI contracts in `tests/test_cli/test_review.py` and `tests/test_cli/test_walkthrough.py` (JSON fields, `brief incomplete`, default no-apply, `--apply` CRITICAL-only) — not by prompt-body substring tests.
|
|
50
|
+
- **`deviate --help` and the README now say which phases commit, spawn, or fail closed.** A short per-phase transparency table covers setup, adhoc, meso, micro, review, and walkthrough (does / commits / debug a fail): RED `git commit --no-verify`, Codex `codex exec --sandbox workspace-write --ask-for-approval never`, nested spawn inside `meso run` / `micro run`, `pre`/`post`, `--profile fast` skips JUDGE **and** REFACTOR, `deviate micro run --review` is a TTY pause (not `/deviate-review`; skill argument `review` is a loop policy), default packs on current main (#134) vs optional packs, `--no-setup --local` Path A, `claim_remote` default false, `/deviate-review` comments-only (not a merge gate; opt-in `--apply` CRITICAL-only), and `/deviate-walkthrough` as the four-look map. Typer help on `setup`, `meso run`, `micro run` (`--profile` / `--review` / `--all`) names the same contract. Help copy is documentation, not a pytest contract.
|
|
51
|
+
- **`deviate setup` default install is execution layers only.** Product is one optional pack (`deviate-flows` + `deviate-architecture` + `deviate-release`, kept together), not a default layer. Bare setup / `--packs none` / TTY default `none` write `macro` + `meso` + `micro` (including `/deviate-init`) plus the shared `deviatdd` skill. The TTY optional-pack selector lists `none`, `all-optional`, then `product` first among named packs, then the existing individuals (`merge`, `pr`, `review`, `walkthrough`, `html`, `hotfix`, `triage`, `prune`, `e2e`). `--packs product` writes all three product commands; `--packs all-optional` includes `product`. Pinned by `tests/test_cli/test_setup.py::TestSetupPacks` and `tests/test_core/test_commands.py::TestCommandPacks`.
|
|
52
|
+
- **`deviate setup` now installs commands and the `deviatdd` skill only to the targeted agent, and auto-detects installed agents when `--agent` is omitted.** `--agent <name>` writes command and skill files only under that agent's directory and fails closed with `AGENT_NOT_INSTALLED` when the named agent is declared but not installed; omitting `--agent` targets exactly the directories `detect_agents` reports as present (`.claude/`, `.opencode/`, `.factory/`, `.pi/`, `.omp/`), instead of the previous install-to-all active-agents dispatch. The `.deviate/` directory is now git-ignored by default through `_ensure_root_gitignore` (new `.deviate/` entry) and the `.deviate/.gitignore`/`.gitignore` provisioning kept intact for existing tracked history (AO-030-01 Boundary). The `[agent].timeout` config field is removed: agent-process and test-command deadlines both resolve from the single consolidated `DeviateConfig.timeout_seconds` (`resolve_agent_deadline`, default 1800s), while the Graphite config surface is fully removed. The `[models]` phase→model resolution order (phase key → `default` → backend-native) is unchanged. Pinned by `tests/test_state/test_config.py`, `tests/test_cli/test_setup.py`, and `tests/test_integration/test_skill_installation.py`.
|
|
53
|
+
- **Push-as-lock is opt-in.** Absent `claim_remote` key, absent `.deviate/config.toml`, and fresh `deviate setup` all mean local claims only (`claim_remote = false`). Fresh setup writes `claim_remote = false`. `deviate setup --claim-remote` writes `true`; `--no-claim-remote` still writes `false`. The TTY prompt “Push claim branches to the remote as a lock” defaults to no. Existing `claim_remote = true` configs keep pushing. `--local` still forces local; effective local remains `--local` OR `claim_remote = false`. Worktrees and local claim commits are unchanged. Specs: `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md`. Pinned by `tests/test_state/test_config.py`, `tests/test_cli/test_setup.py`, `tests/test_cli/test_init.py`, `tests/test_cli/test_meso.py`, `tests/test_meso/test_specify.py`, `tests/test_meso/test_meso_orchestration.py`, and `tests/test_cli/test_top_level_run.py`.
|
|
54
|
+
- **Micro execution profiles are only `full` and `fast`.** `full` runs JUDGE + REFACTOR; `fast` skips both. `--no-judge` / `--no-refactor` remain composable overrides. The old `secure` name is not a public profile and is not a security mode — it stays an undocumented internal alias that keeps JUDGE and skips REFACTOR (`profile = "secure"` / `--profile secure`). Legacy `profile = "default"` still coerces to `full`. `TaskRecord.security_profile` is unchanged. Pinned by `tests/test_core/test_profile.py` and `tests/test_state/test_config.py`.
|
|
55
|
+
- **New-user path is `deviate setup` then `/deviate-init`.** README Quickstart and the Workflow Bootstrap rows no longer claim that setup scaffolds `specs/constitution.md` "in one shot". Setup writes `.deviate/`, persists the agent, and installs default packs (including `deviate-init`) plus the shared `deviatdd` skill. `/deviate-init` (Codex: the `deviate-init` skill) scaffolds `specs/constitution.md`, `mise.toml`, and `specs/issues.jsonl`, skipping files already present. Successful `deviate setup` prints a next-step hint to run `/deviate-init` (no-op if the repo is already scaffolded). Pinned by `tests/test_cli/test_setup.py::TestSetupNextStepHint` and `::TestReadmeNewUserPath`.
|
|
56
|
+
- **`deviate setup` is pack-aware and writes a production-clean `config.toml`.** Omitted `--packs` installs the default layer set (macro + meso + micro plus the shared `deviatdd` skill) and leaves optional packs (`product`, `merge`, `pr`, `review`, `walkthrough`, `html`, `hotfix`, `triage`, `prune`, `e2e`) uninstalled. `--packs none|all-optional|pr,review` (or a TTY prompt) selects optional packs. Generated config always keeps `base_branch` and `claim_remote`; `profile` is now `full`/`fast` (default `full`) and is the `deviate micro run` default when `--profile` is omitted (legacy `profile = "default"` coerces to `full`; legacy `profile = "secure"` remains an undocumented alias that keeps JUDGE and skips REFACTOR). `--libref` is the only libref opt-in — without it there is no `use_libref` key and no libref mention in generated config, governance seeds, or installed command/skill bodies. `[agent]` persists the chosen backend (`codex` when Codex is picked) and writes `transport` only for `pi`/`omp`; `pi_rpc` is never written. Codex Luna + high-reasoning if-empty defaults are unchanged. Pinned by `tests/test_cli/test_setup.py`, `tests/test_core/test_commands.py`, and `tests/test_state/test_config.py`.
|
|
57
|
+
- **The `deviatdd` skill accepts an optional `review` argument that pauses after each successful task.** Default invoke (no argument) still auto-continues `deviate micro run` until `NO_PENDING_TASKS`. When `$ARGUMENTS` contains `review` (or `/deviatdd review` / "deviatdd with review"), the agent stops after each exit 0, shows the task id and the commits just made, and waits for the human before the next bare `deviate micro run`. Does not pass `--review` or `--all` to the runner. Prompt-only change to `src/deviate/prompts/skills/deviatdd/SKILL.md`.
|
|
58
|
+
- **`/deviate-prune` is a manual honeycomb test-thinning surface.** Prefer pytest marks and name tags (`spy` / `impl` drop, `behavioral` / `ac` keep). Untagged tests are classified from the body (drop internal spies/mocks/private state; keep public input-to-output / AC) and must not auto-keep. RED stamps `@pytest.mark.behavioral` | `@pytest.mark.spy` | `@pytest.mark.impl` on each new test (most are behavioral). Prune is manual invoke only — not hooked into micro COMPLETED, `--all`, or the `deviatdd` skill success loop. `apply_prune` / READY never unlinks `plan.md`, `tasks.md`, `explore.md`, `prd.md`, issue md, leftover cycle markdown, or JSONL ledgers. Compact / squash / rewrite is rejected. Thin `deviate prune pre` / `post` CLI is registered. One issue per invocation. Pinned by `tests/test_core/test_prune.py` and `tests/test_cli/test_prune.py`.
|
|
34
59
|
- **The `deviatdd` skill checks for an existing OPEN GitHub issue before filing a harness-bug issue.** The "Filing deviatdd issues" section now runs `gh issue list --repo wernerbisschoff/deviatdd --state open --search` first; when an open issue already matches the same harness failure, it directs commenting the new evidence/task context onto that issue (`gh issue comment`) instead of `gh issue create`, and only creates a new issue when no match exists. Prompt-only change to the packaged skill (`src/deviate/prompts/skills/deviatdd/SKILL.md`), re-installed to all agent skill dirs.
|
|
35
60
|
- **`deviate refactor pre` now scopes `files_to_refactor` to the RED+GREEN production set (GH-98).** The command no longer glob()s every `src/**/*.py` or discards `_resolve_task_context`. It lists production files from `HEAD~2..HEAD`, falling back to the task `Files:` list minus tests when that git range is empty or unavailable. Test files are never included. The JSON contract now emits the documented handover fields (`status`, `task_id`, `task_title`, `task_type`, `test_command`, `lint_command`, `spec_dir`, `verification`, `repo_root`, `git_branch`, `timestamp`) alongside `files_to_refactor`. Auto `_build_auto_prompt("refactor")` injects the same scoped list and keeps the `git log -2` / `git diff HEAD~2..HEAD` inspect step. Pinned by `tests/test_micro/test_refactor.py`.
|
|
36
61
|
- **GREEN / REFACTOR / review prompts fold smallest-change into existing lines** (reuse stdlib or an already-installed dep; in-place refactor; Opportunities do not extract helpers). Prompt-only. Pinned by `tests/test_meso/test_auto_prompt_templates.py::TestSmallestChangeFoldedIntoExistingPrompts`.
|
|
@@ -55,6 +80,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
55
80
|
- **`/deviate-merge` no longer auto-pushes after the squash-merge commit; the push gate runs inline and the network push is opt-in.** The slash command previously ran `git push` as its final step. As of v2.4.0 the squash-merge commit lands on `main`, then a new `push_gate` step inlines the body of `.githooks/pre-push` (lint + format-check + testmon-driven affected tests with the warm-cache / full-suite fallback — bash 3.2 portable, `GIT_DIR` reset + trap preserved) so the safety net fires even though no `git push` happens yet. After the gate passes the skill asks the operator whether to `git push` (which fires the real `pre-push` hook and re-runs the same gate) or stop and push manually. The squash-merge commit and the ledger transition inside it are durable on `main` regardless of the push outcome — only the network push is deferred. New failure states: `Push_Gate_Failed` (inline gate non-zero), `Push_Failed` (`git push` non-zero, raw stderr surfaced), `Push_Deferred` (user chose "Stop — I'll push manually"). Inline gate body and `.githooks/pre-push` body must stay byte-equivalent; divergence is pinned by `tests/test_meso/test_auto_prompt_templates.py::TestMergePromptPushGate::test_hook_and_prompt_agree_on_gate_body` (which compares non-blank non-comment lines in both bodies and fails on drift in either direction) plus 3 supporting assertions on the upstream-first logic, the testmon fallback, and the prompt structure. Prompt: `src/deviate/prompts/commands/deviate-merge.md` (v2.3.0 → v2.4.0). Spec mirrors updated: `specs/DeviaTDD-architecture.md` (Merge bullet, new `**Push gate + opt-in push (v2.4.0)**` sub-bullet) and `specs/DeviaTDD-api.md` (new `**/deviate-merge push behavior (v2.4.0)**` entry under the `deviate merge` reference).
|
|
56
81
|
|
|
57
82
|
### Fixed
|
|
83
|
+
- **Pi spawn no longer passes `--no-extensions` (ISS-ADH-033 fallout, regression from ISS-ADH-026).** Extension-registered providers (such as `pi-commandcode-provider`) must load so pi resolves a saved default model from that provider. With `--no-extensions`, headless `deviate meso/micro` runs silently fell back to the first model the operator's env keys authenticate (observed: `deepseek/deepseek-v4-pro` → HTTP 402 `Insufficient Balance`). Lean tool policy is unchanged: `--tools read,bash,edit,write`, `--no-skills`, optional `--skill`. Specs: `specs/DeviaTDD-api.md` (7), `specs/DeviaTDD-architecture.md` §10. Pinned by `tests/test_core/test_agent.py`, `tests/core/test_agent.py`, and `tests/e2e/test_pi_spawn_lean_tool_schema.bats`.
|
|
84
|
+
- **`deviate setup` on a TTY always picks exactly one agent and lists every optional pack.** Bare `setup` no longer fans out to every leftover `.claude/` / `.opencode/` / … directory via `detect_agents`, and no longer silently uses `detected[0]`. `--agent <name>` pins that one target without prompting (including when leftover dirs exist). On a TTY, omitted `--agent` always shows the Rich agent menu (`AGENT_CHOICES`; existing `[agent].backend` is the default highlight, not an auto-skip). Non-TTY without `--agent` reuses a persisted backend or fail-closes with `NO_AGENT_SELECTED`. The optional-pack prompt is a Rich selector that names every pack (`merge`, `pr`, `review`, `walkthrough`, `html`, `hotfix`, `triage`, `prune`, `e2e`) plus `none` / `all-optional` (default `none`); `--packs` stays for scripts. `_resolve_install_agents` always returns a one-element list. Pinned by `tests/test_cli/test_setup.py::TestSetupPacks` / `::TestSetupPerAgentInstall` and `tests/e2e/test_setup_config_rework.bats`.
|
|
85
|
+
- **`deviate setup --agent X` now installs slash commands and the packaged `deviatdd` skill only for X.** Previously setup wrote into all five of `.claude/`, `.opencode/`, `.factory/`, `.pi/`, and `.omp/` regardless of `--agent`; the flag only persisted `[agent].backend`. `droid` still writes `.factory/` only. Unknown `--agent` values stay fail-closed. Pinned by `tests/test_cli/test_setup.py::TestSetupSelectedAgentIsolation`.
|
|
58
86
|
- **JUDGE prompt injection now uses the Judge-Feedback-stripped task card (GH-118).** `_build_auto_prompt("judge")` still reads the raw card via `_task_card_text`, then applies the same `_strip_judge_feedback` pass that `resolve_task_ac_tokens` already used for token resolution (#89). Prior-round `**Judge Feedback**` bullets and their continuation lines (including false ownership claims such as "AC-PLAN-003 belongs to a later task") no longer appear in `<task_card>`. Token resolution, `_append_judge_feedback`, and the COMPLETED evidence gate are unchanged. Specs: `specs/DeviaTDD-api.md`, `specs/DeviaTDD-architecture.md`. Pinned by `tests/test_micro/test_judge.py::TestJudgePromptStripsJudgeFeedback`.
|
|
59
87
|
- **Ledger appenders no longer concatenate a new JSONL record onto a last line that lacks a trailing newline (GH-117).** `_append_record` and `_append_with_compound_key` insert a leading `\n` when the file is non-empty and does not already end in a newline, so `_read_ledger` cannot skip a fused BACKLOG+SPECIFIED line. `claim_issue` now writes through `append_issue_transition` instead of a raw `"a"` append. Pinned by `tests/test_state/test_ledger.py` and `tests/test_core/test_issues.py`.
|
|
60
88
|
- **JUDGE handover parse now recovers unescaped `"` inside evidence quote fields (GH-116).** `AgentBackend.parse_output` retries `yaml.safe_load` after rewriting broken `quote` / `test_quote` / `impl_quote` double-quoted scalars as `|` block scalars, so citations such as `assert "YAGNI" in text` and `== ["AC-PLAN-002"]` no longer raise `MalformedHandoverManifestError`. Well-formed YAML is unchanged. The auto judge prompt also prefers `|` block scalars when a quote contains `"`. Pinned by `tests/test_core/test_agent.py`.
|
|
@@ -92,7 +92,3 @@ Prefer `libref query <lib> "<topic>"` over web fetching. Workflow: `libref list`
|
|
|
92
92
|
## 📝 Prompt Edit Discipline
|
|
93
93
|
|
|
94
94
|
Edit skill/prompt templates in `src/deviate/prompts/` only. `~/.config/opencode/skills/` is a read-only install mirror.
|
|
95
|
-
|
|
96
|
-
## 🌳 Graphite (when `graphite = true` in `.deviate/config.toml`)
|
|
97
|
-
|
|
98
|
-
`gt create -am "msg"` (or `-m` if tree clean) → `gt submit --stack` → `gt sync`. Never mix `git checkout -b` with `gt`, never `gh pr create` when Graphite is on.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: deviatdd
|
|
3
|
-
Version: 2.
|
|
3
|
+
Version: 2.24.0
|
|
4
4
|
Summary: DeviaTDD CLI — agent orchestration framework
|
|
5
5
|
Project-URL: Homepage, https://github.com/wernerbisschoff/deviatdd
|
|
6
6
|
Project-URL: Repository, https://github.com/wernerbisschoff/deviatdd
|
|
@@ -42,7 +42,7 @@ Description-Content-Type: text/markdown
|
|
|
42
42
|
|
|
43
43
|
# DeviaTDD
|
|
44
44
|
|
|
45
|
-
> **An agent-orchestration framework that runs your entire TDD loop — explore, spec, red, green, refactor — with
|
|
45
|
+
> **An agent-orchestration framework that runs your entire TDD loop — explore, spec, red, green, refactor — with two hard human-in-the-loop gates (design/contract review and merge review); shard is a soft review.**
|
|
46
46
|
|
|
47
47
|
DeviaTDD is a CLI that coordinates AI coding agents across the full Test-Driven Development lifecycle, from problem framing through documentation. It ships with a four-layer architecture (Product · Macro · Meso · Micro), append-only ledgers, worktree isolation, and path-scoped GREEN writes. The system is **agent-agnostic** — Claude Code, OpenCode, Pi, Droid, the Factory Droid IDE, and Oh-My-Pi are first-class backends today.
|
|
48
48
|
|
|
@@ -54,7 +54,7 @@ Most AI coding agents stop at "write code that passes." DeviaTDD goes further
|
|
|
54
54
|
|
|
55
55
|
| Without DeviaTDD | With DeviaTDD |
|
|
56
56
|
|------------------|---------------|
|
|
57
|
-
| Agent writes code, you review after |
|
|
57
|
+
| Agent writes code, you review after | Two hard human gates (design/contract review, merge); shard is a soft review |
|
|
58
58
|
| Test edits slip in silently during "GREEN" | JUDGE flags out-of-scope writes to `tests/`, `specs/`, or protected modules as `COMPLIANCE_VIOLATION` |
|
|
59
59
|
| Lost track of which task is in which state | Append-only JSONL ledgers derive canonical state |
|
|
60
60
|
| Branch drift between parallel features | Worktree isolation + append-only ledger merge driver |
|
|
@@ -72,16 +72,19 @@ Most AI coding agents stop at "write code that passes." DeviaTDD goes further
|
|
|
72
72
|
uv tool install deviatdd
|
|
73
73
|
deviate --version # confirm install
|
|
74
74
|
|
|
75
|
-
#
|
|
76
|
-
#
|
|
77
|
-
#
|
|
78
|
-
#
|
|
79
|
-
#
|
|
80
|
-
#
|
|
81
|
-
|
|
75
|
+
# Step 1 — bootstrap DeviaTDD into this repo.
|
|
76
|
+
# Writes .deviate/, persists the agent, installs default packs
|
|
77
|
+
# (including deviate-init) and the shared deviatdd skill for the
|
|
78
|
+
# selected agent only. Does not write specs/constitution.md,
|
|
79
|
+
# mise.toml, or specs/issues.jsonl — that is /deviate-init.
|
|
80
|
+
# A TTY session prompts for the agent; --agent skips the prompt
|
|
81
|
+
# (`droid` writes `.factory/`; `codex` writes `.agents/skills/`).
|
|
82
|
+
deviate setup --agent claude # or: opencode | pi | droid | factory | omp | codex
|
|
82
83
|
```
|
|
83
84
|
|
|
84
|
-
Once setup
|
|
85
|
+
Once setup finishes, open your agent and run **`/deviate-init` as the first prompt**. That scaffolds `specs/constitution.md`, `mise.toml`, and `specs/issues.jsonl`, skipping anything already present. Codex uses the same prompt, installed as the `deviate-init` skill.
|
|
86
|
+
|
|
87
|
+
Then drive the rest of the lifecycle from inside your agent. Each phase emits a single artifact; **`post` often commits it** (you did not type `git commit`). At the two gates the workflow pauses for human review. See [Phase transparency](#phase-transparency) for which commands commit, spawn, or fail closed.
|
|
85
88
|
|
|
86
89
|
**Product layer** *(optional, for cross-product framing — skip if your repo only ships single features):*
|
|
87
90
|
|
|
@@ -104,7 +107,7 @@ Once setup is done, drive the entire lifecycle from inside your agent. Each phas
|
|
|
104
107
|
/deviate-adhoc "Add a /healthz endpoint" # condenses explore+research+prd+shard into one issue
|
|
105
108
|
```
|
|
106
109
|
|
|
107
|
-
**Meso — claim the issue and enter its worktree.** Meso slash commands run inside the per-issue worktree; claim first, then `cd` in:
|
|
110
|
+
**Meso — claim the issue and enter its worktree.** Default meso uses a worktree. `claim_remote` defaults **false** (local claim only; no push lock). Path A coworker flow stays in this clone: `deviate meso run --no-setup --local`. Meso slash commands run inside the per-issue worktree; claim first, then `cd` in:
|
|
108
111
|
|
|
109
112
|
```
|
|
110
113
|
# From the main checkout (NOT inside a worktree):
|
|
@@ -113,6 +116,9 @@ deviate specify # auto-claim the next unblocked BACKL
|
|
|
113
116
|
deviate specify ISS-001-007 # claim that exact issue; same worktree creation
|
|
114
117
|
cd $(deviate specify ISS-001-007 2>&1 | grep '^WORKTREE' | awk '{print $2}')
|
|
115
118
|
# then re-open the agent inside the worktree
|
|
119
|
+
|
|
120
|
+
# Path A (this clone, no worktree, no remote lock):
|
|
121
|
+
deviate meso run --no-setup --local --issue ISS-001-007
|
|
116
122
|
```
|
|
117
123
|
|
|
118
124
|
**Meso** — with the worktree active, decompose into tasks. `tasks.md` is the human's execution blueprint:
|
|
@@ -140,11 +146,12 @@ cd $(deviate specify ISS-001-007 2>&1 | grep '^WORKTREE' | awk '{print $2}')
|
|
|
140
146
|
/deviate-execute T002 # skips the TDD cycle; still has its own JUDGE pass
|
|
141
147
|
```
|
|
142
148
|
|
|
143
|
-
**Release** — close the loop:
|
|
149
|
+
**Release** — close the loop. `/deviate-pr`, `/deviate-review`, and `/deviate-walkthrough` are **optional packs** (not in default setup on current main):
|
|
144
150
|
|
|
145
151
|
```
|
|
146
152
|
/deviate-pr T001 # conventional-commit PR; merge appends COMPLETED
|
|
147
|
-
/deviate-review # ← Gate 3:
|
|
153
|
+
/deviate-review # ← Gate 3: comments-only PR scan (not a merge gate)
|
|
154
|
+
/deviate-walkthrough # four-look map (brief, tests, production vs checks, command)
|
|
148
155
|
```
|
|
149
156
|
|
|
150
157
|
**Or, run the unattended one-shot pipeline** — the top-level
|
|
@@ -228,7 +235,8 @@ style Ex fill:#f5e1e1
|
|
|
228
235
|
|
|
229
236
|
| Phase | Slash command | Artifact committed | What the human reviews / decides |
|
|
230
237
|
|-------|---------------|--------------------|----------------------------------|
|
|
231
|
-
| **Bootstrap** | `deviate setup --agent <name
|
|
238
|
+
| **Bootstrap · Setup** | `deviate setup [--agent <name>]` | `.deviate/config.toml`, default execution-layer packs (macro + meso + micro, including `deviate-init`), shared `deviatdd` skill, selected-agent `/deviate-*` commands | Confirm the one agent install. TTY always shows the agent selector (existing backend is the default highlight) and a checkbox list of optional packs (`product` first; Space toggles, Enter confirms; default none). |
|
|
239
|
+
| **Bootstrap · Init** | `/deviate-init` | `specs/constitution.md`, `mise.toml`, `specs/issues.jsonl` (skips files already present) | **First prompt after setup.** Codex: the `deviate-init` skill. No-op if the repo is already scaffolded. |
|
|
232
240
|
| **Product · Flows** | `/deviate-flows` | `specs/_product/flows/flows-<domain>.md` + updated `specs/_product/flows/index.md` | Confirm the actor, job-to-be-done, and trigger are right; commit the flow file when asked. |
|
|
233
241
|
| **Product · Architecture** | `/deviate-architecture` | `specs/_product/architecture.md`, `specs/_product/domain-model.md` | Reads existing flows; classify the change as Local / Context-Bridging / Context-Creating; commit when satisfied. |
|
|
234
242
|
| **Product · Release** | `/deviate-release` | `specs/_product/release-next.md` (overrides previous) | Supply a release-goal sentence; confirm the Included Flows / Included Work / Acceptance tables reflect that goal; commit. |
|
|
@@ -236,7 +244,7 @@ style Ex fill:#f5e1e1
|
|
|
236
244
|
| **Macro · Research** *(Gate 1)* | `/deviate-research` | `specs/{epic}/design.md`, `specs/{epic}/data-model.md` | **Gate 1**: approve the design + data-model before PRD synthesis. |
|
|
237
245
|
| **Macro · PRD** | `/deviate-prd` | `specs/{epic}/prd.md` (FR list + acceptance criteria) | Verify each FR is testable; commit. |
|
|
238
246
|
| **Macro · Shard** | `/deviate-shard` | `specs/{epic}/issues/ISS-NNN-*.md` (one file per vertical slice), with `flow_refs:` frontmatter and embedded `## User Stories Ledger` / `## ATDD Acceptance Criteria` sections | Review every sharded issue for completeness, edge cases, and scope (soft review — the system auto-advances to Meso and does not block). Issues are born as full specs — the user-facing *spec content* is embedded here, but **claiming and worktree creation is a separate CLI step (`deviate specify`)** that runs after `/deviate-shard` and before the meso slash commands below. |
|
|
239
|
-
| **Meso · Specify** | `deviate specify [ISS-NNN-NNN]` | A git worktree at `.worktrees/<branch
|
|
247
|
+
| **Meso · Specify** | `deviate specify [ISS-NNN-NNN]` | A git worktree at `.worktrees/<branch>/` and a claim entry appended to `specs/issues.jsonl`. `claim_remote` defaults **false** (local only). Push-as-lock is opt-in (`--claim-remote` / `claim_remote = true`). | The setup step before plan/tasks. With no argument, auto-claims the next unblocked BACKLOG issue; with an explicit ID, claims that issue. Stops after the worktree is created — does NOT advance session state and does NOT run plan or tasks. `cd` into the printed worktree path before running any other meso slash command. Path A: `deviate meso run --no-setup --local` stays in this clone. |
|
|
240
248
|
| **Run** *(full pipeline, end-to-end)* | `deviate run` | Worktree at `.worktrees/<branch>/`, `tasks.md`, `tasks.jsonl`, then completed task commits | The canonical "go do the next thing" command. Discovers the next BACKLOG issue, claims it (creating a per-issue worktree), runs SPECIFY → PLAN → TASKS in that worktree, then drains every PENDING task through the TDD cycle. Forwards `--profile` / `--no-judge` / `--no-refactor` / `--agent` / `--json` to the micro drain. Internally calls `deviate meso run` then `deviate micro run --all` inside the created worktree. |
|
|
241
249
|
| **Meso · Plan** | `/deviate-plan` | `specs/{epic}/issues/ISS-NNN/plan.md` (per-issue localized research, workstation file structure) | **Must be invoked inside the worktree that `deviate specify` created.** Review the workstation mapping and the integration surface listed; commit. Optional when shard already embedded spec sections. |
|
|
242
250
|
| **Meso · Tasks** | `/deviate-tasks` | `specs/{epic}/issues/ISS-NNN/tasks.md` + `specs/{epic}/tasks.jsonl` (append-only ledger) | **Must be invoked inside the same worktree.** The `tasks.md` artifact is the human's execution blueprint. Verify: 4–8 tasks per issue, every task has a Verification CLI command, each task declares a Mode (`TDD` or `IMMEDIATE`) and Type, DAG `blocked_by` deps are right. TDD tasks flow to red→green→judge→refactor; IMMEDIATE tasks route to `/deviate-execute`. |
|
|
@@ -246,24 +254,43 @@ style Ex fill:#f5e1e1
|
|
|
246
254
|
| **Micro · Refactor** | `/deviate-refactor <task-id>` | Polished, behavior-preserving code (only on `JUDGE_PASS`) | If the refactor breaks tests, the CLI discards it and the task completes on the verified GREEN. |
|
|
247
255
|
| **Micro · Execute** | `/deviate-execute <task-id>` | A targeted change for `direct` / `e2e` tasks | Skips the TDD cycle; still has its own JUDGE pass. |
|
|
248
256
|
| **Micro · Run** *(agent-internal drain)* | `deviate micro run [task-id] --all` | Completed task commits per the cycle | Agent-internal dispatch — `deviate micro run <task-id>` runs a single task; `deviate micro run --all` drains every PENDING task. Top-level `deviate run` invokes this with `--all` inside the worktree the meso step just created. Forwards `--profile` / `--no-judge` / `--no-refactor` / `--agent` / `--json`. |
|
|
249
|
-
| **Release** | `/deviate-pr <task-id>` | A conventional-commit PR | Open the PR; on merge, the issue ledger is appended with `COMPLETED`. |
|
|
250
|
-
| **Release** *(Gate 3)* | `/deviate-review` |
|
|
257
|
+
| **Release** | `/deviate-pr <task-id>` | A conventional-commit PR | Optional pack. Open the PR; on merge, the issue ledger is appended with `COMPLETED`. |
|
|
258
|
+
| **Release** *(Gate 3)* | `/deviate-review` | Comments-only PR scan (optional pack) | **Gate 3**: comments only by default (stdout / GitHub COMMENT). Not a merge gate. Opt-in `--apply` is CRITICAL-only. Auto-apply of CRITICAL+SUGGESTION is old PyPI 2.23.1 — not current main. |
|
|
259
|
+
| **Walkthrough** | `/deviate-walkthrough` | Four-look map (optional pack; no commit) | Brief location, test hunks, production hunks vs named checks, command to run those checks. Does not approve or auto-edit. |
|
|
260
|
+
| **Cleanup** | `/deviate-prune` | Spy/impl tests thinned; `plan.md` / `tasks.md` and JSONL ledgers unchanged | Manual honeycomb pass for **one** issue. Drops `spy` / `impl` (marks, name tags, or untagged internal probes); keeps `behavioral` / `ac` and public input-to-output. Never deletes `plan.md`, `tasks.md`, `explore.md`, `prd.md`, or `issues/*.md`. Never touches `issues.jsonl`, `tasks.jsonl`, or `flows.jsonl`. Manual invoke only — not hooked into COMPLETED, `--all`, or the skill success loop. |
|
|
261
|
+
|
|
262
|
+
Operational tools (no gate): `/deviate-triage`, `/deviate-constitution`, `/deviate-hotfix`. `/deviate-prune` is the manual honeycomb test-thinning surface (thin CLI `deviate prune pre` / `post`; the slash command commits the cleanup).
|
|
263
|
+
|
|
264
|
+
---
|
|
265
|
+
|
|
266
|
+
## Phase transparency
|
|
267
|
+
|
|
268
|
+
`--help` and this table say which phases **commit**, **spawn an agent**, or **fail closed**. `pre` injects a JSON contract to the agent; `post` validates, writes, and often commits (you did not type `git commit`). Slash prompts tell the agent to run `pre`/`post`; auto prompts tell the agent not to — the orchestrator does. Codex spawn is `codex exec --sandbox workspace-write --ask-for-approval never` (`src/deviate/core/agent.py`). `meso run` and `micro run` nest that spawn.
|
|
269
|
+
|
|
270
|
+
Default setup on current main (#134) installs **macro + meso + micro** plus the shared `deviatdd` skill. Optional packs stay off until selected: `product` (bundle), `merge`, `pr`, `review`, `walkthrough`, `html`, `hotfix`, `triage`, `prune`, `e2e`. PyPI 2.23.1 still installs all slash files — that is not current main.
|
|
251
271
|
|
|
252
|
-
|
|
272
|
+
| Phase | Does | Commits | Debug a fail |
|
|
273
|
+
|-------|------|---------|--------------|
|
|
274
|
+
| **setup** | Writes `.deviate/`, persists one agent, installs default packs + `deviatdd`. `claim_remote` defaults **false**. | No. | Non-TTY without `--agent` → `NO_AGENT_SELECTED`. Unknown `--packs` → fail closed. Named agent missing → `AGENT_NOT_INSTALLED`. |
|
|
275
|
+
| **adhoc** | One spec-enriched issue + `FR-ADHOC-NNN` + a BACKLOG ledger row. | Yes — `post` commits artifacts. Record stays BACKLOG. | Missing problem statement; complexity HIGH without `--force`. |
|
|
276
|
+
| **meso** | Default: worktree + claim, then PLAN → TASKS (spawns the agent). Path A: `deviate meso run --no-setup --local` stays in this clone and skips the remote lock. | Yes — claim, `plan post`, `tasks post`. | `MESO_PLAN_INVALID`, `MESO_TASKS_INVALID`, `NO_CLAIMABLE_ISSUES`. |
|
|
277
|
+
| **micro** | RED → GREEN → JUDGE → REFACTOR (or EXECUTE). Spawns the agent each phase. `--profile fast` skips **JUDGE and REFACTOR**. `deviate micro run --review` is a **TTY pause before the phase commit**, not `/deviate-review`. Skill argument `review` is an agent loop policy — never pass `--review` from the skill. | Yes — each phase. RED uses `git commit --no-verify`. | `REVIEW_REQUIRES_TTY`, `TRAIN_EXHAUSTED`, `COMMIT_FAILED`. `NO_PENDING_TASKS` (exit 1) means the queue is empty. |
|
|
278
|
+
| **review** | Optional pack. `/deviate-review` is **comments-only** by default. Not a merge gate. | No (unless opt-in `--apply` landed a CRITICAL fix). | Missing named checks → `brief incomplete`. Unclaimed plan ACs stay comment input (`uncovered`); not a fail-close. |
|
|
279
|
+
| **walkthrough** | Optional pack. Four-look map: brief location, test hunks, production hunks vs named checks, command to run those checks. | No. | Missing brief / named checks: stop. |
|
|
253
280
|
|
|
254
281
|
---
|
|
255
282
|
|
|
256
283
|
## Why Each Phase Exists
|
|
257
284
|
|
|
258
|
-
DeviaTDD's phase structure is not arbitrary. Each phase exists because the alternative — an agent that skips it — produces a documented failure mode. The rationale below is split into five parts: why the four layers exist, why each Product / Macro / Meso phase exists, why the
|
|
285
|
+
DeviaTDD's phase structure is not arbitrary. Each phase exists because the alternative — an agent that skips it — produces a documented failure mode. The rationale below is split into five parts: why the four layers exist, why each Product / Macro / Meso phase exists, why the two hard human gates exist (design/contract review and merge; shard is soft), why the append-only ledgers exist, and why the TDD micro-loop is `Red → Green → Judge/Train → Refactor`. Direct article citations appear inline in italics; consolidated URLs are listed under [References](#references) below.
|
|
259
286
|
|
|
260
287
|
### Why the four layers
|
|
261
288
|
|
|
262
|
-
- **Each layer matches a different model strength and a different cost profile.** Spec authoring, issue decomposition, and isolated judgement are high-judgment, low-frequency tasks best suited to a strong model. Test writing, implementation, and refactor are high-frequency, low-judgment tasks that a cheap model can perform with the right context. Splitting them into layers routes each turn to the appropriate model
|
|
263
|
-
- **Layering converts a monolithic chat into a chain of accountable artifacts.** Each phase commits a single artifact (an exploration note, a design doc, an issue file, a test, a commit). A chain of small, committed artifacts is auditable, recoverable, and parallelizable; a single long conversation is none of those *(Agile-V
|
|
289
|
+
- **Each layer matches a different model strength and a different cost profile.** Spec authoring, issue decomposition, and isolated judgement are high-judgment, low-frequency tasks best suited to a strong model. Test writing, implementation, and refactor are high-frequency, low-judgment tasks that a cheap model can perform with the right context. Splitting them into layers routes each turn to the appropriate model. *_(UCCI's 31% cost cut is a NER cascade, not coding-agent layer routing; RoBatch is batch prompting for data-management ICL. Neither paper is evidence that Product/Macro/Meso/Micro exist.)_*
|
|
290
|
+
- **Layering converts a monolithic chat into a chain of accountable artifacts.** Each phase commits a single artifact (an exploration note, a design doc, an issue file, a test, a commit). A chain of small, committed artifacts is auditable, recoverable, and parallelizable; a single long conversation is none of those *(SDD stages Specify → Plan → Implement → Validate; Spec Kit stages Specify/Plan/Tasks/Implement. Agile-V has two layers — Agile-V lifecycle + SCOPE-V task loop — and argues for a reviewed brief rather than a long chat. None of these papers define Product/Macro/Meso/Micro)*.
|
|
264
291
|
- **The Product layer is optional because most repos only ship one feature stream at a time.** A team maintaining ten products needs cross-product framing; a team shipping one web app does not. Making the Product layer optional means DeviaTDD does not impose ceremony on teams that do not need it. *_(Design proposal — closest supporting evidence is Agile-V's R0–R3 risk-adaptive framing, which is about gate strictness, not layer optionality.)_*
|
|
265
|
-
- **The Macro / Meso / Micro split separates *what to build* from *how to build it* from *how to verify it*.** Macro is intent. Meso is structure. Micro is verification.
|
|
266
|
-
- **Escalation between layers is risk-gated, not always-on.** The strong model intervenes as judge or verifier when the repair budget exhausts or when the change touches protected modules. Routine work stays in the cheap layer; high-risk work escalates automatically *(Agile-V
|
|
292
|
+
- **The Macro / Meso / Micro split separates *what to build* from *how to build it* from *how to verify it*.** Macro is intent. Meso is structure. Micro is verification. SDD separates *what* from *how* at the feature level (Specify vs Plan); Spec Kit uses the same feature stages. That is not a documented finding that conflating what/how/verify in one prompt is an agent failure — the Survey is a neural-vs-symbolic architecture review and does not make that claim.
|
|
293
|
+
- **Escalation between layers is risk-gated, not always-on.** The strong model intervenes as judge or verifier when the repair budget exhausts or when the change touches protected modules. Routine work stays in the cheap layer; high-risk work escalates automatically *(Agile-V — R0–R3 risk-adaptive acceptance. TDDev studies protocol–model fit for web-app TDD, not layer escalation; UCCI is NER cascade routing.)*
|
|
267
294
|
|
|
268
295
|
### Why the Product layer phases exist
|
|
269
296
|
|
|
@@ -274,41 +301,40 @@ DeviaTDD's phase structure is not arbitrary. Each phase exists because the alter
|
|
|
274
301
|
|
|
275
302
|
### Why the Macro layer phases exist
|
|
276
303
|
|
|
277
|
-
- **The Macro layer is the only place where a business goal is decomposed into spec-enriched issues at the right granularity.** Without a dedicated decomposition layer, the agent either ships the goal as one monolithic change (too large to review) or as ad-hoc task lists (too small to be independently testable). The Macro layer is where the granularity decision is made *(Agile-V
|
|
278
|
-
- **Splitting Explore and Research is what lets a cheap model do the cheap work and a strong model do the strong work.** Explore is a factual codebase scan; Research is an architectural reasoning task. A combined phase would either over-pay for trivial scans or under-pay for critical design decisions. The split routes each turn to the appropriate model and gives the human a cheaper artifact to review at the cheap stage *(Spec Kit)
|
|
279
|
-
- **Keeping Explore a *what exists* artifact, not a *what to do* artifact, is what lets the research phase build on it without inheriting the scan's biases.** A factual scan is the only input a research phase can build on without smuggling in design recommendations. Conflated "scan and recommend" phases lock the research phase into a direction before the human has reviewed anything *(Spec Kit;
|
|
280
|
-
- **Splitting the PRD, the design, and the data-model into three artifacts is what lets each be reviewed by a different lens and revised on a different cadence.** The PRD is *what* the system must do (requirements); the design is *how* (architecture); the data-model is *what shape the information takes* (entities, relations). Conflated artifacts force joint review and weaken every review *(SDD; Agile-V)*.
|
|
281
|
-
- **Testable acceptance criteria are the requirement for a requirement to enter the PRD.** A requirement without criteria is a wish; a PRD full of wishes cannot be sharded because there is nothing for the issues to test against. The criterion test is what turns a wishlist into a PRD *(SDD — "passing spec tests only guarantee the code matches the spec";
|
|
282
|
-
- **Decomposing a feature into vertical slices, not horizontal layers, is what makes each issue independently shippable
|
|
283
|
-
- **Decomposing a feature into vertical slices, not horizontal layers, is what makes each issue independently shippable.** A vertical slice is a complete, testable behavior end-to-end; a horizontal layer is a file or module. Vertical slices can be reviewed for missing behavior; horizontal layers hide integration risk until merge *(TDFlow; TDDev)*.
|
|
304
|
+
- **The Macro layer is the only place where a business goal is decomposed into spec-enriched issues at the right granularity.** Without a dedicated decomposition layer, the agent either ships the goal as one monolithic change (too large to review) or as ad-hoc task lists (too small to be independently testable). The Macro layer is where the granularity decision is made *(SDD decomposes a feature through Specify → Plan → Implement → Validate; Spec Kit through Specify/Plan/Tasks/Implement — neither is Product/Macro/Meso/Micro; Agile-V has two layers — lifecycle + SCOPE-V)*.
|
|
305
|
+
- **Splitting Explore and Research is what lets a cheap model do the cheap work and a strong model do the strong work.** Explore is a factual codebase scan; Research is an architectural reasoning task. A combined phase would either over-pay for trivial scans or under-pay for critical design decisions. The split routes each turn to the appropriate model and gives the human a cheaper artifact to review at the cheap stage. *(Spec Kit's discovery/validation hooks ground Specify/Plan/Tasks/Implement in repository evidence; they are not an Explore/Research cheap-vs-strong model split. Spec Kit plan-review is optional; a 40-minute arm skips intermediate artifacts.)*
|
|
306
|
+
- **Keeping Explore a *what exists* artifact, not a *what to do* artifact, is what lets the research phase build on it without inheriting the scan's biases.** A factual scan is the only input a research phase can build on without smuggling in design recommendations. Conflated "scan and recommend" phases lock the research phase into a direction before the human has reviewed anything. *(Spec Kit discovery hooks collect repository evidence before a stage; they do not name an Explore-as-inventory phase. Agile-V treats conversation as discovery and structured artifacts as the implementation contract.)*
|
|
307
|
+
- **Splitting the PRD, the design, and the data-model into three artifacts is what lets each be reviewed by a different lens and revised on a different cadence.** The PRD is *what* the system must do (requirements); the design is *how* (architecture); the data-model is *what shape the information takes* (entities, relations). Conflated artifacts force joint review and weaken every review *(SDD separates Specify (what) from Plan (how) at the feature level; Agile-V's two layers are lifecycle + SCOPE-V, not PRD/design/data-model)*.
|
|
308
|
+
- **Testable acceptance criteria are the requirement for a requirement to enter the PRD.** A requirement without criteria is a wish; a PRD full of wishes cannot be sharded because there is nothing for the issues to test against. The criterion test is what turns a wishlist into a PRD *(SDD — "passing spec tests only guarantee the code matches the spec"; BCMS vendor blog — "~3–10× higher first-pass success rate … according to early adopter reports," not the SDD paper; Acceptance Test Gen — LLM-generated acceptance tests are usable in production at 60% as-generated, 92% after fixes; LLM BDD)*.
|
|
309
|
+
- **Decomposing a feature into vertical slices, not horizontal layers, is what makes each issue independently shippable.** A vertical slice is a complete, testable behavior end-to-end; a horizontal layer is a file or module. Vertical slices can be reviewed for missing behavior; horizontal layers hide integration risk until merge. *_(TDFlow and TDDev do not use the word "vertical." TDFlow is forced-decoupled sub-agents on SWE-bench given human-written tests; TDDev compares TDD protocols by model generation style. The vertical-slice wording is DeviaTDD's.)_*
|
|
284
310
|
- **A complexity gate (low / medium → proceed, high → reject) is what makes Adhoc safe to expose as a shortcut.** Without the gate, "adhoc" becomes a workaround for skipping ceremony on work that needs the full Macro chain. The gate is the structural mechanism that prevents the shortcut from being misused. *_(Adaptive-enforcement concept is parallel to TDDev's protocol-model fit and TDD Governance's N=3 repair cap; the specific "low/medium → proceed, high → reject" classifier is DeviaTDD-original.)_*
|
|
285
311
|
|
|
286
312
|
### Why the Meso layer phases exist
|
|
287
313
|
|
|
288
|
-
- **The Meso layer is the only place where a spec-enriched issue is decomposed into TDD-executable tasks at the right granularity.** A spec is too coarse for a single TDD cycle
|
|
289
|
-
- **Re-running Plan per issue is what keeps the "what exists now" context fresh within a sprint.** Epic-level Explore becomes stale within days — by the time the fifth issue of a feature is being planned, the codebase has changed and prior issues have shipped. Per-issue Plan reads what prior issues implemented via the issues ledger, so the context reflects the current state, not the state at the start of the epic. *_(Parallel: Mise en Place's
|
|
290
|
-
- **Human review of decomposition and machine execution of decomposition require different formats and must be separate files.** `tasks.md` is the only surface a human can read, amend, and approve task decomposition against; `tasks.jsonl` is the only surface a CLI can parse deterministically and replay across parallel branches. Combining them forces one format to compromise on both readers *(TDAD — `test_map.txt` vs `SKILL.md` separation
|
|
291
|
-
- **The 4–8 tasks-per-issue target
|
|
292
|
-
- **Explicit DAG `blocked_by` dependencies are what make the parallel work graph visible to both the CLI and the human reviewer.** A flat list of tasks with implicit order is invisible to a parallelizing CLI and uninspectable to a human looking for the critical path. The DAG is the only structure that supports parallel-execution scheduling and critical-path reasoning *(Runtime Decomp — 80.5% lower retry cost vs. static
|
|
314
|
+
- **The Meso layer is the only place where a spec-enriched issue is decomposed into TDD-executable tasks at the right granularity.** A spec is too coarse for a single TDD cycle; a task list is too fine for a human to review coherently. The Meso layer is where the granularity decision is made. *(TDAID describes Plan → Red → Green → Refactor → Validate. It does not give a 15–60 minute cycle time.)*
|
|
315
|
+
- **Re-running Plan per issue is what keeps the "what exists now" context fresh within a sprint.** Epic-level Explore becomes stale within days — by the time the fifth issue of a feature is being planned, the codebase has changed and prior issues have shipped. Per-issue Plan reads what prior issues implemented via the issues ledger, so the context reflects the current state, not the state at the start of the epic. *_(Parallel: Mise en Place's three preparation phases — contextual grounding, collaborative specification, task decomposition — and Runtime Decomp's runtime branching. Mise en Place does not name a "fresh-context per task" phase. The specific "per-issue Plan reads prior issues via the issues ledger" pattern is DeviaTDD-original.)_*
|
|
316
|
+
- **Human review of decomposition and machine execution of decomposition require different formats and must be separate files.** `tasks.md` is the only surface a human can read, amend, and approve task decomposition against; `tasks.jsonl` is the only surface a CLI can parse deterministically and replay across parallel branches. Combining them forces one format to compromise on both readers *(TDAD — `test_map.txt` vs `SKILL.md` separation: shrinking `SKILL.md` 107 → 20 lines raised resolution 12% → 50%; the 70% regression cut (6.08% → 1.82%) is an impact-map / `test_map.txt` result, not a red-first result)*.
|
|
317
|
+
- **The 4–8 tasks-per-issue target is the granularity DeviaTDD uses so each task stays reviewable and independently testable — outside that range, decomposition becomes either fragmented or bloated.** More tasks force micro-decomposition that fragments the acceptance criteria; fewer tasks hide integration risk. *_(The specific 4–8 count is DeviaTDD-original. TDAID does not state a 15–60 minute cycle.)_*
|
|
318
|
+
- **Explicit DAG `blocked_by` dependencies are what make the parallel work graph visible to both the CLI and the human reviewer.** A flat list of tasks with implicit order is invisible to a parallelizing CLI and uninspectable to a human looking for the critical path. The DAG is the only structure that supports parallel-execution scheduling and critical-path reasoning *(Runtime Decomp — on the Kubernetes RCA workload (N=10, simulated failure), static decomposition's retry cost exceeded monolithic by 80.5%; runtime-branched vs monolithic was 51.7% lower retry cost and vs static 73.2%. The 80.5% figure is static *worse than monolithic*, not a win for runtime vs static)*.
|
|
293
319
|
- **Treating the GitHub PR as a structural merge boundary, not a code-formatting step, is what gives reviewers a single artifact to review (title, body, diff, review surface).** A list of commits is a history; a PR is the unit of *what we are about to merge* *(Agile-V — SCOPE-V's verify step at "before / during / before-merge / after-deployment", treating the merge boundary as a discrete verification point)*.
|
|
294
320
|
- **Gate 3 (final PR review) is the only audit over the full atomic git history of a feature.** Per-task review sees a slice; Gate 3 sees the whole. A final human audit catches the long-tail issues — integration regressions, doc drift, scope creep — that escaped per-task validation *(extension of Agile-V's verify-step pattern)*.
|
|
295
321
|
|
|
296
322
|
### Why two non-bypassable human gates
|
|
297
323
|
|
|
298
|
-
- **Spec errors are the most expensive to fix downstream.** A bug in a contract caught at the post-shard review saves the plan, the tasks, and every TDD cycle that would have implemented the bug. The same bug caught after merge costs the bug report, the rollback, the post-mortem, and the customer trust. Task decomposition is cheap to regenerate; cascades of implemented tasks are not *(Agile-V — SCOPE-V's evidence-based acceptance
|
|
299
|
-
- **An LLM cannot self-verify its own output.** Every frontier model is a stochastic generator with zero internal semantic verification capability — the tool is irrelevant, the process is determinative. The same agent that produces a plausible design or plausible code will produce a plausible-looking review of them. A human gate at design and merge is the only verification mechanism with the necessary independence *(IACDM — "verification gap";
|
|
324
|
+
- **Spec errors are the most expensive to fix downstream.** A bug in a contract caught at the post-shard review saves the plan, the tasks, and every TDD cycle that would have implemented the bug. The same bug caught after merge costs the bug report, the rollback, the post-mortem, and the customer trust. Task decomposition is cheap to regenerate; cascades of implemented tasks are not *(Agile-V — SCOPE-V's evidence-based acceptance. The 3–10× first-pass figure is a BCMS early-adopter report, not a finding in the SDD paper)*.
|
|
325
|
+
- **An LLM cannot self-verify its own output.** Every frontier model is a stochastic generator with zero internal semantic verification capability — the tool is irrelevant, the process is determinative. The same agent that produces a plausible design or plausible code will produce a plausible-looking review of them. A human gate at design and merge is the only verification mechanism with the necessary independence *(IACDM — "verification gap"; State Contamination — memory laundering can preserve adversarial influence below classifier threshold. PRIME's executor/verifier agents are for algorithmic reasoning — sorting, automata — not TDD judge wiring)*.
|
|
300
326
|
- **Gates are cheap; the work that gates prevent is expensive.** A five-minute human check at Gate 1 prevents a multi-day agent cycle that would have built the wrong thing. The economics favor verification early (Gate 1, design) and at the merge boundary (Gate 3, final audit) — but not in between, where a working agent loop is already verifiable on its own *(Agile-V — risk-adaptive acceptance at discrete levels rather than everywhere, always)*.
|
|
301
|
-
- **"Do not let an agent implement from a long chat
|
|
327
|
+
- **"Do not let an agent implement from a long chat. Let it implement from a reviewed brief."** Gates are the mechanism that converts a long conversation into a reviewed brief. The contract is what the agent implements against; the chat is at most a source of the contract. Without gates, the implementation drifts away from the original intent as the chat lengthens *(Agile-V / SCOPE-V — direct quote from the paper)*.
|
|
302
328
|
- **Two gates, not one and not ten.** One gate at the end is too late — errors have already cascaded. A gate at every micro-step is bureaucracy. Two gates correspond to the two failure modes that compound across the lifecycle: bad design (caught at Gate 1) and bad merge (caught at Gate 3). Each gate catches the class of error that the prior phases are most likely to produce. *_(The risk-adaptive framing is supported by Agile-V's R0–R3 acceptance levels; the specific count of two is DeviaTDD-original.)_
|
|
303
329
|
|
|
304
330
|
### Why the append-only ledgers exist
|
|
305
331
|
|
|
306
|
-
- **Append-only is the merge strategy DeviaTDD uses to let parallel feature branches share state without coordination or a database.** Mutable state files would require lock-step coordination; a `.jsonl` file with `merge=union` declared in `.gitattributes` lets concurrent appends on parallel branches merge without conflict markers. The state machine scales beyond a single branch because the state format is append-only. *_(Git's `merge=union` is the structural basis.
|
|
332
|
+
- **Append-only is the merge strategy DeviaTDD uses to let parallel feature branches share state without coordination or a database.** Mutable state files would require lock-step coordination; a `.jsonl` file with `merge=union` declared in `.gitattributes` lets concurrent appends on parallel branches merge without conflict markers. The state machine scales beyond a single branch because the state format is append-only. *_(Git's `merge=union` is the structural basis. TDFlow does not discuss worktrees, isolation, immutability, or branches — it is forced-decoupled sub-agents on SWE-bench given human-written tests. The "only viable strategy" claim is a software-engineering argument, not a research finding — see References §Gaps.)_*
|
|
307
333
|
- **Deriving CLI state from the ledger, rather than caching it in a separate file, is what prevents state drift between the CLI and the repo.** The CLI's view of current task, active issue, completed work, and FR traceability is computed by sequential parsing of the ledger on demand. A separate state file would be a cache that could disagree with the source; a derived state cannot disagree with itself. *_(Design proposal — derived-state-from-log is a general software-engineering principle; no direct source supports the specific "sequential parse on demand, no cache" pattern.)_*
|
|
308
|
-
- **Recording transitions as events, rather than mutating state in place, is what makes the ledger re-derivable from history.** An event can be replayed; a mutation cannot. A corrupted state file can be reconstructed by re-running the event stream, and the canonical state can always be recomputed by re-parsing the ledger *(
|
|
334
|
+
- **Recording transitions as events, rather than mutating state in place, is what makes the ledger re-derivable from history.** An event can be replayed; a mutation cannot. A corrupted state file can be reconstructed by re-running the event stream, and the canonical state can always be recomputed by re-parsing the ledger. *_(Design proposal — PRIME's opened abstract is algorithmic reasoning with an executor/verifier/coordinator, not an event-replayable git State Stack.)_*
|
|
309
335
|
- **Deriving issue IDs from the ledger, rather than assigning them externally, is what makes them collision-free across parallel branches.** Externally-assigned IDs require coordination to avoid duplicates; ledger-derived IDs compute the next ID from the current ledger state and encode the issue's lineage by construction. The `next_issue_id` field on each shard contract is computed by parsing the ledger, not from a counter file. *_(Design proposal — the collision-free argument is a software-engineering claim, not a research finding.)_*
|
|
310
336
|
- **`flow_refs:` in each issue's frontmatter is the only mechanism that connects a code change back to the customer flow that motivated it.** Without the trace, a refactor that "improves" a vertical slice may break the flow that motivated it without anyone noticing. The trace is what makes the issue→flow→release chain auditable. *_(The traceability principle is supported by SDD's spec-first case studies; the specific `flow_refs:` frontmatter convention is DeviaTDD-original.)_*
|
|
311
|
-
- **A review surface and an execution surface serve different readers and must be separate files.** `tasks.md` is the only artifact a human can read, amend, and approve task decomposition against; `tasks.jsonl` is the only artifact a CLI can parse deterministically and replay across parallel branches. Combining them forces one format to compromise on both readers — a human-readable markdown becomes hard to parse, or a parseable JSONL becomes hard to review *(TDAD — `test_map.txt` vs `SKILL.md` separation
|
|
337
|
+
- **A review surface and an execution surface serve different readers and must be separate files.** `tasks.md` is the only artifact a human can read, amend, and approve task decomposition against; `tasks.jsonl` is the only artifact a CLI can parse deterministically and replay across parallel branches. Combining them forces one format to compromise on both readers — a human-readable markdown becomes hard to parse, or a parseable JSONL becomes hard to review *(TDAD — `test_map.txt` vs `SKILL.md` separation)*.
|
|
312
338
|
|
|
313
339
|
### Why the TDD micro-loop is `Red → Green → Judge/Train → Refactor`
|
|
314
340
|
|
|
@@ -316,7 +342,7 @@ The TDD micro-loop is what makes agent-written code trustworthy. Each phase exis
|
|
|
316
342
|
|
|
317
343
|
#### Why Red (write a failing test first)
|
|
318
344
|
|
|
319
|
-
- **Tests written after implementation tend to reflect what the code does, not what it should do.** When the same agent writes both test and implementation in one session, the implementation bleeds into the test — a failure mode known as *context pollution*. The test ends up passing trivially because it asserts whatever the implementation does, not whatever the spec requires. Forcing the test to be written first, in a session with no implementation,
|
|
345
|
+
- **Tests written after implementation tend to reflect what the code does, not what it should do.** When the same agent writes both test and implementation in one session, the implementation bleeds into the test — a failure mode known as *context pollution*. The test ends up passing trivially because it asserts whatever the implementation does, not whatever the spec requires. Forcing the test to be written first, in a session with no implementation, is DeviaTDD's structural counter to that bleed *(TDD Agent Dev; Refactor Pattern — context pollution when test and implementation share a session)*. TDAD is not evidence for red-first: its 70% regression cut (6.08% → 1.82%) and `SKILL.md` 107 → 20 / 12% → 50% are impact-map / `test_map.txt` results. The same paper's TDD Prompting Paradox: TDD instructions without a targeted test map *raised* regressions to 9.94%.
|
|
320
346
|
- **A test that passes immediately is not a test.** Confirming the test fails before any production code exists verifies that the test actually exercises the new behavior. Skipping this step risks shipping a test that exercises nothing — green by construction, useless by construction.
|
|
321
347
|
- **The test is the agent's only objective specification.** An agent given "make this work" produces plausible-looking code; an agent given "make this test pass" produces code whose correctness is mechanically checkable. The test is the only artifact in the loop that the agent cannot rationalize its way past *(TDD Agent Dev — tests as spec and guardrail)*.
|
|
322
348
|
|
|
@@ -328,11 +354,11 @@ The TDD micro-loop is what makes agent-written code trustworthy. Each phase exis
|
|
|
328
354
|
|
|
329
355
|
#### Why Judge / Train (the Green → Judge → Green loop)
|
|
330
356
|
|
|
331
|
-
- **The same agent that wrote the green code cannot reliably review it.** A self-review inherits the biases and blind spots of the producer — the same hallucinations, the same shortcuts, the same rationalizations. The Judge phase runs in an isolated session with a fresh context, breaking the recursive subjectivity of "did I do what I would have approved?" *(
|
|
357
|
+
- **The same agent that wrote the green code cannot reliably review it.** A self-review inherits the biases and blind spots of the producer — the same hallucinations, the same shortcuts, the same rationalizations. The Judge phase runs in an isolated session with a fresh context, breaking the recursive subjectivity of "did I do what I would have approved?" *(IACDM — external verification agents at discrete gates. PRIME's executor/verifier split is algorithmic-reasoning constraint checking, not a TDD judge)*.
|
|
332
358
|
- **Tests passing is necessary but not sufficient.** A green test suite verifies the implementation matches the test; it does not verify the implementation matches the spec, the architecture, the security model, the performance budget, or the protected-module list. Judge evaluates the production diff against the contract for invariant, security, and structural violations the test cannot express. A separate human-style validation step at the end of an agentic session is required even when tests are green *(IACDM — verification gap; State Contamination — even classifier-cleaned memory can carry adversarial influence)*.
|
|
333
|
-
- **Bounded repair with feedback injection converts failure into a learning signal.** When Judge rejects, the CLI rolls the task back to the RED commit (a known-good state), injects the failure feedback into the next GREEN prompt, and retries — up to three times. The next attempt has the same context plus the explicit feedback of why the previous attempt failed. This is the "Train" half: the failure is preserved as a constraint on the next attempt, not discarded as a dead end *(TDD Governance — N=3 repair cap
|
|
359
|
+
- **Bounded repair with feedback injection converts failure into a learning signal.** When Judge rejects, the CLI rolls the task back to the RED commit (a known-good state), injects the failure feedback into the next GREEN prompt, and retries — up to three times. The next attempt has the same context plus the explicit feedback of why the previous attempt failed. This is the "Train" half: the failure is preserved as a constraint on the next attempt, not discarded as a dead end *(TDD Governance — N=3 GREEN repair cap, a chosen bound; the paper has four *stages* — planning, generation, repair, validation — not "4 validation gates." Parallel: TDAD treats regression as a first-class metric via a targeted test map, not via red-first prompting)*.
|
|
334
360
|
- **Three retries is enough; more would be a sign of a wrong test or a wrong spec.** A green implementation that fails Judge three times is unlikely to converge on the fourth. The bound forces escalation back to the human (an amendment, a new plan, or a spec revision) rather than burning compute on a fundamentally misaligned task. The empirical cap `N=3` is consistent with documented multi-agent TDD governance practice *(TDD Governance)*.
|
|
335
|
-
- **Rollback is to a known-good commit, not a fresh start.** The RED commit is the verified-good test boundary. Resetting to it discards the suspect GREEN cleanly, but preserves all prior work — the test, the spec, the agent session, the audit trail. Starting from a fresh checkout would also discard the test, which is the only artifact whose correctness has actually been confirmed *(
|
|
361
|
+
- **Rollback is to a known-good commit, not a fresh start.** The RED commit is the verified-good test boundary. Resetting to it discards the suspect GREEN cleanly, but preserves all prior work — the test, the spec, the agent session, the audit trail. Starting from a fresh checkout would also discard the test, which is the only artifact whose correctness has actually been confirmed *(TDD Governance — proposal-execution separation. TDFlow does not discuss worktrees, isolation, or immutable branch state)*.
|
|
336
362
|
|
|
337
363
|
#### Why Refactor (behavior-preserving improvement)
|
|
338
364
|
|
|
@@ -349,34 +375,34 @@ The rationale above grounds each architectural choice in published agentic-engin
|
|
|
349
375
|
|
|
350
376
|
### Methodology (de jure) — frameworks and governance
|
|
351
377
|
|
|
352
|
-
- [**Agile-V / SCOPE-V**](https://arxiv.org/abs/2605.20456) — Agentic-Agile vs Vibe-Coding: Verified Engineering. Defines R0–R3 risk-adaptive acceptance; SCOPE-V's verify step at "before / during / before-merge / after-deployment"; "
|
|
378
|
+
- [**Agile-V / SCOPE-V**](https://arxiv.org/abs/2605.20456) — Agentic-Agile vs Vibe-Coding: Verified Engineering. Two layers (Agile-V lifecycle + SCOPE-V task loop), not Product/Macro/Meso/Micro. Defines R0–R3 risk-adaptive acceptance; SCOPE-V's verify step at "before / during / before-merge / after-deployment"; "Do not let an agent implement from a long chat. Let it implement from a reviewed brief."
|
|
353
379
|
- [**IACDM**](https://arxiv.org/abs/2604.16399) — Interactive Adversarial Convergence Development Methodology. 8-phase framework with external verification agents at discrete gates; "the tool is irrelevant, the process is determinative"; foundational source for the "verification gap" rationale.
|
|
354
|
-
- [**PRIME**](https://doi.org/10.20944/preprints202601.1479.v1) — Policy-Reinforced Iterative Multi-Agent Execution.
|
|
380
|
+
- [**PRIME**](https://doi.org/10.20944/preprints202601.1479.v1) — Policy-Reinforced Iterative Multi-Agent Execution for Algorithmic Reasoning. Crossref abstract: executor / verifier / coordinator agents for sorting, automata, and state-machine tasks. Full text 403 on preprints.org. Not a git ledger or TDD judge paper.
|
|
355
381
|
- [**State Contamination**](https://arxiv.org/abs/2605.16746) — State Contamination in Memory-Augmented LLM Agents. Memory laundering can preserve adversarial influence below classifier thresholds.
|
|
356
|
-
- [**Survey**](https://doi.org/10.1007/s10462-
|
|
357
|
-
- [**UCCI**](https://arxiv.org/abs/2605.18796) — Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing. 31% cost reduction at
|
|
358
|
-
- [**RoBatch**](https://doi.org/10.14778/3734839.3734853) — Cost-Effective LLMs
|
|
382
|
+
- [**Survey**](https://doi.org/10.1007/s10462-025-11422-4) — Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions ([arXiv:2510.25445](https://arxiv.org/abs/2510.25445)). Dual-paradigm review of neural vs symbolic agent architectures. Not a source for "conflating what/how/verify in one prompt."
|
|
383
|
+
- [**UCCI**](https://arxiv.org/abs/2605.18796) — Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing. 31% cost reduction at micro-F1 = 0.91 on a production NER workload (4B/12B cascade). Not coding-agent layer routing.
|
|
384
|
+
- [**RoBatch**](https://doi.org/10.14778/3734839.3734853) — Optimized Batch Prompting for Cost-Effective LLMs. Batch prompting for data-management in-context learning. Not cascade routing of spec vs implement.
|
|
359
385
|
|
|
360
386
|
### Methodology (de jure) — process and decomposition
|
|
361
387
|
|
|
362
|
-
- [**SDD**](https://arxiv.org/abs/2602.00180) — Spec-Driven Development: From Code to Contract.
|
|
363
|
-
- [**Spec Kit**](https://arxiv.org/abs/2604.05278) — Spec-Kit Agents: Context-Grounded Agentic Workflows.
|
|
364
|
-
- [**Mise en Place**](https://arxiv.org/abs/2605.05400) — Mise en Place for Agentic Coding. Three
|
|
365
|
-
- [**Runtime Decomp**](https://arxiv.org/abs/2605.15425) — Runtime-Structured Task Decomposition for Agentic Coding Systems.
|
|
388
|
+
- [**SDD**](https://arxiv.org/abs/2602.00180) — Spec-Driven Development: From Code to Contract. Feature stages Specify → Plan → Implement → Validate (tools often add Tasks). "A passing spec test doesn't guarantee correct software—it only guarantees that the software matches the spec." Does **not** contain a 3–10× first-pass figure.
|
|
389
|
+
- [**Spec Kit**](https://arxiv.org/abs/2604.05278) — Spec-Kit Agents: Context-Grounded Agentic Workflows. Feature stages Specify/Plan/Tasks/Implement with discovery/validation hooks; SPEC.md + PLAN.md + TASKS.md intermediate artifacts. Plan-review is optional (auto-approved in the study); a 40-minute arm skips those artifacts.
|
|
390
|
+
- [**Mise en Place**](https://arxiv.org/abs/2605.05400) — Mise en Place for Agentic Coding. Three named phases: contextual grounding, collaborative specification, task decomposition. Does not name a "fresh-context per task" phase.
|
|
391
|
+
- [**Runtime Decomp**](https://arxiv.org/abs/2605.15425) — Runtime-Structured Task Decomposition for Agentic Coding Systems. N=10, simulated failure: static retry cost exceeded monolithic by 80.5% (RCA); runtime-branched vs monolithic 51.7% lower retry cost, vs static 73.2%.
|
|
366
392
|
|
|
367
393
|
### Implementation (de facto) — TDD, agents, and refactor
|
|
368
394
|
|
|
369
|
-
- [**TDAD**](https://arxiv.org/abs/2603.17973) — Test-Driven Agentic Development.
|
|
370
|
-
- [**TDFlow**](https://arxiv.org/abs/2510.23761) — TDFlow: Agentic Workflows for Test-Driven Development. Forced-
|
|
371
|
-
- [**TDDev**](https://arxiv.org/abs/2605.17242) — From Runnable to Shippable: Multi-Agent TDD. Protocol-model fit
|
|
372
|
-
- [**TDD Governance**](https://arxiv.org/abs/2604.26615) — TDD Governance for Multi-Agent Code Generation. N=3 repair cap
|
|
373
|
-
- [**TDAID**](https://www.awesome-testing.com/2025/10/test-driven-ai-development-tdaid) — Test-Driven AI Development. Plan → Red → Green → Refactor → Validate
|
|
395
|
+
- [**TDAD**](https://arxiv.org/abs/2603.17973) — Test-Driven Agentic Development. Impact-map / `test_map.txt` results: 70% regression reduction (6.08% → 1.82%); shrinking `SKILL.md` 107 → 20 lines raised resolution 12% → 50%. TDD Prompting Paradox: TDD instructions without a targeted test map raised regressions to 9.94%. Not a red-first proof.
|
|
396
|
+
- [**TDFlow**](https://arxiv.org/abs/2510.23761) — TDFlow: Agentic Workflows for Test-Driven Development. Forced-decoupled sub-agents (propose / debug / revise / optional test generation) on SWE-bench given human-written tests. Does not discuss worktrees, isolation, immutability, branches, or append-only ledgers.
|
|
397
|
+
- [**TDDev**](https://arxiv.org/abs/2605.17242) — From Runnable to Shippable: Multi-Agent TDD. Protocol-model fit: holistic models benefit from agentic TDD; conservative read-then-extend models benefit from incremental TDD. Does not use the word "vertical."
|
|
398
|
+
- [**TDD Governance**](https://arxiv.org/abs/2604.26615) — TDD Governance for Multi-Agent Code Generation. N=3 GREEN repair cap (chosen bound); four *stages* (planning, generation, repair, validation); design hygiene (refactor continuously while green); proposal-execution separation.
|
|
399
|
+
- [**TDAID**](https://www.awesome-testing.com/2025/10/test-driven-ai-development-tdaid) — Test-Driven AI Development. Plan → Red → Green → Refactor → Validate; local commits after each TDD phase. Does not state a 15–60 minute cycle.
|
|
374
400
|
- [**Refactor Pattern**](https://agentpatterns.ai/verification/red-green-refactor-agents/) — Red-Green-Refactor with Agents: Tests as the Spec. Refactor must be behavior-preserving, never test-preserving; failed refactors are discarded.
|
|
375
401
|
- [**TDD Agent Dev**](https://agentpatterns.ai/verification/tdd-agent-development/) — Test-Driven Agent Development: Tests as Spec and Guardrail. Tests-as-spec-guardrail pattern; structural enforcement against test-hacking.
|
|
376
402
|
|
|
377
403
|
### Specification and acceptance
|
|
378
404
|
|
|
379
|
-
- [**Definitive SDD**](https://thebcms.com/blog/spec-driven-development) — Spec-Driven Development: The Definitive 2026 Guide. EARS notation (Ubiquitous / Event-driven / State-driven / Unwanted / Optional).
|
|
405
|
+
- [**Definitive SDD**](https://thebcms.com/blog/spec-driven-development) — Spec-Driven Development: The Definitive 2026 Guide. EARS notation (Ubiquitous / Event-driven / State-driven / Unwanted / Optional). "~3–10× higher first-pass success rate from AI agents on non-trivial tasks, according to early adopter reports from GitHub and AWS."
|
|
380
406
|
- [**Acceptance Test Gen**](https://arxiv.org/abs/2504.07244) — Acceptance Test Generation with LLMs (Industrial Case Study). 95% helpfulness, 92% semantic relevance, 60% directly usable as generated.
|
|
381
407
|
- [**LLM BDD**](https://arxiv.org/abs/2403.14965) — Comprehensive Evaluation: LLMs for BDD Acceptance Test Formulation. Comprehensive evaluation of GPT-3.5/4, Llama-2, PaLM-2 on BDD generation.
|
|
382
408
|
|
|
@@ -388,11 +414,11 @@ The rationale above grounds each architectural choice in published agentic-engin
|
|
|
388
414
|
|
|
389
415
|
Claims in this README flagged with an italic _design proposal_ note have **no direct source** in the agentic-engineering literature at the time of writing. They are explicit gaps in the evidence chain; treat them as DeviaTDD design choices, not research findings:
|
|
390
416
|
|
|
391
|
-
- **Append-only JSONL over mutable state** — the "only viable merge strategy across parallel branches" claim is a software-engineering argument supported by git's `merge=union` semantics
|
|
417
|
+
- **Append-only JSONL over mutable state** — the "only viable merge strategy across parallel branches" claim is a software-engineering argument supported by git's `merge=union` semantics, not a research finding (TDFlow does not discuss isolation or ledgers).
|
|
392
418
|
- **Product layer optionality; Flows / Architecture / Release triad; single-sentence release goal** — DeviaTDD-original; closest support is SDD's spec-from-plan-from-implementation separation at the feature level.
|
|
393
|
-
- **4–8 tasks per issue** —
|
|
419
|
+
- **4–8 tasks per issue** — DeviaTDD-original; TDAID does not state a 15–60 minute cycle.
|
|
394
420
|
- **Per-issue Plan cadence; Adhoc complexity classifier; ledger-derived issue IDs; `flow_refs:` frontmatter convention; deriving CLI state from the ledger** — DeviaTDD-original; parallel support from adjacent work exists but does not directly cover these patterns.
|
|
395
|
-
- **Two gates, not one and not ten** — the risk-adaptive framing is supported (Agile-V R0–R3); the specific count of two is DeviaTDD-original.
|
|
421
|
+
- **Two gates, not one and not ten** — the risk-adaptive framing is supported (Agile-V R0–R3); the specific count of two is DeviaTDD-original. Design/contract review and merge review are the hard gates; shard is a soft review.
|
|
396
422
|
|
|
397
423
|
---
|
|
398
424
|
|