easy-coding-harness 0.10.0-beta.1 → 0.10.0-beta.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/CHANGELOG.md +151 -0
  2. package/README.md +50 -20
  3. package/dist/cli.js +478 -47
  4. package/dist/cli.js.map +1 -1
  5. package/package.json +1 -1
  6. package/templates/claude/agents/ec-implementer.md +11 -0
  7. package/templates/claude/agents/ec-reviewer.md +6 -1
  8. package/templates/codex/agents/ec-implementer.toml +11 -0
  9. package/templates/codex/agents/ec-reviewer.toml +6 -1
  10. package/templates/common/bundled-skills/ec-init/SKILL.md +18 -6
  11. package/templates/common/bundled-skills/ec-meta/references/local-architecture/README.md +27 -12
  12. package/templates/common/bundled-skills/ec-meta/references/platform-files/README.md +1 -1
  13. package/templates/common/skills/ec-analysis/SKILL.md +141 -35
  14. package/templates/common/skills/ec-config/SKILL.md +24 -2
  15. package/templates/common/skills/ec-git/SKILL.md +7 -1
  16. package/templates/common/skills/ec-implementing/SKILL.md +61 -1
  17. package/templates/common/skills/ec-memory/SKILL.md +76 -6
  18. package/templates/common/skills/ec-reviewing/SKILL.md +21 -3
  19. package/templates/common/skills/ec-task-close/SKILL.md +4 -0
  20. package/templates/common/skills/ec-task-management/SKILL.md +7 -1
  21. package/templates/common/skills/ec-tdd-init/SKILL.md +101 -0
  22. package/templates/common/skills/ec-verification/SKILL.md +68 -26
  23. package/templates/common/skills/ec-workflow/SKILL.md +81 -22
  24. package/templates/main-constraint/AGENTS.md.tpl +49 -14
  25. package/templates/main-constraint/CLAUDE.md.tpl +46 -14
  26. package/templates/qoder/agents/ec-implementer.md +11 -0
  27. package/templates/qoder/agents/ec-reviewer.md +6 -1
  28. package/templates/runtime/templates/dev-spec-skeleton.md +8 -1
  29. package/templates/runtime/tools/easy_coding_tdd_readiness.py +306 -0
  30. package/templates/shared-hooks/easy_coding_state.py +3878 -259
  31. package/templates/shared-hooks/easy_dev_spec.py +444 -30
  32. package/templates/shared-hooks/easy_dev_spec_execution.py +1014 -0
  33. package/templates/shared-hooks/easy_dev_spec_protocol.py +1426 -18
@@ -36,6 +36,37 @@ Communicate with the user in the user's language.
36
36
  enters REVIEW.
37
37
  7. Read-only `doc` / `analysis` / `report` tasks remain `single` with `files:[]`, make no writes,
38
38
  return a non-empty `deliverable`, then follow the mode-aware IMPLEMENT -> COMPLETE edge.
39
+ 8. When a project template, local convention, or new source header uses author attribution, the
40
+ author value must be `<Current Agent Name> with Easy Coding`, for example
41
+ `Codex with Easy Coding`. `Current Agent Name` means the user-facing host Agent (for example,
42
+ Codex, Claude, or Qoder), never an implementation sub-agent role such as `ec-implementer`.
43
+ This value is display attribution only: never pass it to the workflow state API's `--agent`,
44
+ which accepts only `claude-code`, `codex`, or `qoder`. Never copy a previous human or Agent name
45
+ into newly authored code.
46
+ 9. Every newly added field in a data-bearing model must have a meaningful field-level comment.
47
+ This includes new or extended entity/DO/DTO/VO/BO, request/response, configuration, and similar
48
+ model types. Every new enum member and every new declared constant requires the same treatment.
49
+ Describe the semantic meaning and, when relevant, units, format, allowed values, nullability,
50
+ default behavior, or compatibility constraints. A type-level comment does not replace comments
51
+ on its fields or members; do not add low-value comments to ordinary local variables.
52
+ 10. Treat the task card's `Local Baseline` as the default implementation shape. Match the nearest
53
+ comparable code's naming, control flow, null/empty and error handling, layering, object model,
54
+ and extraction granularity unless correctness, security, an explicit requirement, or a hard
55
+ project rule requires a deviation. Do not add defensive null checks solely because they are a
56
+ generic best practice when the evidenced local contract intentionally omits them.
57
+ 11. Implement the smallest coherent design. Do not add speculative abstractions, wrappers,
58
+ factories, layers, or extension points, and do not fragment one readable flow into many
59
+ single-use micro-methods. Extract code only for a clear semantic boundary, real reuse,
60
+ independent testability, or a material reduction in complexity.
61
+ 12. Literals and magic values are allowed when they are obvious, local, and consistent with the
62
+ surrounding code. Introduce a constant for repeated use, stable domain/config/protocol
63
+ semantics, or an established project convention—not merely to hold the single return value of
64
+ a getter.
65
+ 13. In a newly added core Java class, every method and field requires meaningful Javadoc. In an
66
+ existing core Java class, every added or materially modified method and field requires it.
67
+ Add focused inline comments to core or complex logic to explain intent, constraints, or
68
+ non-obvious tradeoffs. Do not mass-retrofit untouched legacy code, and for non-Java code
69
+ follow the language's doc-comment form plus the evidenced project convention.
39
70
 
40
71
  ## Choose the execution owner
41
72
 
@@ -81,26 +112,49 @@ Sub-agents never dispatch other sub-agents or read `.easy-coding` workflow asset
81
112
  ## Test Points {unit.test_points and exact targeted commands}
82
113
  ## Contracts {inputs, outputs, invariants shared with other units}
83
114
  ## Risks {known edge cases and compatibility risks}
115
+ ## Local Baseline {nearest comparable code conventions and evidence paths}
116
+ ## Code Comments {resolved host Agent author value; model-field, enum-member, and constant rules}
84
117
  ## Coding Rules {pre-digested RULES sections}
85
118
  ## Architecture {pre-digested ABSTRACT sections}
86
119
  ## Output
87
120
  status:"completed", repo_id|null, source_task_id|null, changed_files[], summary,
88
- deliverable|null, issues:[], needs_attention:[]
121
+ deliverable|null, checks:[{command,passed,failures:[]}], issues:[], needs_attention:[]
89
122
  ```
90
123
 
91
124
  ## Dispatch and result loop
92
125
 
93
126
  1. Append a `dispatch` record before work begins. Canonical-backed records include `repo_id` and
94
127
  `source_task_id`; resolve every file relative to `task.repo_paths[repo_id]` before dispatch.
128
+ Before dispatching a selected task that is not already `in_progress`, call
129
+ `writeback-spec-task --status in_progress` with a key stable for that dispatch/recovery
130
+ attempt but distinct from any earlier accepted `in_progress` event. Do this only after its
131
+ hard/contract dependencies are ready; do not batch-start dependent tasks at the initial
132
+ IMPLEMENT boundary.
133
+ Populate `Code Comments` on every code task card with the resolved user-facing host Agent
134
+ author value, the field/member/constant rules, and the core Java Javadoc rule above. Populate
135
+ `Local Baseline` from the Unit's analyzed evidence; sub-agents do not read this Skill.
95
136
  2. Execute according to dependency order and selected owner.
96
137
  3. Run targeted unit tests and self-audit scope, contracts, TODOs, and introduced warnings.
138
+ Also audit new author attributions and every new model field, enum member, and constant against
139
+ the comment requirements above, then check local-style deviations, unnecessary abstractions,
140
+ one-use constant extraction, and affected core Java Javadoc before recording success.
97
141
  4. Append one `result` record. Only a successful unit uses `status:"completed"`; include
98
142
  unresolved issues rather than hiding them, and do not advance while `issues` or
99
143
  `needs_attention` is non-empty.
144
+ For Canonical-backed success, write each owned source Step `completed` through
145
+ `writeback-spec-step`, with passed evidence for every bound Canonical Test ID and a stable key.
146
+ After every source Step for that task is complete, write the task `implemented`. On failure,
147
+ write the affected Step `failed`; the shared writer moves its task to `blocked`. Local evidence
148
+ is appended first, shared projection second, and the returned acknowledgment last.
100
149
  5. If a result changes a cross-unit contract, stop dependent units and return to ANALYSIS.
101
150
  6. For parallel units, detect overlapping writes before advancing.
102
151
  7. If implementation needs a file, symbol, repository, or source step outside the mapped
103
152
  Canonical change set, stop and return to ANALYSIS instead of expanding scope implicitly.
153
+ 8. If a static Canonical change is confirmed, revise the original design by exactly one revision
154
+ and use `sync-spec-design`; never edit the machine-owned execution block. If a writeback was
155
+ interrupted, run `reconcile-spec-execution` with the stored idempotent pending action.
156
+ Reconciliation only consumes dispatch/result evidence created after the current `in_progress`
157
+ acknowledgment; it never opens a new repair attempt or reuses an earlier attempt's result.
104
158
 
105
159
  Do not emit a progress message for every trivial edit. Report at unit boundaries to reduce
106
160
  conversation overhead while keeping work observable.
@@ -119,4 +173,10 @@ conversation overhead while keeping work observable.
119
173
  - [ ] Every unit has a dispatch/result pair and satisfied its acceptance criteria.
120
174
  - [ ] Targeted tests ran or a concrete blocker is recorded.
121
175
  - [ ] Cross-unit contracts still match.
176
+ - [ ] The implementation follows the evidenced Local Baseline or records a required deviation.
177
+ - [ ] No speculative layer, fragmented micro-method set, or single-use getter constant was added.
178
+ - [ ] New author attributions use the user-facing host `<Current Agent Name> with Easy Coding`.
179
+ - [ ] Every new model field, enum member, and constant has a meaningful field-level comment.
180
+ - [ ] Every method/field in a new core Java class, and every added or materially modified one in
181
+ an existing core Java class, has Javadoc.
122
182
  - [ ] Code tasks enter REVIEW, regardless of workflow mode.
@@ -3,10 +3,17 @@ name: ec-memory
3
3
  description: MEMORY-stage skill. Creates a workflow-mode-aware schema-v2 checkpoint from existing task evidence and performs conditional long-memory distillation.
4
4
  ---
5
5
 
6
- # ec-memory — evidence-derived checkpoint
6
+ # ec-memory — evidence-derived checkpoint and knowledge governance
7
7
 
8
- MEMORY remains mandatory for code tasks. It must not re-analyze the repository or repeat the
9
- entire conversation. Generate from `task.json`, `dev-spec.md`, and `execution.jsonl`.
8
+ MEMORY remains mandatory for code tasks. Daily task processing and architecture maintenance are
9
+ separate responsibilities: every completed code task produces one immutable short-memory fact;
10
+ only a long-memory distillation, or the explicit missing-ABSTRACT startup exception, may open an
11
+ architecture assessment. Never update architecture merely because MEMORY was entered.
12
+
13
+ The short-memory checkpoint must not re-analyze the repository or repeat the entire conversation.
14
+ Generate it only from the verified evidence already stored in `task.json`, `dev-spec.md`, and
15
+ `execution.jsonl`. The bounded repository reads described below belong only to a required
16
+ `backfill` or `update` architecture assessment.
10
17
 
11
18
  ## Depth by workflow mode
12
19
 
@@ -28,13 +35,76 @@ Name it `{memory_id}_{YYYYMMDD}_{smart_name}.md` and set
28
35
  `.easy-coding/memory/short/`, then register it with
29
36
  `memory-short-complete`. Never invent test results or commit hashes.
30
37
 
31
- When frozen TDD is enabled, add its threshold, lifecycle evidence, local changed-line result,
32
- and remote CI status to the short memory's execution evidence. When TDD is off, omit TDD fields
33
- entirely so ordinary tasks incur no additional memory work.
38
+ Copy the final `acceptance` record from `execution.jsonl` into the checkpoint as a concise
39
+ decision fact: authorization source, decision summary, `diff_sha256`, review policy, verification
40
+ policy, changed files, and any Canonical source tasks that required targeted verification.
41
+ `memory-short-complete` rejects a checkpoint that omits any of those decision fields. This records
42
+ the user's accepted exception without re-reviewing or re-analyzing the code. Canonical writeback
43
+ already carries the same digest and authorization as shared `acceptance` evidence.
44
+
45
+ When frozen TDD is enabled, add its threshold, lifecycle evidence, passed local unit-test result,
46
+ and local changed-line result to the short memory's execution evidence. Remote CI status is not
47
+ part of Harness acceptance or task memory. When TDD is off, omit TDD fields entirely so ordinary
48
+ tasks incur no additional memory work.
34
49
 
35
50
  Ask the state API for `memory-instruction`. Distill only when it returns `action:distill`;
36
51
  otherwise record `no-op`. Long memory receives reusable facts only, not file dumps, transient
37
52
  logs, routine command output, or speculation.
38
53
 
54
+ ## Architecture assessment
55
+
56
+ Read the frozen `architecture_assessment` contract returned by `memory-instruction`.
57
+
58
+ - `required:false`: do not read the repository for architecture purposes and do not modify
59
+ `.easy-coding/ABSTRACT.md` or `.easy-coding/CHANGELOG.md`.
60
+ - `trigger:distillation`: finish classifying the frozen `candidate_files`, then assess whether
61
+ their stable, reusable facts make the current architecture cognition stale. Default to
62
+ `no-op`.
63
+ - `trigger:missing-abstract`: use `backfill` after the first substantive startup task even when
64
+ long-memory action is `no-op`. This is the only non-distillation architecture exception.
65
+
66
+ An architecture `update` is justified only by evidence of at least one of these changes:
67
+
68
+ - a module was added, removed, split, or merged;
69
+ - module responsibility, ownership, or dependency direction changed;
70
+ - a core request, data, state, or event flow changed;
71
+ - the technology stack, runtime, build, or deployment infrastructure changed;
72
+ - the existing ABSTRACT conflicts with verified current facts.
73
+
74
+ Do not update for a bug fix, local implementation detail, DTO/field-only change, local refactor,
75
+ temporary workaround, routine dependency patch, or the mere fact that distillation ran. Stable
76
+ new coding conventions belong in `TECHNICAL.md` as explicit RULES update candidates; never
77
+ silently edit `RULES.md`, `SOUL.md`, or `TEST_STRATEGY.md` from MEMORY.
78
+
79
+ For `no-op`, use only the frozen memory evidence and give a concrete reason. For `backfill` or
80
+ `update`, read only candidate-related modules, entrypoints, dependencies, and affected ABSTRACT
81
+ sections. Do not perform an unbounded repository re-analysis. Create or edit only the affected
82
+ sections of `.easy-coding/ABSTRACT.md`, and create or append `.easy-coding/CHANGELOG.md`; never
83
+ regenerate the whole ABSTRACT when a bounded edit is sufficient.
84
+
85
+ Whenever `required:true`, record the decision through the command below. For distillation this
86
+ must succeed before deleting any candidate; for `missing-abstract` it must succeed before the
87
+ `no-op` long-memory action can complete:
88
+
89
+ ```bash
90
+ {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py memory-architecture-assessment \
91
+ --session-file <P> --action <no-op|backfill|update> --reason <reason> \
92
+ --evidence <frozen-memory-file> [--evidence <frozen-memory-file> ...] \
93
+ [--affected-section <section> ...] --agent <agent-id>
94
+ ```
95
+
96
+ `backfill` and `update` require affected sections; `no-op` must not declare them. Evidence must
97
+ come from the frozen candidate set, or from the current checkpoint for the missing-ABSTRACT
98
+ exception. If assessment or architecture-file validation fails, keep every candidate file and
99
+ remain in MEMORY.
100
+
101
+ After the assessment succeeds, a distillation may delete all frozen `candidate_files` while
102
+ preserving every `kept_file`. Then call `memory-complete`. The state API rechecks that no-op
103
+ assets stayed unchanged, changed assets still match the recorded assessment, all candidates were
104
+ consumed, and all retained memories still exist.
105
+
39
106
  Complete processing with `memory-complete`. When the state API reports
40
107
  `memory_progress.completed:true`, call `auto-transition --stage COMPLETE`.
108
+ For Canonical-backed tasks this automatic edge first writes every verified selected source task
109
+ to shared `completed`; pending integration dependencies or failed writeback keep the task in
110
+ MEMORY. Do not claim COMPLETE from local memory state alone.
@@ -12,7 +12,17 @@ Every new code task enters REVIEW. Read-only tasks do not. Obtain the current fi
12
12
  ```
13
13
 
14
14
  Review the final diff against `dev-spec.md`, RULES, unit acceptance criteria, tests, contracts,
15
- and obvious security risks. Every finding cites `file:line`.
15
+ the Unit's Local Baseline, and obvious security risks. Every finding cites `file:line`.
16
+
17
+ Review local fit before recommending generic cleanup. Flag an implementation when it departs from
18
+ the nearest comparable naming, control flow, null/error handling, layering, modeling, or method
19
+ granularity without a correctness, security, requirement, or hard-rule reason. Also flag
20
+ speculative layers, fragmented one-use micro-methods, constants created only for one getter
21
+ return, and missing Javadoc in core Java code: check every method/field in a new core class and
22
+ each added or materially modified one in an existing core class. Do not demand defensive null
23
+ checks, abstraction, constant extraction, or legacy-wide comment retrofits merely because they
24
+ are generic best practices. A violation of an explicit task-card coding/comment contract is a
25
+ contract defect, not optional stylistic advice.
16
26
 
17
27
  For Canonical-backed tasks, group evidence by `repo_id` and `source_task_id`. Every selected
18
28
  Spec task needs an implementation result and source test evidence; file references remain
@@ -24,8 +34,10 @@ review dimension. A global record without source ownership cannot satisfy the ga
24
34
  When frozen TDD is enabled, add a passed review dimension named exactly `tdd` for each source
25
35
  task. Review whether RED/GREEN/REFACTOR (or characterization GREEN for pure refactors) is genuine,
26
36
  tests exercise changed behavior and boundaries, mocks do not merely mirror implementation, and
27
- the local/CI changed-line coverage gates share the frozen threshold. When TDD is off, do not add
28
- this dimension or raise the ordinary review depth.
37
+ the local unit-test command genuinely passes while the changed-line coverage command uses the
38
+ frozen baseline and threshold. Generated CI configuration may be reviewed when it changed, but
39
+ remote CI status is never a review or acceptance dependency. When TDD is off, do not add this
40
+ dimension or raise the ordinary review depth.
29
41
 
30
42
  ## Depth by workflow mode
31
43
 
@@ -67,6 +79,12 @@ Verdict:
67
79
  In-scope defects are fixed automatically. Ask the user only for a new design choice, changed
68
80
  public contract, or contradiction with a confirmed decision.
69
81
 
82
+ For Canonical-backed review, append the local review record first. Any blocking finding then
83
+ writes the owning source task `blocked` through `writeback-spec-task`, referencing the local
84
+ record. A passed review does not change shared task status. A `replan` verdict returns to ANALYSIS;
85
+ confirmed static Spec changes use revision + READY + `sync-spec-design` rather than edits to the
86
+ derived plan or machine execution block.
87
+
70
88
  ## Evidence record
71
89
 
72
90
  Append one final record per executed dimension for the current implementation fingerprint.
@@ -22,6 +22,9 @@ when you recognize abandonment intent in the user's message.
22
22
  `{{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py close-current --session-file <P> --reason "<reason>" --agent <agent-id>`.
23
23
  This sets `task.json.status` to `CLOSED`, records `closed_reason`, updates history, and
24
24
  clears session `current_task` so the next hook injection returns to Ready.
25
+ For Canonical-backed work, the same command first projects every unfinished selected task to
26
+ shared `cancelled` (using an intermediate `blocked` transition when required). A shared writer
27
+ failure leaves the Harness task open; never close only the local side.
25
28
  Use the returned `status_context` as the authoritative status source for the rest of the
26
29
  current turn.
27
30
  4. **No memory flow.** Do not run MEMORY. An incomplete task's memory is dirty data.
@@ -34,6 +37,7 @@ when you recognize abandonment intent in the user's message.
34
37
 
35
38
  - Never delete task folders — CLOSED tasks stay as a record.
36
39
  - Never run the memory/archive flow.
40
+ - Never hand-edit the Canonical execution region to force cancellation.
37
41
  - This skill closes the `current_task`. If the user wants to close a different (suspended)
38
42
  task, they should first switch to it via ec-workflow, then invoke ec-task-close.
39
43
  - Division of labor: ec-task-management lists/creates (read-only panel), ec-workflow runs the
@@ -16,7 +16,8 @@ Call the state API snapshot and show:
16
16
  - task `concrete_workflow_mode` and frozen TDD state when present;
17
17
  - harness enabled/disabled state;
18
18
  - active and resumable tasks.
19
- - for Canonical-backed tasks: source Spec ID/revision/SHA, selected task IDs, repository
19
+ - for Canonical-backed tasks: source locator/path mode, Spec ID/design revision/design digest,
20
+ document digest, execution revision, writeback status, selected task IDs, repository
20
21
  bindings/baseline status, and pending dependency evidence.
21
22
 
22
23
  Mode inspection and configuration belongs to `ec-config`. If the user asks to change Approval,
@@ -32,3 +33,8 @@ When creating from a Canonical Spec, call `inspect-dev-spec`, display the comple
32
33
  dependency selection, then call `select-dev-spec-scope` and `create-task-from-spec` only after
33
34
  explicit user selection. Multiple selected Spec tasks still create one Harness task, while the
34
35
  selector returns one deterministic consumption closure per selected repository.
36
+ Initialize missing shared execution before creation. Support `rebind-spec-source` only when the
37
+ new file matches schema + spec_id + design revision + design_sha256 and does not roll execution
38
+ revision backward. A pending writeback is repaired with `reconcile-spec-execution`, never by
39
+ editing the execution JSON block or starting a different writeback. A deterministic rejected
40
+ action is cleared with `status:error`; correct its input instead of replaying it.
@@ -0,0 +1,101 @@
1
+ ---
2
+ name: ec-tdd-init
3
+ description: Initialize or refresh Java changed-line TDD coverage infrastructure before TDD can be enabled.
4
+ ---
5
+
6
+ # ec-tdd-init — Java changed-line gate initialization
7
+
8
+ Communicate in the user's language. This skill owns TDD infrastructure readiness, not historical
9
+ test-debt cleanup. It must never bulk-generate tests for existing business code, require
10
+ repository-wide coverage, or modify production behavior merely to raise coverage.
11
+
12
+ ## Non-circular ordering
13
+
14
+ The only legal order is:
15
+
16
+ ```text
17
+ TDD off -> initialize infrastructure -> readiness ready -> user explicitly enables TDD
18
+ ```
19
+
20
+ Run this skill as a dedicated code task with `type=tdd-init`. The state API always freezes that
21
+ task with `tdd_enabled=false`, even when a legacy project/session setting or a suspended task has
22
+ TDD enabled. Never offer "enable now and initialize later". Never enable TDD automatically after
23
+ initialization.
24
+
25
+ ## Read-only preflight
26
+
27
+ First run:
28
+
29
+ ```bash
30
+ python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . check
31
+ ```
32
+
33
+ If it returns `ready`, report the recorded build/CI contract and stop without creating a task.
34
+ The user may then use `ec-config` or `easy-coding config` to enable TDD.
35
+
36
+ If it returns `needs_init`, inspect only the infrastructure needed to form a confirmed plan:
37
+
38
+ - Maven/Gradle files and the existing JUnit runner;
39
+ - JaCoCo XML generation configuration;
40
+ - `.gitlab-ci.yml` and its repository-local include chain;
41
+ - TEST-stage job, JUnit/JaCoCo artifacts, and invocation of
42
+ `.easy-coding/tools/easy_coding_java_coverage.py` with task-supplied baseline/threshold values.
43
+
44
+ Do not measure current whole-project coverage. A project with no historical business tests may
45
+ still become ready when the test runner, JaCoCo reporting, and parameterized changed-line gate
46
+ are functional.
47
+
48
+ ## Initialization task
49
+
50
+ After the user confirms the exact infrastructure scope, create one task and route it through the
51
+ ordinary workflow:
52
+
53
+ ```bash
54
+ {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py create-task \
55
+ --task-id <safe-unique-id> --type tdd-init \
56
+ --title "Initialize Java changed-line TDD infrastructure" \
57
+ --agent <agent-id> --session-file <P>
58
+ ```
59
+
60
+ ANALYSIS may plan build files, GitLab CI files, common scripts, and the readiness receipt. It must
61
+ state `historical coverage required: no` and `coverage scope: changed production lines since each
62
+ future task baseline`. IMPLEMENT changes infrastructure only. It must not add tests whose sole
63
+ purpose is to cover unchanged production code.
64
+
65
+ The reusable GitLab job must consume a baseline SHA and threshold supplied for the future task;
66
+ do not hardcode the initialization commit or the default 90% threshold. The same Python coverage
67
+ tool must be usable locally and remotely. This generated job is remote automation infrastructure,
68
+ not a Harness task acceptance dependency: later tasks require local unit-test and coverage
69
+ evidence only, and never wait for a pipeline URL, job identity, or remote success status.
70
+
71
+ ## Readiness receipt and verification
72
+
73
+ At the end of IMPLEMENT, after the infrastructure files are stable, record their fingerprints.
74
+ The receipt is part of the implementation and must exist before REVIEW so review/verification
75
+ fingerprints do not change after review. The recorder automatically includes the harness-managed
76
+ `.easy-coding/tools/easy_coding_java_coverage.py` fingerprint:
77
+
78
+ ```bash
79
+ python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . record \
80
+ --build-file <pom.xml-or-build.gradle> [--build-file <included-build-file>]... \
81
+ --ci-file .gitlab-ci.yml [--ci-file <repository-local-include>]... \
82
+ --coverage-report <jacoco-xml-pattern> [--coverage-report <pattern>]... \
83
+ --gate-command "python3 .easy-coding/tools/easy_coding_java_coverage.py check --base \$EASY_CODING_TDD_BASE_SHA --threshold \$EASY_CODING_TDD_THRESHOLD" \
84
+ --agent <agent-id>
85
+ ```
86
+
87
+ REVIEW includes the receipt and its declared infrastructure boundary. VERIFICATION runs the
88
+ frozen Workflow Mode's applicable build/test/CI syntax checks, then performs only the read-only
89
+ readiness check:
90
+
91
+ ```bash
92
+ python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . check
93
+ ```
94
+
95
+ The `VERIFICATION -> MEMORY` gate requires the final check to return `ready`. If any recorded
96
+ build or CI file changes after the receipt was created, readiness becomes `needs_init`; return to
97
+ IMPLEMENT, refresh the receipt, and repeat REVIEW before verifying again. Rerun this skill when
98
+ the same drift occurs after task completion.
99
+
100
+ After completion, tell the user that TDD remains off and provide the explicit project/session
101
+ enable route. Do not treat readiness as consent to enable it.
@@ -15,7 +15,8 @@ Read-only tasks never enter this stage. Obtain fresh fingerprints before running
15
15
 
16
16
  - No completion claim without executed verification evidence.
17
17
  - Evidence is reusable only while both returned fingerprints remain unchanged.
18
- - Relevant code or config changes invalidate old evidence automatically.
18
+ - Relevant code or config changes invalidate old evidence automatically unless the exact
19
+ post-verification code diff is explicitly accepted under the checkpoint protocol below.
19
20
  - Failed or missing evidence never becomes acceptance because of approval mode.
20
21
 
21
22
  ## Verification depth
@@ -26,10 +27,12 @@ Read-only tasks never enter this stage. Obtain fresh fingerprints before running
26
27
  - `strict`: run the project's full applicable lint, typecheck, test, and build gates.
27
28
 
28
29
  These rules remain unchanged when frozen TDD is off: do not discover JaCoCo reports, run the
29
- coverage tool, inspect GitLab, or add a coverage record.
30
+ coverage tool, inspect GitLab, or add a coverage record. The explicit `type=tdd-init` task is an
31
+ infrastructure exception: run its planned build/CI syntax checks and readiness tool, but do not
32
+ measure repository-wide coverage or append TDD coverage evidence for unchanged production code.
30
33
 
31
- When frozen TDD is on, first run the planned Java unit command and generate JaCoCo XML, then run
32
- the same deterministic gate intended for GitLab:
34
+ When frozen TDD is on, first run the planned local Java unit command and generate JaCoCo XML,
35
+ then run the deterministic local acceptance gate:
33
36
 
34
37
  ```bash
35
38
  python3 .easy-coding/tools/easy_coding_java_coverage.py check \
@@ -41,10 +44,10 @@ The tool measures covered added/modified production Java executable lines only.
41
44
  comment, blank, import, and test-source lines are excluded by diff/JaCoCo intersection. Missing
42
45
  or ambiguous source files and reports older than their modified source fail; zero modified
43
46
  executable lines is explicit N/A. Always regenerate JaCoCo XML after the final source change.
44
- Record CI as pending until the remote pipeline actually passes; local green is not remote green.
45
47
  Never substitute `HEAD`, a mutable ref, project defaults, or current session settings for the
46
- task-frozen baseline SHA and threshold. GitLab must invoke the tool with the same two frozen
47
- values; a session override therefore requires the CI command/variable for this task to match it.
48
+ task-frozen baseline SHA and threshold. `ec-tdd-init` still generates a GitLab job that can run
49
+ the same tool, but remote pipeline execution and status are outside Harness acceptance. Never
50
+ request an intermediate commit or push merely to obtain CI evidence.
48
51
 
49
52
  The main Agent may run commands inline. Dispatch verifier sub-agents only when checks are
50
53
  independent and parallel execution materially saves time or isolates specialist environments.
@@ -75,7 +78,10 @@ and a non-empty `not_applicable_reason`; it does not count as the required appli
75
78
  check, and must not be represented by an invented successful command.
76
79
 
77
80
  Record failures in `failures[]`. If any current-fingerprint record fails, return to IMPLEMENT;
78
- do not append a later synthetic pass without rerunning the failed command.
81
+ do not append a later synthetic pass without rerunning the failed command. For a Canonical-backed
82
+ failure, append the local verify record first, then write the owning source task `blocked` with a
83
+ concise reference to that record. The repair transition automatically reopens blocked source tasks
84
+ only; unaffected implemented tasks retain their latest shared conclusion.
79
85
 
80
86
  ## Coverage and acceptance
81
87
 
@@ -84,23 +90,17 @@ For TDD coverage, copy the tool output into `coverage`: `baseline_sha`, `covered
84
90
  `applicable:false` plus the tool's reason only for zero executable modified lines. A percentage
85
91
  below the frozen threshold fails even when ordinary tests pass.
86
92
 
87
- Append two coverage records per repository (and per Canonical source task): one with
88
- `coverage_scope:"local"`, and one with `coverage_scope:"gitlab"`. The GitLab record may be
89
- appended only after the remote job succeeds and must also include:
93
+ Append one coverage record with `coverage_scope:"local"` per repository (and per Canonical
94
+ source task). The state gate also requires a passed local `check_type:"test"` record for the same
95
+ owner. The coverage record preserves the task-frozen baseline and threshold. Do not append or
96
+ wait for GitLab pipeline evidence; historical remote coverage records are ignored by acceptance
97
+ without modifying or deleting the stored records.
90
98
 
91
- ```json
92
- {
93
- "ci": {
94
- "provider": "gitlab",
95
- "pipeline_url": "https://gitlab.example/.../pipelines/123",
96
- "job_name": "changed-line-coverage",
97
- "status": "success"
98
- }
99
- }
100
- ```
101
-
102
- Both records must preserve the same task-frozen baseline and threshold. A local-only result,
103
- pending/failed pipeline, missing job identity, or synthetic remote pass cannot satisfy MEMORY.
99
+ For `type=tdd-init`, the infrastructure receipt must already have been recorded during IMPLEMENT
100
+ and reviewed with the rest of the implementation. Run only `easy_coding_tdd_readiness.py check`
101
+ here. If it reports drift, return to IMPLEMENT to refresh the receipt and repeat REVIEW; never
102
+ rewrite it inside VERIFICATION. The state gate requires `ready` before MEMORY. This does not
103
+ enable TDD; report the explicit `ec-config`/`easy-coding config` next step.
104
104
 
105
105
  - Every must-test item has an executed check.
106
106
  - Bug fixes include a regression test when project infrastructure exists.
@@ -110,6 +110,38 @@ pending/failed pipeline, missing job identity, or synthetic remote pass cannot s
110
110
  user wait.
111
111
  - A reported in-scope problem returns to IMPLEMENT; out-of-scope work becomes a separate task.
112
112
 
113
+ After the final green evidence is recorded, freeze the acceptance baseline before presenting the
114
+ result or applying the boundary:
115
+
116
+ ```bash
117
+ {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py verification-checkpoint \
118
+ --agent <agent-id> --session-file <P>
119
+ ```
120
+
121
+ Then request or auto-apply VERIFICATION -> MEMORY according to `approval_mode`. `auto` remains
122
+ automatic when the checkpoint is unchanged. If any in-scope code changed after the checkpoint,
123
+ the state API returns `action:"acceptance-drift"`, keeps the task in VERIFICATION, and includes
124
+ the exact unified patches or binary/mode-change descriptions plus a stable `diff_sha256`. This
125
+ exceptional drift pauses every approval mode, including `auto`; it does not permanently change
126
+ the configured mode.
127
+
128
+ Show the complete returned diff and ask whether to accept that exact digest. Do not re-enter
129
+ IMPLEMENT or rerun REVIEW merely because this drift exists. On acceptance, call
130
+ `confirm-transition --stage MEMORY --diff-sha256 <digest>` with exactly one policy:
131
+
132
+ - `carry-forward`: only when every changed hunk is confidently non-executable and existing
133
+ verification remains applicable;
134
+ - `targeted`: executable behavior changed; append passed current-fingerprint targeted verification
135
+ before confirming. Canonical tasks must cover every affected source task reported by the
136
+ acceptance record, without rerunning checks for unaffected source tasks;
137
+ - `waived`: the user explicitly accepts the stated unverified risk.
138
+
139
+ Include `--decision-summary` with the user's decision. A changed digest invalidates the pending
140
+ confirmation and must be shown again. Behavior config, execution plan, workflow, Canonical
141
+ design, or nested-repository metadata drift cannot use this shortcut; return to ANALYSIS or
142
+ IMPLEMENT as reported by the state API. The acceptance record bridges only the accepted
143
+ implementation fingerprints, so prior REVIEW evidence remains valid without a second REVIEW.
144
+
113
145
  For Canonical-backed tasks, run each repository's commands from `task.repo_paths[repo_id]` and
114
146
  cover every selected task's source test IDs. Report pending integration edges separately from
115
147
  local green checks. They do not block local implementation evidence, but the state API blocks
@@ -120,6 +152,16 @@ tasks remain separate evidence records. In `strict`, every involved repository i
120
152
  records all four check types; a repository-specific non-applicable record still needs its reason
121
153
  and source ownership.
122
154
 
155
+ After implementation and local checks, each selected Canonical source task remains
156
+ `implemented`. Do not call `writeback-spec-task --status verified` from VERIFICATION. Applying
157
+ VERIFICATION -> MEMORY is the authoritative acceptance boundary: the state API writes each
158
+ still-implemented source task to `verified` through CAS/idempotent recoverable events with its
159
+ accepted test evidence and acceptance digest, then enters MEMORY only after every write is
160
+ confirmed. For `approve`/`guard`, that authority is the explicit boundary
161
+ confirmation; for `confirm`/`auto`, it is the standing approval-mode authorization when no new
162
+ drift exists. If writeback is interrupted, run `reconcile-spec-execution` before retrying the
163
+ transition. Remote CI remains outside this acceptance gate.
164
+
123
165
  Record the exact integration edge only after its evidence exists:
124
166
 
125
167
  ```bash
@@ -131,5 +173,5 @@ Record the exact integration edge only after its evidence exists:
131
173
  --agent <agent>
132
174
  ```
133
175
 
134
- The state API rejects VERIFICATION -> MEMORY unless all evidence for the current implementation
135
- and config fingerprints is green.
176
+ The state API rejects VERIFICATION -> MEMORY unless all effective evidence is green and the
177
+ checkpoint is either unchanged or bound to an exact accepted diff.