easy-coding-harness 0.10.0-beta.0 → 0.10.0-beta.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. package/CHANGELOG.md +166 -0
  2. package/README.md +53 -19
  3. package/dist/cli.js +573 -53
  4. package/dist/cli.js.map +1 -1
  5. package/package.json +1 -1
  6. package/templates/claude/agents/ec-implementer.md +11 -0
  7. package/templates/claude/agents/ec-reviewer.md +6 -1
  8. package/templates/codex/agents/ec-implementer.toml +11 -0
  9. package/templates/codex/agents/ec-reviewer.toml +6 -1
  10. package/templates/common/bundled-skills/ec-init/SKILL.md +20 -2
  11. package/templates/common/bundled-skills/ec-meta/references/local-architecture/README.md +33 -11
  12. package/templates/common/bundled-skills/ec-meta/references/platform-files/README.md +1 -1
  13. package/templates/common/skills/ec-analysis/SKILL.md +162 -26
  14. package/templates/common/skills/ec-config/SKILL.md +76 -0
  15. package/templates/common/skills/ec-git/SKILL.md +7 -1
  16. package/templates/common/skills/ec-implementing/SKILL.md +72 -1
  17. package/templates/common/skills/ec-memory/SKILL.md +77 -3
  18. package/templates/common/skills/ec-reviewing/SKILL.md +25 -1
  19. package/templates/common/skills/ec-task-close/SKILL.md +4 -0
  20. package/templates/common/skills/ec-task-management/SKILL.md +13 -32
  21. package/templates/common/skills/ec-tdd-init/SKILL.md +101 -0
  22. package/templates/common/skills/ec-verification/SKILL.md +91 -5
  23. package/templates/common/skills/ec-workflow/SKILL.md +86 -21
  24. package/templates/main-constraint/AGENTS.md.tpl +54 -12
  25. package/templates/main-constraint/CLAUDE.md.tpl +51 -12
  26. package/templates/qoder/agents/ec-implementer.md +11 -0
  27. package/templates/qoder/agents/ec-reviewer.md +6 -1
  28. package/templates/runtime/templates/dev-spec-skeleton.md +8 -1
  29. package/templates/runtime/tools/easy_coding_java_coverage.py +317 -0
  30. package/templates/runtime/tools/easy_coding_tdd_readiness.py +306 -0
  31. package/templates/shared-hooks/easy_coding_state.py +4782 -574
  32. package/templates/shared-hooks/easy_dev_spec.py +444 -30
  33. package/templates/shared-hooks/easy_dev_spec_execution.py +1014 -0
  34. package/templates/shared-hooks/easy_dev_spec_protocol.py +1426 -18
  35. package/templates/shared-hooks/inject-subagent-context.py +5 -0
@@ -9,6 +9,16 @@ Use only after ANALYSIS has frozen `task.json.workflow_mode` to `fast`, `standar
9
9
  `strict`. Read `dev-spec.md`, the latest `plan` record in `execution.jsonl`, relevant RULES
10
10
  and ABSTRACT sections, and `test-strategy.md` for code tasks.
11
11
 
12
+ If frozen `task.tdd_enabled` is not `true`, preserve the existing shift-left behavior exactly;
13
+ do not load the Java coverage tool, require RED/GREEN/REFACTOR, inspect CI, or run extra test
14
+ commands. TDD is an independent opt-in mode, not an implicit consequence of strict workflow.
15
+
16
+ When frozen TDD is enabled, every feature/bug unit must capture a meaningful failing unit test
17
+ before production code (RED), the smallest passing implementation (GREEN), and a green refactor.
18
+ Pure refactors instead capture a passing characterization test before the change and rerun it
19
+ afterward. Never fake RED evidence. Keep tests deterministic, boundary-focused, and minimally
20
+ mocked, and design changed production code toward 100% unit coverage.
21
+
12
22
  Communicate with the user in the user's language.
13
23
 
14
24
  ## Non-negotiable gates
@@ -26,6 +36,37 @@ Communicate with the user in the user's language.
26
36
  enters REVIEW.
27
37
  7. Read-only `doc` / `analysis` / `report` tasks remain `single` with `files:[]`, make no writes,
28
38
  return a non-empty `deliverable`, then follow the mode-aware IMPLEMENT -> COMPLETE edge.
39
+ 8. When a project template, local convention, or new source header uses author attribution, the
40
+ author value must be `<Current Agent Name> with Easy Coding`, for example
41
+ `Codex with Easy Coding`. `Current Agent Name` means the user-facing host Agent (for example,
42
+ Codex, Claude, or Qoder), never an implementation sub-agent role such as `ec-implementer`.
43
+ This value is display attribution only: never pass it to the workflow state API's `--agent`,
44
+ which accepts only `claude-code`, `codex`, or `qoder`. Never copy a previous human or Agent name
45
+ into newly authored code.
46
+ 9. Every newly added field in a data-bearing model must have a meaningful field-level comment.
47
+ This includes new or extended entity/DO/DTO/VO/BO, request/response, configuration, and similar
48
+ model types. Every new enum member and every new declared constant requires the same treatment.
49
+ Describe the semantic meaning and, when relevant, units, format, allowed values, nullability,
50
+ default behavior, or compatibility constraints. A type-level comment does not replace comments
51
+ on its fields or members; do not add low-value comments to ordinary local variables.
52
+ 10. Treat the task card's `Local Baseline` as the default implementation shape. Match the nearest
53
+ comparable code's naming, control flow, null/empty and error handling, layering, object model,
54
+ and extraction granularity unless correctness, security, an explicit requirement, or a hard
55
+ project rule requires a deviation. Do not add defensive null checks solely because they are a
56
+ generic best practice when the evidenced local contract intentionally omits them.
57
+ 11. Implement the smallest coherent design. Do not add speculative abstractions, wrappers,
58
+ factories, layers, or extension points, and do not fragment one readable flow into many
59
+ single-use micro-methods. Extract code only for a clear semantic boundary, real reuse,
60
+ independent testability, or a material reduction in complexity.
61
+ 12. Literals and magic values are allowed when they are obvious, local, and consistent with the
62
+ surrounding code. Introduce a constant for repeated use, stable domain/config/protocol
63
+ semantics, or an established project convention—not merely to hold the single return value of
64
+ a getter.
65
+ 13. In a newly added core Java class, every method and field requires meaningful Javadoc. In an
66
+ existing core Java class, every added or materially modified method and field requires it.
67
+ Add focused inline comments to core or complex logic to explain intent, constraints, or
68
+ non-obvious tradeoffs. Do not mass-retrofit untouched legacy code, and for non-Java code
69
+ follow the language's doc-comment form plus the evidenced project convention.
29
70
 
30
71
  ## Choose the execution owner
31
72
 
@@ -59,6 +100,7 @@ Sub-agents never dispatch other sub-agents or read `.easy-coding` workflow asset
59
100
  # Task Card
60
101
  ## Identity Easy Coding implementation unit
61
102
  ## Workflow Mode {fast|standard|strict}
103
+ ## TDD {off | on, frozen changed-line threshold N%}
62
104
  ## Task {unit description}
63
105
  ## Source Spec {spec_id@revision + sha256 | NONE}
64
106
  ## Source Task {source_task_id | NONE}
@@ -70,26 +112,49 @@ Sub-agents never dispatch other sub-agents or read `.easy-coding` workflow asset
70
112
  ## Test Points {unit.test_points and exact targeted commands}
71
113
  ## Contracts {inputs, outputs, invariants shared with other units}
72
114
  ## Risks {known edge cases and compatibility risks}
115
+ ## Local Baseline {nearest comparable code conventions and evidence paths}
116
+ ## Code Comments {resolved host Agent author value; model-field, enum-member, and constant rules}
73
117
  ## Coding Rules {pre-digested RULES sections}
74
118
  ## Architecture {pre-digested ABSTRACT sections}
75
119
  ## Output
76
120
  status:"completed", repo_id|null, source_task_id|null, changed_files[], summary,
77
- deliverable|null, issues:[], needs_attention:[]
121
+ deliverable|null, checks:[{command,passed,failures:[]}], issues:[], needs_attention:[]
78
122
  ```
79
123
 
80
124
  ## Dispatch and result loop
81
125
 
82
126
  1. Append a `dispatch` record before work begins. Canonical-backed records include `repo_id` and
83
127
  `source_task_id`; resolve every file relative to `task.repo_paths[repo_id]` before dispatch.
128
+ Before dispatching a selected task that is not already `in_progress`, call
129
+ `writeback-spec-task --status in_progress` with a key stable for that dispatch/recovery
130
+ attempt but distinct from any earlier accepted `in_progress` event. Do this only after its
131
+ hard/contract dependencies are ready; do not batch-start dependent tasks at the initial
132
+ IMPLEMENT boundary.
133
+ Populate `Code Comments` on every code task card with the resolved user-facing host Agent
134
+ author value, the field/member/constant rules, and the core Java Javadoc rule above. Populate
135
+ `Local Baseline` from the Unit's analyzed evidence; sub-agents do not read this Skill.
84
136
  2. Execute according to dependency order and selected owner.
85
137
  3. Run targeted unit tests and self-audit scope, contracts, TODOs, and introduced warnings.
138
+ Also audit new author attributions and every new model field, enum member, and constant against
139
+ the comment requirements above, then check local-style deviations, unnecessary abstractions,
140
+ one-use constant extraction, and affected core Java Javadoc before recording success.
86
141
  4. Append one `result` record. Only a successful unit uses `status:"completed"`; include
87
142
  unresolved issues rather than hiding them, and do not advance while `issues` or
88
143
  `needs_attention` is non-empty.
144
+ For Canonical-backed success, write each owned source Step `completed` through
145
+ `writeback-spec-step`, with passed evidence for every bound Canonical Test ID and a stable key.
146
+ After every source Step for that task is complete, write the task `implemented`. On failure,
147
+ write the affected Step `failed`; the shared writer moves its task to `blocked`. Local evidence
148
+ is appended first, shared projection second, and the returned acknowledgment last.
89
149
  5. If a result changes a cross-unit contract, stop dependent units and return to ANALYSIS.
90
150
  6. For parallel units, detect overlapping writes before advancing.
91
151
  7. If implementation needs a file, symbol, repository, or source step outside the mapped
92
152
  Canonical change set, stop and return to ANALYSIS instead of expanding scope implicitly.
153
+ 8. If a static Canonical change is confirmed, revise the original design by exactly one revision
154
+ and use `sync-spec-design`; never edit the machine-owned execution block. If a writeback was
155
+ interrupted, run `reconcile-spec-execution` with the stored idempotent pending action.
156
+ Reconciliation only consumes dispatch/result evidence created after the current `in_progress`
157
+ acknowledgment; it never opens a new repair attempt or reuses an earlier attempt's result.
93
158
 
94
159
  Do not emit a progress message for every trivial edit. Report at unit boundaries to reduce
95
160
  conversation overhead while keeping work observable.
@@ -108,4 +173,10 @@ conversation overhead while keeping work observable.
108
173
  - [ ] Every unit has a dispatch/result pair and satisfied its acceptance criteria.
109
174
  - [ ] Targeted tests ran or a concrete blocker is recorded.
110
175
  - [ ] Cross-unit contracts still match.
176
+ - [ ] The implementation follows the evidenced Local Baseline or records a required deviation.
177
+ - [ ] No speculative layer, fragmented micro-method set, or single-use getter constant was added.
178
+ - [ ] New author attributions use the user-facing host `<Current Agent Name> with Easy Coding`.
179
+ - [ ] Every new model field, enum member, and constant has a meaningful field-level comment.
180
+ - [ ] Every method/field in a new core Java class, and every added or materially modified one in
181
+ an existing core Java class, has Javadoc.
111
182
  - [ ] Code tasks enter REVIEW, regardless of workflow mode.
@@ -3,10 +3,17 @@ name: ec-memory
3
3
  description: MEMORY-stage skill. Creates a workflow-mode-aware schema-v2 checkpoint from existing task evidence and performs conditional long-memory distillation.
4
4
  ---
5
5
 
6
- # ec-memory — evidence-derived checkpoint
6
+ # ec-memory — evidence-derived checkpoint and knowledge governance
7
7
 
8
- MEMORY remains mandatory for code tasks. It must not re-analyze the repository or repeat the
9
- entire conversation. Generate from `task.json`, `dev-spec.md`, and `execution.jsonl`.
8
+ MEMORY remains mandatory for code tasks. Daily task processing and architecture maintenance are
9
+ separate responsibilities: every completed code task produces one immutable short-memory fact;
10
+ only a long-memory distillation, or the explicit missing-ABSTRACT startup exception, may open an
11
+ architecture assessment. Never update architecture merely because MEMORY was entered.
12
+
13
+ The short-memory checkpoint must not re-analyze the repository or repeat the entire conversation.
14
+ Generate it only from the verified evidence already stored in `task.json`, `dev-spec.md`, and
15
+ `execution.jsonl`. The bounded repository reads described below belong only to a required
16
+ `backfill` or `update` architecture assessment.
10
17
 
11
18
  ## Depth by workflow mode
12
19
 
@@ -28,9 +35,76 @@ Name it `{memory_id}_{YYYYMMDD}_{smart_name}.md` and set
28
35
  `.easy-coding/memory/short/`, then register it with
29
36
  `memory-short-complete`. Never invent test results or commit hashes.
30
37
 
38
+ Copy the final `acceptance` record from `execution.jsonl` into the checkpoint as a concise
39
+ decision fact: authorization source, decision summary, `diff_sha256`, review policy, verification
40
+ policy, changed files, and any Canonical source tasks that required targeted verification.
41
+ `memory-short-complete` rejects a checkpoint that omits any of those decision fields. This records
42
+ the user's accepted exception without re-reviewing or re-analyzing the code. Canonical writeback
43
+ already carries the same digest and authorization as shared `acceptance` evidence.
44
+
45
+ When frozen TDD is enabled, add its threshold, lifecycle evidence, passed local unit-test result,
46
+ and local changed-line result to the short memory's execution evidence. Remote CI status is not
47
+ part of Harness acceptance or task memory. When TDD is off, omit TDD fields entirely so ordinary
48
+ tasks incur no additional memory work.
49
+
31
50
  Ask the state API for `memory-instruction`. Distill only when it returns `action:distill`;
32
51
  otherwise record `no-op`. Long memory receives reusable facts only, not file dumps, transient
33
52
  logs, routine command output, or speculation.
34
53
 
54
+ ## Architecture assessment
55
+
56
+ Read the frozen `architecture_assessment` contract returned by `memory-instruction`.
57
+
58
+ - `required:false`: do not read the repository for architecture purposes and do not modify
59
+ `.easy-coding/ABSTRACT.md` or `.easy-coding/CHANGELOG.md`.
60
+ - `trigger:distillation`: finish classifying the frozen `candidate_files`, then assess whether
61
+ their stable, reusable facts make the current architecture cognition stale. Default to
62
+ `no-op`.
63
+ - `trigger:missing-abstract`: use `backfill` after the first substantive startup task even when
64
+ long-memory action is `no-op`. This is the only non-distillation architecture exception.
65
+
66
+ An architecture `update` is justified only by evidence of at least one of these changes:
67
+
68
+ - a module was added, removed, split, or merged;
69
+ - module responsibility, ownership, or dependency direction changed;
70
+ - a core request, data, state, or event flow changed;
71
+ - the technology stack, runtime, build, or deployment infrastructure changed;
72
+ - the existing ABSTRACT conflicts with verified current facts.
73
+
74
+ Do not update for a bug fix, local implementation detail, DTO/field-only change, local refactor,
75
+ temporary workaround, routine dependency patch, or the mere fact that distillation ran. Stable
76
+ new coding conventions belong in `TECHNICAL.md` as explicit RULES update candidates; never
77
+ silently edit `RULES.md`, `SOUL.md`, or `TEST_STRATEGY.md` from MEMORY.
78
+
79
+ For `no-op`, use only the frozen memory evidence and give a concrete reason. For `backfill` or
80
+ `update`, read only candidate-related modules, entrypoints, dependencies, and affected ABSTRACT
81
+ sections. Do not perform an unbounded repository re-analysis. Create or edit only the affected
82
+ sections of `.easy-coding/ABSTRACT.md`, and create or append `.easy-coding/CHANGELOG.md`; never
83
+ regenerate the whole ABSTRACT when a bounded edit is sufficient.
84
+
85
+ Whenever `required:true`, record the decision through the command below. For distillation this
86
+ must succeed before deleting any candidate; for `missing-abstract` it must succeed before the
87
+ `no-op` long-memory action can complete:
88
+
89
+ ```bash
90
+ {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py memory-architecture-assessment \
91
+ --session-file <P> --action <no-op|backfill|update> --reason <reason> \
92
+ --evidence <frozen-memory-file> [--evidence <frozen-memory-file> ...] \
93
+ [--affected-section <section> ...] --agent <agent-id>
94
+ ```
95
+
96
+ `backfill` and `update` require affected sections; `no-op` must not declare them. Evidence must
97
+ come from the frozen candidate set, or from the current checkpoint for the missing-ABSTRACT
98
+ exception. If assessment or architecture-file validation fails, keep every candidate file and
99
+ remain in MEMORY.
100
+
101
+ After the assessment succeeds, a distillation may delete all frozen `candidate_files` while
102
+ preserving every `kept_file`. Then call `memory-complete`. The state API rechecks that no-op
103
+ assets stayed unchanged, changed assets still match the recorded assessment, all candidates were
104
+ consumed, and all retained memories still exist.
105
+
35
106
  Complete processing with `memory-complete`. When the state API reports
36
107
  `memory_progress.completed:true`, call `auto-transition --stage COMPLETE`.
108
+ For Canonical-backed tasks this automatic edge first writes every verified selected source task
109
+ to shared `completed`; pending integration dependencies or failed writeback keep the task in
110
+ MEMORY. Do not claim COMPLETE from local memory state alone.
@@ -12,7 +12,17 @@ Every new code task enters REVIEW. Read-only tasks do not. Obtain the current fi
12
12
  ```
13
13
 
14
14
  Review the final diff against `dev-spec.md`, RULES, unit acceptance criteria, tests, contracts,
15
- and obvious security risks. Every finding cites `file:line`.
15
+ the Unit's Local Baseline, and obvious security risks. Every finding cites `file:line`.
16
+
17
+ Review local fit before recommending generic cleanup. Flag an implementation when it departs from
18
+ the nearest comparable naming, control flow, null/error handling, layering, modeling, or method
19
+ granularity without a correctness, security, requirement, or hard-rule reason. Also flag
20
+ speculative layers, fragmented one-use micro-methods, constants created only for one getter
21
+ return, and missing Javadoc in core Java code: check every method/field in a new core class and
22
+ each added or materially modified one in an existing core class. Do not demand defensive null
23
+ checks, abstraction, constant extraction, or legacy-wide comment retrofits merely because they
24
+ are generic best practices. A violation of an explicit task-card coding/comment contract is a
25
+ contract defect, not optional stylistic advice.
16
26
 
17
27
  For Canonical-backed tasks, group evidence by `repo_id` and `source_task_id`. Every selected
18
28
  Spec task needs an implementation result and source test evidence; file references remain
@@ -21,6 +31,14 @@ a blocking correctness finding. Every Canonical review record includes its `repo
21
31
  `source_task_id`; emit at least one current-fingerprint record per selected task and required
22
32
  review dimension. A global record without source ownership cannot satisfy the gate.
23
33
 
34
+ When frozen TDD is enabled, add a passed review dimension named exactly `tdd` for each source
35
+ task. Review whether RED/GREEN/REFACTOR (or characterization GREEN for pure refactors) is genuine,
36
+ tests exercise changed behavior and boundaries, mocks do not merely mirror implementation, and
37
+ the local unit-test command genuinely passes while the changed-line coverage command uses the
38
+ frozen baseline and threshold. Generated CI configuration may be reviewed when it changed, but
39
+ remote CI status is never a review or acceptance dependency. When TDD is off, do not add this
40
+ dimension or raise the ordinary review depth.
41
+
24
42
  ## Depth by workflow mode
25
43
 
26
44
  - `fast`: main Agent performs one final-diff self-review across correctness, scope, tests, and
@@ -61,6 +79,12 @@ Verdict:
61
79
  In-scope defects are fixed automatically. Ask the user only for a new design choice, changed
62
80
  public contract, or contradiction with a confirmed decision.
63
81
 
82
+ For Canonical-backed review, append the local review record first. Any blocking finding then
83
+ writes the owning source task `blocked` through `writeback-spec-task`, referencing the local
84
+ record. A passed review does not change shared task status. A `replan` verdict returns to ANALYSIS;
85
+ confirmed static Spec changes use revision + READY + `sync-spec-design` rather than edits to the
86
+ derived plan or machine execution block.
87
+
64
88
  ## Evidence record
65
89
 
66
90
  Append one final record per executed dimension for the current implementation fingerprint.
@@ -22,6 +22,9 @@ when you recognize abandonment intent in the user's message.
22
22
  `{{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py close-current --session-file <P> --reason "<reason>" --agent <agent-id>`.
23
23
  This sets `task.json.status` to `CLOSED`, records `closed_reason`, updates history, and
24
24
  clears session `current_task` so the next hook injection returns to Ready.
25
+ For Canonical-backed work, the same command first projects every unfinished selected task to
26
+ shared `cancelled` (using an intermediate `blocked` transition when required). A shared writer
27
+ failure leaves the Harness task open; never close only the local side.
25
28
  Use the returned `status_context` as the authoritative status source for the rest of the
26
29
  current turn.
27
30
  4. **No memory flow.** Do not run MEMORY. An incomplete task's memory is dirty data.
@@ -34,6 +37,7 @@ when you recognize abandonment intent in the user's message.
34
37
 
35
38
  - Never delete task folders — CLOSED tasks stay as a record.
36
39
  - Never run the memory/archive flow.
40
+ - Never hand-edit the Canonical execution region to force cancellation.
37
41
  - This skill closes the `current_task`. If the user wants to close a different (suspended)
38
42
  task, they should first switch to it via ec-workflow, then invoke ec-task-close.
39
43
  - Division of labor: ec-task-management lists/creates (read-only panel), ec-workflow runs the
@@ -1,9 +1,9 @@
1
1
  ---
2
2
  name: ec-task-management
3
- description: View and manage Easy Coding tasks plus project/session approval and workflow-mode settings.
3
+ description: View and manage Easy Coding task lifecycle, ownership, handoff, and closure.
4
4
  ---
5
5
 
6
- # ec-task-management — tasks and session modes
6
+ # ec-task-management — task lifecycle
7
7
 
8
8
  Communicate with the user in the user's language. A bare invocation is read-only: show the
9
9
  panel and available actions, but do not mutate a session without an explicit choice.
@@ -13,47 +13,28 @@ panel and available actions, but do not mutate a session without an explicit cho
13
13
  Call the state API snapshot and show:
14
14
 
15
15
  - current task, stage, last Agent, and pending transition;
16
- - `project_approval_mode`, `session_approval_mode`, `effective_approval_mode`;
17
- - `project_workflow_mode`, `session_workflow_mode`, `configured_workflow_mode`;
18
- - task `concrete_workflow_mode` or ANALYSIS proposal when present;
16
+ - task `concrete_workflow_mode` and frozen TDD state when present;
19
17
  - harness enabled/disabled state;
20
18
  - active and resumable tasks.
21
- - for Canonical-backed tasks: source Spec ID/revision/SHA, selected task IDs, repository
19
+ - for Canonical-backed tasks: source locator/path mode, Spec ID/design revision/design digest,
20
+ document digest, execution revision, writeback status, selected task IDs, repository
22
21
  bindings/baseline status, and pending dependency evidence.
23
22
 
24
- Explain precedence:
25
-
26
- `session override > project config > approval:guard / workflow:adaptive`
27
-
28
- ## Session settings
29
-
30
- After explicit user selection:
31
-
32
- ```bash
33
- # approval
34
- {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py set-approval-mode --mode approve|guard|confirm|auto --agent <agent-id> --session-file <P>
35
- {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py clear-approval-mode --agent <agent-id> --session-file <P>
36
-
37
- # workflow
38
- {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py set-workflow-mode --mode adaptive|fast|standard|strict --agent <agent-id> --session-file <P>
39
- {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py clear-workflow-mode --agent <agent-id> --session-file <P>
40
- ```
41
-
42
- Changing a session setting affects future ANALYSIS proposals. It does not silently rewrite a
43
- mode already frozen on an active task. During ANALYSIS, regenerate and show the proposal. During
44
- IMPLEMENT or REVIEW, use `raise-workflow-mode` for a justified increase; lowering is forbidden.
45
- From VERIFICATION, return to IMPLEMENT first so the raised mode receives fresh REVIEW evidence.
46
-
47
- Project settings are changed with `easy-coding config`, which edits both dimensions in one
48
- confirmed interaction.
23
+ Mode inspection and configuration belongs to `ec-config`. If the user asks to change Approval,
24
+ Workflow, TDD, or the TDD coverage threshold, route there and do not mutate those fields here.
49
25
 
50
26
  ## Task actions
51
27
 
52
28
  Support listing, creating, selecting, claiming, handing off, and closing tasks through the
53
- state API. Preserve pending transitions when merely changing approval mode. Never infer user
29
+ state API. Preserve pending transitions when inspecting tasks. Never infer user
54
30
  acceptance from opening this panel.
55
31
 
56
32
  When creating from a Canonical Spec, call `inspect-dev-spec`, display the complete task and
57
33
  dependency selection, then call `select-dev-spec-scope` and `create-task-from-spec` only after
58
34
  explicit user selection. Multiple selected Spec tasks still create one Harness task, while the
59
35
  selector returns one deterministic consumption closure per selected repository.
36
+ Initialize missing shared execution before creation. Support `rebind-spec-source` only when the
37
+ new file matches schema + spec_id + design revision + design_sha256 and does not roll execution
38
+ revision backward. A pending writeback is repaired with `reconcile-spec-execution`, never by
39
+ editing the execution JSON block or starting a different writeback. A deterministic rejected
40
+ action is cleared with `status:error`; correct its input instead of replaying it.
@@ -0,0 +1,101 @@
1
+ ---
2
+ name: ec-tdd-init
3
+ description: Initialize or refresh Java changed-line TDD coverage infrastructure before TDD can be enabled.
4
+ ---
5
+
6
+ # ec-tdd-init — Java changed-line gate initialization
7
+
8
+ Communicate in the user's language. This skill owns TDD infrastructure readiness, not historical
9
+ test-debt cleanup. It must never bulk-generate tests for existing business code, require
10
+ repository-wide coverage, or modify production behavior merely to raise coverage.
11
+
12
+ ## Non-circular ordering
13
+
14
+ The only legal order is:
15
+
16
+ ```text
17
+ TDD off -> initialize infrastructure -> readiness ready -> user explicitly enables TDD
18
+ ```
19
+
20
+ Run this skill as a dedicated code task with `type=tdd-init`. The state API always freezes that
21
+ task with `tdd_enabled=false`, even when a legacy project/session setting or a suspended task has
22
+ TDD enabled. Never offer "enable now and initialize later". Never enable TDD automatically after
23
+ initialization.
24
+
25
+ ## Read-only preflight
26
+
27
+ First run:
28
+
29
+ ```bash
30
+ python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . check
31
+ ```
32
+
33
+ If it returns `ready`, report the recorded build/CI contract and stop without creating a task.
34
+ The user may then use `ec-config` or `easy-coding config` to enable TDD.
35
+
36
+ If it returns `needs_init`, inspect only the infrastructure needed to form a confirmed plan:
37
+
38
+ - Maven/Gradle files and the existing JUnit runner;
39
+ - JaCoCo XML generation configuration;
40
+ - `.gitlab-ci.yml` and its repository-local include chain;
41
+ - TEST-stage job, JUnit/JaCoCo artifacts, and invocation of
42
+ `.easy-coding/tools/easy_coding_java_coverage.py` with task-supplied baseline/threshold values.
43
+
44
+ Do not measure current whole-project coverage. A project with no historical business tests may
45
+ still become ready when the test runner, JaCoCo reporting, and parameterized changed-line gate
46
+ are functional.
47
+
48
+ ## Initialization task
49
+
50
+ After the user confirms the exact infrastructure scope, create one task and route it through the
51
+ ordinary workflow:
52
+
53
+ ```bash
54
+ {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py create-task \
55
+ --task-id <safe-unique-id> --type tdd-init \
56
+ --title "Initialize Java changed-line TDD infrastructure" \
57
+ --agent <agent-id> --session-file <P>
58
+ ```
59
+
60
+ ANALYSIS may plan build files, GitLab CI files, common scripts, and the readiness receipt. It must
61
+ state `historical coverage required: no` and `coverage scope: changed production lines since each
62
+ future task baseline`. IMPLEMENT changes infrastructure only. It must not add tests whose sole
63
+ purpose is to cover unchanged production code.
64
+
65
+ The reusable GitLab job must consume a baseline SHA and threshold supplied for the future task;
66
+ do not hardcode the initialization commit or the default 90% threshold. The same Python coverage
67
+ tool must be usable locally and remotely. This generated job is remote automation infrastructure,
68
+ not a Harness task acceptance dependency: later tasks require local unit-test and coverage
69
+ evidence only, and never wait for a pipeline URL, job identity, or remote success status.
70
+
71
+ ## Readiness receipt and verification
72
+
73
+ At the end of IMPLEMENT, after the infrastructure files are stable, record their fingerprints.
74
+ The receipt is part of the implementation and must exist before REVIEW so review/verification
75
+ fingerprints do not change after review. The recorder automatically includes the harness-managed
76
+ `.easy-coding/tools/easy_coding_java_coverage.py` fingerprint:
77
+
78
+ ```bash
79
+ python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . record \
80
+ --build-file <pom.xml-or-build.gradle> [--build-file <included-build-file>]... \
81
+ --ci-file .gitlab-ci.yml [--ci-file <repository-local-include>]... \
82
+ --coverage-report <jacoco-xml-pattern> [--coverage-report <pattern>]... \
83
+ --gate-command "python3 .easy-coding/tools/easy_coding_java_coverage.py check --base \$EASY_CODING_TDD_BASE_SHA --threshold \$EASY_CODING_TDD_THRESHOLD" \
84
+ --agent <agent-id>
85
+ ```
86
+
87
+ REVIEW includes the receipt and its declared infrastructure boundary. VERIFICATION runs the
88
+ frozen Workflow Mode's applicable build/test/CI syntax checks, then performs only the read-only
89
+ readiness check:
90
+
91
+ ```bash
92
+ python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . check
93
+ ```
94
+
95
+ The `VERIFICATION -> MEMORY` gate requires the final check to return `ready`. If any recorded
96
+ build or CI file changes after the receipt was created, readiness becomes `needs_init`; return to
97
+ IMPLEMENT, refresh the receipt, and repeat REVIEW before verifying again. Rerun this skill when
98
+ the same drift occurs after task completion.
99
+
100
+ After completion, tell the user that TDD remains off and provide the explicit project/session
101
+ enable route. Do not treat readiness as consent to enable it.
@@ -15,7 +15,8 @@ Read-only tasks never enter this stage. Obtain fresh fingerprints before running
15
15
 
16
16
  - No completion claim without executed verification evidence.
17
17
  - Evidence is reusable only while both returned fingerprints remain unchanged.
18
- - Relevant code or config changes invalidate old evidence automatically.
18
+ - Relevant code or config changes invalidate old evidence automatically unless the exact
19
+ post-verification code diff is explicitly accepted under the checkpoint protocol below.
19
20
  - Failed or missing evidence never becomes acceptance because of approval mode.
20
21
 
21
22
  ## Verification depth
@@ -25,6 +26,29 @@ Read-only tasks never enter this stage. Obtain fresh fingerprints before running
25
26
  - `standard`: run impacted lint/typecheck/test scopes and every must-test item.
26
27
  - `strict`: run the project's full applicable lint, typecheck, test, and build gates.
27
28
 
29
+ These rules remain unchanged when frozen TDD is off: do not discover JaCoCo reports, run the
30
+ coverage tool, inspect GitLab, or add a coverage record. The explicit `type=tdd-init` task is an
31
+ infrastructure exception: run its planned build/CI syntax checks and readiness tool, but do not
32
+ measure repository-wide coverage or append TDD coverage evidence for unchanged production code.
33
+
34
+ When frozen TDD is on, first run the planned local Java unit command and generate JaCoCo XML,
35
+ then run the deterministic local acceptance gate:
36
+
37
+ ```bash
38
+ python3 .easy-coding/tools/easy_coding_java_coverage.py check \
39
+ --base <task.tdd_baselines[repo-id-or-project]> \
40
+ --threshold <task.tdd_coverage_threshold> [--report <jacoco.xml>]...
41
+ ```
42
+
43
+ The tool measures covered added/modified production Java executable lines only. Deleted,
44
+ comment, blank, import, and test-source lines are excluded by diff/JaCoCo intersection. Missing
45
+ or ambiguous source files and reports older than their modified source fail; zero modified
46
+ executable lines is explicit N/A. Always regenerate JaCoCo XML after the final source change.
47
+ Never substitute `HEAD`, a mutable ref, project defaults, or current session settings for the
48
+ task-frozen baseline SHA and threshold. `ec-tdd-init` still generates a GitLab job that can run
49
+ the same tool, but remote pipeline execution and status are outside Harness acceptance. Never
50
+ request an intermediate commit or push merely to obtain CI evidence.
51
+
28
52
  The main Agent may run commands inline. Dispatch verifier sub-agents only when checks are
29
53
  independent and parallel execution materially saves time or isolates specialist environments.
30
54
  Platform spawn rule: {{platform_spawn_instruction}}
@@ -48,16 +72,36 @@ Append one record per executed check:
48
72
  }
49
73
  ```
50
74
 
51
- `check_type` is one of `lint`, `typecheck`, `test`, or `build`. In `strict`, append current
75
+ `check_type` is one of `lint`, `typecheck`, `test`, `build`, or (TDD only) `coverage`. In `strict`, append current
52
76
  evidence for all four types. When a type genuinely does not apply, record `applicable: false`
53
77
  and a non-empty `not_applicable_reason`; it does not count as the required applicable executed
54
78
  check, and must not be represented by an invented successful command.
55
79
 
56
80
  Record failures in `failures[]`. If any current-fingerprint record fails, return to IMPLEMENT;
57
- do not append a later synthetic pass without rerunning the failed command.
81
+ do not append a later synthetic pass without rerunning the failed command. For a Canonical-backed
82
+ failure, append the local verify record first, then write the owning source task `blocked` with a
83
+ concise reference to that record. The repair transition automatically reopens blocked source tasks
84
+ only; unaffected implemented tasks retain their latest shared conclusion.
58
85
 
59
86
  ## Coverage and acceptance
60
87
 
88
+ For TDD coverage, copy the tool output into `coverage`: `baseline_sha`, `covered_lines`,
89
+ `total_lines`, `percentage`, frozen `threshold`, `report_paths`, and `report_sha256`. Set
90
+ `applicable:false` plus the tool's reason only for zero executable modified lines. A percentage
91
+ below the frozen threshold fails even when ordinary tests pass.
92
+
93
+ Append one coverage record with `coverage_scope:"local"` per repository (and per Canonical
94
+ source task). The state gate also requires a passed local `check_type:"test"` record for the same
95
+ owner. The coverage record preserves the task-frozen baseline and threshold. Do not append or
96
+ wait for GitLab pipeline evidence; historical remote coverage records are ignored by acceptance
97
+ without modifying or deleting the stored records.
98
+
99
+ For `type=tdd-init`, the infrastructure receipt must already have been recorded during IMPLEMENT
100
+ and reviewed with the rest of the implementation. Run only `easy_coding_tdd_readiness.py check`
101
+ here. If it reports drift, return to IMPLEMENT to refresh the receipt and repeat REVIEW; never
102
+ rewrite it inside VERIFICATION. The state gate requires `ready` before MEMORY. This does not
103
+ enable TDD; report the explicit `ec-config`/`easy-coding config` next step.
104
+
61
105
  - Every must-test item has an executed check.
62
106
  - Bug fixes include a regression test when project infrastructure exists.
63
107
  - Present changed scope, commands, results, and unverified items.
@@ -66,6 +110,38 @@ do not append a later synthetic pass without rerunning the failed command.
66
110
  user wait.
67
111
  - A reported in-scope problem returns to IMPLEMENT; out-of-scope work becomes a separate task.
68
112
 
113
+ After the final green evidence is recorded, freeze the acceptance baseline before presenting the
114
+ result or applying the boundary:
115
+
116
+ ```bash
117
+ {{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py verification-checkpoint \
118
+ --agent <agent-id> --session-file <P>
119
+ ```
120
+
121
+ Then request or auto-apply VERIFICATION -> MEMORY according to `approval_mode`. `auto` remains
122
+ automatic when the checkpoint is unchanged. If any in-scope code changed after the checkpoint,
123
+ the state API returns `action:"acceptance-drift"`, keeps the task in VERIFICATION, and includes
124
+ the exact unified patches or binary/mode-change descriptions plus a stable `diff_sha256`. This
125
+ exceptional drift pauses every approval mode, including `auto`; it does not permanently change
126
+ the configured mode.
127
+
128
+ Show the complete returned diff and ask whether to accept that exact digest. Do not re-enter
129
+ IMPLEMENT or rerun REVIEW merely because this drift exists. On acceptance, call
130
+ `confirm-transition --stage MEMORY --diff-sha256 <digest>` with exactly one policy:
131
+
132
+ - `carry-forward`: only when every changed hunk is confidently non-executable and existing
133
+ verification remains applicable;
134
+ - `targeted`: executable behavior changed; append passed current-fingerprint targeted verification
135
+ before confirming. Canonical tasks must cover every affected source task reported by the
136
+ acceptance record, without rerunning checks for unaffected source tasks;
137
+ - `waived`: the user explicitly accepts the stated unverified risk.
138
+
139
+ Include `--decision-summary` with the user's decision. A changed digest invalidates the pending
140
+ confirmation and must be shown again. Behavior config, execution plan, workflow, Canonical
141
+ design, or nested-repository metadata drift cannot use this shortcut; return to ANALYSIS or
142
+ IMPLEMENT as reported by the state API. The acceptance record bridges only the accepted
143
+ implementation fingerprints, so prior REVIEW evidence remains valid without a second REVIEW.
144
+
69
145
  For Canonical-backed tasks, run each repository's commands from `task.repo_paths[repo_id]` and
70
146
  cover every selected task's source test IDs. Report pending integration edges separately from
71
147
  local green checks. They do not block local implementation evidence, but the state API blocks
@@ -76,6 +152,16 @@ tasks remain separate evidence records. In `strict`, every involved repository i
76
152
  records all four check types; a repository-specific non-applicable record still needs its reason
77
153
  and source ownership.
78
154
 
155
+ After implementation and local checks, each selected Canonical source task remains
156
+ `implemented`. Do not call `writeback-spec-task --status verified` from VERIFICATION. Applying
157
+ VERIFICATION -> MEMORY is the authoritative acceptance boundary: the state API writes each
158
+ still-implemented source task to `verified` through CAS/idempotent recoverable events with its
159
+ accepted test evidence and acceptance digest, then enters MEMORY only after every write is
160
+ confirmed. For `approve`/`guard`, that authority is the explicit boundary
161
+ confirmation; for `confirm`/`auto`, it is the standing approval-mode authorization when no new
162
+ drift exists. If writeback is interrupted, run `reconcile-spec-execution` before retrying the
163
+ transition. Remote CI remains outside this acceptance gate.
164
+
79
165
  Record the exact integration edge only after its evidence exists:
80
166
 
81
167
  ```bash
@@ -87,5 +173,5 @@ Record the exact integration edge only after its evidence exists:
87
173
  --agent <agent>
88
174
  ```
89
175
 
90
- The state API rejects VERIFICATION -> MEMORY unless all evidence for the current implementation
91
- and config fingerprints is green.
176
+ The state API rejects VERIFICATION -> MEMORY unless all effective evidence is green and the
177
+ checkpoint is either unchanged or bound to an exact accepted diff.