easy-coding-harness 0.10.0-beta.1 → 0.10.0-beta.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +151 -0
- package/README.md +50 -20
- package/dist/cli.js +478 -47
- package/dist/cli.js.map +1 -1
- package/package.json +1 -1
- package/templates/claude/agents/ec-implementer.md +11 -0
- package/templates/claude/agents/ec-reviewer.md +6 -1
- package/templates/codex/agents/ec-implementer.toml +11 -0
- package/templates/codex/agents/ec-reviewer.toml +6 -1
- package/templates/common/bundled-skills/ec-init/SKILL.md +18 -6
- package/templates/common/bundled-skills/ec-meta/references/local-architecture/README.md +27 -12
- package/templates/common/bundled-skills/ec-meta/references/platform-files/README.md +1 -1
- package/templates/common/skills/ec-analysis/SKILL.md +141 -35
- package/templates/common/skills/ec-config/SKILL.md +24 -2
- package/templates/common/skills/ec-git/SKILL.md +7 -1
- package/templates/common/skills/ec-implementing/SKILL.md +61 -1
- package/templates/common/skills/ec-memory/SKILL.md +76 -6
- package/templates/common/skills/ec-reviewing/SKILL.md +21 -3
- package/templates/common/skills/ec-task-close/SKILL.md +4 -0
- package/templates/common/skills/ec-task-management/SKILL.md +7 -1
- package/templates/common/skills/ec-tdd-init/SKILL.md +101 -0
- package/templates/common/skills/ec-verification/SKILL.md +68 -26
- package/templates/common/skills/ec-workflow/SKILL.md +81 -22
- package/templates/main-constraint/AGENTS.md.tpl +49 -14
- package/templates/main-constraint/CLAUDE.md.tpl +46 -14
- package/templates/qoder/agents/ec-implementer.md +11 -0
- package/templates/qoder/agents/ec-reviewer.md +6 -1
- package/templates/runtime/templates/dev-spec-skeleton.md +8 -1
- package/templates/runtime/tools/easy_coding_tdd_readiness.py +306 -0
- package/templates/shared-hooks/easy_coding_state.py +3878 -259
- package/templates/shared-hooks/easy_dev_spec.py +444 -30
- package/templates/shared-hooks/easy_dev_spec_execution.py +1014 -0
- package/templates/shared-hooks/easy_dev_spec_protocol.py +1426 -18
|
@@ -36,6 +36,37 @@ Communicate with the user in the user's language.
|
|
|
36
36
|
enters REVIEW.
|
|
37
37
|
7. Read-only `doc` / `analysis` / `report` tasks remain `single` with `files:[]`, make no writes,
|
|
38
38
|
return a non-empty `deliverable`, then follow the mode-aware IMPLEMENT -> COMPLETE edge.
|
|
39
|
+
8. When a project template, local convention, or new source header uses author attribution, the
|
|
40
|
+
author value must be `<Current Agent Name> with Easy Coding`, for example
|
|
41
|
+
`Codex with Easy Coding`. `Current Agent Name` means the user-facing host Agent (for example,
|
|
42
|
+
Codex, Claude, or Qoder), never an implementation sub-agent role such as `ec-implementer`.
|
|
43
|
+
This value is display attribution only: never pass it to the workflow state API's `--agent`,
|
|
44
|
+
which accepts only `claude-code`, `codex`, or `qoder`. Never copy a previous human or Agent name
|
|
45
|
+
into newly authored code.
|
|
46
|
+
9. Every newly added field in a data-bearing model must have a meaningful field-level comment.
|
|
47
|
+
This includes new or extended entity/DO/DTO/VO/BO, request/response, configuration, and similar
|
|
48
|
+
model types. Every new enum member and every new declared constant requires the same treatment.
|
|
49
|
+
Describe the semantic meaning and, when relevant, units, format, allowed values, nullability,
|
|
50
|
+
default behavior, or compatibility constraints. A type-level comment does not replace comments
|
|
51
|
+
on its fields or members; do not add low-value comments to ordinary local variables.
|
|
52
|
+
10. Treat the task card's `Local Baseline` as the default implementation shape. Match the nearest
|
|
53
|
+
comparable code's naming, control flow, null/empty and error handling, layering, object model,
|
|
54
|
+
and extraction granularity unless correctness, security, an explicit requirement, or a hard
|
|
55
|
+
project rule requires a deviation. Do not add defensive null checks solely because they are a
|
|
56
|
+
generic best practice when the evidenced local contract intentionally omits them.
|
|
57
|
+
11. Implement the smallest coherent design. Do not add speculative abstractions, wrappers,
|
|
58
|
+
factories, layers, or extension points, and do not fragment one readable flow into many
|
|
59
|
+
single-use micro-methods. Extract code only for a clear semantic boundary, real reuse,
|
|
60
|
+
independent testability, or a material reduction in complexity.
|
|
61
|
+
12. Literals and magic values are allowed when they are obvious, local, and consistent with the
|
|
62
|
+
surrounding code. Introduce a constant for repeated use, stable domain/config/protocol
|
|
63
|
+
semantics, or an established project convention—not merely to hold the single return value of
|
|
64
|
+
a getter.
|
|
65
|
+
13. In a newly added core Java class, every method and field requires meaningful Javadoc. In an
|
|
66
|
+
existing core Java class, every added or materially modified method and field requires it.
|
|
67
|
+
Add focused inline comments to core or complex logic to explain intent, constraints, or
|
|
68
|
+
non-obvious tradeoffs. Do not mass-retrofit untouched legacy code, and for non-Java code
|
|
69
|
+
follow the language's doc-comment form plus the evidenced project convention.
|
|
39
70
|
|
|
40
71
|
## Choose the execution owner
|
|
41
72
|
|
|
@@ -81,26 +112,49 @@ Sub-agents never dispatch other sub-agents or read `.easy-coding` workflow asset
|
|
|
81
112
|
## Test Points {unit.test_points and exact targeted commands}
|
|
82
113
|
## Contracts {inputs, outputs, invariants shared with other units}
|
|
83
114
|
## Risks {known edge cases and compatibility risks}
|
|
115
|
+
## Local Baseline {nearest comparable code conventions and evidence paths}
|
|
116
|
+
## Code Comments {resolved host Agent author value; model-field, enum-member, and constant rules}
|
|
84
117
|
## Coding Rules {pre-digested RULES sections}
|
|
85
118
|
## Architecture {pre-digested ABSTRACT sections}
|
|
86
119
|
## Output
|
|
87
120
|
status:"completed", repo_id|null, source_task_id|null, changed_files[], summary,
|
|
88
|
-
deliverable|null, issues:[], needs_attention:[]
|
|
121
|
+
deliverable|null, checks:[{command,passed,failures:[]}], issues:[], needs_attention:[]
|
|
89
122
|
```
|
|
90
123
|
|
|
91
124
|
## Dispatch and result loop
|
|
92
125
|
|
|
93
126
|
1. Append a `dispatch` record before work begins. Canonical-backed records include `repo_id` and
|
|
94
127
|
`source_task_id`; resolve every file relative to `task.repo_paths[repo_id]` before dispatch.
|
|
128
|
+
Before dispatching a selected task that is not already `in_progress`, call
|
|
129
|
+
`writeback-spec-task --status in_progress` with a key stable for that dispatch/recovery
|
|
130
|
+
attempt but distinct from any earlier accepted `in_progress` event. Do this only after its
|
|
131
|
+
hard/contract dependencies are ready; do not batch-start dependent tasks at the initial
|
|
132
|
+
IMPLEMENT boundary.
|
|
133
|
+
Populate `Code Comments` on every code task card with the resolved user-facing host Agent
|
|
134
|
+
author value, the field/member/constant rules, and the core Java Javadoc rule above. Populate
|
|
135
|
+
`Local Baseline` from the Unit's analyzed evidence; sub-agents do not read this Skill.
|
|
95
136
|
2. Execute according to dependency order and selected owner.
|
|
96
137
|
3. Run targeted unit tests and self-audit scope, contracts, TODOs, and introduced warnings.
|
|
138
|
+
Also audit new author attributions and every new model field, enum member, and constant against
|
|
139
|
+
the comment requirements above, then check local-style deviations, unnecessary abstractions,
|
|
140
|
+
one-use constant extraction, and affected core Java Javadoc before recording success.
|
|
97
141
|
4. Append one `result` record. Only a successful unit uses `status:"completed"`; include
|
|
98
142
|
unresolved issues rather than hiding them, and do not advance while `issues` or
|
|
99
143
|
`needs_attention` is non-empty.
|
|
144
|
+
For Canonical-backed success, write each owned source Step `completed` through
|
|
145
|
+
`writeback-spec-step`, with passed evidence for every bound Canonical Test ID and a stable key.
|
|
146
|
+
After every source Step for that task is complete, write the task `implemented`. On failure,
|
|
147
|
+
write the affected Step `failed`; the shared writer moves its task to `blocked`. Local evidence
|
|
148
|
+
is appended first, shared projection second, and the returned acknowledgment last.
|
|
100
149
|
5. If a result changes a cross-unit contract, stop dependent units and return to ANALYSIS.
|
|
101
150
|
6. For parallel units, detect overlapping writes before advancing.
|
|
102
151
|
7. If implementation needs a file, symbol, repository, or source step outside the mapped
|
|
103
152
|
Canonical change set, stop and return to ANALYSIS instead of expanding scope implicitly.
|
|
153
|
+
8. If a static Canonical change is confirmed, revise the original design by exactly one revision
|
|
154
|
+
and use `sync-spec-design`; never edit the machine-owned execution block. If a writeback was
|
|
155
|
+
interrupted, run `reconcile-spec-execution` with the stored idempotent pending action.
|
|
156
|
+
Reconciliation only consumes dispatch/result evidence created after the current `in_progress`
|
|
157
|
+
acknowledgment; it never opens a new repair attempt or reuses an earlier attempt's result.
|
|
104
158
|
|
|
105
159
|
Do not emit a progress message for every trivial edit. Report at unit boundaries to reduce
|
|
106
160
|
conversation overhead while keeping work observable.
|
|
@@ -119,4 +173,10 @@ conversation overhead while keeping work observable.
|
|
|
119
173
|
- [ ] Every unit has a dispatch/result pair and satisfied its acceptance criteria.
|
|
120
174
|
- [ ] Targeted tests ran or a concrete blocker is recorded.
|
|
121
175
|
- [ ] Cross-unit contracts still match.
|
|
176
|
+
- [ ] The implementation follows the evidenced Local Baseline or records a required deviation.
|
|
177
|
+
- [ ] No speculative layer, fragmented micro-method set, or single-use getter constant was added.
|
|
178
|
+
- [ ] New author attributions use the user-facing host `<Current Agent Name> with Easy Coding`.
|
|
179
|
+
- [ ] Every new model field, enum member, and constant has a meaningful field-level comment.
|
|
180
|
+
- [ ] Every method/field in a new core Java class, and every added or materially modified one in
|
|
181
|
+
an existing core Java class, has Javadoc.
|
|
122
182
|
- [ ] Code tasks enter REVIEW, regardless of workflow mode.
|
|
@@ -3,10 +3,17 @@ name: ec-memory
|
|
|
3
3
|
description: MEMORY-stage skill. Creates a workflow-mode-aware schema-v2 checkpoint from existing task evidence and performs conditional long-memory distillation.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# ec-memory — evidence-derived checkpoint
|
|
6
|
+
# ec-memory — evidence-derived checkpoint and knowledge governance
|
|
7
7
|
|
|
8
|
-
MEMORY remains mandatory for code tasks.
|
|
9
|
-
|
|
8
|
+
MEMORY remains mandatory for code tasks. Daily task processing and architecture maintenance are
|
|
9
|
+
separate responsibilities: every completed code task produces one immutable short-memory fact;
|
|
10
|
+
only a long-memory distillation, or the explicit missing-ABSTRACT startup exception, may open an
|
|
11
|
+
architecture assessment. Never update architecture merely because MEMORY was entered.
|
|
12
|
+
|
|
13
|
+
The short-memory checkpoint must not re-analyze the repository or repeat the entire conversation.
|
|
14
|
+
Generate it only from the verified evidence already stored in `task.json`, `dev-spec.md`, and
|
|
15
|
+
`execution.jsonl`. The bounded repository reads described below belong only to a required
|
|
16
|
+
`backfill` or `update` architecture assessment.
|
|
10
17
|
|
|
11
18
|
## Depth by workflow mode
|
|
12
19
|
|
|
@@ -28,13 +35,76 @@ Name it `{memory_id}_{YYYYMMDD}_{smart_name}.md` and set
|
|
|
28
35
|
`.easy-coding/memory/short/`, then register it with
|
|
29
36
|
`memory-short-complete`. Never invent test results or commit hashes.
|
|
30
37
|
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
38
|
+
Copy the final `acceptance` record from `execution.jsonl` into the checkpoint as a concise
|
|
39
|
+
decision fact: authorization source, decision summary, `diff_sha256`, review policy, verification
|
|
40
|
+
policy, changed files, and any Canonical source tasks that required targeted verification.
|
|
41
|
+
`memory-short-complete` rejects a checkpoint that omits any of those decision fields. This records
|
|
42
|
+
the user's accepted exception without re-reviewing or re-analyzing the code. Canonical writeback
|
|
43
|
+
already carries the same digest and authorization as shared `acceptance` evidence.
|
|
44
|
+
|
|
45
|
+
When frozen TDD is enabled, add its threshold, lifecycle evidence, passed local unit-test result,
|
|
46
|
+
and local changed-line result to the short memory's execution evidence. Remote CI status is not
|
|
47
|
+
part of Harness acceptance or task memory. When TDD is off, omit TDD fields entirely so ordinary
|
|
48
|
+
tasks incur no additional memory work.
|
|
34
49
|
|
|
35
50
|
Ask the state API for `memory-instruction`. Distill only when it returns `action:distill`;
|
|
36
51
|
otherwise record `no-op`. Long memory receives reusable facts only, not file dumps, transient
|
|
37
52
|
logs, routine command output, or speculation.
|
|
38
53
|
|
|
54
|
+
## Architecture assessment
|
|
55
|
+
|
|
56
|
+
Read the frozen `architecture_assessment` contract returned by `memory-instruction`.
|
|
57
|
+
|
|
58
|
+
- `required:false`: do not read the repository for architecture purposes and do not modify
|
|
59
|
+
`.easy-coding/ABSTRACT.md` or `.easy-coding/CHANGELOG.md`.
|
|
60
|
+
- `trigger:distillation`: finish classifying the frozen `candidate_files`, then assess whether
|
|
61
|
+
their stable, reusable facts make the current architecture cognition stale. Default to
|
|
62
|
+
`no-op`.
|
|
63
|
+
- `trigger:missing-abstract`: use `backfill` after the first substantive startup task even when
|
|
64
|
+
long-memory action is `no-op`. This is the only non-distillation architecture exception.
|
|
65
|
+
|
|
66
|
+
An architecture `update` is justified only by evidence of at least one of these changes:
|
|
67
|
+
|
|
68
|
+
- a module was added, removed, split, or merged;
|
|
69
|
+
- module responsibility, ownership, or dependency direction changed;
|
|
70
|
+
- a core request, data, state, or event flow changed;
|
|
71
|
+
- the technology stack, runtime, build, or deployment infrastructure changed;
|
|
72
|
+
- the existing ABSTRACT conflicts with verified current facts.
|
|
73
|
+
|
|
74
|
+
Do not update for a bug fix, local implementation detail, DTO/field-only change, local refactor,
|
|
75
|
+
temporary workaround, routine dependency patch, or the mere fact that distillation ran. Stable
|
|
76
|
+
new coding conventions belong in `TECHNICAL.md` as explicit RULES update candidates; never
|
|
77
|
+
silently edit `RULES.md`, `SOUL.md`, or `TEST_STRATEGY.md` from MEMORY.
|
|
78
|
+
|
|
79
|
+
For `no-op`, use only the frozen memory evidence and give a concrete reason. For `backfill` or
|
|
80
|
+
`update`, read only candidate-related modules, entrypoints, dependencies, and affected ABSTRACT
|
|
81
|
+
sections. Do not perform an unbounded repository re-analysis. Create or edit only the affected
|
|
82
|
+
sections of `.easy-coding/ABSTRACT.md`, and create or append `.easy-coding/CHANGELOG.md`; never
|
|
83
|
+
regenerate the whole ABSTRACT when a bounded edit is sufficient.
|
|
84
|
+
|
|
85
|
+
Whenever `required:true`, record the decision through the command below. For distillation this
|
|
86
|
+
must succeed before deleting any candidate; for `missing-abstract` it must succeed before the
|
|
87
|
+
`no-op` long-memory action can complete:
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
{{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py memory-architecture-assessment \
|
|
91
|
+
--session-file <P> --action <no-op|backfill|update> --reason <reason> \
|
|
92
|
+
--evidence <frozen-memory-file> [--evidence <frozen-memory-file> ...] \
|
|
93
|
+
[--affected-section <section> ...] --agent <agent-id>
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
`backfill` and `update` require affected sections; `no-op` must not declare them. Evidence must
|
|
97
|
+
come from the frozen candidate set, or from the current checkpoint for the missing-ABSTRACT
|
|
98
|
+
exception. If assessment or architecture-file validation fails, keep every candidate file and
|
|
99
|
+
remain in MEMORY.
|
|
100
|
+
|
|
101
|
+
After the assessment succeeds, a distillation may delete all frozen `candidate_files` while
|
|
102
|
+
preserving every `kept_file`. Then call `memory-complete`. The state API rechecks that no-op
|
|
103
|
+
assets stayed unchanged, changed assets still match the recorded assessment, all candidates were
|
|
104
|
+
consumed, and all retained memories still exist.
|
|
105
|
+
|
|
39
106
|
Complete processing with `memory-complete`. When the state API reports
|
|
40
107
|
`memory_progress.completed:true`, call `auto-transition --stage COMPLETE`.
|
|
108
|
+
For Canonical-backed tasks this automatic edge first writes every verified selected source task
|
|
109
|
+
to shared `completed`; pending integration dependencies or failed writeback keep the task in
|
|
110
|
+
MEMORY. Do not claim COMPLETE from local memory state alone.
|
|
@@ -12,7 +12,17 @@ Every new code task enters REVIEW. Read-only tasks do not. Obtain the current fi
|
|
|
12
12
|
```
|
|
13
13
|
|
|
14
14
|
Review the final diff against `dev-spec.md`, RULES, unit acceptance criteria, tests, contracts,
|
|
15
|
-
and obvious security risks. Every finding cites `file:line`.
|
|
15
|
+
the Unit's Local Baseline, and obvious security risks. Every finding cites `file:line`.
|
|
16
|
+
|
|
17
|
+
Review local fit before recommending generic cleanup. Flag an implementation when it departs from
|
|
18
|
+
the nearest comparable naming, control flow, null/error handling, layering, modeling, or method
|
|
19
|
+
granularity without a correctness, security, requirement, or hard-rule reason. Also flag
|
|
20
|
+
speculative layers, fragmented one-use micro-methods, constants created only for one getter
|
|
21
|
+
return, and missing Javadoc in core Java code: check every method/field in a new core class and
|
|
22
|
+
each added or materially modified one in an existing core class. Do not demand defensive null
|
|
23
|
+
checks, abstraction, constant extraction, or legacy-wide comment retrofits merely because they
|
|
24
|
+
are generic best practices. A violation of an explicit task-card coding/comment contract is a
|
|
25
|
+
contract defect, not optional stylistic advice.
|
|
16
26
|
|
|
17
27
|
For Canonical-backed tasks, group evidence by `repo_id` and `source_task_id`. Every selected
|
|
18
28
|
Spec task needs an implementation result and source test evidence; file references remain
|
|
@@ -24,8 +34,10 @@ review dimension. A global record without source ownership cannot satisfy the ga
|
|
|
24
34
|
When frozen TDD is enabled, add a passed review dimension named exactly `tdd` for each source
|
|
25
35
|
task. Review whether RED/GREEN/REFACTOR (or characterization GREEN for pure refactors) is genuine,
|
|
26
36
|
tests exercise changed behavior and boundaries, mocks do not merely mirror implementation, and
|
|
27
|
-
the local
|
|
28
|
-
|
|
37
|
+
the local unit-test command genuinely passes while the changed-line coverage command uses the
|
|
38
|
+
frozen baseline and threshold. Generated CI configuration may be reviewed when it changed, but
|
|
39
|
+
remote CI status is never a review or acceptance dependency. When TDD is off, do not add this
|
|
40
|
+
dimension or raise the ordinary review depth.
|
|
29
41
|
|
|
30
42
|
## Depth by workflow mode
|
|
31
43
|
|
|
@@ -67,6 +79,12 @@ Verdict:
|
|
|
67
79
|
In-scope defects are fixed automatically. Ask the user only for a new design choice, changed
|
|
68
80
|
public contract, or contradiction with a confirmed decision.
|
|
69
81
|
|
|
82
|
+
For Canonical-backed review, append the local review record first. Any blocking finding then
|
|
83
|
+
writes the owning source task `blocked` through `writeback-spec-task`, referencing the local
|
|
84
|
+
record. A passed review does not change shared task status. A `replan` verdict returns to ANALYSIS;
|
|
85
|
+
confirmed static Spec changes use revision + READY + `sync-spec-design` rather than edits to the
|
|
86
|
+
derived plan or machine execution block.
|
|
87
|
+
|
|
70
88
|
## Evidence record
|
|
71
89
|
|
|
72
90
|
Append one final record per executed dimension for the current implementation fingerprint.
|
|
@@ -22,6 +22,9 @@ when you recognize abandonment intent in the user's message.
|
|
|
22
22
|
`{{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py close-current --session-file <P> --reason "<reason>" --agent <agent-id>`.
|
|
23
23
|
This sets `task.json.status` to `CLOSED`, records `closed_reason`, updates history, and
|
|
24
24
|
clears session `current_task` so the next hook injection returns to Ready.
|
|
25
|
+
For Canonical-backed work, the same command first projects every unfinished selected task to
|
|
26
|
+
shared `cancelled` (using an intermediate `blocked` transition when required). A shared writer
|
|
27
|
+
failure leaves the Harness task open; never close only the local side.
|
|
25
28
|
Use the returned `status_context` as the authoritative status source for the rest of the
|
|
26
29
|
current turn.
|
|
27
30
|
4. **No memory flow.** Do not run MEMORY. An incomplete task's memory is dirty data.
|
|
@@ -34,6 +37,7 @@ when you recognize abandonment intent in the user's message.
|
|
|
34
37
|
|
|
35
38
|
- Never delete task folders — CLOSED tasks stay as a record.
|
|
36
39
|
- Never run the memory/archive flow.
|
|
40
|
+
- Never hand-edit the Canonical execution region to force cancellation.
|
|
37
41
|
- This skill closes the `current_task`. If the user wants to close a different (suspended)
|
|
38
42
|
task, they should first switch to it via ec-workflow, then invoke ec-task-close.
|
|
39
43
|
- Division of labor: ec-task-management lists/creates (read-only panel), ec-workflow runs the
|
|
@@ -16,7 +16,8 @@ Call the state API snapshot and show:
|
|
|
16
16
|
- task `concrete_workflow_mode` and frozen TDD state when present;
|
|
17
17
|
- harness enabled/disabled state;
|
|
18
18
|
- active and resumable tasks.
|
|
19
|
-
- for Canonical-backed tasks: source Spec ID/revision/
|
|
19
|
+
- for Canonical-backed tasks: source locator/path mode, Spec ID/design revision/design digest,
|
|
20
|
+
document digest, execution revision, writeback status, selected task IDs, repository
|
|
20
21
|
bindings/baseline status, and pending dependency evidence.
|
|
21
22
|
|
|
22
23
|
Mode inspection and configuration belongs to `ec-config`. If the user asks to change Approval,
|
|
@@ -32,3 +33,8 @@ When creating from a Canonical Spec, call `inspect-dev-spec`, display the comple
|
|
|
32
33
|
dependency selection, then call `select-dev-spec-scope` and `create-task-from-spec` only after
|
|
33
34
|
explicit user selection. Multiple selected Spec tasks still create one Harness task, while the
|
|
34
35
|
selector returns one deterministic consumption closure per selected repository.
|
|
36
|
+
Initialize missing shared execution before creation. Support `rebind-spec-source` only when the
|
|
37
|
+
new file matches schema + spec_id + design revision + design_sha256 and does not roll execution
|
|
38
|
+
revision backward. A pending writeback is repaired with `reconcile-spec-execution`, never by
|
|
39
|
+
editing the execution JSON block or starting a different writeback. A deterministic rejected
|
|
40
|
+
action is cleared with `status:error`; correct its input instead of replaying it.
|
|
@@ -0,0 +1,101 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ec-tdd-init
|
|
3
|
+
description: Initialize or refresh Java changed-line TDD coverage infrastructure before TDD can be enabled.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# ec-tdd-init — Java changed-line gate initialization
|
|
7
|
+
|
|
8
|
+
Communicate in the user's language. This skill owns TDD infrastructure readiness, not historical
|
|
9
|
+
test-debt cleanup. It must never bulk-generate tests for existing business code, require
|
|
10
|
+
repository-wide coverage, or modify production behavior merely to raise coverage.
|
|
11
|
+
|
|
12
|
+
## Non-circular ordering
|
|
13
|
+
|
|
14
|
+
The only legal order is:
|
|
15
|
+
|
|
16
|
+
```text
|
|
17
|
+
TDD off -> initialize infrastructure -> readiness ready -> user explicitly enables TDD
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
Run this skill as a dedicated code task with `type=tdd-init`. The state API always freezes that
|
|
21
|
+
task with `tdd_enabled=false`, even when a legacy project/session setting or a suspended task has
|
|
22
|
+
TDD enabled. Never offer "enable now and initialize later". Never enable TDD automatically after
|
|
23
|
+
initialization.
|
|
24
|
+
|
|
25
|
+
## Read-only preflight
|
|
26
|
+
|
|
27
|
+
First run:
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . check
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
If it returns `ready`, report the recorded build/CI contract and stop without creating a task.
|
|
34
|
+
The user may then use `ec-config` or `easy-coding config` to enable TDD.
|
|
35
|
+
|
|
36
|
+
If it returns `needs_init`, inspect only the infrastructure needed to form a confirmed plan:
|
|
37
|
+
|
|
38
|
+
- Maven/Gradle files and the existing JUnit runner;
|
|
39
|
+
- JaCoCo XML generation configuration;
|
|
40
|
+
- `.gitlab-ci.yml` and its repository-local include chain;
|
|
41
|
+
- TEST-stage job, JUnit/JaCoCo artifacts, and invocation of
|
|
42
|
+
`.easy-coding/tools/easy_coding_java_coverage.py` with task-supplied baseline/threshold values.
|
|
43
|
+
|
|
44
|
+
Do not measure current whole-project coverage. A project with no historical business tests may
|
|
45
|
+
still become ready when the test runner, JaCoCo reporting, and parameterized changed-line gate
|
|
46
|
+
are functional.
|
|
47
|
+
|
|
48
|
+
## Initialization task
|
|
49
|
+
|
|
50
|
+
After the user confirms the exact infrastructure scope, create one task and route it through the
|
|
51
|
+
ordinary workflow:
|
|
52
|
+
|
|
53
|
+
```bash
|
|
54
|
+
{{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py create-task \
|
|
55
|
+
--task-id <safe-unique-id> --type tdd-init \
|
|
56
|
+
--title "Initialize Java changed-line TDD infrastructure" \
|
|
57
|
+
--agent <agent-id> --session-file <P>
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
ANALYSIS may plan build files, GitLab CI files, common scripts, and the readiness receipt. It must
|
|
61
|
+
state `historical coverage required: no` and `coverage scope: changed production lines since each
|
|
62
|
+
future task baseline`. IMPLEMENT changes infrastructure only. It must not add tests whose sole
|
|
63
|
+
purpose is to cover unchanged production code.
|
|
64
|
+
|
|
65
|
+
The reusable GitLab job must consume a baseline SHA and threshold supplied for the future task;
|
|
66
|
+
do not hardcode the initialization commit or the default 90% threshold. The same Python coverage
|
|
67
|
+
tool must be usable locally and remotely. This generated job is remote automation infrastructure,
|
|
68
|
+
not a Harness task acceptance dependency: later tasks require local unit-test and coverage
|
|
69
|
+
evidence only, and never wait for a pipeline URL, job identity, or remote success status.
|
|
70
|
+
|
|
71
|
+
## Readiness receipt and verification
|
|
72
|
+
|
|
73
|
+
At the end of IMPLEMENT, after the infrastructure files are stable, record their fingerprints.
|
|
74
|
+
The receipt is part of the implementation and must exist before REVIEW so review/verification
|
|
75
|
+
fingerprints do not change after review. The recorder automatically includes the harness-managed
|
|
76
|
+
`.easy-coding/tools/easy_coding_java_coverage.py` fingerprint:
|
|
77
|
+
|
|
78
|
+
```bash
|
|
79
|
+
python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . record \
|
|
80
|
+
--build-file <pom.xml-or-build.gradle> [--build-file <included-build-file>]... \
|
|
81
|
+
--ci-file .gitlab-ci.yml [--ci-file <repository-local-include>]... \
|
|
82
|
+
--coverage-report <jacoco-xml-pattern> [--coverage-report <pattern>]... \
|
|
83
|
+
--gate-command "python3 .easy-coding/tools/easy_coding_java_coverage.py check --base \$EASY_CODING_TDD_BASE_SHA --threshold \$EASY_CODING_TDD_THRESHOLD" \
|
|
84
|
+
--agent <agent-id>
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
REVIEW includes the receipt and its declared infrastructure boundary. VERIFICATION runs the
|
|
88
|
+
frozen Workflow Mode's applicable build/test/CI syntax checks, then performs only the read-only
|
|
89
|
+
readiness check:
|
|
90
|
+
|
|
91
|
+
```bash
|
|
92
|
+
python3 .easy-coding/tools/easy_coding_tdd_readiness.py --cwd . check
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
The `VERIFICATION -> MEMORY` gate requires the final check to return `ready`. If any recorded
|
|
96
|
+
build or CI file changes after the receipt was created, readiness becomes `needs_init`; return to
|
|
97
|
+
IMPLEMENT, refresh the receipt, and repeat REVIEW before verifying again. Rerun this skill when
|
|
98
|
+
the same drift occurs after task completion.
|
|
99
|
+
|
|
100
|
+
After completion, tell the user that TDD remains off and provide the explicit project/session
|
|
101
|
+
enable route. Do not treat readiness as consent to enable it.
|
|
@@ -15,7 +15,8 @@ Read-only tasks never enter this stage. Obtain fresh fingerprints before running
|
|
|
15
15
|
|
|
16
16
|
- No completion claim without executed verification evidence.
|
|
17
17
|
- Evidence is reusable only while both returned fingerprints remain unchanged.
|
|
18
|
-
- Relevant code or config changes invalidate old evidence automatically
|
|
18
|
+
- Relevant code or config changes invalidate old evidence automatically unless the exact
|
|
19
|
+
post-verification code diff is explicitly accepted under the checkpoint protocol below.
|
|
19
20
|
- Failed or missing evidence never becomes acceptance because of approval mode.
|
|
20
21
|
|
|
21
22
|
## Verification depth
|
|
@@ -26,10 +27,12 @@ Read-only tasks never enter this stage. Obtain fresh fingerprints before running
|
|
|
26
27
|
- `strict`: run the project's full applicable lint, typecheck, test, and build gates.
|
|
27
28
|
|
|
28
29
|
These rules remain unchanged when frozen TDD is off: do not discover JaCoCo reports, run the
|
|
29
|
-
coverage tool, inspect GitLab, or add a coverage record.
|
|
30
|
+
coverage tool, inspect GitLab, or add a coverage record. The explicit `type=tdd-init` task is an
|
|
31
|
+
infrastructure exception: run its planned build/CI syntax checks and readiness tool, but do not
|
|
32
|
+
measure repository-wide coverage or append TDD coverage evidence for unchanged production code.
|
|
30
33
|
|
|
31
|
-
When frozen TDD is on, first run the planned Java unit command and generate JaCoCo XML,
|
|
32
|
-
the
|
|
34
|
+
When frozen TDD is on, first run the planned local Java unit command and generate JaCoCo XML,
|
|
35
|
+
then run the deterministic local acceptance gate:
|
|
33
36
|
|
|
34
37
|
```bash
|
|
35
38
|
python3 .easy-coding/tools/easy_coding_java_coverage.py check \
|
|
@@ -41,10 +44,10 @@ The tool measures covered added/modified production Java executable lines only.
|
|
|
41
44
|
comment, blank, import, and test-source lines are excluded by diff/JaCoCo intersection. Missing
|
|
42
45
|
or ambiguous source files and reports older than their modified source fail; zero modified
|
|
43
46
|
executable lines is explicit N/A. Always regenerate JaCoCo XML after the final source change.
|
|
44
|
-
Record CI as pending until the remote pipeline actually passes; local green is not remote green.
|
|
45
47
|
Never substitute `HEAD`, a mutable ref, project defaults, or current session settings for the
|
|
46
|
-
task-frozen baseline SHA and threshold.
|
|
47
|
-
|
|
48
|
+
task-frozen baseline SHA and threshold. `ec-tdd-init` still generates a GitLab job that can run
|
|
49
|
+
the same tool, but remote pipeline execution and status are outside Harness acceptance. Never
|
|
50
|
+
request an intermediate commit or push merely to obtain CI evidence.
|
|
48
51
|
|
|
49
52
|
The main Agent may run commands inline. Dispatch verifier sub-agents only when checks are
|
|
50
53
|
independent and parallel execution materially saves time or isolates specialist environments.
|
|
@@ -75,7 +78,10 @@ and a non-empty `not_applicable_reason`; it does not count as the required appli
|
|
|
75
78
|
check, and must not be represented by an invented successful command.
|
|
76
79
|
|
|
77
80
|
Record failures in `failures[]`. If any current-fingerprint record fails, return to IMPLEMENT;
|
|
78
|
-
do not append a later synthetic pass without rerunning the failed command.
|
|
81
|
+
do not append a later synthetic pass without rerunning the failed command. For a Canonical-backed
|
|
82
|
+
failure, append the local verify record first, then write the owning source task `blocked` with a
|
|
83
|
+
concise reference to that record. The repair transition automatically reopens blocked source tasks
|
|
84
|
+
only; unaffected implemented tasks retain their latest shared conclusion.
|
|
79
85
|
|
|
80
86
|
## Coverage and acceptance
|
|
81
87
|
|
|
@@ -84,23 +90,17 @@ For TDD coverage, copy the tool output into `coverage`: `baseline_sha`, `covered
|
|
|
84
90
|
`applicable:false` plus the tool's reason only for zero executable modified lines. A percentage
|
|
85
91
|
below the frozen threshold fails even when ordinary tests pass.
|
|
86
92
|
|
|
87
|
-
Append
|
|
88
|
-
|
|
89
|
-
|
|
93
|
+
Append one coverage record with `coverage_scope:"local"` per repository (and per Canonical
|
|
94
|
+
source task). The state gate also requires a passed local `check_type:"test"` record for the same
|
|
95
|
+
owner. The coverage record preserves the task-frozen baseline and threshold. Do not append or
|
|
96
|
+
wait for GitLab pipeline evidence; historical remote coverage records are ignored by acceptance
|
|
97
|
+
without modifying or deleting the stored records.
|
|
90
98
|
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
"job_name": "changed-line-coverage",
|
|
97
|
-
"status": "success"
|
|
98
|
-
}
|
|
99
|
-
}
|
|
100
|
-
```
|
|
101
|
-
|
|
102
|
-
Both records must preserve the same task-frozen baseline and threshold. A local-only result,
|
|
103
|
-
pending/failed pipeline, missing job identity, or synthetic remote pass cannot satisfy MEMORY.
|
|
99
|
+
For `type=tdd-init`, the infrastructure receipt must already have been recorded during IMPLEMENT
|
|
100
|
+
and reviewed with the rest of the implementation. Run only `easy_coding_tdd_readiness.py check`
|
|
101
|
+
here. If it reports drift, return to IMPLEMENT to refresh the receipt and repeat REVIEW; never
|
|
102
|
+
rewrite it inside VERIFICATION. The state gate requires `ready` before MEMORY. This does not
|
|
103
|
+
enable TDD; report the explicit `ec-config`/`easy-coding config` next step.
|
|
104
104
|
|
|
105
105
|
- Every must-test item has an executed check.
|
|
106
106
|
- Bug fixes include a regression test when project infrastructure exists.
|
|
@@ -110,6 +110,38 @@ pending/failed pipeline, missing job identity, or synthetic remote pass cannot s
|
|
|
110
110
|
user wait.
|
|
111
111
|
- A reported in-scope problem returns to IMPLEMENT; out-of-scope work becomes a separate task.
|
|
112
112
|
|
|
113
|
+
After the final green evidence is recorded, freeze the acceptance baseline before presenting the
|
|
114
|
+
result or applying the boundary:
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
{{PYTHON_CMD}} {{platform_config_dir}}/hooks/easy_coding_state.py verification-checkpoint \
|
|
118
|
+
--agent <agent-id> --session-file <P>
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Then request or auto-apply VERIFICATION -> MEMORY according to `approval_mode`. `auto` remains
|
|
122
|
+
automatic when the checkpoint is unchanged. If any in-scope code changed after the checkpoint,
|
|
123
|
+
the state API returns `action:"acceptance-drift"`, keeps the task in VERIFICATION, and includes
|
|
124
|
+
the exact unified patches or binary/mode-change descriptions plus a stable `diff_sha256`. This
|
|
125
|
+
exceptional drift pauses every approval mode, including `auto`; it does not permanently change
|
|
126
|
+
the configured mode.
|
|
127
|
+
|
|
128
|
+
Show the complete returned diff and ask whether to accept that exact digest. Do not re-enter
|
|
129
|
+
IMPLEMENT or rerun REVIEW merely because this drift exists. On acceptance, call
|
|
130
|
+
`confirm-transition --stage MEMORY --diff-sha256 <digest>` with exactly one policy:
|
|
131
|
+
|
|
132
|
+
- `carry-forward`: only when every changed hunk is confidently non-executable and existing
|
|
133
|
+
verification remains applicable;
|
|
134
|
+
- `targeted`: executable behavior changed; append passed current-fingerprint targeted verification
|
|
135
|
+
before confirming. Canonical tasks must cover every affected source task reported by the
|
|
136
|
+
acceptance record, without rerunning checks for unaffected source tasks;
|
|
137
|
+
- `waived`: the user explicitly accepts the stated unverified risk.
|
|
138
|
+
|
|
139
|
+
Include `--decision-summary` with the user's decision. A changed digest invalidates the pending
|
|
140
|
+
confirmation and must be shown again. Behavior config, execution plan, workflow, Canonical
|
|
141
|
+
design, or nested-repository metadata drift cannot use this shortcut; return to ANALYSIS or
|
|
142
|
+
IMPLEMENT as reported by the state API. The acceptance record bridges only the accepted
|
|
143
|
+
implementation fingerprints, so prior REVIEW evidence remains valid without a second REVIEW.
|
|
144
|
+
|
|
113
145
|
For Canonical-backed tasks, run each repository's commands from `task.repo_paths[repo_id]` and
|
|
114
146
|
cover every selected task's source test IDs. Report pending integration edges separately from
|
|
115
147
|
local green checks. They do not block local implementation evidence, but the state API blocks
|
|
@@ -120,6 +152,16 @@ tasks remain separate evidence records. In `strict`, every involved repository i
|
|
|
120
152
|
records all four check types; a repository-specific non-applicable record still needs its reason
|
|
121
153
|
and source ownership.
|
|
122
154
|
|
|
155
|
+
After implementation and local checks, each selected Canonical source task remains
|
|
156
|
+
`implemented`. Do not call `writeback-spec-task --status verified` from VERIFICATION. Applying
|
|
157
|
+
VERIFICATION -> MEMORY is the authoritative acceptance boundary: the state API writes each
|
|
158
|
+
still-implemented source task to `verified` through CAS/idempotent recoverable events with its
|
|
159
|
+
accepted test evidence and acceptance digest, then enters MEMORY only after every write is
|
|
160
|
+
confirmed. For `approve`/`guard`, that authority is the explicit boundary
|
|
161
|
+
confirmation; for `confirm`/`auto`, it is the standing approval-mode authorization when no new
|
|
162
|
+
drift exists. If writeback is interrupted, run `reconcile-spec-execution` before retrying the
|
|
163
|
+
transition. Remote CI remains outside this acceptance gate.
|
|
164
|
+
|
|
123
165
|
Record the exact integration edge only after its evidence exists:
|
|
124
166
|
|
|
125
167
|
```bash
|
|
@@ -131,5 +173,5 @@ Record the exact integration edge only after its evidence exists:
|
|
|
131
173
|
--agent <agent>
|
|
132
174
|
```
|
|
133
175
|
|
|
134
|
-
The state API rejects VERIFICATION -> MEMORY unless all evidence
|
|
135
|
-
|
|
176
|
+
The state API rejects VERIFICATION -> MEMORY unless all effective evidence is green and the
|
|
177
|
+
checkpoint is either unchanged or bound to an exact accepted diff.
|