@michelj/context-guard 0.4.3 → 0.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/Coordinator.md +88 -0
- package/Executor.md +53 -0
- package/README.md +89 -224
- package/README.zh-CN.md +89 -224
- package/SKILL.md +26 -684
- package/THIRD_PARTY_NOTICES.md +47 -0
- package/Tester.md +53 -0
- package/agents/openai.yaml +2 -2
- package/bin/build-runtime.mjs +96 -0
- package/bin/context-guard-skill.js +399 -78
- package/bin/postinstall.js +2 -2
- package/hooks.json +89 -13
- package/licenses/JSONParse-MIT.txt +24 -0
- package/licenses/Marked-MIT.txt +44 -0
- package/licenses/Portless-Apache-2.0.txt +201 -0
- package/package.json +35 -6
- package/prototype/LICENSES/Marked-MIT.txt +44 -0
- package/prototype/LICENSES/Ready-redistribution.txt +14 -0
- package/prototype/attachments.mjs +75 -0
- package/prototype/coordinator-markdown.mjs +283 -0
- package/prototype/coordinator-working-blot.mjs +124 -0
- package/prototype/vendor/marked.mjs +2189 -0
- package/prototype/workbench-app.js +5197 -0
- package/prototype/workbench-data.js +33 -0
- package/prototype/workbench-sync.mjs +898 -0
- package/prototype/workbench.css +1050 -0
- package/prototype/workbench.html +211 -0
- package/prototype/working-blot-atlas.png +0 -0
- package/references/agent-handoff.md +40 -0
- package/references/claude-runtime.md +120 -0
- package/references/cloud-sync-interface.md +66 -0
- package/references/design-current.md +14 -0
- package/references/map-mount.md +41 -0
- package/references/map-read.md +50 -0
- package/references/memory-definition.md +120 -0
- package/references/memory-filesystem-v2/Bug.en.md +162 -0
- package/references/memory-filesystem-v2/Bug.md +162 -0
- package/references/memory-filesystem-v2/Bug_Coordinater.md +8 -0
- package/references/memory-filesystem-v2/Bug_Executor.md +8 -0
- package/references/memory-filesystem-v2/Bug_Tester.md +7 -0
- package/references/memory-filesystem-v2/Idea.en.md +36 -0
- package/references/memory-filesystem-v2/Idea.md +36 -0
- package/references/memory-filesystem-v2/Node_Module_Index.en.md +88 -0
- package/references/memory-filesystem-v2/Node_Module_Index.md +88 -0
- package/references/memory-filesystem-v2/README.md +60 -0
- package/references/memory-filesystem-v2/Todo.en.md +137 -0
- package/references/memory-filesystem-v2/Todo.md +137 -0
- package/references/memory-filesystem-v2/Todo_Coordinater.md +7 -0
- package/references/memory-filesystem-v2/Todo_Executor.md +7 -0
- package/references/memory-filesystem-v2/Todo_Tester.md +7 -0
- package/references/named-workbench.md +124 -0
- package/references/plan-review.md +12 -0
- package/references/server-memory.md +276 -0
- package/references/test-check.md +7 -0
- package/references/user-reply.md +38 -0
- package/references/workbench-interface.md +531 -0
- package/roles.md +13 -0
- package/scripts/context_guard.py +1366 -7602
- package/scripts/context_guard_hook.py +1960 -711
- package/scripts/map_owns.py +699 -0
- package/scripts/shared/LICENSES/JSONParse-MIT.txt +24 -0
- package/scripts/shared/filesystem-v2.mjs +430 -0
- package/scripts/shared/io.mjs +117 -0
- package/scripts/shared/map-model.mjs +506 -0
- package/scripts/shared/memory-schema.mjs +13 -0
- package/scripts/shared/protocol-blobs.mjs +112 -0
- package/scripts/shared/protocol-map.mjs +146 -0
- package/scripts/shared/protocol-snapshots.mjs +84 -0
- package/scripts/shared/protocol-store.mjs +624 -0
- package/scripts/shared/protocol-workflow.mjs +226 -0
- package/scripts/shared/protocol.mjs +125 -0
- package/scripts/shared/vendor/jsonparse.cjs +413 -0
- package/scripts/workbench/access.mjs +496 -0
- package/scripts/workbench/attachments.mjs +92 -0
- package/scripts/workbench/browser-login.mjs +78 -0
- package/scripts/workbench/claude-runtime.mjs +372 -0
- package/scripts/workbench/cli.mjs +980 -0
- package/scripts/workbench/device-heartbeat.mjs +72 -0
- package/scripts/workbench/hook-status.mjs +38 -0
- package/scripts/workbench/inbox.mjs +155 -0
- package/scripts/workbench/journal.mjs +56 -0
- package/scripts/workbench/memory-merge.mjs +65 -0
- package/scripts/workbench/memory.mjs +252 -0
- package/scripts/workbench/named-proxy.mjs +108 -0
- package/scripts/workbench/named.mjs +152 -0
- package/scripts/workbench/portless-routes.mjs +51 -0
- package/scripts/workbench/project.mjs +327 -0
- package/scripts/workbench/projections.mjs +68 -0
- package/scripts/workbench/protocol-client.mjs +165 -0
- package/scripts/workbench/protocol-delivery.mjs +133 -0
- package/scripts/workbench/protocol-device.mjs +316 -0
- package/scripts/workbench/protocol-events.mjs +53 -0
- package/scripts/workbench/protocol-repository.mjs +58 -0
- package/scripts/workbench/reconcile.mjs +244 -0
- package/scripts/workbench/registry.mjs +111 -0
- package/scripts/workbench/runtime.mjs +54 -0
- package/scripts/workbench/server.mjs +1171 -0
- package/scripts/workbench/store.mjs +243 -0
- package/scripts/workbench/sync-coordinator.mjs +518 -0
- package/scripts/workbench/sync.mjs +86 -0
- package/references/context-template.md +0 -341
- package/references/feature-chain-methodology.md +0 -228
- package/references/register-template.md +0 -85
- package/references/task-case-template.md +0 -63
- package/tests/BC-20260618-063.sh +0 -116
- package/tests/BC-20260618-065.sh +0 -66
- package/tests/BC-20260626-080.sh +0 -48
- package/tests/BC-20260626-081.sh +0 -40
- package/tests/BC-20260626-082.sh +0 -32
- package/tests/BC-20260626-083.sh +0 -66
- package/tests/BC-20260627-084.sh +0 -74
- package/tests/BC-20260630-086.sh +0 -50
- package/tests/BC-20260630-087.sh +0 -103
- package/tests/BC-20260630-088.sh +0 -32
- package/tests/BC-20260630-089.sh +0 -63
- package/tests/BC-20260701-090.sh +0 -84
- package/tests/BC-20260702-096.sh +0 -48
- package/tests/BC-20260706-098.sh +0 -66
- package/tests/BC-20260707-099.sh +0 -47
- package/tests/BC-20260707-100.sh +0 -46
- package/tests/BC-20260707-101.sh +0 -47
- package/tests/BC-20260707-102.sh +0 -68
- package/tests/BC-20260707-103.sh +0 -59
- package/tests/BC-20260707-104.sh +0 -103
- package/tests/BC-20260707-105.sh +0 -109
- package/tests/BC-20260707-106.sh +0 -80
- package/tests/BC-20260707-107.sh +0 -74
- package/tests/BC-20260707-108.sh +0 -48
- package/tests/BC-20260707-109.sh +0 -56
- package/tests/BC-20260707-110.sh +0 -71
- package/tests/BC-20260707-111.sh +0 -70
- package/tests/BC-20260707-112.sh +0 -45
- package/tests/BC-20260707-113.sh +0 -73
- package/tests/BC-20260707-115.sh +0 -77
- package/tests/BC-20260707-116.sh +0 -77
- package/tests/BC-20260707-118.sh +0 -115
- package/tests/BC-20260707-119.sh +0 -47
- package/tests/BC-20260707-120.sh +0 -60
- package/tests/BC-20260707-121.sh +0 -66
- package/tests/BC-20260707-122.sh +0 -48
- package/tests/BC-20260707-123.sh +0 -43
- package/tests/BC-20260707-124.sh +0 -56
- package/tests/BC-20260707-125.sh +0 -64
- package/tests/BC-20260707-126.sh +0 -80
- package/tests/BC-20260707-127.sh +0 -88
- package/tests/BC-20260707-129.sh +0 -59
- package/tests/BC-20260707-130.sh +0 -69
- package/tests/BC-20260707-131.sh +0 -140
- package/tests/BC-20260707-132.sh +0 -150
- package/tests/BC-20260707-133.sh +0 -70
- package/tests/BC-20260708-136.sh +0 -210
- package/tests/BC-20260708-137.sh +0 -106
- package/tests/BC-20260708-138.sh +0 -168
- package/tests/BC-20260708-139.sh +0 -79
- package/tests/BC-20260709-002.sh +0 -63
- package/tests/BC-20260709-003.sh +0 -239
- package/tests/BC-20260709-006.sh +0 -76
- package/tests/BC-20260709-008.sh +0 -168
- package/tests/BC-20260710-001.sh +0 -61
- package/tests/BC-20260710-002.sh +0 -111
- package/tests/npm-install-smoke.sh +0 -53
|
@@ -1,341 +0,0 @@
|
|
|
1
|
-
# Context Folder Template
|
|
2
|
-
|
|
3
|
-
Use this structure for project context:
|
|
4
|
-
|
|
5
|
-
```text
|
|
6
|
-
<opened Codex project root>/
|
|
7
|
-
.codex/context/
|
|
8
|
-
├── index.md
|
|
9
|
-
├── user-messages.md
|
|
10
|
-
├── roadmap.md
|
|
11
|
-
├── bad-cases.md
|
|
12
|
-
├── preferences.json
|
|
13
|
-
├── private/
|
|
14
|
-
│ └── secrets.local.json
|
|
15
|
-
├── roadmap/
|
|
16
|
-
│ ├── roadmap.html
|
|
17
|
-
│ ├── roadmap-details.html
|
|
18
|
-
│ ├── roadmap.md
|
|
19
|
-
│ └── roadmap.json
|
|
20
|
-
├── tasks/
|
|
21
|
-
│ └── CTX-YYYYMMDD-short-slug/
|
|
22
|
-
│ ├── context.md
|
|
23
|
-
│ └── bad-cases.md
|
|
24
|
-
├── task-cases/
|
|
25
|
-
├── test-hub/
|
|
26
|
-
│ ├── registry.json
|
|
27
|
-
│ ├── feature-chains.json
|
|
28
|
-
│ ├── last-run.json
|
|
29
|
-
│ └── runs/
|
|
30
|
-
├── bad-case-tests/
|
|
31
|
-
└── archive/
|
|
32
|
-
```
|
|
33
|
-
|
|
34
|
-
The opened Codex project root is the local folder selected in Codex or the local workspace root for the current thread. Do not place this folder inside a skill installation directory, remote SSH path, chat/thread-specific folder, or temporary execution directory unless the user explicitly asks that location to own its own context.
|
|
35
|
-
|
|
36
|
-
The `private/` folder is local-only sensitive memory. It must be covered by `.codex/.gitignore`, use restrictive local permissions where possible, and never be projected into Roadmap HTML, agent-readable exports, README files, logs, or final answers.
|
|
37
|
-
|
|
38
|
-
## preferences.json
|
|
39
|
-
|
|
40
|
-
```json
|
|
41
|
-
{
|
|
42
|
-
"record_language": "unset",
|
|
43
|
-
"display_language": "auto",
|
|
44
|
-
"last_updated": "YYYY-MM-DD",
|
|
45
|
-
"note": "Set with: context_guard.py set-language --language <language>"
|
|
46
|
-
}
|
|
47
|
-
```
|
|
48
|
-
|
|
49
|
-
Ask the user for a context record language the first time `record_language` is `unset`, then update this file. Use that language for future source context records. Keep literal code identifiers, paths, commands, logs, API names, and exact error text unchanged. If the user changes language later, update this file and use the new language going forward; do not bulk-translate history unless asked.
|
|
50
|
-
|
|
51
|
-
## index.md
|
|
52
|
-
|
|
53
|
-
```md
|
|
54
|
-
# Context Index
|
|
55
|
-
|
|
56
|
-
This is a dynamic queue of active and recently parked folder context. Keep it short enough to scan in seconds.
|
|
57
|
-
|
|
58
|
-
## Quick Scan
|
|
59
|
-
|
|
60
|
-
- Current: CTX-YYYYMMDD-short-slug
|
|
61
|
-
- Latest roadmap node: NODE-YYYYMMDD-001
|
|
62
|
-
- Hot bad-case tags: #hot-ui, #flaky-test
|
|
63
|
-
- Resume candidate: CTX-YYYYMMDD-other-slug
|
|
64
|
-
|
|
65
|
-
Keep Quick Scan to these four lines unless the user explicitly asks for a fuller view.
|
|
66
|
-
|
|
67
|
-
## Current
|
|
68
|
-
|
|
69
|
-
- ID: CTX-YYYYMMDD-short-slug
|
|
70
|
-
- Title: short task title
|
|
71
|
-
- State: current
|
|
72
|
-
- Folder: `.codex/context/tasks/CTX-YYYYMMDD-short-slug/`
|
|
73
|
-
- Last updated: YYYY-MM-DD
|
|
74
|
-
- Summary: one sentence of the current direction
|
|
75
|
-
- Next step: the next useful action
|
|
76
|
-
|
|
77
|
-
## Parked / Resume Candidates
|
|
78
|
-
|
|
79
|
-
### CTX-YYYYMMDD-other-slug
|
|
80
|
-
|
|
81
|
-
- Title: short task title
|
|
82
|
-
- State: parked | resume-candidate
|
|
83
|
-
- Folder: `.codex/context/tasks/CTX-YYYYMMDD-other-slug/`
|
|
84
|
-
- Parked because: urgent bug, unrelated request, waiting on user, etc.
|
|
85
|
-
- Resume prompt: concise question to ask when the interruption is done
|
|
86
|
-
- Last updated: YYYY-MM-DD
|
|
87
|
-
|
|
88
|
-
## Archived
|
|
89
|
-
|
|
90
|
-
Keep only concise summaries here. Move detailed stale context to `.codex/context/archive/`.
|
|
91
|
-
```
|
|
92
|
-
|
|
93
|
-
## user-messages.md
|
|
94
|
-
|
|
95
|
-
```md
|
|
96
|
-
# User Message Memory
|
|
97
|
-
|
|
98
|
-
This file preserves concise user wording that future Codex turns may need. It is agent-readable context, not a public transcript.
|
|
99
|
-
|
|
100
|
-
## Recent User Signals
|
|
101
|
-
|
|
102
|
-
### YYYY-MM-DD HH:MM:SS
|
|
103
|
-
|
|
104
|
-
- Mode: verbatim | summary | redacted | ephemeral
|
|
105
|
-
- User message: short original user wording, concise summary for large inputs, or redacted message when secrets are present
|
|
106
|
-
- Secret pointer: USER-SECRET-YYYYMMDD-HHMMSS | ephemeral-not-stored | omitted when not relevant
|
|
107
|
-
- Use: Preserve this wording when deciding task direction, constraints, credentials, preferences, bad-case intake, or roadmap `User request` fields.
|
|
108
|
-
|
|
109
|
-
## Durable User Constraints
|
|
110
|
-
|
|
111
|
-
- User preference or rule that should affect future turns.
|
|
112
|
-
|
|
113
|
-
## Secret Pointers
|
|
114
|
-
|
|
115
|
-
- USER-SECRET-YYYYMMDD-HHMMSS: redacted purpose only; raw value lives in `.codex/context/private/secrets.local.json`.
|
|
116
|
-
```
|
|
117
|
-
|
|
118
|
-
Rules:
|
|
119
|
-
|
|
120
|
-
- Record short user prompts near-verbatim when they contain requirements, preferences, constraints, route changes, bad-case reports, server details, or credentials needed for the task.
|
|
121
|
-
- Summarize large pasted files, logs, attachments, or generated blobs instead of copying them wholesale.
|
|
122
|
-
- Never store raw secrets outside `.codex/context/private/` or a secure OS credential store.
|
|
123
|
-
- Do not persist OTP or short-lived one-time verification codes; record only `ephemeral-not-stored`.
|
|
124
|
-
- Promote durable user requirements into task context or roadmap `User request:` fields, then keep this file as a concise source of recent user wording.
|
|
125
|
-
|
|
126
|
-
## roadmap.md
|
|
127
|
-
|
|
128
|
-
```md
|
|
129
|
-
# Context Roadmap
|
|
130
|
-
|
|
131
|
-
This is the route map through the task. It may contain one mainline, forked side routes, or multiple parallel mainlines. Keep nodes concise. Do not record every tiny action or chat turn.
|
|
132
|
-
|
|
133
|
-
## Nodes
|
|
134
|
-
|
|
135
|
-
### NODE-YYYYMMDD-001: Short node title
|
|
136
|
-
|
|
137
|
-
- Date: YYYY-MM-DD
|
|
138
|
-
- Status: planned | active | done | superseded
|
|
139
|
-
- Level: major | checkpoint
|
|
140
|
-
- Branch: Main | short branch name
|
|
141
|
-
- Parent: NODE-YYYYMMDD-000 when this branch forks, otherwise none
|
|
142
|
-
- Task: `CTX-YYYYMMDD-short-slug`
|
|
143
|
-
- Display title: short human-facing card title; use clear language close to what the user cares about
|
|
144
|
-
- User request: concise summary of the user's actual request, using the user's wording as much as possible
|
|
145
|
-
- Progress summary: short human-facing current progress; omit if Outcome already reads naturally
|
|
146
|
-
- Method summary: short human-facing method; omit if Decision / reason already reads naturally
|
|
147
|
-
- Outcome: one-line result
|
|
148
|
-
- Decision / reason: why this node exists, one line
|
|
149
|
-
- Avoid going back: rejected path or lesson, only if it prevents backtracking
|
|
150
|
-
- Next: next useful node or action
|
|
151
|
-
- Linked bad cases: BC-YYYYMMDD-001, BC-YYYYMMDD-002
|
|
152
|
-
- Test chain: compact checkpoint evidence only; user-facing recurrence checks come from linked bad-case guards
|
|
153
|
-
- End-of-work self-check: changed behavior checked; for frontend/layout work include browser/plugin/screenshot evidence or the exact blocker
|
|
154
|
-
```
|
|
155
|
-
|
|
156
|
-
Use the `### NODE-...` section form as the canonical editable source. If a session accidentally records loose bullet blocks such as `- ID: NODE-...`, `- Title: ...`, `- Level: ...`, the renderer should still project them, but future edits should normalize them back into formal node sections.
|
|
157
|
-
|
|
158
|
-
## tasks/<task-id>/context.md
|
|
159
|
-
|
|
160
|
-
```md
|
|
161
|
-
# Task Context: short task title
|
|
162
|
-
|
|
163
|
-
- ID: CTX-YYYYMMDD-short-slug
|
|
164
|
-
- State: current | parked | resume-candidate | done | archived
|
|
165
|
-
- Created: YYYY-MM-DD
|
|
166
|
-
- Last updated: YYYY-MM-DD
|
|
167
|
-
|
|
168
|
-
## Objective
|
|
169
|
-
|
|
170
|
-
One sentence describing what the user is trying to accomplish.
|
|
171
|
-
|
|
172
|
-
## Key Points
|
|
173
|
-
|
|
174
|
-
- Important ideas, constraints, and decisions only.
|
|
175
|
-
- Rejected approaches only when they prevent repeating a wrong route.
|
|
176
|
-
- Product, design, architecture, or implementation notes only when needed to resume.
|
|
177
|
-
|
|
178
|
-
## Open Questions
|
|
179
|
-
|
|
180
|
-
- Questions that need user input or future investigation.
|
|
181
|
-
|
|
182
|
-
## Files / Areas
|
|
183
|
-
|
|
184
|
-
- Relevant files, modules, commands, screenshots, or external references.
|
|
185
|
-
|
|
186
|
-
## Bad Cases
|
|
187
|
-
|
|
188
|
-
- Link to shared `.codex/context/bad-cases.md` entries or task-local `bad-cases.md`.
|
|
189
|
-
|
|
190
|
-
## Roadmap Nodes
|
|
191
|
-
|
|
192
|
-
- Link to `NODE-...` entries in `.codex/context/roadmap.md`.
|
|
193
|
-
|
|
194
|
-
## Verification / Self-Check
|
|
195
|
-
|
|
196
|
-
- Behavior checked before final answer.
|
|
197
|
-
- Frontend/layout artifacts opened with a browser/plugin or inspected via screenshot when possible.
|
|
198
|
-
- If visual inspection was blocked, record the blocker and residual risk.
|
|
199
|
-
|
|
200
|
-
## Next Step
|
|
201
|
-
|
|
202
|
-
The smallest useful action to resume this task.
|
|
203
|
-
```
|
|
204
|
-
|
|
205
|
-
## task-cases/<task-case-id>.md
|
|
206
|
-
|
|
207
|
-
Use task cases for realistic multi-step verification flows. They should catch bugs by simulating a real task, not by testing one isolated bug at a time. Task-case design is human-owned: Codex may draft and structure a proposal, but durable task cases must stay `proposed` until the user confirms them.
|
|
208
|
-
|
|
209
|
-
```md
|
|
210
|
-
# Task Case: short realistic workflow title
|
|
211
|
-
|
|
212
|
-
- ID: TC-YYYYMMDD-short-slug
|
|
213
|
-
- Status: proposed | approved | active | stable | deferred | obsolete
|
|
214
|
-
- Route/task: `CTX-...` or branch name
|
|
215
|
-
- Scope: feature, service, UI flow, agent workflow, or subsystem
|
|
216
|
-
- Last checked: YYYY-MM-DD
|
|
217
|
-
- Design confirmation: pending | user-approved YYYY-MM-DD
|
|
218
|
-
- Run policy: every-dev-completion | relevant-only | manual | release-only | goal-final | disabled-with-reason | user-defined cadence
|
|
219
|
-
- Automation entry: native command | script path | prompt/manual runner | none
|
|
220
|
-
- Artifact policy: cleanup-on-pass | preserve-on-fail | manual-preserve
|
|
221
|
-
- Linked roadmap nodes: NODE-...
|
|
222
|
-
- Linked bad cases: BC-..., BC-...
|
|
223
|
-
- Entry command/prompt: command, prompt, manual setup, or fixture
|
|
224
|
-
- Not covered: explicit exclusions to avoid fake confidence
|
|
225
|
-
- Stop condition: what means the workflow is complete
|
|
226
|
-
- Cleanup: required cleanup or none
|
|
227
|
-
|
|
228
|
-
## Phases
|
|
229
|
-
|
|
230
|
-
### Phase 1: setup or trigger
|
|
231
|
-
|
|
232
|
-
- Action: one realistic action
|
|
233
|
-
- Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
|
|
234
|
-
- Covers bad cases: BC-...
|
|
235
|
-
- Failure localization: what this phase failure usually means
|
|
236
|
-
- Log note: what the script/agent should record
|
|
237
|
-
|
|
238
|
-
### Phase 2: transition or recovery
|
|
239
|
-
|
|
240
|
-
- Action: one realistic action
|
|
241
|
-
- Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
|
|
242
|
-
- Covers bad cases: BC-...
|
|
243
|
-
- Failure localization: what this phase failure usually means
|
|
244
|
-
- Log note: what the script/agent should record
|
|
245
|
-
|
|
246
|
-
## Result Log
|
|
247
|
-
|
|
248
|
-
- YYYY-MM-DD: pass/fail, failed phase/checkpoint if any, evidence path or command output summary
|
|
249
|
-
```
|
|
250
|
-
|
|
251
|
-
## Maintenance Rules
|
|
252
|
-
|
|
253
|
-
- Keep `index.md` small and useful, not exhaustive.
|
|
254
|
-
- Keep `roadmap.md` as the route map. It should show progress as nodes, not a raw transcript.
|
|
255
|
-
- Use `Level: major` for significant milestones shown as main route cards; use `Level: checkpoint` for minor progress that should live in details.
|
|
256
|
-
- Use `Branch:` for forked or parallel routes. Missing `Branch:` means `Main`; use `Parent:` to point to the node where a branch forked.
|
|
257
|
-
- In the human overview, visible card numbers should be consecutive per route group after checkpoint filtering, not source node numbers with gaps.
|
|
258
|
-
- If the human overview has multiple route groups, show all route lines together with parent/fork markers; keep tests and bad cases in clicked node details, Test Hub, and agent-readable exports by default.
|
|
259
|
-
- Each node should be concise enough for Codex to scan quickly: outcome, decision, next step, linked bad cases.
|
|
260
|
-
- Link nodes to bad cases and test-chain notes instead of duplicating full details.
|
|
261
|
-
- Treat human-facing test coverage as human-designed bad-case recurrence detection. Single-route overview hides the test lane; multi-route compact test routes should be generated only from user-approved tests with explicit `Run policy`, approved task-case checkpoints, or approved test registry entries, not from ordinary linked bad-case guards or roadmap node checkpoint logs.
|
|
262
|
-
- In branch or multi-route views, align each visible test item to the roadmap node whose approved bad-case test or approved task-case checkpoint it covers. If a route has no approved tests, do not show a test route for it. Empty test slots should be subtle timeline placeholders only when the route has at least one approved test elsewhere.
|
|
263
|
-
- Prefer a task-oriented case in `.codex/context/task-cases/` when realistic workflow phases matter more than isolated bug checks. Bad-case guards should often point to a task-case checkpoint that covers them.
|
|
264
|
-
- Prefer a feature chain in `.codex/context/test-hub/feature-chains.json` when one user-visible feature or workflow can cover several bad cases through checkpoint coverage. Attach new bad cases to existing chain checkpoints before proposing a new chain.
|
|
265
|
-
- Before proposing a new feature chain for a bad case, run `context_guard.py feature-chain-plan --root <project> --query "<bad case or feature text>"`; this is a read-only intake to either review an existing chain or propose a `feature-chain-propose` skeleton with checkpoint coverage state. Do not use `feature-chain-add` as the planner skeleton.
|
|
266
|
-
- Before approving automation, or when proposed feature chains sound similar, run `context_guard.py feature-chain-overlap --root <project>`; this is a read-only duplicate-chain audit to help merge or extend an existing workflow instead of creating another always-run test.
|
|
267
|
-
- `feature-chain-add` creates proposed chains by default. After the user confirms the business flow and test design, use `feature-chain-approve --chain-id <id> --command-text "<approved command>"` to promote the same chain. Do not hand-edit `feature-chains.json` or create a duplicate approved chain to bypass the approval gate.
|
|
268
|
-
- If the user says an approved feature chain should not run every time, use `feature-chain-set-policy --chain-id <id> --run-policy <policy> --reason <reason>` to update the same chain. Do not delete or duplicate feature chains just to change cadence.
|
|
269
|
-
- After editing feature chains, run `context_guard.py validate-feature-chains --root <project>` to catch missing entries, exits, approved automation, checkpoint checks, linked bad-case coverage, and artifact policy issues. This validates structure only; it does not approve business test design.
|
|
270
|
-
- Feature-chain commands may emit `CG_CHECKPOINT:<checkpoint title or id>:PASS` or `CG_CHECKPOINT:<checkpoint title or id>:FAIL:<short reason>`; Test Hub treats any `FAIL` marker as a failed chain and reports that checkpoint. Marker names must match registered checkpoint titles or ids; unknown markers are test-chain failures.
|
|
271
|
-
- Task-case scripts or agents should log the phase/checkpoint that failed, so Codex can locate the broken workflow step without re-debugging the whole task.
|
|
272
|
-
- Before writing any new durable task-case script or active task case, ask the user to confirm with only the business path: from what state to what state, the main task, and the major risk. Keep technical phases/checkpoints/logs inside the task-case file, not in the confirmation prompt. If confirmation is unavailable, keep the case `proposed` and avoid broad new scripts.
|
|
273
|
-
- When the user explicitly asks to create, write, generate, design, or add a test/test task/task case, start the user-visible response with `测试创建识别:...` or the folder-language equivalent, then summarize the test target from what state to what state and the main risk it catches.
|
|
274
|
-
- When the user creates or approves a test, register it with `Run policy: every-dev-completion` by default. At the end of every development turn, run all approved tests with that policy or record the exact blocker.
|
|
275
|
-
- Use `.codex/context/test-hub/registry.json` as the explicit Test Hub registry for user-approved automated tests. Do not populate it from ordinary bad-case guards or roadmap node `Test chain:` notes.
|
|
276
|
-
- Keep Test Hub as a simple control layer: registry, `dev-complete`, `last-run.json`, and lightweight commands to list, enable, disable, change policy, or remove registry tests.
|
|
277
|
-
- At development completion, prefer `context_guard.py dev-complete --root <project>` so the hub runs the approved always-run set, handles parallel workers when safe, cleans success artifacts, and preserves failure evidence under `.codex/context/test-hub/runs/`.
|
|
278
|
-
- After approval, automate a test when it can be safely scripted or run as a native command. Future Codex turns should execute the registered entry with minimal reinterpretation.
|
|
279
|
-
- Automated tests should clean temporary files after full success and preserve concise diagnostic artifacts on failure.
|
|
280
|
-
- Failed approved tests become a bad-case analysis loop: inspect the preserved evidence, fix the in-scope cause, rerun the same approved test, and stop only after pass or a non-actionable blocker.
|
|
281
|
-
- If blocked by credentials, unavailable external service, permissions, hardware/resource limits, network, destructive-risk confirmation, or user-only judgment, ask or warn the user with the exact blocker and evidence path.
|
|
282
|
-
- Change a test to `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or another cadence only when the user explicitly asks. Record the user's reason beside the policy.
|
|
283
|
-
- During goal mode, task cases should act as phase gates: select an approved case or propose one for confirmation, log phase progress during continuations, and run the smallest human-approved path before claiming the goal complete.
|
|
284
|
-
- Keep multilingual display as an HTML projection concern; do not duplicate source context by language. When supported, localize human-facing record titles, summaries, bad-case labels, and test-chain snippets in the projection.
|
|
285
|
-
- Keep source records in the configured `.codex/context/preferences.json` record language. The HTML roadmap should follow that preference and should not show a visible language selector by default.
|
|
286
|
-
- During goal mode or long-running autonomous work, keep the active goal aligned to the current task, add compact goal checkpoints during meaningful phase changes, and record bad cases as soon as they appear.
|
|
287
|
-
- Treat `.codex/context/index.md`, `.codex/context/roadmap.md`, `.codex/context/bad-cases.md`, and task context files as the source of truth.
|
|
288
|
-
- Treat `.codex/context/user-messages.md` as the source for recent user wording and durable user signals. Use it to populate roadmap `User request:` fields and avoid asking the user to repeat short instructions.
|
|
289
|
-
- Treat `.codex/context/private/` as local-only sensitive memory. Never expose raw secrets in human-facing HTML or git-tracked context.
|
|
290
|
-
- Treat `.codex/context/roadmap/roadmap.html` as a human-facing view only. Codex should not use it for context intake or bad-case management.
|
|
291
|
-
- Treat `.codex/context/roadmap/roadmap.md` and `.codex/context/roadmap/roadmap.json` as stable agent-readable exports for quick scanning, route lookup, bad-case lookup, and recurrence-guard lookup, not as primary editable sources.
|
|
292
|
-
- Keep `NODE-...`, `BC-...`, and `CTX-...` IDs in source files for linking, but hide them in the default human-facing HTML. Show short natural-language node and bad-case labels instead.
|
|
293
|
-
- In human-facing HTML, prefer color, symbols, and compact visual markers over labels like `Status:`, `Nodes:`, `Frequency:`, or fallback text such as `untagged`.
|
|
294
|
-
- Show meaningful `#tags` as compact colored chips with small emoji cues in human-facing HTML. Limit overview tags; show full tags on the detail page; omit the tag row when no tags exist.
|
|
295
|
-
- A sharp task direction change should park the current task before starting a new one.
|
|
296
|
-
- If the user explicitly says a task is a branch/side route/fork/支线/分支, run or emulate `scripts/context_guard.py create-branch-task --title <task title> --branch <branch name> --parent-node <parent NODE id>` before implementation so the task folder, current index entry, and `Branch:`/`Parent:` roadmap node all exist.
|
|
297
|
-
- If work significantly drifts from the mainline architecture without an explicit branch request, ask whether to create a branch before silently continuing.
|
|
298
|
-
- When an interruption finishes, ask whether to resume the most relevant parked task.
|
|
299
|
-
- Do not let parked items grow endlessly. Mark stale items `archived` and compress them to a short summary.
|
|
300
|
-
- Do not delete unresolved user intent unless the user explicitly discards it.
|
|
301
|
-
- Use `scripts/context_guard.py show-roadmap` to generate and display the stable human-friendly overview at `.codex/context/roadmap/roadmap.html`, with details at `.codex/context/roadmap/roadmap-details.html`, agent-readable Markdown at `.codex/context/roadmap/roadmap.md`, and structured lookup at `.codex/context/roadmap/roadmap.json`. Use `export-roadmap --format md` only for Markdown-only export.
|
|
302
|
-
- Do not accumulate timestamped HTML roadmap files. Showing the roadmap overwrites the same stable HTML files.
|
|
303
|
-
- With one route group, the HTML roadmap overview should show only the main route cards. Keep linked bad cases and recurrence checks in clicked node details, source context, and agent-readable exports.
|
|
304
|
-
- In one-route HTML, do not render bad-case/test-chain lanes or a left lane-label column. The route board should size to real content, and main route summaries should remain readable rather than being clipped after a very short fragment.
|
|
305
|
-
- With multiple route groups, the overview should show all route lines as a branch map. If a route has user-approved tests, also show a compact node-aligned test route under that route line. Route selection may affect details, but the default view should not invent tests from ordinary bad-case context.
|
|
306
|
-
- Parent/fork markers should appear only on side routes whose parent node is outside that route. Main route should not show a fork marker just because a later main node references an earlier main node.
|
|
307
|
-
- Side routes should visually start near their parent node's visible position on the parent route, not all from the first column.
|
|
308
|
-
- Branch route labels, parent chips, and checkpoint text should sit near the branch's first visible card by reusing the same spacer/grid coordinate as the branch cards.
|
|
309
|
-
- Branch overview should use one shared horizontal route canvas. Route alignment should use grid spacer columns, not padding that shifts or clips the whole route section.
|
|
310
|
-
- Branch connector lines should use the same offset coordinate as the route's spacer columns, not a fixed left-edge position.
|
|
311
|
-
- Branch connector endpoints should be anchored to the status dots inside the source and target node cards; do not draw connector lines from the whole route section or unrelated card edges.
|
|
312
|
-
- Route progression connectors should be card-to-card through card gaps; branch connectors should be dot-to-dot through an empty branch corridor and must not cross node cards or text.
|
|
313
|
-
- Side routes may drift right from exact column alignment when that creates a cleaner non-crossing branch path.
|
|
314
|
-
- Connector layers should render behind route cards so cards mask any line segment that would otherwise pass over content.
|
|
315
|
-
- Hide heavy native horizontal scrollbar chrome in the roadmap overview while preserving horizontal scroll interaction.
|
|
316
|
-
- Human-facing node detail cards should show only one concise summary sentence, linked bad cases, and linked bad-case recurrence tests. Do not show a standalone status dot under the node detail title. Keep route, parent, decision, avoid-going-back, and next-step source fields in agent-readable context, not in the human detail card.
|
|
317
|
-
- Human-facing bad-case details should localize phenomenon, trigger, root cause, fix, and guard prose to the folder language preference while preserving technical identifiers, commands, and paths.
|
|
318
|
-
- Route color should encode branch depth: main route green, first-level branch cool cyan/teal, deeper branch levels progressively colder toward blue and indigo.
|
|
319
|
-
- Before finalizing frontend, roadmap HTML, or visual layout work, open or render the artifact with an available browser/plugin or screenshot path and inspect for obvious visual bugs. Do not rely only on string assertions for layout changes.
|
|
320
|
-
- Treat Stop hook output as a completion reliability gate. Before claiming fixed/done/passing, record real verification evidence for the changed artifact or workflow and rerun relevant bad-case guards.
|
|
321
|
-
- For UI/browser/binding/frontend work, verify the original user-visible symptom, not only build success or process restart.
|
|
322
|
-
- User-facing projected text should follow the folder language preference; avoid untranslated English prose in Chinese overview output except for intentional technical strings.
|
|
323
|
-
- For a single route group, do not show lane titles or a left label column; the overview is only the main route.
|
|
324
|
-
- Keep overview cards sparse. Put full Outcome, Decision, Next, and guard details in same-file detail anchors and the stable `roadmap-details.html` sidecar.
|
|
325
|
-
- In multi-route branch overview, route cards should read as a compact map skeleton: number, title, date/status cue, and no visible outcome paragraph. Keep route summaries in details and source context.
|
|
326
|
-
- Default overview links should target same-file `#node-*` and `#case-*` anchors, not `roadmap-details.html#...`, to avoid `file://` access-denied navigation.
|
|
327
|
-
|
|
328
|
-
## Pruning Rules
|
|
329
|
-
|
|
330
|
-
- Do not record normal implementation chatter.
|
|
331
|
-
- Do not drop short user messages that contain requirements, constraints, preferences, credentials, route changes, or bad-case reports.
|
|
332
|
-
- Do not paste huge files or logs into user-message memory; summarize them.
|
|
333
|
-
- Do not store raw secrets in public context files, roadmap HTML, exports, logs, or final answers.
|
|
334
|
-
- Do not record every command; record only commands that prove a checkpoint or guard a bad case.
|
|
335
|
-
- Do not let roadmap node `Test chain:` history replace bad-case recurrence guards in user-facing roadmap output.
|
|
336
|
-
- Do not split a real workflow into many unrelated bug-level tests when one task case with checkpoints would reveal the failure location more clearly.
|
|
337
|
-
- Do not silently enter test creation. If the user explicitly asks to create a test, acknowledge the test-creation intake first so the user can see the skill activated.
|
|
338
|
-
- Do not wait until goal completion to record important roadmap progress or bad cases.
|
|
339
|
-
- Merge tiny adjacent updates into one roadmap node.
|
|
340
|
-
- Archive stale parked tasks as a one-sentence summary.
|
|
341
|
-
- If a reader cannot use a detail to resume, decide, verify, or avoid recurrence, remove it.
|
|
@@ -1,228 +0,0 @@
|
|
|
1
|
-
# Feature Chain Methodology
|
|
2
|
-
|
|
3
|
-
Use feature chains when the goal is to prevent fixed bad cases from reappearing without creating one durable test per bad case.
|
|
4
|
-
|
|
5
|
-
## Core Idea
|
|
6
|
-
|
|
7
|
-
The durable test unit is a user-visible feature or workflow. A bad case is coverage attached to one checkpoint inside that workflow.
|
|
8
|
-
|
|
9
|
-
```text
|
|
10
|
-
Feature chain
|
|
11
|
-
Entry: the real trigger users or Codex will perform
|
|
12
|
-
Checkpoint 1: expected intermediate state
|
|
13
|
-
Covers: BC-...
|
|
14
|
-
Checkpoint 2: expected transition or output
|
|
15
|
-
Covers: BC-..., BC-...
|
|
16
|
-
Exit check: strict final green condition
|
|
17
|
-
```
|
|
18
|
-
|
|
19
|
-
## Creation Rule
|
|
20
|
-
|
|
21
|
-
When a bad case appears:
|
|
22
|
-
|
|
23
|
-
1. Identify the feature entry that can reproduce or guard the symptom.
|
|
24
|
-
2. Search existing feature chains for the same entry, workflow, component, route, or service. Use the read-only planning helper first:
|
|
25
|
-
|
|
26
|
-
```bash
|
|
27
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-plan --root <project> --query "<bad case or feature text>"
|
|
28
|
-
```
|
|
29
|
-
|
|
30
|
-
The query can be natural language or a `BC-...` ID. The planner is read-only: it either says `action: review-existing-chain` with match evidence and an after-confirmation attach-command skeleton, or `action: propose-new-chain` with a compact user-confirmation prompt and a `feature-chain-propose` command skeleton. The proposed skeleton must contain a checkpoint and explicit coverage state: use `--bad-cases` for a real bad-case seed, or `--coverage-pending-reason` for a user-described test target that has no concrete bad case yet. If the input is only a `BC-...` ID, the prompt should describe the bad case by title or summary instead of asking the user to confirm an opaque ID. It must not create a feature chain, attach a bad case, or approve automation; run the skeleton only after user confirmation.
|
|
31
|
-
|
|
32
|
-
3. Use `feature-chain-suggest` when you only need raw candidate chains/checkpoints. If only a bad-case ID is available, the helper expands it from `bad-cases.md` first so matching uses the case title, summary, phenomenon, trigger, cause, and tags rather than the opaque ID alone.
|
|
33
|
-
4. Use `feature-chain-coverage` when you need the whole register view. It should show covered cases, unassigned candidates, and possible existing chains for visible candidates without changing the registry. Strong suggestions should include short match evidence terms; if the evidence is weak or unreadable, treat the suggestion as planning noise rather than coverage.
|
|
34
|
-
5. Use `feature-chain-candidates` when the unassigned list is too large. It groups unassigned bad cases by shared feature tags, prefers more specific tag combinations over broad single tags, suppresses repeated groups that add little new coverage, and proposes a small set of candidate feature chains. Treat `new coverage` as the key signal: a candidate with high total count but low new coverage may be a subcase of an earlier chain. It must not create, attach, or approve anything; it only helps choose which user-visible flow is worth designing.
|
|
35
|
-
6. Use `feature-chain-overlap` before approving automation, or whenever several proposed chains sound similar. It is read-only and flags pairs that likely describe the same workflow. If it reports overlap, merge the intent or extend one chain before creating another always-run test.
|
|
36
|
-
7. If a chain exists and the match is semantically correct, attach the bad case to the closest checkpoint and tighten that checkpoint.
|
|
37
|
-
8. If no chain exists, propose a new chain in one short business-facing sentence and wait for user confirmation before approving or automating it.
|
|
38
|
-
|
|
39
|
-
For natural-language test requests, keep the user's business intent but remove the request wrapper before writing the confirmation prompt. For example, `写一个测试,检验每次开发完成后 Markdown 编辑器里的单行、多行和矩阵公式都能正常渲染` should become a compact subject such as `Markdown 编辑器里的单行、多行和矩阵公式能正常渲染`, not a verbatim copy of the whole chat sentence. This keeps the prompt useful for human confirmation while avoiding agent-invented workflow details.
|
|
40
|
-
|
|
41
|
-
When the user already gives a workflow shape, preserve it. A request like `创建一个测试任务:从编辑器输入 Markdown 到预览正确渲染,主要验证公式渲染回归` should be confirmed as `从「编辑器输入 Markdown」到「预览正确渲染」,主要验证「公式渲染回归」`. Do not replace explicit entry/exit/risk wording with generic "相关入口到正确结果" language.
|
|
42
|
-
|
|
43
|
-
For this explicit shape, `feature-chain-plan` may prefill the after-confirmation `feature-chain-propose` skeleton with the stated entry and exit check, and may print the stated risk as a suggested checkpoint. This is still only a confirmation aid: it must not create the chain, approve automation, or invent missing checkpoint details before the user confirms the business flow.
|
|
44
|
-
|
|
45
|
-
CLI rule: `feature-chain-add` creates `status: proposed` by default. Do not treat this as an approved test. `feature-chain-add --test-status approved` is not allowed for `every-dev-completion` chains because it skips the user confirmation and approval dry-run gates. Use `feature-chain-approve` on the same proposed chain instead.
|
|
46
|
-
|
|
47
|
-
After the user confirms a candidate flow, use `feature-chain-propose` when you need to record a safe draft with seed bad-case coverage:
|
|
48
|
-
|
|
49
|
-
```bash
|
|
50
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-propose \
|
|
51
|
-
--root <project> \
|
|
52
|
-
--title "<confirmed feature title>" \
|
|
53
|
-
--entry "<confirmed user-visible entry>" \
|
|
54
|
-
--exit-check "<confirmed strict final green condition>" \
|
|
55
|
-
--node-title "<confirmed checkpoint>" \
|
|
56
|
-
--bad-cases "BC-..., BC-..." \
|
|
57
|
-
--check "<checkpoint recurrence check>"
|
|
58
|
-
```
|
|
59
|
-
|
|
60
|
-
This command creates only a `proposed` chain. It records the user-confirmed shape and seed bad cases, but it does not add an executable command and must not enter `dev-complete` until `feature-chain-approve` is run with user-approved automation.
|
|
61
|
-
|
|
62
|
-
If the user confirms the feature flow before any concrete bad case exists, keep the same command but replace `--bad-cases ...` with `--coverage-pending-reason "<why there is no linked bad case yet>"`. This records the draft so it is not lost, but it is not recurrence coverage and cannot be approved for `every-dev-completion` until a real bad case is attached to a checkpoint.
|
|
63
|
-
|
|
64
|
-
When a later bad case appears, run `feature-chain-plan` first. If it points to a coverage-pending chain/checkpoint, attach the bad case there with `feature-chain-attach-bc` and tighten the checkpoint text. The attach step should clear the pending-coverage note, because the checkpoint now has real bad-case coverage.
|
|
65
|
-
|
|
66
|
-
The expected lifecycle is:
|
|
67
|
-
|
|
68
|
-
1. `feature-chain-plan` turns user wording or a bad-case ID into a read-only confirmation prompt.
|
|
69
|
-
2. After user confirmation, `feature-chain-propose` records a non-executable draft with either linked bad cases or a coverage-pending reason.
|
|
70
|
-
3. Later bad cases are routed through `feature-chain-plan` and attached to the nearest existing checkpoint when semantically correct.
|
|
71
|
-
4. `feature-chain-summary` gives the fast coverage map; `feature-chain-overlap` checks duplicate workflow coverage before approval.
|
|
72
|
-
5. `feature-chain-approve` is the only path into the always-run set and must pass the checkpoint dry run.
|
|
73
|
-
6. `dev-complete` runs the approved chain with structured checkpoint markers and cleans success artifacts.
|
|
74
|
-
|
|
75
|
-
This lifecycle is the core experiment: fewer feature chains should cover more bad-case recurrence checks without Codex inventing a broad test suite.
|
|
76
|
-
|
|
77
|
-
Multi-project trials are the sanity check for this method. These trials must start from fresh sandbox projects instead of reusing an existing project, existing context folder, or previous test registry; otherwise the result may only prove that old context happened to work. The sandbox themes should also be genuinely different, such as a life utility, a creative tool, and a game or interaction, not merely three variants of the same engineering workflow. In small linear flows, one feature chain with two to four checkpoints can cover about three related bad cases and localize the failed phase. Treat planner checkpoint suggestions as hints, not business truth: Codex or the user must attach each bad case to the real phase where it can recur. When the workflow has queues, retries, multiple workers, recovery branches, or cross-process cleanup, upgrade the design to a task case instead of stretching a simple feature chain.
|
|
78
|
-
|
|
79
|
-
A single-chain trial is not enough to validate Context Guard itself. A system-level regression should include at least two independent approved feature chains in one fresh project, then prove that `dev-complete` runs both, reports one chain failure without hiding the other chain's pass result, preserves the failing evidence, and returns to all-pass after the same chain is fixed. This checks the Test Hub orchestration layer rather than only the lifecycle of one feature chain.
|
|
80
|
-
|
|
81
|
-
Before approving automation, use a dry run when the proposed command or checkpoint markers need validation:
|
|
82
|
-
|
|
83
|
-
```bash
|
|
84
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-dry-run \
|
|
85
|
-
--root <project> \
|
|
86
|
-
--chain-id FC-YYYYMMDD-001 \
|
|
87
|
-
--command-text "<candidate command>"
|
|
88
|
-
```
|
|
89
|
-
|
|
90
|
-
Dry run executes the candidate command against the proposed chain's registered checkpoints, reports missing/failed/unknown checkpoint markers, cleans success artifacts, and preserves failure evidence under `.codex/context/test-hub/dry-runs/`. It does not approve the chain, does not write the command into `feature-chains.json`, and does not add the chain to `dev-complete`.
|
|
91
|
-
|
|
92
|
-
Dry-run evidence paths must be unique per run. Fast repeated or parallel dry runs must not reuse or overwrite a previous failure directory, because the preserved evidence is what lets Codex locate the failed checkpoint without reinterpreting the whole task.
|
|
93
|
-
|
|
94
|
-
## Approval Rule
|
|
95
|
-
|
|
96
|
-
After the user confirms the feature flow and test design, promote the existing proposed chain with:
|
|
97
|
-
|
|
98
|
-
```bash
|
|
99
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-approve \
|
|
100
|
-
--root <project> \
|
|
101
|
-
--chain-id FC-YYYYMMDD-001 \
|
|
102
|
-
--command-text "<approved command>"
|
|
103
|
-
```
|
|
104
|
-
|
|
105
|
-
Approval is a safety gate and the only supported path from proposed feature-chain automation to `every-dev-completion`. It refuses a chain that has no checkpoint, no checkpoint check text, no linked bad-case coverage, or no automated command when the run policy is `every-dev-completion`. For `every-dev-completion` automation, approval must also run a dry-run preflight before mutating the registry. If the command misses a required checkpoint marker, emits an unknown marker, emits a `FAIL` marker, times out, or hits a blocker, approval fails and the chain stays `proposed`. Do not bypass this by hand-editing `feature-chains.json`, using `feature-chain-add --test-status approved`, or creating a second approved chain.
|
|
106
|
-
|
|
107
|
-
## Policy Rule
|
|
108
|
-
|
|
109
|
-
After approval, the user's cadence still wins. If the user says a feature chain should not run after every development turn, update the existing chain instead of deleting or duplicating it:
|
|
110
|
-
|
|
111
|
-
```bash
|
|
112
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-set-policy \
|
|
113
|
-
--root <project> \
|
|
114
|
-
--chain-id FC-YYYYMMDD-001 \
|
|
115
|
-
--run-policy relevant-only \
|
|
116
|
-
--reason "User said this chain is only needed when touching GPU monitor flows."
|
|
117
|
-
```
|
|
118
|
-
|
|
119
|
-
Keep the reason short and human-readable. Use `disabled-with-reason` only when the user asks to disable the chain or it cannot be run safely.
|
|
120
|
-
|
|
121
|
-
## Design Heuristic
|
|
122
|
-
|
|
123
|
-
A good chain has:
|
|
124
|
-
|
|
125
|
-
- one clear entry point
|
|
126
|
-
- one realistic flow, not a pile of unrelated checks
|
|
127
|
-
- two to five checkpoints
|
|
128
|
-
- strict red and green conditions
|
|
129
|
-
- failure localization, so the runner reports which checkpoint broke
|
|
130
|
-
- cleanup-on-pass and preserve-on-fail behavior
|
|
131
|
-
|
|
132
|
-
Avoid:
|
|
133
|
-
|
|
134
|
-
- one script per bad case when one feature flow can cover them
|
|
135
|
-
- broad suites that run unrelated product areas
|
|
136
|
-
- checks that only prove code executed but cannot catch the old symptom
|
|
137
|
-
- durable tests written from agent guesses without user confirmation
|
|
138
|
-
|
|
139
|
-
## Storage
|
|
140
|
-
|
|
141
|
-
Store chain metadata in:
|
|
142
|
-
|
|
143
|
-
```text
|
|
144
|
-
.codex/context/test-hub/feature-chains.json
|
|
145
|
-
```
|
|
146
|
-
|
|
147
|
-
Store large scenario specs in `.codex/context/task-cases/` only when the workflow needs richer phases, logs, or human-readable execution notes.
|
|
148
|
-
|
|
149
|
-
## Execution
|
|
150
|
-
|
|
151
|
-
Approved chains with `status: approved | active | stable` and `run_policy: every-dev-completion` are part of Test Hub. At development completion, run them through:
|
|
152
|
-
|
|
153
|
-
```bash
|
|
154
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py dev-complete --root <project>
|
|
155
|
-
```
|
|
156
|
-
|
|
157
|
-
The runner should execute the approved command with minimal Codex reinterpretation, clean success artifacts, preserve failure evidence, and report the failed checkpoint or blocker.
|
|
158
|
-
|
|
159
|
-
A feature-chain experiment is not complete just because one happy path passes. It should prove the closed loop: one chain covers multiple bad cases, a failing checkpoint preserves evidence with an actionable reason, and the fixed path passes while cleaning temporary artifacts.
|
|
160
|
-
|
|
161
|
-
When an approved chain fails, do not design a new test to prove the same workflow. Treat the failed checkpoint as the recurrence signal, fix the cause, and rerun the same approved chain. A good runner makes this loop cheap by emitting readable checkpoint markers, preserving only the useful failure evidence, and cleaning success artifacts after the rerun passes.
|
|
162
|
-
|
|
163
|
-
Feature-chain commands can report phase-level status with lightweight markers:
|
|
164
|
-
|
|
165
|
-
```text
|
|
166
|
-
CG_CHECKPOINT:<checkpoint title or id>:PASS
|
|
167
|
-
CG_CHECKPOINT:<checkpoint title or id>:FAIL:<short reason>
|
|
168
|
-
```
|
|
169
|
-
|
|
170
|
-
Test Hub treats any `FAIL` marker as a failed feature chain, even if the command exits 0. Prefer these markers when one command covers several checkpoints, because the preserved result will point to the broken workflow step instead of only saying that the whole command failed.
|
|
171
|
-
|
|
172
|
-
Marker names must match registered checkpoint titles or ids. Unknown markers fail the chain because they usually mean the script no longer matches the approved workflow. Keep non-English checkpoint titles readable and distinct; do not collapse them into generic ids.
|
|
173
|
-
|
|
174
|
-
Approved feature-chain commands must report every registered checkpoint unless a checkpoint is explicitly optional (`optional: true` or `required: false`). Missing markers fail the chain, because an unreported checkpoint was not proven to run.
|
|
175
|
-
|
|
176
|
-
If the user or business flow says one checkpoint should not be required every time, change that checkpoint explicitly instead of weakening the whole chain:
|
|
177
|
-
|
|
178
|
-
```bash
|
|
179
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-set-checkpoint \
|
|
180
|
-
--root <project> \
|
|
181
|
-
--chain-id FC-YYYYMMDD-001 \
|
|
182
|
-
--node-title "前端打开监控页" \
|
|
183
|
-
--required optional \
|
|
184
|
-
--reason "Only runs in browser integration environment."
|
|
185
|
-
```
|
|
186
|
-
|
|
187
|
-
Use `--required required` to restore the checkpoint to every-run coverage. This keeps the chain strict by default while allowing intentional, documented exceptions.
|
|
188
|
-
|
|
189
|
-
Audit required/optional coverage without opening JSON:
|
|
190
|
-
|
|
191
|
-
```bash
|
|
192
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-list --root <project> --verbose
|
|
193
|
-
```
|
|
194
|
-
|
|
195
|
-
Audit the compact coverage map before creating new coverage:
|
|
196
|
-
|
|
197
|
-
```bash
|
|
198
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-summary --root <project>
|
|
199
|
-
```
|
|
200
|
-
|
|
201
|
-
This is the quickest way to see whether a small number of feature chains already covers several bad cases, and which checkpoints are still waiting for a real bad case before approval.
|
|
202
|
-
Treat the `coverage density`, `reuse signal`, and `next:` lines as decision aids: they should push Codex toward reusing or extending an existing workflow when possible, not toward creating another standalone test.
|
|
203
|
-
|
|
204
|
-
Audit possible duplicate feature chains before approval:
|
|
205
|
-
|
|
206
|
-
```bash
|
|
207
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-overlap --root <project>
|
|
208
|
-
```
|
|
209
|
-
|
|
210
|
-
This is a route-choice guard, not a test runner. It compares existing chain wording and linked bad cases, then prints pairs that may be the same workflow. Use it to avoid turning one business flow into multiple always-run tests.
|
|
211
|
-
|
|
212
|
-
Audit bad-case coverage across chains without mutating records:
|
|
213
|
-
|
|
214
|
-
```bash
|
|
215
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-coverage --root <project>
|
|
216
|
-
```
|
|
217
|
-
|
|
218
|
-
Use this to decide whether a new bad case should attach to an existing chain or remain a standalone guard. Treat unassigned cases as candidates, not as required new tests.
|
|
219
|
-
|
|
220
|
-
## Quality Gate
|
|
221
|
-
|
|
222
|
-
After editing feature chains, run:
|
|
223
|
-
|
|
224
|
-
```bash
|
|
225
|
-
python3 ~/.agents/skills/context-guard/scripts/context_guard.py validate-feature-chains --root <project>
|
|
226
|
-
```
|
|
227
|
-
|
|
228
|
-
This gate checks structure, not business judgment. It catches approved chains that are missing an entry point, exit check, automated command, checkpoint nodes, checkpoint check text, linked bad-case coverage, or a clear artifact policy. It does not create tests, approve tests, or decide whether the user's workflow deserves a durable chain.
|