@michelj/context-guard 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,290 @@
1
+ # Context Folder Template
2
+
3
+ Use this structure for project context:
4
+
5
+ ```text
6
+ <opened Codex project root>/
7
+ .codex/context/
8
+ ├── index.md
9
+ ├── roadmap.md
10
+ ├── bad-cases.md
11
+ ├── preferences.json
12
+ ├── roadmap/
13
+ │ ├── roadmap.html
14
+ │ ├── roadmap-details.html
15
+ │ ├── roadmap.md
16
+ │ └── roadmap.json
17
+ ├── tasks/
18
+ │ └── CTX-YYYYMMDD-short-slug/
19
+ │ ├── context.md
20
+ │ └── bad-cases.md
21
+ ├── task-cases/
22
+ ├── test-hub/
23
+ │ ├── registry.json
24
+ │ ├── last-run.json
25
+ │ └── runs/
26
+ ├── bad-case-tests/
27
+ └── archive/
28
+ ```
29
+
30
+ The opened Codex project root is the local folder selected in Codex or the local workspace root for the current thread. Do not place this folder inside a skill installation directory, remote SSH path, chat/thread-specific folder, or temporary execution directory unless the user explicitly asks that location to own its own context.
31
+
32
+ ## preferences.json
33
+
34
+ ```json
35
+ {
36
+ "record_language": "unset",
37
+ "display_language": "auto",
38
+ "last_updated": "YYYY-MM-DD",
39
+ "note": "Set with: context_guard.py set-language --language <language>"
40
+ }
41
+ ```
42
+
43
+ Ask the user for a context record language the first time `record_language` is `unset`, then update this file. Use that language for future source context records. Keep literal code identifiers, paths, commands, logs, API names, and exact error text unchanged. If the user changes language later, update this file and use the new language going forward; do not bulk-translate history unless asked.
44
+
45
+ ## index.md
46
+
47
+ ```md
48
+ # Context Index
49
+
50
+ This is a dynamic queue of active and recently parked folder context. Keep it short enough to scan in seconds.
51
+
52
+ ## Quick Scan
53
+
54
+ - Current: CTX-YYYYMMDD-short-slug
55
+ - Latest roadmap node: NODE-YYYYMMDD-001
56
+ - Hot bad-case tags: #hot-ui, #flaky-test
57
+ - Resume candidate: CTX-YYYYMMDD-other-slug
58
+
59
+ Keep Quick Scan to these four lines unless the user explicitly asks for a fuller view.
60
+
61
+ ## Current
62
+
63
+ - ID: CTX-YYYYMMDD-short-slug
64
+ - Title: short task title
65
+ - State: current
66
+ - Folder: `.codex/context/tasks/CTX-YYYYMMDD-short-slug/`
67
+ - Last updated: YYYY-MM-DD
68
+ - Summary: one sentence of the current direction
69
+ - Next step: the next useful action
70
+
71
+ ## Parked / Resume Candidates
72
+
73
+ ### CTX-YYYYMMDD-other-slug
74
+
75
+ - Title: short task title
76
+ - State: parked | resume-candidate
77
+ - Folder: `.codex/context/tasks/CTX-YYYYMMDD-other-slug/`
78
+ - Parked because: urgent bug, unrelated request, waiting on user, etc.
79
+ - Resume prompt: concise question to ask when the interruption is done
80
+ - Last updated: YYYY-MM-DD
81
+
82
+ ## Archived
83
+
84
+ Keep only concise summaries here. Move detailed stale context to `.codex/context/archive/`.
85
+ ```
86
+
87
+ ## roadmap.md
88
+
89
+ ```md
90
+ # Context Roadmap
91
+
92
+ This is the route map through the task. It may contain one mainline, forked side routes, or multiple parallel mainlines. Keep nodes concise. Do not record every tiny action or chat turn.
93
+
94
+ ## Nodes
95
+
96
+ ### NODE-YYYYMMDD-001: Short node title
97
+
98
+ - Date: YYYY-MM-DD
99
+ - Status: planned | active | done | superseded
100
+ - Level: major | checkpoint
101
+ - Branch: Main | short branch name
102
+ - Parent: NODE-YYYYMMDD-000 when this branch forks, otherwise none
103
+ - Task: `CTX-YYYYMMDD-short-slug`
104
+ - Display title: short human-facing card title; use clear language close to what the user cares about
105
+ - User request: concise summary of the user's actual request, using the user's wording as much as possible
106
+ - Progress summary: short human-facing current progress; omit if Outcome already reads naturally
107
+ - Method summary: short human-facing method; omit if Decision / reason already reads naturally
108
+ - Outcome: one-line result
109
+ - Decision / reason: why this node exists, one line
110
+ - Avoid going back: rejected path or lesson, only if it prevents backtracking
111
+ - Next: next useful node or action
112
+ - Linked bad cases: BC-YYYYMMDD-001, BC-YYYYMMDD-002
113
+ - Test chain: compact checkpoint evidence only; user-facing recurrence checks come from linked bad-case guards
114
+ - End-of-work self-check: changed behavior checked; for frontend/layout work include browser/plugin/screenshot evidence or the exact blocker
115
+ ```
116
+
117
+ Use the `### NODE-...` section form as the canonical editable source. If a session accidentally records loose bullet blocks such as `- ID: NODE-...`, `- Title: ...`, `- Level: ...`, the renderer should still project them, but future edits should normalize them back into formal node sections.
118
+
119
+ ## tasks/<task-id>/context.md
120
+
121
+ ```md
122
+ # Task Context: short task title
123
+
124
+ - ID: CTX-YYYYMMDD-short-slug
125
+ - State: current | parked | resume-candidate | done | archived
126
+ - Created: YYYY-MM-DD
127
+ - Last updated: YYYY-MM-DD
128
+
129
+ ## Objective
130
+
131
+ One sentence describing what the user is trying to accomplish.
132
+
133
+ ## Key Points
134
+
135
+ - Important ideas, constraints, and decisions only.
136
+ - Rejected approaches only when they prevent repeating a wrong route.
137
+ - Product, design, architecture, or implementation notes only when needed to resume.
138
+
139
+ ## Open Questions
140
+
141
+ - Questions that need user input or future investigation.
142
+
143
+ ## Files / Areas
144
+
145
+ - Relevant files, modules, commands, screenshots, or external references.
146
+
147
+ ## Bad Cases
148
+
149
+ - Link to shared `.codex/context/bad-cases.md` entries or task-local `bad-cases.md`.
150
+
151
+ ## Roadmap Nodes
152
+
153
+ - Link to `NODE-...` entries in `.codex/context/roadmap.md`.
154
+
155
+ ## Verification / Self-Check
156
+
157
+ - Behavior checked before final answer.
158
+ - Frontend/layout artifacts opened with a browser/plugin or inspected via screenshot when possible.
159
+ - If visual inspection was blocked, record the blocker and residual risk.
160
+
161
+ ## Next Step
162
+
163
+ The smallest useful action to resume this task.
164
+ ```
165
+
166
+ ## task-cases/<task-case-id>.md
167
+
168
+ Use task cases for realistic multi-step verification flows. They should catch bugs by simulating a real task, not by testing one isolated bug at a time. Task-case design is human-owned: Codex may draft and structure a proposal, but durable task cases must stay `proposed` until the user confirms them.
169
+
170
+ ```md
171
+ # Task Case: short realistic workflow title
172
+
173
+ - ID: TC-YYYYMMDD-short-slug
174
+ - Status: proposed | approved | active | stable | deferred | obsolete
175
+ - Route/task: `CTX-...` or branch name
176
+ - Scope: feature, service, UI flow, agent workflow, or subsystem
177
+ - Last checked: YYYY-MM-DD
178
+ - Design confirmation: pending | user-approved YYYY-MM-DD
179
+ - Run policy: every-dev-completion | relevant-only | manual | release-only | goal-final | disabled-with-reason | user-defined cadence
180
+ - Automation entry: native command | script path | prompt/manual runner | none
181
+ - Artifact policy: cleanup-on-pass | preserve-on-fail | manual-preserve
182
+ - Linked roadmap nodes: NODE-...
183
+ - Linked bad cases: BC-..., BC-...
184
+ - Entry command/prompt: command, prompt, manual setup, or fixture
185
+ - Not covered: explicit exclusions to avoid fake confidence
186
+ - Stop condition: what means the workflow is complete
187
+ - Cleanup: required cleanup or none
188
+
189
+ ## Phases
190
+
191
+ ### Phase 1: setup or trigger
192
+
193
+ - Action: one realistic action
194
+ - Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
195
+ - Covers bad cases: BC-...
196
+ - Failure localization: what this phase failure usually means
197
+ - Log note: what the script/agent should record
198
+
199
+ ### Phase 2: transition or recovery
200
+
201
+ - Action: one realistic action
202
+ - Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
203
+ - Covers bad cases: BC-...
204
+ - Failure localization: what this phase failure usually means
205
+ - Log note: what the script/agent should record
206
+
207
+ ## Result Log
208
+
209
+ - YYYY-MM-DD: pass/fail, failed phase/checkpoint if any, evidence path or command output summary
210
+ ```
211
+
212
+ ## Maintenance Rules
213
+
214
+ - Keep `index.md` small and useful, not exhaustive.
215
+ - Keep `roadmap.md` as the route map. It should show progress as nodes, not a raw transcript.
216
+ - Use `Level: major` for significant milestones shown as main route cards; use `Level: checkpoint` for minor progress that should live in details.
217
+ - Use `Branch:` for forked or parallel routes. Missing `Branch:` means `Main`; use `Parent:` to point to the node where a branch forked.
218
+ - In the human overview, visible card numbers should be consecutive per route group after checkpoint filtering, not source node numbers with gaps.
219
+ - If the human overview has multiple route groups, show all route lines together with parent/fork markers and a compact test route aligned under visible roadmap nodes.
220
+ - Each node should be concise enough for Codex to scan quickly: outcome, decision, next step, linked bad cases.
221
+ - Link nodes to bad cases and test-chain notes instead of duplicating full details.
222
+ - Treat human-facing test coverage as human-designed bad-case recurrence detection. Single-route overview hides the test lane; multi-route compact test routes should be generated only from user-approved tests with explicit `Run policy`, approved task-case checkpoints, or approved test registry entries, not from ordinary linked bad-case guards or roadmap node checkpoint logs.
223
+ - In branch or multi-route views, align each visible test item to the roadmap node whose approved bad-case test or approved task-case checkpoint it covers. If a route has no approved tests, do not show a test route for it. Empty test slots should be subtle timeline placeholders only when the route has at least one approved test elsewhere.
224
+ - Prefer a task-oriented case in `.codex/context/task-cases/` when realistic workflow phases matter more than isolated bug checks. Bad-case guards should often point to a task-case checkpoint that covers them.
225
+ - Task-case scripts or agents should log the phase/checkpoint that failed, so Codex can locate the broken workflow step without re-debugging the whole task.
226
+ - Before writing any new durable task-case script or active task case, ask the user to confirm with only the business path: from what state to what state, the main task, and the major risk. Keep technical phases/checkpoints/logs inside the task-case file, not in the confirmation prompt. If confirmation is unavailable, keep the case `proposed` and avoid broad new scripts.
227
+ - When the user explicitly asks to create, write, generate, design, or add a test/test task/task case, start the user-visible response with `测试创建识别:...` or the folder-language equivalent, then summarize the test target from what state to what state and the main risk it catches.
228
+ - When the user creates or approves a test, register it with `Run policy: every-dev-completion` by default. At the end of every development turn, run all approved tests with that policy or record the exact blocker.
229
+ - Use `.codex/context/test-hub/registry.json` as the explicit Test Hub registry for user-approved automated tests. Do not populate it from ordinary bad-case guards or roadmap node `Test chain:` notes.
230
+ - Keep Test Hub as a simple control layer: registry, `dev-complete`, `last-run.json`, and lightweight commands to list, enable, disable, change policy, or remove registry tests.
231
+ - At development completion, prefer `context_guard.py dev-complete --root <project>` so the hub runs the approved always-run set, handles parallel workers when safe, cleans success artifacts, and preserves failure evidence under `.codex/context/test-hub/runs/`.
232
+ - After approval, automate a test when it can be safely scripted or run as a native command. Future Codex turns should execute the registered entry with minimal reinterpretation.
233
+ - Automated tests should clean temporary files after full success and preserve concise diagnostic artifacts on failure.
234
+ - Failed approved tests become a bad-case analysis loop: inspect the preserved evidence, fix the in-scope cause, rerun the same approved test, and stop only after pass or a non-actionable blocker.
235
+ - If blocked by credentials, unavailable external service, permissions, hardware/resource limits, network, destructive-risk confirmation, or user-only judgment, ask or warn the user with the exact blocker and evidence path.
236
+ - Change a test to `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or another cadence only when the user explicitly asks. Record the user's reason beside the policy.
237
+ - During goal mode, task cases should act as phase gates: select an approved case or propose one for confirmation, log phase progress during continuations, and run the smallest human-approved path before claiming the goal complete.
238
+ - Keep multilingual display as an HTML projection concern; do not duplicate source context by language. When supported, localize human-facing record titles, summaries, bad-case labels, and test-chain snippets in the projection.
239
+ - Keep source records in the configured `.codex/context/preferences.json` record language. The HTML roadmap should follow that preference and should not show a visible language selector by default.
240
+ - During goal mode or long-running autonomous work, keep the active goal aligned to the current task, add compact goal checkpoints during meaningful phase changes, and record bad cases as soon as they appear.
241
+ - Treat `.codex/context/index.md`, `.codex/context/roadmap.md`, `.codex/context/bad-cases.md`, and task context files as the source of truth.
242
+ - Treat `.codex/context/roadmap/roadmap.html` as a human-facing view only. Codex should not use it for context intake or bad-case management.
243
+ - Treat `.codex/context/roadmap/roadmap.md` and `.codex/context/roadmap/roadmap.json` as stable agent-readable exports for quick scanning, route lookup, bad-case lookup, and recurrence-guard lookup, not as primary editable sources.
244
+ - Keep `NODE-...`, `BC-...`, and `CTX-...` IDs in source files for linking, but hide them in the default human-facing HTML. Show short natural-language node and bad-case labels instead.
245
+ - In human-facing HTML, prefer color, symbols, and compact visual markers over labels like `Status:`, `Nodes:`, `Frequency:`, or fallback text such as `untagged`.
246
+ - Show meaningful `#tags` as compact colored chips with small emoji cues in human-facing HTML. Limit overview tags; show full tags on the detail page; omit the tag row when no tags exist.
247
+ - A sharp task direction change should park the current task before starting a new one.
248
+ - If the user explicitly says a task is a branch/side route/fork/支线/分支, run or emulate `scripts/context_guard.py create-branch-task --title <task title> --branch <branch name> --parent-node <parent NODE id>` before implementation so the task folder, current index entry, and `Branch:`/`Parent:` roadmap node all exist.
249
+ - If work significantly drifts from the mainline architecture without an explicit branch request, ask whether to create a branch before silently continuing.
250
+ - When an interruption finishes, ask whether to resume the most relevant parked task.
251
+ - Do not let parked items grow endlessly. Mark stale items `archived` and compress them to a short summary.
252
+ - Do not delete unresolved user intent unless the user explicitly discards it.
253
+ - Use `scripts/context_guard.py show-roadmap` to generate and display the stable human-friendly overview at `.codex/context/roadmap/roadmap.html`, with details at `.codex/context/roadmap/roadmap-details.html`, agent-readable Markdown at `.codex/context/roadmap/roadmap.md`, and structured lookup at `.codex/context/roadmap/roadmap.json`. Use `export-roadmap --format md` only for Markdown-only export.
254
+ - Do not accumulate timestamped HTML roadmap files. Showing the roadmap overwrites the same stable HTML files.
255
+ - With one route group, the HTML roadmap overview should show only the main route cards. Keep linked bad cases and recurrence checks in clicked node details, source context, and agent-readable exports.
256
+ - In one-route HTML, do not render bad-case/test-chain lanes or a left lane-label column. The route board should size to real content, and main route summaries should remain readable rather than being clipped after a very short fragment.
257
+ - With multiple route groups, the overview should show all route lines as a branch map. If a route has user-approved tests, also show a compact node-aligned test route under that route line. Route selection may affect details, but the default view should not invent tests from ordinary bad-case context.
258
+ - Parent/fork markers should appear only on side routes whose parent node is outside that route. Main route should not show a fork marker just because a later main node references an earlier main node.
259
+ - Side routes should visually start near their parent node's visible position on the parent route, not all from the first column.
260
+ - Branch route labels, parent chips, and checkpoint text should sit near the branch's first visible card by reusing the same spacer/grid coordinate as the branch cards.
261
+ - Branch overview should use one shared horizontal route canvas. Route alignment should use grid spacer columns, not padding that shifts or clips the whole route section.
262
+ - Branch connector lines should use the same offset coordinate as the route's spacer columns, not a fixed left-edge position.
263
+ - Branch connector endpoints should be anchored to the status dots inside the source and target node cards; do not draw connector lines from the whole route section or unrelated card edges.
264
+ - Route progression connectors should be card-to-card through card gaps; branch connectors should be dot-to-dot through an empty branch corridor and must not cross node cards or text.
265
+ - Side routes may drift right from exact column alignment when that creates a cleaner non-crossing branch path.
266
+ - Connector layers should render behind route cards so cards mask any line segment that would otherwise pass over content.
267
+ - Hide heavy native horizontal scrollbar chrome in the roadmap overview while preserving horizontal scroll interaction.
268
+ - Human-facing node detail cards should show only one concise summary sentence, linked bad cases, and linked bad-case recurrence tests. Do not show a standalone status dot under the node detail title. Keep route, parent, decision, avoid-going-back, and next-step source fields in agent-readable context, not in the human detail card.
269
+ - Human-facing bad-case details should localize phenomenon, trigger, root cause, fix, and guard prose to the folder language preference while preserving technical identifiers, commands, and paths.
270
+ - Route color should encode branch depth: main route green, first-level branch cool cyan/teal, deeper branch levels progressively colder toward blue and indigo.
271
+ - Before finalizing frontend, roadmap HTML, or visual layout work, open or render the artifact with an available browser/plugin or screenshot path and inspect for obvious visual bugs. Do not rely only on string assertions for layout changes.
272
+ - Treat Stop hook output as a completion reliability gate. Before claiming fixed/done/passing, record real verification evidence for the changed artifact or workflow and rerun relevant bad-case guards.
273
+ - For UI/browser/binding/frontend work, verify the original user-visible symptom, not only build success or process restart.
274
+ - User-facing projected text should follow the folder language preference; avoid untranslated English prose in Chinese overview output except for intentional technical strings.
275
+ - For a single route group, do not show lane titles or a left label column; the overview is only the main route.
276
+ - Keep overview cards sparse. Put full Outcome, Decision, Next, and guard details in same-file detail anchors and the stable `roadmap-details.html` sidecar.
277
+ - In multi-route branch overview, route cards should read as a compact map skeleton: number, title, date/status cue, and no visible outcome paragraph. Keep route summaries in details and source context.
278
+ - Default overview links should target same-file `#node-*` and `#case-*` anchors, not `roadmap-details.html#...`, to avoid `file://` access-denied navigation.
279
+
280
+ ## Pruning Rules
281
+
282
+ - Do not record normal implementation chatter.
283
+ - Do not record every command; record only commands that prove a checkpoint or guard a bad case.
284
+ - Do not let roadmap node `Test chain:` history replace bad-case recurrence guards in user-facing roadmap output.
285
+ - Do not split a real workflow into many unrelated bug-level tests when one task case with checkpoints would reveal the failure location more clearly.
286
+ - Do not silently enter test creation. If the user explicitly asks to create a test, acknowledge the test-creation intake first so the user can see the skill activated.
287
+ - Do not wait until goal completion to record important roadmap progress or bad cases.
288
+ - Merge tiny adjacent updates into one roadmap node.
289
+ - Archive stale parked tasks as a one-sentence summary.
290
+ - If a reader cannot use a detail to resume, decide, verify, or avoid recurrence, remove it.
@@ -0,0 +1,85 @@
1
+ # Bad Case Register Template
2
+
3
+ Use this format for `.codex/context/bad-cases.md`, task-local `.codex/context/tasks/<task-id>/bad-cases.md`, or the existing project context register.
4
+
5
+ ```md
6
+ # Bad Case Register
7
+
8
+ This register tracks bad cases found during development and the guards that prevent them from recurring.
9
+
10
+ Record only bad cases that are user-visible, recurring, risky, fixed, deferred, or needed to explain a guard. Do not turn the register into a defect diary.
11
+
12
+ ## Active Cases
13
+
14
+ ### BC-YYYYMMDD-001: Short descriptive title
15
+
16
+ - Status: open | resolved | recurred | deferred | superseded-by-route-change
17
+ - First observed: YYYY-MM-DD
18
+ - Last checked: YYYY-MM-DD
19
+ - Scope: feature, files, tests, route, UI flow, API, or subsystem
20
+ - Context task: `CTX-...` folder or shared
21
+ - Roadmap nodes: `NODE-...`
22
+ - Tags: #hot | #flaky | #ui | #data-loss | #route-risk | custom tags
23
+ - Frequency: first-seen | repeated-N | high-frequency
24
+ - Phenomenon: one-line user-visible behavior or failing output
25
+ - Trigger / reproduction: shortest command, step, input, environment, or precondition
26
+ - Root cause: confirmed cause, suspected cause, or unknown, one line
27
+ - Fix method: code/test/config/documentation change that fixed it, one line
28
+ - Guard type: script | native-test | manual | browser-screenshot | browser-dom | curl | cli | prompt | log-invariant | fixture | unit | integration | e2e | custom
29
+ - Guard / verification: native test, command, reusable script, manual check, screenshot, log, invariant, or reproduction note, one line
30
+ - Run policy: every-dev-completion | relevant-only | manual | release-only | goal-final | disabled-with-reason | user-defined cadence
31
+ - Artifact policy: cleanup-on-pass | preserve-on-fail | manual-preserve | none
32
+ - Blocker handling: credentials | external-service | permissions | resource-limits | network | destructive-confirmation | user-judgment | none
33
+ - Red condition: exact output, visual state, assertion, or symptom that means this bad case has recurred
34
+ - Green condition: exact evidence that means this bad case is absent
35
+ - Expected failure reason: why the guard should fail for the old symptom, not for a broken test or unrelated environment issue
36
+ - Reusable guard path: project test file, `.codex/context/task-cases/...#phase-name`, `.codex/context/bad-case-tests/...`, or none
37
+ - Covered by task case: TC-YYYYMMDD-short-slug phase/checkpoint, or none
38
+ - Test-chain issue: false-positive | false-negative | wrong-granularity | missing-phase | wrong-assertion | unrealistic-setup | missing-cleanup | unclear-localization | none
39
+ - Guard reuse rule: reuse this recorded guard before creating any new test or script for this case
40
+ - Test chain: ordered checks only when multiple checks are genuinely needed
41
+ - High-frequency note: warning text to show Codex when this pattern repeats often
42
+ - Recurrence analysis: why it came back, if it ever did
43
+ - Route-change note: only when an approved technical route change intentionally changes expected behavior
44
+ - Evidence: links to tests, commands run, PRs, commits, screenshots, or logs
45
+
46
+ ## Resolved History
47
+
48
+ Move old resolved entries here only if the active section becomes noisy. Keep enough detail to replay the guard.
49
+ ```
50
+
51
+ Use the `### BC-...` section form as the canonical editable source. If a session accidentally records loose bullet blocks such as `- ID: BC-...`, `- Title: ...`, `- Status: ...`, or `- Nodes: ...`, the renderer should still project them, but future edits should normalize them back into formal case sections.
52
+
53
+ ## Status Rules
54
+
55
+ - `open`: bad case is known and not fixed.
56
+ - `resolved`: fix is implemented and verification passed.
57
+ - `recurred`: bad case came back after resolution; must be analyzed and fixed before completion.
58
+ - `deferred`: intentionally not fixed in the current task; requires reason and owner/next step.
59
+ - `superseded-by-route-change`: old behavior is no longer expected because an approved technical route changed it.
60
+
61
+ ## Context Guard Rules
62
+
63
+ - Use `.codex/context/` as the project folder for bad-case memory. Do not introduce a separate bad-case folder for new projects.
64
+ - Use the configured `.codex/context/preferences.json` `record_language` for bad-case titles, phenomenon, root cause, fix method, guard summaries, and test-chain notes.
65
+ - Preserve exact commands, paths, code identifiers, logs, API names, and error messages in their original language.
66
+ - Prefer existing recorded context, user-approved commands, native tests, screenshots, logs, or manual checks over newly invented tests.
67
+ - When the user creates or approves a test, default its `Run policy` to `every-dev-completion`; Codex must run it at every development completion unless the user sets another cadence.
68
+ - Only the user can demote an approved test to `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or a custom cadence. Record why.
69
+ - After user approval, automate the check when feasible; successful automated checks should clean temporary files, while failed checks should preserve concise diagnostic evidence for Codex to analyze and rerun after fixing.
70
+ - Store reusable user-approved automated tests in `.codex/context/test-hub/registry.json` or approved task-case files; do not auto-register ordinary bad-case guards or roadmap `Test chain:` notes as always-run tests.
71
+ - At development completion, prefer `context_guard.py dev-complete --root <project>` so Test Hub runs the approved always-run set and writes `.codex/context/test-hub/last-run.json`.
72
+ - If an approved automated check is blocked by credentials, external services, permissions, resource limits, network, destructive confirmation, or user-only judgment, record the blocker and ask or warn the user instead of looping.
73
+ - For resolved or recurred cases, `Guard / verification`, `Guard type`, `Red condition`, `Green condition`, and `Expected failure reason` are required.
74
+ - The guard must be red-capable: it should fail if the same user-visible symptom returns.
75
+ - When the bad case is part of a longer workflow, attach it to a human-approved task-case checkpoint instead of creating a separate isolated script. The bad-case entry should say which task case phase covers it.
76
+ - If the test chain itself is wrong, record that as a bad case and classify the test-chain issue. Fix the test-chain design before trusting its result.
77
+ - Do not script every bad case. Store bad-case-specific scripts under `.codex/context/bad-case-tests/` only when the user approved the test design, the script is genuinely reusable, and it does not belong in the native test suite.
78
+ - Name any guard script with the bad case ID so it is easy to find and reuse.
79
+ - Update existing context when expected behavior changes; do not create parallel guards for the same case unless the old one is explicitly obsolete.
80
+ - If a guard is manual-only, list the exact manual check and why that is acceptable for now.
81
+ - Link bad cases to roadmap nodes so Codex can quickly see which mainline decisions created or fixed them.
82
+ - Keep record/display linkage explicit: use `Roadmap nodes:` or `Nodes:` on the bad case, or `Linked bad cases:` on the roadmap node.
83
+ - Add tags and frequency notes when a bad case repeats often; high-frequency cases should stand out during quick scanning.
84
+ - Promote high-frequency cases into fixed pressure checks and rerun them whenever related code, UI, context, or hooks change.
85
+ - Keep entries compact. If the same information appears in a roadmap node, link to it instead of duplicating it.
@@ -0,0 +1,63 @@
1
+ # Task Case Template
2
+
3
+ Use task cases for realistic multi-step verification. A task case should simulate a full user or agent workflow and log phase-level checkpoints so failures identify the broken step. Test design is human-owned: Codex can draft a proposal, but a durable task case stays `proposed` until the user confirms it.
4
+
5
+ ```md
6
+ # Task Case: short realistic workflow title
7
+
8
+ - ID: TC-YYYYMMDD-short-slug
9
+ - Status: proposed | approved | active | stable | deferred | obsolete
10
+ - Route/task: `CTX-...` or branch name
11
+ - Scope: feature, service, UI flow, agent workflow, or subsystem
12
+ - Last checked: YYYY-MM-DD
13
+ - Design confirmation: pending | user-approved YYYY-MM-DD
14
+ - Run policy: every-dev-completion | relevant-only | manual | release-only | goal-final | disabled-with-reason | user-defined cadence
15
+ - Automation entry: native command | script path | prompt/manual runner | none
16
+ - Artifact policy: cleanup-on-pass | preserve-on-fail | manual-preserve
17
+ - Linked roadmap nodes: NODE-...
18
+ - Linked bad cases: BC-..., BC-...
19
+ - Entry command/prompt: command, prompt, manual setup, or fixture
20
+ - Not covered: explicit exclusions to avoid fake confidence
21
+ - Stop condition: what means the workflow is complete
22
+ - Cleanup: required cleanup or none
23
+ - Blocker handling: credentials | external service | permissions | resource limits | network | destructive confirmation | user judgment | none
24
+
25
+ ## Phases
26
+
27
+ ### Phase 1: setup or trigger
28
+
29
+ - Action: one realistic action
30
+ - Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
31
+ - Covers bad cases: BC-...
32
+ - Failure localization: what this phase failure usually means
33
+ - Log note: what the script/agent should record
34
+
35
+ ### Phase 2: transition or recovery
36
+
37
+ - Action: one realistic action
38
+ - Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
39
+ - Covers bad cases: BC-...
40
+ - Failure localization: what this phase failure usually means
41
+ - Log note: what the script/agent should record
42
+
43
+ ## Result Log
44
+
45
+ - YYYY-MM-DD: pass/fail, failed phase/checkpoint if any, evidence path or command output summary
46
+ ```
47
+
48
+ Rules:
49
+
50
+ - When the user explicitly asks to create, write, generate, design, or add a test/task case, begin the user-visible response with a compact intake line such as `测试创建识别:...` before proposing or implementing the case.
51
+ - Prefer one task case with clear checkpoints over many disconnected bug-level scripts when the same workflow is being exercised.
52
+ - Link bad cases to the checkpoint that catches them.
53
+ - Keep the task case realistic enough to match actual product or agent usage.
54
+ - Once a test is approved, automate it when it can be safely encapsulated. Future Codex turns should run the registered entry instead of reinterpreting the test design.
55
+ - Automated approved tests should be registered in `.codex/context/test-hub/registry.json` or represented by this approved task-case file with `Run policy: every-dev-completion` and an executable `Entry command/prompt`.
56
+ - At development completion, run the approved always-run set through `context_guard.py dev-complete --root <project>` instead of manually reconstructing each test.
57
+ - Automated task cases should clean temporary files after full success and preserve the smallest useful evidence on failure.
58
+ - If the automated case fails, Codex should analyze the failed phase/checkpoint, record or update the bad case, fix in scope, and rerun the same approved test until it passes or a non-actionable blocker is reached.
59
+ - If blocked by credentials, unavailable services, permissions, hardware/resource limits, network, destructive-risk confirmation, or user-only judgment, stop and ask or warn the user with the blocker and evidence path.
60
+ - Keep checkpoint logs concise and useful for localizing failure.
61
+ - Keep the case in `proposed` state until the user confirms the design.
62
+ - Once the user confirms the design, default `Run policy` to `every-dev-completion`; lower the cadence only when the user explicitly asks.
63
+ - In goal mode, use human-approved task cases as phase gates and log phase progress; do not create broad new scripts or active task cases without confirmation unless the user explicitly asked for that exact test.