@michelj/context-guard 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (72) hide show
  1. package/README.md +74 -6
  2. package/README.zh-CN.md +74 -6
  3. package/{skills/context-guard/SKILL.md → SKILL.md} +216 -77
  4. package/bin/context-guard-skill.js +88 -3
  5. package/hooks.json +25 -1
  6. package/package.json +7 -3
  7. package/{skills/context-guard/references → references}/context-template.md +52 -1
  8. package/references/feature-chain-methodology.md +228 -0
  9. package/{skills/context-guard/scripts → scripts}/context_guard.py +3119 -121
  10. package/{skills/context-guard/scripts → scripts}/context_guard_hook.py +344 -12
  11. package/{skills/context-guard/tests → tests}/BC-20260701-090.sh +8 -7
  12. package/{skills/context-guard/tests → tests}/BC-20260702-096.sh +6 -3
  13. package/tests/BC-20260706-098.sh +66 -0
  14. package/tests/BC-20260707-099.sh +47 -0
  15. package/tests/BC-20260707-100.sh +46 -0
  16. package/tests/BC-20260707-101.sh +47 -0
  17. package/tests/BC-20260707-102.sh +68 -0
  18. package/tests/BC-20260707-103.sh +59 -0
  19. package/tests/BC-20260707-104.sh +103 -0
  20. package/tests/BC-20260707-105.sh +109 -0
  21. package/tests/BC-20260707-106.sh +80 -0
  22. package/tests/BC-20260707-107.sh +74 -0
  23. package/tests/BC-20260707-108.sh +48 -0
  24. package/tests/BC-20260707-109.sh +56 -0
  25. package/tests/BC-20260707-110.sh +71 -0
  26. package/tests/BC-20260707-111.sh +70 -0
  27. package/tests/BC-20260707-112.sh +45 -0
  28. package/tests/BC-20260707-113.sh +73 -0
  29. package/tests/BC-20260707-115.sh +77 -0
  30. package/tests/BC-20260707-116.sh +77 -0
  31. package/tests/BC-20260707-118.sh +115 -0
  32. package/tests/BC-20260707-119.sh +47 -0
  33. package/tests/BC-20260707-120.sh +60 -0
  34. package/tests/BC-20260707-121.sh +66 -0
  35. package/tests/BC-20260707-122.sh +48 -0
  36. package/tests/BC-20260707-123.sh +43 -0
  37. package/tests/BC-20260707-124.sh +56 -0
  38. package/tests/BC-20260707-125.sh +64 -0
  39. package/tests/BC-20260707-126.sh +80 -0
  40. package/tests/BC-20260707-127.sh +88 -0
  41. package/tests/BC-20260707-129.sh +59 -0
  42. package/tests/BC-20260707-130.sh +69 -0
  43. package/tests/BC-20260707-131.sh +140 -0
  44. package/tests/BC-20260707-132.sh +150 -0
  45. package/tests/BC-20260707-133.sh +70 -0
  46. package/tests/BC-20260708-136.sh +210 -0
  47. package/tests/BC-20260708-137.sh +106 -0
  48. package/tests/BC-20260708-138.sh +168 -0
  49. package/tests/BC-20260708-139.sh +79 -0
  50. package/tests/BC-20260709-002.sh +63 -0
  51. package/tests/BC-20260709-003.sh +239 -0
  52. package/tests/BC-20260709-006.sh +76 -0
  53. package/tests/BC-20260709-008.sh +168 -0
  54. package/tests/BC-20260710-001.sh +61 -0
  55. package/tests/BC-20260710-002.sh +111 -0
  56. package/tests/npm-install-smoke.sh +53 -0
  57. package/skills/context-guard/README.md +0 -234
  58. package/skills/context-guard/README.zh-CN.md +0 -234
  59. /package/{skills/context-guard/agents → agents}/openai.yaml +0 -0
  60. /package/{skills/context-guard/references → references}/register-template.md +0 -0
  61. /package/{skills/context-guard/references → references}/task-case-template.md +0 -0
  62. /package/{skills/context-guard/tests → tests}/BC-20260618-063.sh +0 -0
  63. /package/{skills/context-guard/tests → tests}/BC-20260618-065.sh +0 -0
  64. /package/{skills/context-guard/tests → tests}/BC-20260626-080.sh +0 -0
  65. /package/{skills/context-guard/tests → tests}/BC-20260626-081.sh +0 -0
  66. /package/{skills/context-guard/tests → tests}/BC-20260626-082.sh +0 -0
  67. /package/{skills/context-guard/tests → tests}/BC-20260626-083.sh +0 -0
  68. /package/{skills/context-guard/tests → tests}/BC-20260627-084.sh +0 -0
  69. /package/{skills/context-guard/tests → tests}/BC-20260630-086.sh +0 -0
  70. /package/{skills/context-guard/tests → tests}/BC-20260630-087.sh +0 -0
  71. /package/{skills/context-guard/tests → tests}/BC-20260630-088.sh +0 -0
  72. /package/{skills/context-guard/tests → tests}/BC-20260630-089.sh +0 -0
@@ -17,8 +17,9 @@ Context is a navigation aid, not a transcript. Record only information that help
17
17
  2. Keep major roadmap nodes for significant changes only; record small implementation updates as `Level: checkpoint`.
18
18
  3. Keep task context to key points only: objective, constraints/decisions, open questions, touched areas, and next step.
19
19
  4. Record a bad case only when it is user-visible, recurring, risky, fixed, deferred, or needed to explain a guard.
20
- 5. Prefer one-line summaries. If a detail is not needed for resume, route choice, or recurrence prevention, omit it.
21
- 6. When exporting or displaying context, show the shortest useful view first and leave secondary details folded or linked.
20
+ 5. Treat the user's own wording as context when it contains requirements, preferences, constraints, credentials, route changes, or bad-case reports. Preserve short user messages without turning context into a full transcript.
21
+ 6. Prefer one-line summaries. If a detail is not needed for resume, route choice, or recurrence prevention, omit it.
22
+ 7. When exporting or displaying context, show the shortest useful view first and leave secondary details folded or linked.
22
23
 
23
24
  ## Context Folder
24
25
 
@@ -46,14 +47,16 @@ Do not use the active chat/thread name, the skill installation directory, a remo
46
47
  3. Maintain the quick-browse index at `.codex/context/index.md`.
47
48
  4. Maintain the main route map at `.codex/context/roadmap.md`.
48
49
  5. Maintain folder preferences at `.codex/context/preferences.json`.
49
- 6. Store task-specific context under `.codex/context/tasks/<task-id>/`.
50
- 7. Store task-oriented evaluation scenarios under `.codex/context/task-cases/` when a reusable long workflow is more useful than isolated bug checks.
51
- 8. Store the Test Hub registry and run evidence under `.codex/context/test-hub/`.
52
- 9. Store shared bad-case and test-chain context at `.codex/context/bad-cases.md` unless a bad case belongs only inside one task folder.
53
- 10. If no canonical context exists, read legacy bad-case locations if present: `.codex/bad-cases.md`, `BAD_CASES.md`, `docs/bad-cases.md`, or `.agents/bad-cases.md`.
54
- 11. If legacy context exists and the task modifies context, migrate or copy it into `.codex/context/` unless the repository clearly standardizes on the legacy path.
55
- 12. Use `references/context-template.md` for index, roadmap, task-folder, and task-case formats.
56
- 13. Use `references/register-template.md` when creating or updating bad-case entries.
50
+ 6. Maintain user-message memory at `.codex/context/user-messages.md`.
51
+ 7. Store local-only sensitive memory under `.codex/context/private/`; keep it gitignored and never project it into HTML.
52
+ 8. Store task-specific context under `.codex/context/tasks/<task-id>/`.
53
+ 9. Store task-oriented evaluation scenarios under `.codex/context/task-cases/` when a reusable long workflow is more useful than isolated bug checks.
54
+ 10. Store the Test Hub registry and run evidence under `.codex/context/test-hub/`.
55
+ 11. Store shared bad-case and test-chain context at `.codex/context/bad-cases.md` unless a bad case belongs only inside one task folder.
56
+ 12. If no canonical context exists, read legacy bad-case locations if present: `.codex/bad-cases.md`, `BAD_CASES.md`, `docs/bad-cases.md`, or `.agents/bad-cases.md`.
57
+ 13. If legacy context exists and the task modifies context, migrate or copy it into `.codex/context/` unless the repository clearly standardizes on the legacy path.
58
+ 14. Use `references/context-template.md` for index, roadmap, task-folder, and task-case formats.
59
+ 15. Use `references/register-template.md` when creating or updating bad-case entries.
57
60
 
58
61
  Do not store project context inside the skill directory. If a command would use `/Users/.../.agents/skills/context-guard` or another installed skill path as the implicit root, stop and rerun it from the opened Codex workspace or pass the workspace with `--root`. Do not create a separate top-level bad-case folder; bad cases are part of `context`.
59
62
 
@@ -80,6 +83,20 @@ Context records need a folder-scoped language preference so Codex does not mix l
80
83
  7. Do not bulk-translate historical records unless the user explicitly asks for migration.
81
84
  8. The HTML roadmap follows the folder language preference by default. Do not show a visible language selector in the human-facing roadmap unless the user explicitly asks for one.
82
85
 
86
+ ## User Message Memory
87
+
88
+ User messages are not disposable chat noise. Short user instructions, corrections, preferences, constraints, route hints, server connection details, credentials needed for the current task, and bad-case reports must be preserved in project context so a later Codex turn does not ask the user to repeat them.
89
+
90
+ 1. At `UserPromptSubmit` or turn start, record the latest user prompt in `.codex/context/user-messages.md` when it contains durable context. For normal short prompts, keep the user's wording as written or near-verbatim.
91
+ 2. Do not store every long pasted file, log, generated artifact, or giant attachment. For large inputs, record the file/source, purpose, first meaningful line or summary, and critical identifiers only.
92
+ 3. Keep `.codex/context/user-messages.md` agent-readable and concise: recent user signals plus durable constraints. Promote stable requirements into the active task context or roadmap `User request:` fields, then archive stale chatter.
93
+ 4. Preserve exact code identifiers, paths, hosts, ports, commands, API names, and error messages unless they are secrets.
94
+ 5. Never put raw secrets in `index.md`, `roadmap.md`, `bad-cases.md`, task context, Roadmap HTML, exported JSON/Markdown, logs, README, git-tracked files, or final answers.
95
+ 6. If the user explicitly provides a credential, token, password, or similar secret that future Codex turns need, store the raw value only in `.codex/context/private/secrets.local.json` with local-only permissions or an OS credential store. In public context, record only a redacted pointer such as `USER-SECRET-...`.
96
+ 7. Do not persist one-time codes such as OTP or short-lived verification codes. Record only a redacted note that a one-time code was supplied.
97
+ 8. If safe private storage is unavailable, record a redacted note and ask the user to re-provide the secret or use a secure store when needed.
98
+ 9. When creating a roadmap node, derive `User request:` from the preserved user wording rather than from implementation logs, `Outcome`, or `Decision / reason`.
99
+
83
100
  ## Dynamic Task Index
84
101
 
85
102
  Use `.codex/context/index.md` as a small, actively maintained queue of work context and `.codex/context/roadmap.md` as the route map. A route map may have one mainline, forked side routes, or multiple parallel mainlines.
@@ -150,15 +167,17 @@ Do not create timestamped HTML roadmap exports for display. The roadmap folder s
150
167
  - Show the roadmap tracks, concise node titles, status/date chips, and at most one short summary line.
151
168
  - Overview node titles must use `Display title:` when present. Keep them close to the user's wording and outcome, not implementation-log language. Use short natural phrases such as "节点详情更容易读" instead of "压缩节点详情长字段".
152
169
  - Show only `Level: major` nodes as main route cards; summarize hidden checkpoints compactly and put checkpoint details in `roadmap-details.html`.
170
+ - Keep two presentations in the same stable `roadmap.html`: a card view for reading a small route and a compact overview for scanning many nodes. Use icon-only controls with accessible labels, preserve the user's selected mode locally, and default to compact overview when more than sixteen major nodes are visible.
171
+ - In compact overview, wrap concise node tiles into the available viewport width, show only the visible number, short title, and status cue, and keep every tile linked to the same node detail. Do not duplicate or discard source context to create the compact view.
172
+ - Compact overview may replace connector curves with route grouping and parent labels to avoid lines crossing wrapped tiles; the card view remains the detailed branch-connector view.
153
173
  - Number visible overview cards consecutively per route group after checkpoint filtering; keep source node IDs and source-order detail anchors hidden from the overview.
154
174
  - When there is only one route group, show only the main route cards in the default overview. Do not show the Bad Cases lane, Test Chain lane, or the left lane-label column. Keep linked bad cases and recurrence checks inside clicked node details, source context, and agent-readable exports.
155
175
  - In single-route overview mode, keep the board content-height compact: main route cards should size to content, and summaries may use up to three readable lines before truncation.
156
176
  - When there are multiple route groups, show all route lines together as a branch overview so users can see where each side route forked. Do not show a separate always-visible bad-case/test-chain drilldown under the route map; keep detailed case/check relationships in source context and agent-readable exports.
157
177
  - In multi-route branch overview, treat route cards as a compact map skeleton: show the number, title, and small date/status cues only. Hide outcome summaries from the visible cards and keep them in same-file details/source context.
158
- - When user-approved tests exist, show a compact test route directly under the route cards in multi-route overview. Align each visible test slot to the corresponding roadmap node column, so users can see which task phase has approved recurrence coverage.
159
- - If no user-approved tests exist for a route, do not show a test route for that route at all.
160
- - Empty test slots should be thin timeline placeholders only when the route has at least one approved test elsewhere; do not create blank cards or "no test" messages.
161
- - Test route items must come from user-approved tests with an explicit `Run policy`, approved task-case checkpoints, or approved test registry entries. Do not treat ordinary linked bad-case guards as tests unless the user approved them as tests and a run policy is recorded. Do not populate the human-facing test route from roadmap node development logs.
178
+ - Do not show a compact test route, bad-case lane, recurrence-check lane, or node-linked test notes in the default overview. Keep the overview focused on route nodes only.
179
+ - Keep user-approved tests, bad cases, and recurrence checks in clicked node details, the Test Hub page, and agent-readable exports. Do not create a second row of cards under route nodes for this information.
180
+ - Do not show empty test slots, "no test" placeholders, or compact test cards under roadmap nodes.
162
181
  - Show parent/fork markers only for side routes whose parent node belongs to another route. Never show a fork marker on the Main route merely because a later main node references an earlier node.
163
182
  - In branch overview, visually align each side route's starting position to the parent node's visible position on its parent route. Do not render every side route from the first column.
164
183
  - Place branch route titles, parent chips, and checkpoint text near that branch's first visible card by using the same spacer/grid coordinate as the branch cards. Do not leave branch labels pinned to the far-left edge when the branch starts later.
@@ -211,6 +230,7 @@ Do not use roadmap.html as a context source. The HTML file is only a human-facin
211
230
  For Codex context intake, checkpointing, bad-case review, and task switching, read the source context files directly:
212
231
 
213
232
  - `.codex/context/index.md`
233
+ - `.codex/context/user-messages.md`
214
234
  - `.codex/context/roadmap.md`
215
235
  - `.codex/context/bad-cases.md`
216
236
  - `.codex/context/tasks/<task-id>/context.md`
@@ -225,8 +245,8 @@ The HTML roadmap is a route-grouped board:
225
245
  1. A roadmap may contain multiple route groups, using `Branch:` on nodes. Missing `Branch:` means `Main`.
226
246
  2. Horizontal movement inside each route group follows that route's nodes over time.
227
247
  3. If there is only one route group, the default overview shows only the main route cards. Bad cases and test-chain notes stay in clicked node details and agent-readable context.
228
- 4. If there are multiple route groups, the overview first shows all route lines as a branch map with parent/fork markers and a compact test line aligned under the visible route nodes. Selecting a route may change the route focus state, but it must not open a separate always-visible bad-case/test-chain drilldown below the map.
229
- 5. Linked bad cases and verification chain should appear only as compact node-aligned route details, while full details stay in source context and agent-readable exports.
248
+ 4. If there are multiple route groups, the overview first shows all route lines as a branch map with parent/fork markers. Selecting a route may change the route focus state, but it must not open a separate always-visible bad-case/test-chain drilldown below the map.
249
+ 5. Linked bad cases and verification chain should appear only in clicked node details, the Test Hub page, source context, and agent-readable exports.
230
250
 
231
251
  Treat a single-route roadmap as the user's mainline summary first, not as a three-lane board. Treat multi-route roadmaps as route navigation first, with compact bad-case/test coverage scoped to the visible route. Use `Parent:` when a branch forks from an earlier node.
232
252
 
@@ -247,29 +267,31 @@ The core artifact is context, not scripts. Record enough context that a future C
247
267
 
248
268
  1. Test design is human-owned. Codex may run existing tests, execute user-provided checks, or draft a short proposed check, but it must not silently design durable tests, task cases, guard scripts, or broad verification plans on the user's behalf.
249
269
  2. When the user explicitly asks to create, write, generate, design, or add a test / testing task / task case, the first user-visible sentence must acknowledge that Context Guard recognized a test-creation request. Use the folder language and a compact style such as `测试创建识别:我会先把测试目标确认成一句话:从 <起点> 到 <终点>,主要验证 <风险/行为>。` This visible intake is required so the user knows the test mechanism activated.
250
- 3. Treat a user-approved test case, user-provided reproduction, native project test, or existing recorded guard as the source of truth. Codex can structure it, link it to bad cases, run it, and log results.
251
- 4. When the user creates or approves a test, register it with `Run policy: every-dev-completion` by default. This means every development turn that changes code, behavior, UI, workflow, docs with behavior rules, or project artifacts must run that test before the final answer.
252
- 5. Only change a registered test's run policy when the user says it does not need to run every time. Supported policies: `every-dev-completion`, `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or a user-defined cadence. Record the user's reason next to the policy.
253
- 6. Prefer the existing `Guard / verification` note on the bad case: it may be a command, native test, manual check, screenshot comparison, log invariant, reproduction note, or script.
254
- 7. Reuse recorded commands, tests, and manual checks before proposing any new check.
255
- 8. After a test design is user-approved, prefer automation when the check can be safely and repeatably encapsulated as a native project test, command, script, or task-case runner. The goal is to reduce future Codex judgment: already-built tests should run automatically and report structured results.
256
- 9. Do not turn every bad case into a script. Create or update a durable script only when the user explicitly asked for it or confirmed the proposed test design, and the check is repeatable, valuable, and cheaper than repeatedly reconstructing it.
257
- 10. If a user-approved script is justified and does not belong in the native test suite, place it under `.codex/context/bad-case-tests/`, for example `.codex/context/bad-case-tests/BC-YYYYMMDD-001.sh`.
258
- 11. Scripted tests and task-case runners should own their temporary workspace. On full success, they should clean up generated temporary files automatically. On failure, they should preserve the smallest useful artifacts, logs, screenshots, fixtures, or temp directory path needed for Codex to diagnose the bad case.
259
- 12. If an approved automated test fails, Codex must analyze the failed phase/checkpoint or artifact, record/update the bad case, fix the cause when within scope, and rerun the same approved test until it passes or an external blocker is reached.
260
- 13. If a test cannot proceed because of a non-actionable blocker such as missing credentials, unavailable external service, permission denial, hardware/resource limits, network outage, destructive-risk confirmation, or user-only domain judgment, stop the loop and ask or warn the user with the exact blocker and the preserved evidence path.
261
- 14. Record why the chosen guard is enough. If the guard is manual-only, record the exact manual steps and why automation is not currently worth it.
262
- 15. Add tags and frequency notes for recurring bad cases, such as `#hot`, `#flaky`, `#ui`, `#data-loss`, or `#route-risk`, so Codex can quickly spot high-risk patterns.
263
- 16. Treat each resolved bad case guard as a red-capable recurrence signal: it must be able to catch the original symptom if it returns, not merely prove that related code ran.
264
- 17. For resolved or recurred cases, record `Guard type`, `Red condition`, `Green condition`, `Expected failure reason`, and `Run policy` in addition to `Guard / verification`.
265
- 18. When a new guard is needed but no human-approved design exists, write it as `proposed` with a short confirmation prompt instead of creating or running a durable test.
266
- 19. Promote repeated or high-frequency bad cases into fixed pressure checks only after the user approves that pressure check.
267
- 20. Keep verification proportional for agent-selected ad hoc checks, but do not use the verification budget to skip user-approved tests whose policy is `every-dev-completion`.
268
- 21. Default verification budget for ordinary turns is one primary check plus the complete human-approved `every-dev-completion` test set, plus at most two extra relevant bad-case guards. Exceed this only when the user-approved always-run set requires it, or for high-risk, shared, release, security/data-loss, or user-requested exhaustive work.
269
- 22. Select extra guards by overlap: changed files, feature area, route branch, bad-case tags, and the original user-visible symptom. Skip unrelated resolved cases unless their `Run policy` says they must run every development completion.
270
- 23. Prefer existing native project tests or one focused symptom check over adding new `.codex/context/bad-case-tests/` scripts. Add a new script only after explicit user approval.
271
- 24. Do not let guard work become an agent-created testing loop. If the user-approved always-run suite is too broad or expensive, ask the user which tests to demote instead of silently skipping or redesigning it.
272
- 25. Do not import a strict test-first workflow into every task. Context Guard requires credible evidence, not always a newly written failing test. For urgent bugs, remote patches, UI polish, or small documentation/skill edits, existing user evidence plus one targeted verification is enough unless the user asks for TDD or the risk is high.
270
+ 3. When a task changes user-visible behavior, fixes a recurring bug, touches a multi-step workflow, changes UI/HTML/browser behavior, updates remote/service operations, or enters goal-mode work, gently remind the user that this may be a good moment to create a reusable test task. Keep the reminder optional and one sentence, for example: `这个流程后续可能会复发,要不要把它沉淀成一个测试任务?`
271
+ 4. Do not ask for a test task on every turn. Skip the reminder when the task is purely conversational, already covered by an approved test, very small, or the user has recently declined/demoted similar coverage.
272
+ 5. Treat a user-approved test case, user-provided reproduction, native project test, or existing recorded guard as the source of truth. Codex can structure it, link it to bad cases, run it, and log results.
273
+ 6. When the user creates or approves a test, register it with `Run policy: every-dev-completion` by default. This means every development turn that changes code, behavior, UI, workflow, docs with behavior rules, or project artifacts must run that test before the final answer.
274
+ 7. Only change a registered test's run policy when the user says it does not need to run every time. Supported policies: `every-dev-completion`, `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or a user-defined cadence. Record the user's reason next to the policy.
275
+ 8. Prefer the existing `Guard / verification` note on the bad case: it may be a command, native test, manual check, screenshot comparison, log invariant, reproduction note, or script.
276
+ 9. Reuse recorded commands, tests, and manual checks before proposing any new check.
277
+ 10. After a test design is user-approved, prefer automation when the check can be safely and repeatably encapsulated as a native project test, command, script, or task-case runner. The goal is to reduce future Codex judgment: already-built tests should run automatically and report structured results.
278
+ 11. Do not turn every bad case into a script. Create or update a durable script only when the user explicitly asked for it or confirmed the proposed test design, and the check is repeatable, valuable, and cheaper than repeatedly reconstructing it.
279
+ 12. If a user-approved script is justified and does not belong in the native test suite, place it under `.codex/context/bad-case-tests/`, for example `.codex/context/bad-case-tests/BC-YYYYMMDD-001.sh`.
280
+ 13. Scripted tests and task-case runners should own their temporary workspace. On full success, they should clean up generated temporary files automatically. On failure, they should preserve the smallest useful artifacts, logs, screenshots, fixtures, or temp directory path needed for Codex to diagnose the bad case.
281
+ 14. If an approved automated test fails, Codex must analyze the failed phase/checkpoint or artifact, record/update the bad case, fix the cause when within scope, and rerun the same approved test until it passes or an external blocker is reached.
282
+ 15. If a test cannot proceed because of a non-actionable blocker such as missing credentials, unavailable external service, permission denial, hardware/resource limits, network outage, destructive-risk confirmation, or user-only domain judgment, stop the loop and ask or warn the user with the exact blocker and the preserved evidence path.
283
+ 16. Record why the chosen guard is enough. If the guard is manual-only, record the exact manual steps and why automation is not currently worth it.
284
+ 17. Add tags and frequency notes for recurring bad cases, such as `#hot`, `#flaky`, `#ui`, `#data-loss`, or `#route-risk`, so Codex can quickly spot high-risk patterns.
285
+ 18. Treat each resolved bad case guard as a red-capable recurrence signal: it must be able to catch the original symptom if it returns, not merely prove that related code ran.
286
+ 19. For resolved or recurred cases, record `Guard type`, `Red condition`, `Green condition`, `Expected failure reason`, and `Run policy` in addition to `Guard / verification`.
287
+ 20. When a new guard is needed but no human-approved design exists, write it as `proposed` with a short confirmation prompt instead of creating or running a durable test.
288
+ 21. Promote repeated or high-frequency bad cases into fixed pressure checks only after the user approves that pressure check.
289
+ 22. Keep verification proportional for agent-selected ad hoc checks, but do not use the verification budget to skip user-approved tests whose policy is `every-dev-completion`.
290
+ 23. Default verification budget for ordinary turns is one primary check plus the complete human-approved `every-dev-completion` test set, plus at most two extra relevant bad-case guards. Exceed this only when the user-approved always-run set requires it, or for high-risk, shared, release, security/data-loss, or user-requested exhaustive work.
291
+ 24. Select extra guards by overlap: changed files, feature area, route branch, bad-case tags, and the original user-visible symptom. Skip unrelated resolved cases unless their `Run policy` says they must run every development completion.
292
+ 25. Prefer existing native project tests or one focused symptom check over adding new `.codex/context/bad-case-tests/` scripts. Add a new script only after explicit user approval.
293
+ 26. Do not let guard work become an agent-created testing loop. If the user-approved always-run suite is too broad or expensive, ask the user which tests to demote instead of silently skipping or redesigning it.
294
+ 27. Do not import a strict test-first workflow into every task. Context Guard requires credible evidence, not always a newly written failing test. For urgent bugs, remote patches, UI polish, or small documentation/skill edits, existing user evidence plus one targeted verification is enough unless the user asks for TDD or the risk is high.
273
295
 
274
296
  Preferred bad-case source format is one `### BC-YYYYMMDD-001: Title` section per case with bullet fields below it. If a legacy or interrupted session wrote loose bullet blocks with fields such as `ID`, `Title`, `Status`, and `Nodes`, the roadmap projector should still recognize those blocks instead of showing "No linked bad cases"; normalize them back to formal sections when editing the source file.
275
297
 
@@ -277,6 +299,8 @@ Recording and display must stay connected. Every bad case that should appear on
277
299
 
278
300
  ### Feature-Oriented Test Chains
279
301
 
302
+ When designing or updating a feature chain, read `references/feature-chain-methodology.md`.
303
+
280
304
  Use as few durable test chains as possible to cover as many bad-case recurrence checks as possible. The durable testing unit is a feature or workflow chain, not an individual bad case.
281
305
 
282
306
  A feature chain has:
@@ -289,9 +313,11 @@ A feature chain has:
289
313
  When a new bad case appears:
290
314
 
291
315
  1. First ask whether it belongs to an existing feature chain.
292
- 2. If it does, attach the bad case to the matching chain node and strengthen that node's checkpoint instead of creating a separate long-lived test.
293
- 3. If no existing chain matches the feature or workflow, propose a new feature chain with the same short human-facing confirmation style as task cases.
294
- 4. Only approve or automate the chain after user confirmation, unless the user explicitly provided the exact test to implement.
316
+ 2. Use `feature-chain-plan --query <bad-case or feature text>` as the default intake before proposing a new chain. This command is read-only; it says whether to review an existing chain or propose a new chain, shows match evidence for strong candidates, and must not approve or mutate coverage.
317
+ 3. Use `feature-chain-suggest --query <bad-case or feature text>` when you only need raw candidate chains/checkpoints.
318
+ 4. If a strong existing-chain match is semantically correct, attach the bad case to the matching chain node and strengthen that node's checkpoint instead of creating a separate long-lived test.
319
+ 5. If no existing chain matches the feature or workflow, propose a new feature chain with the same short human-facing confirmation style as task cases.
320
+ 6. Only approve or automate the chain after user confirmation, unless the user explicitly provided the exact test to implement.
295
321
 
296
322
  Store feature chains in `.codex/context/test-hub/feature-chains.json`. Use:
297
323
 
@@ -300,8 +326,7 @@ python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-ad
300
326
  --root <project> \
301
327
  --title "GPU 监控按钮" \
302
328
  --entry "点击 GPU 监控按钮" \
303
- --exit-check "打开包含有效 grafana_url 的监控页" \
304
- --command-text "<approved command>"
329
+ --exit-check "打开包含有效 grafana_url 的监控页"
305
330
 
306
331
  python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-attach-bc \
307
332
  --root <project> \
@@ -309,9 +334,66 @@ python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-at
309
334
  --node-title "后端返回监控 URL" \
310
335
  --bad-case BC-YYYYMMDD-001 \
311
336
  --check "grafana_url 不为空且前端没有卡住"
337
+
338
+ python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-suggest \
339
+ --root <project> \
340
+ --query "GPU 监控点击后没有打开 grafana_url"
341
+
342
+ python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-plan \
343
+ --root <project> \
344
+ --query BC-YYYYMMDD-001
312
345
  ```
313
346
 
314
- Approved feature chains with `Run policy: every-dev-completion` are included in `dev-complete` and the Stop hook. Relevant-only or proposed chains remain context until the user approves or asks to run them.
347
+ `feature-chain-plan` is the safest first step for Codex. If it finds a strong existing match, it prints `action: review-existing-chain`, the candidate chain/checkpoint, short `match evidence`, and an after-confirmation attach-command skeleton. If it does not find a strong match, it prints `action: propose-new-chain` plus a compact confirmation prompt and a `feature-chain-propose` command skeleton, not `feature-chain-add`. The skeleton must include one concrete checkpoint (`--node-title` and `--check`) plus either `--bad-cases` when the query came from real bad-case coverage, or `--coverage-pending-reason` when it is only a user-described test target. When the query is a `BC-...` ID, the confirmation prompt must use the bad-case title or display summary instead of the opaque ID so the user can judge the business flow. It must not create feature chains, attach bad cases, or approve automation; the skeleton is only for use after user confirmation.
348
+
349
+ For natural-language test requests, the confirmation prompt should strip request scaffolding such as "写一个测试", "创建测试任务", "检验", "验证", or "每次开发完成后" and keep the short business behavior/risk phrase. If the user explicitly describes a workflow as "从 A 到 B,主要验证 C", preserve that entry/exit/risk shape in the confirmation prompt instead of replacing it with generic wording, and use the stated entry/exit to prefill the after-confirmation command skeleton. The stated risk may be printed as a suggested checkpoint, but it remains a confirmation aid only. Preserve the original query in debug output if useful, but the user-facing confirmation prompt should read like a compact test goal rather than a copied chat sentence.
350
+
351
+ `feature-chain-suggest` may receive a natural-language symptom or a `BC-...` ID. When a `BC-...` ID is provided, it should expand the query from `bad-cases.md` using that case's title, summary, phenomenon, trigger, root cause, and tags before scoring candidate chains. The command remains read-only; it suggests where to attach coverage but must not mutate the registry.
352
+
353
+ `feature-chain-add` defaults to `status: proposed` even when a command is supplied. A proposed chain is non-executable context and must not enter the always-run suite. Do not use `feature-chain-add --test-status approved` for `every-dev-completion` chains; the command must refuse that path because it skips the human confirmation and approval dry-run gates. Create or keep the chain as `proposed`, then use `feature-chain-approve` after the user confirms the flow and automation.
354
+
355
+ After the user confirms a candidate flow but before approving automation, prefer `feature-chain-propose --title ... --entry ... --exit-check ... --node-title ... --bad-cases ... --check ...`. It creates one proposed chain with seed bad-case coverage and a checkpoint, but writes no executable command and must not run in `dev-complete`.
356
+
357
+ If the user confirms a feature-oriented test target before any concrete bad case exists, use `feature-chain-propose` with `--coverage-pending-reason <why coverage is pending>` instead of inventing a fake bad case. This records a proposed chain and checkpoint for future attachment, but it is not active coverage, cannot be approved for `every-dev-completion`, and should be revisited when a real bad case or user-provided recurrence risk is available.
358
+
359
+ When a later bad case matches a `coverage_pending_reason` chain or checkpoint, prefer attaching it to that proposed checkpoint with `feature-chain-attach-bc` instead of creating a new chain. After the first real bad case is attached, the pending-coverage note should be removed from that checkpoint because it now has concrete recurrence coverage.
360
+
361
+ Before approving an automation command for a proposed chain, use `feature-chain-dry-run --chain-id <id> --command-text "<candidate command>"` when the command or checkpoint markers are not yet proven. Dry run checks the proposed chain's checkpoint markers, cleans success artifacts, preserves failure evidence, and must not mutate the chain, approve it, or add it to `dev-complete`.
362
+
363
+ After the user confirms the business flow and test design, promote the same chain with `feature-chain-approve --chain-id <id> --command-text "<approved command>"`. This is the only supported path for turning a proposed `every-dev-completion` feature chain into an approved automated test. The approval gate must refuse chains that have no checkpoint node, no checkpoint check text, no linked bad-case coverage, or no automated command. It must also run an approval dry-run before mutating the registry; if required checkpoint markers are missing, unknown, failing, blocked, or timed out, the chain must remain `proposed`. Do not bypass approval by hand-editing `feature-chains.json`, by using `feature-chain-add --test-status approved`, or by creating a duplicate approved chain.
364
+
365
+ Feature-chain commands may emit checkpoint markers so Test Hub can localize failures without Codex reinterpreting the whole log:
366
+
367
+ ```text
368
+ CG_CHECKPOINT:<checkpoint title or id>:PASS
369
+ CG_CHECKPOINT:<checkpoint title or id>:FAIL:<short reason>
370
+ ```
371
+
372
+ For feature-chain tests, any `FAIL` marker is a failed test even if the command exits 0. Use these markers for important workflow phases so the final report says which checkpoint broke, not only that the chain failed.
373
+
374
+ Checkpoint marker names must match a registered feature-chain checkpoint title or id. An unknown marker is a test-chain failure because it means the automated script has drifted from the approved feature-chain design. For non-English checkpoint titles, preserve the readable title; do not collapse distinct checkpoints into generic IDs such as `test`.
375
+
376
+ Approved feature-chain commands must emit a marker for every registered checkpoint unless that checkpoint is explicitly marked optional (`optional: true` or `required: false`). A missing marker is a test-chain failure because the run did not prove that the approved workflow step was exercised.
377
+
378
+ If a registered checkpoint should not be required on every run, update the existing checkpoint with `feature-chain-set-checkpoint --chain-id <id> --node-title <checkpoint> --required optional --reason <short reason>`. Do not hand-edit `feature-chains.json`, remove the checkpoint, or silently ignore missing markers. Use `--required required` when the user later wants that checkpoint restored to every-run coverage.
379
+
380
+ Use `feature-chain-list --verbose` to audit each chain's required/optional checkpoint counts and optional reasons without opening `feature-chains.json`.
381
+
382
+ Use `feature-chain-summary` as the fast coverage view before creating or updating test coverage. It shows the small map of feature chain -> checkpoint -> covered bad-case titles, plus any pending checkpoints waiting for a real bad case. It also prints lightweight reuse signals, such as coverage density and whether one workflow already covers multiple bad cases. Prefer this summary when deciding whether one existing chain can absorb a new bad case instead of creating another test.
383
+
384
+ Use `feature-chain-overlap` before approving automation or when several proposed chains look similar. It is a read-only duplicate-chain audit: it compares chain entries, exits, checkpoints, and linked bad cases, then prints candidate pairs that may belong to the same workflow. If it reports overlap, review whether one chain should absorb the other before creating or approving another always-run test.
385
+
386
+ Use `feature-chain-coverage` to see which registered bad cases are already covered by feature-chain checkpoints and which remain unassigned candidates. For each visible unassigned candidate, it may show the most likely existing chain and checkpoint based on the bad-case semantics. Strong suggestions should include short `match evidence` terms so Codex and the user can judge why the candidate was suggested. This is a planning aid, not a mandate to create tests for every unassigned bad case, and it must not automatically attach coverage or approve tests.
387
+
388
+ Use `feature-chain-candidates` when many bad cases remain unassigned and Codex needs a small set of feature-chain candidates instead of a long case list. It groups unassigned bad cases by shared feature tags, prefers more specific tag combinations over broad single tags, hides cases already covered by existing feature chains, suppresses low-value repeated candidate groups, and prints compact confirmation prompts with `new coverage` counts. Treat the output as planning hints only: ask the user which candidate flow is real before creating a proposed chain or automation.
389
+
390
+ When an approved feature-chain test fails or is blocked, the completion report must be actionable without rereading the full log: include the feature-chain title, the failed or missing checkpoint, the short reason, and the preserved evidence path. Keep long logs as evidence only.
391
+
392
+ If the user later says an approved feature chain should not run every time, use `feature-chain-set-policy --chain-id <id> --run-policy <policy> --reason <short reason>` on the existing chain. Do not delete the chain or create a duplicate just to change cadence.
393
+
394
+ Use `validate-feature-chains` after editing feature-chain records or before relying on a new chain as durable coverage. This validation is a quality gate only: it checks structural readiness, checkpoint coverage, approved automation, and artifact policy, but it does not decide whether the business test should exist.
395
+
396
+ Approved feature chains with `Run policy: every-dev-completion` are included in `dev-complete` and the Stop/SubagentStop hooks. Relevant-only or proposed chains remain context until the user approves or asks to run them.
315
397
 
316
398
  ### Task-Oriented Test Cases
317
399
 
@@ -411,13 +493,67 @@ Use the Test Hub as the automation control plane for approved tests. The hub col
411
493
  - A registry entry with `status: approved | active | stable` and `run_policy: every-dev-completion` is part of the always-run set.
412
494
  - Manage registry tests with `test-hub-list`, `test-hub-enable`, `test-hub-disable`, `test-hub-set-policy`, and `test-hub-remove`.
413
495
  - Use `show-test-hub` to write the stable read-only human-facing page at `.codex/context/test-hub/test-hub.html`.
414
- - The Test Hub HTML is a status page only. Do not make users start tests from HTML buttons; approved tests run from the Stop hook or `dev-complete`.
496
+ - The Test Hub HTML is a status page only. Do not make users start tests from HTML buttons; approved tests run from the Stop/SubagentStop hooks or `dev-complete`.
497
+ - In the Test Hub HTML, show required/optional checkpoint policy for approved feature-chain tests inside that test card. Do not show proposed feature chains, empty checkpoint placeholders, or ordinary bad-case guards as tests.
415
498
  - Task cases in `.codex/context/task-cases/` may also join the always-run set only when they are `approved | active | stable`, have `Run policy: every-dev-completion`, and include an automated entry command.
416
- - At development completion, the Stop hook should invoke `scripts/context_guard.py dev-complete --root <project>` so registered tests run automatically. If running manually, use `dev-complete` over hand-running tests one by one. Use `--jobs <n>` only when parallel execution is safe for the registered tests.
499
+ - At development completion, the Stop and SubagentStop hooks should invoke `scripts/context_guard.py dev-complete --root <project>` so registered tests run automatically. If running manually, use `dev-complete` over hand-running tests one by one. Use `--jobs <n>` only when parallel execution is safe for the registered tests.
417
500
  - `dev-complete` must report passed, failed, and blocked tests. On full success it should clean the run artifacts; on failure or blocker it should preserve evidence under `.codex/context/test-hub/runs/`.
418
501
  - If the user says a test should not run every time, update the registry or task case run policy instead of silently skipping it.
419
502
  - If no approved every-dev-completion tests exist, the hub should report that clearly and exit successfully; Codex should not invent tests to fill the gap.
420
503
 
504
+ ### Subagent Completion Fallback
505
+
506
+ Some Codex subagent transports may not fire project-local `SubagentStart` or `SubagentStop` hooks, and a subagent launched from a parent workspace may inherit the parent's working directory even when its product files live in a child folder. Treat hooks as supplemental and bind every spawned subagent to its real local project root immediately after `spawn_agent` returns:
507
+
508
+ ```bash
509
+ python3 ~/.agents/skills/context-guard/scripts/context_guard.py subagent-register \
510
+ --root <main/control workspace root> \
511
+ --agent-id <agent id> \
512
+ --project-root <subagent project root> \
513
+ --task "<ordinary product task>"
514
+ ```
515
+
516
+ Do not initialize or update Context Guard on an SSH server merely because the subagent edits remote files. `--project-root` is always the opened local Codex project folder; remote host/path belongs in task metadata.
517
+
518
+ If a completed agent is reused with `send_input` for another development round, run `subagent-register` again with the same agent ID and project root before sending or immediately after sending the new request. This marks the assignment active again without losing prior completion history.
519
+
520
+ When the main agent uses `wait_agent`, receives a completion notification, or closes a completed agent, pass the exact completion output back through the registered assignment:
521
+
522
+ ```bash
523
+ python3 ~/.agents/skills/context-guard/scripts/context_guard.py subagent-complete \
524
+ --root <main/control workspace root> \
525
+ --agent-id <agent id> \
526
+ --summary "<short completion summary>" \
527
+ --evidence-file <file containing the exact subagent final output>
528
+ ```
529
+
530
+ For simple integrations that cannot write an evidence file, `--summary` remains valid, but it loses detail and should not be preferred. The completion command resolves the registered child root, initializes folder context there, records one idempotent handoff/checkpoint, runs Test Hub `dev-complete`, archives concrete repair evidence, and attempts feature-chain auto-proposal from recorded bad cases. Replaying the same successful completion evidence must not duplicate roadmap nodes or bad cases. If Test Hub fails or blocks, the same completion may be retried after the fix.
531
+
532
+ When a subagent actually found or fixed a problem, its final output should include this compact evidence block in the folder language. Include it only for observed or reproducible problems, not speculative risks:
533
+
534
+ ```text
535
+ CG_BAD_CASE: <short user-visible problem>
536
+ CG_PHENOMENON: <what actually happened>
537
+ CG_TRIGGER: <minimal reproduction>
538
+ CG_CAUSE: <known cause, if confirmed>
539
+ CG_FIX: <what changed>
540
+ CG_VERIFICATION: <real post-fix evidence>
541
+ CG_SCOPE: <feature/workflow>
542
+ ```
543
+
544
+ The main agent must pass this exact output to `subagent-complete`; do not rewrite it into a generic "implemented X and smoke passed" summary. Context Guard records verified evidence as a concrete resolved bad case, or as open when verification is absent. It never turns that evidence into approved automation without user confirmation.
545
+
546
+ On Codex clients that emit the documented `SubagentStop` payload, the hook reads `agent_id` and `last_assistant_message` and invokes this same completion command automatically. The explicit main-agent call remains mandatory when the transport omits the hook, the completion notification arrives without a local hook record, or a resumed agent disappears before its final message is delivered. Native and fallback paths share the same completion fingerprint, so a successful event is archived once.
547
+
548
+ The completion-risk audit is not a license to invent a test suite. It exists to catch the failure mode where a subagent only reports "done / smoke passed" while a stateful, persistent, reset, replay, copy/export, or input-validation workflow has no bad-case input at all. When a completed subagent summary contains multiple high-risk workflow cues and the project has no parsable bad cases, `subagent-complete` may write one `risk-audit` bad-case candidate. Keep it conservative:
549
+
550
+ - Do not create a bad case for every small concern.
551
+ - Prefer real user-reported, observed, or reproducible bad cases when they exist.
552
+ - Treat risk-audit entries as open candidates until a later run confirms, merges, or closes them.
553
+ - Let `feature-chain-auto-propose` group those candidates into proposed feature chains; do not approve or execute them without user confirmation.
554
+
555
+ Do not rely on the subagent's final prose as proof that Context Guard ran. Verify the project has `.codex/context/` and, when approved tests exist, a current `.codex/context/test-hub/last-run.json`.
556
+
421
557
  Each reusable check should answer four questions:
422
558
 
423
559
  - Red condition: what output, visual state, error, or assertion means the bad case has recurred
@@ -426,7 +562,7 @@ Each reusable check should answer four questions:
426
562
  - Guard type: script, native-test, manual, browser-screenshot, browser-dom, curl, cli, prompt, log-invariant, fixture, unit, integration, e2e, or another concise type
427
563
  - Artifact policy: cleanup-on-pass, preserve-on-fail, or manual-preserve, including the log/temp path convention when relevant
428
564
 
429
- In roadmap overview, single-route pages hide the Test Chain lane by default, while multi-route compact test routes must be generated only from user-approved tests with explicit `Run policy`, approved task-case checkpoints, or approved test registry entries. Ordinary linked bad-case guards are recurrence context, not user-approved tests, unless the user approved that guard as a test and the source record has a run policy. Do not fill user-facing test coverage from roadmap node `Test chain:` history. Roadmap node `Test chain:` may keep compact checkpoint evidence in source/details, but it is not the primary bad-case recurrence chain.
565
+ In roadmap overview, both single-route and multi-route pages hide Test Chain lanes, compact test routes, and bad-case lanes by default. Ordinary linked bad-case guards are recurrence context, not user-approved tests, unless the user approved that guard as a test and the source record has a run policy. Do not fill user-facing overview coverage from roadmap node `Test chain:` history. Roadmap node `Test chain:` may keep compact checkpoint evidence in source/details, but it is not the primary bad-case recurrence chain.
430
566
 
431
567
  When a task-oriented case exists for the changed workflow, prefer running or following that scenario and its checkpoint logs over running several disconnected bad-case scripts. Stay within the verification budget by selecting the smallest relevant task case and only the checkpoint guards that overlap the current change.
432
568
 
@@ -446,15 +582,16 @@ Run this before any substantive answer or action.
446
582
 
447
583
  1. Ensure the folder-scoped context skeleton exists when this is the first task in a Codex folder.
448
584
  2. Read `.codex/context/preferences.json`. If the record language is unset, ask the user to choose the context record language and store it before adding substantive context.
449
- 3. Decide whether the user's latest message continues the current task, starts a substantially different task, reports a bad case, or changes expected behavior.
450
- 4. Locate and read `.codex/context/index.md`, `.codex/context/roadmap.md`, `.codex/context/bad-cases.md`, and the relevant task folder if they exist.
451
- 5. If the request changes direction, park the previous task context before switching.
452
- 6. If the user reports a bad case, add or update the matching bad-case entry before fixing it.
453
- 7. Identify context entries relevant to the files, features, tests, or workflows likely to be touched.
454
- 8. Keep relevant context in mind while planning and editing.
455
- 9. Do not use the generated HTML roadmap as the context source.
456
- 10. In goal mode, call `get_goal` when available and align the active task with the goal objective before continuing work.
457
- 11. At the start of the user-visible answer, include a compact intake statement when useful: `Context intake: continuing <task>`, `Context intake: parked <task>, starting <task>`, `Bad-case intake: recorded BC-...`, `测试创建识别:...`, or `Context intake: no active context`.
585
+ 3. Preserve the latest user prompt in `.codex/context/user-messages.md` when it contains durable context; if it contains a secret, store only a redacted pointer in public context and keep raw values local-only under `.codex/context/private/`.
586
+ 4. Decide whether the user's latest message continues the current task, starts a substantially different task, reports a bad case, or changes expected behavior.
587
+ 5. Locate and read `.codex/context/index.md`, `.codex/context/user-messages.md`, `.codex/context/roadmap.md`, `.codex/context/bad-cases.md`, and the relevant task folder if they exist.
588
+ 6. If the request changes direction, park the previous task context before switching.
589
+ 7. If the user reports a bad case, add or update the matching bad-case entry before fixing it.
590
+ 8. Identify context entries relevant to the files, features, tests, or workflows likely to be touched.
591
+ 9. Keep relevant context in mind while planning and editing.
592
+ 10. Do not use the generated HTML roadmap as the context source.
593
+ 11. In goal mode, call `get_goal` when available and align the active task with the goal objective before continuing work.
594
+ 12. At the start of the user-visible answer, include a compact intake statement when useful: `Context intake: continuing <task>`, `Context intake: parked <task>, starting <task>`, `Bad-case intake: recorded BC-...`, `测试创建识别:...`, `User-message memory: saved`, or `Context intake: no active context`.
458
595
 
459
596
  ### During Work
460
597
 
@@ -498,7 +635,7 @@ Run this before the final answer whenever Codex changed code, generated artifact
498
635
  5. Do not end with only string/DOM assertions when the risk is visual. If visual inspection is blocked, say exactly what was blocked, record the residual risk, and avoid claiming visual polish was verified.
499
636
  6. If the self-check reveals a new or recurring bad case, record it immediately, fix it before the final answer unless the user pauses, and rerun the self-check.
500
637
  7. Record the self-check evidence in the relevant roadmap node, bad-case entry, or task context using the folder language preference.
501
- 8. Treat the Stop hook as a completion reliability gate, not a decorative reminder. If the hook asks for verification evidence, branch-task handling, or BC summary, satisfy it before finalizing.
638
+ 8. Treat the Stop/SubagentStop hook as a completion reliability gate, not a decorative reminder. If the hook asks for verification evidence, branch-task handling, or BC summary, satisfy it before finalizing.
502
639
  9. Do not claim a bug is fixed because a build passed or a helper restarted. Verify the original user-visible symptom with the smallest real check that could falsify the claim.
503
640
  10. If the work touched frontend, browser, UI binding, routing, HTML/CSS, or visual state, the self-check must include Browser/plugin/screenshot/DOM evidence tied to the original symptom, or an explicit blocker and residual risk.
504
641
  11. Keep the self-check inside the verification budget unless risk is high. If the budget would be exceeded, prefer the original symptom check and the highest-risk relevant guard, then record the skipped checks as unrelated or deferred.
@@ -509,40 +646,42 @@ Run this before the final answer whenever Codex changed code, generated artifact
509
646
  Run this before every final answer.
510
647
 
511
648
  1. Re-read the project context index and relevant task folder.
512
- 2. Re-read the route map and make a roadmap checkpoint decision. If this turn changed direction, made a durable decision, fixed a problem, created a branch/fork, reached a user-visible milestone, or refreshed a stale route, create or update one concise node. If none of those apply, do not create a node; mention that no roadmap node was needed.
513
- 3. Update the active task summary with key decisions, bad cases, open questions, and next step.
514
- 4. If the task direction changed this turn, ensure the previous task is parked and the new task is current.
515
- 5. Run the End-of-Work Self-Check for the changed behavior or artifact before claiming success.
516
- 6. Select only the bad-case entries whose scope clearly overlaps the changed code, feature, route, or user-visible symptom. Do not select all resolved cases merely because they have recorded guards.
517
- 7. Run the full human-approved test registry whose `Run policy` is `every-dev-completion`. This includes approved task cases, approved native test commands, approved guard scripts, and approved manual/prompt checks. If any cannot run, record the blocker and residual risk; do not imply the always-run suite passed.
649
+ 2. Re-read `.codex/context/user-messages.md` and ensure the latest durable user wording has been promoted into the active task context, bad-case entry, or roadmap `User request:` when relevant.
650
+ 3. Re-read the route map and make a roadmap checkpoint decision. If this turn changed direction, made a durable decision, fixed a problem, created a branch/fork, reached a user-visible milestone, or refreshed a stale route, create or update one concise node. If none of those apply, do not create a node; mention that no roadmap node was needed.
651
+ 4. Update the active task summary with key decisions, bad cases, open questions, and next step.
652
+ 5. If the task direction changed this turn, ensure the previous task is parked and the new task is current.
653
+ 6. Run the End-of-Work Self-Check for the changed behavior or artifact before claiming success.
654
+ 7. Select only the bad-case entries whose scope clearly overlaps the changed code, feature, route, or user-visible symptom. Do not select all resolved cases merely because they have recorded guards.
655
+ 8. Run the full human-approved test registry whose `Run policy` is `every-dev-completion`. This includes approved task cases, approved native test commands, approved guard scripts, and approved manual/prompt checks. If any cannot run, record the blocker and residual risk; do not imply the always-run suite passed.
518
656
  - For automated approved tests, run the registered command/script rather than reconstructing the test manually.
519
657
  - If all approved automated tests pass, ensure their temp files were cleaned or note the script's cleanup policy.
520
658
  - If any approved test fails, preserve evidence, analyze the failed phase/checkpoint as a bad case, fix what is in scope, and rerun until the test passes or a non-actionable blocker is reached.
521
659
  - If blocked by credentials, unavailable services, permissions, hardware/resource limits, network, destructive-risk confirmation, or user-only judgment, ask or warn the user instead of looping.
522
- 8. For tests whose policy is `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or a user-defined cadence, follow that policy exactly and mention skipped items only when they overlap the current change or affect confidence.
523
- 9. If a relevant task-oriented case exists, use its phase/checkpoint flow as the primary verification and note which checkpoints covered the linked bad cases. Otherwise re-run or re-perform the recorded guard for the highest-risk selected resolved entries, staying within the default budget of one primary check plus the complete always-run test set plus at most two extra relevant bad-case guards unless this is high-risk or the user requested exhaustive verification. Use the existing context, command, native test, script, screenshot/manual check, or visual inspection first.
524
- 10. If no recorded human-approved guard exists, do not invent active coverage. Use the lightest credible evidence for this turn, record any new durable guard as `proposed`, and ask for user confirmation with the short business-facing format before writing durable scripts or marking it approved. Existing user screenshots/logs/reproductions may serve as the red condition; a new failing test is optional, not mandatory.
525
- 11. If a resolved bad case recurs:
660
+ 9. For tests whose policy is `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or a user-defined cadence, follow that policy exactly and mention skipped items only when they overlap the current change or affect confidence.
661
+ 10. If a relevant task-oriented case exists, use its phase/checkpoint flow as the primary verification and note which checkpoints covered the linked bad cases. Otherwise re-run or re-perform the recorded guard for the highest-risk selected resolved entries, staying within the default budget of one primary check plus the complete always-run test set plus at most two extra relevant bad-case guards unless this is high-risk or the user requested exhaustive verification. Use the existing context, command, native test, script, screenshot/manual check, or visual inspection first.
662
+ 11. If no recorded human-approved guard exists, do not invent active coverage. Use the lightest credible evidence for this turn, record any new durable guard as `proposed`, and ask for user confirmation with the short business-facing format before writing durable scripts or marking it approved. Existing user screenshots/logs/reproductions may serve as the red condition; a new failing test is optional, not mandatory.
663
+ 12. If a resolved bad case recurs:
526
664
  - Mark it `recurred`.
527
665
  - Explain why it recurred: missed guard, incomplete fix, route conflict, test gap, refactor side effect, environment drift, or unknown.
528
666
  - Fix it immediately unless the user explicitly pauses the work or the recurrence is due to an approved technical route change.
529
667
  - Add or update the context and guard so the recurrence is easier to catch next time.
530
668
  - Re-run the verification and update the entry back to `resolved` only when evidence passes.
531
- 12. If a case is exempt because of a technical route change, mark it `superseded-by-route-change` and document the approved change.
532
- 13. If a bad case becomes frequent, add or update a high-frequency tag and warning note.
533
- 14. In goal mode, finish this checkpoint before calling `update_goal` to mark the goal complete or blocked.
534
- 15. If urgent or unrelated work is complete and a parked task exists, ask the user whether to resume the most relevant parked task.
535
- 16. If the Stop hook detects an explicit branch request, ensure the branch task and `Branch:`/`Parent:` roadmap node exist before finalizing.
536
- 17. If the Stop hook detects possible drift from the mainline architecture and no explicit branch exists, ask the user whether this should become a branch instead of silently continuing the mainline.
537
- 18. When a roadmap node is needed at turn end, prefer `scripts/context_guard.py checkpoint-roadmap-node --title <source title> --display-title <short human title> --user-request <short summary of the user's actual request> --progress-summary <readable current progress> --method-summary <readable method> --branch <Main or route> --level <major|checkpoint> --outcome <one-line source progress> --next-step <next>` instead of hand-editing. Use `create-branch-task` first when the user explicitly asks for a new branch.
538
- 19. Run `scripts/context_guard.py validate-bad-cases` only after updating bad-case entries, changing bad-case schema/renderer/hook behavior, or intentionally auditing the register; do not run it on unrelated code turns. Historical resolved cases without the new fields may remain warnings until touched; use `--strict` only when intentionally migrating or auditing all resolved cases.
539
- 20. Run `scripts/context_guard.py validate-roadmap-maintenance` only after adding route nodes, changing roadmap maintenance rules, or before showing the roadmap; if it reports too many hidden checkpoints after a route's latest visible node, promote or add a major node before finalizing.
669
+ 13. If a case is exempt because of a technical route change, mark it `superseded-by-route-change` and document the approved change.
670
+ 14. If a bad case becomes frequent, add or update a high-frequency tag and warning note.
671
+ 15. In goal mode, finish this checkpoint before calling `update_goal` to mark the goal complete or blocked.
672
+ 16. If urgent or unrelated work is complete and a parked task exists, ask the user whether to resume the most relevant parked task.
673
+ 17. If the Stop hook detects an explicit branch request, ensure the branch task and `Branch:`/`Parent:` roadmap node exist before finalizing.
674
+ 18. If the Stop hook detects possible drift from the mainline architecture and no explicit branch exists, ask the user whether this should become a branch instead of silently continuing the mainline.
675
+ 19. When a roadmap node is needed at turn end, prefer `scripts/context_guard.py checkpoint-roadmap-node --title <source title> --display-title <short human title> --user-request <short summary of the user's actual request> --progress-summary <readable current progress> --method-summary <readable method> --branch <Main or route> --level <major|checkpoint> --outcome <one-line source progress> --next-step <next>` instead of hand-editing. Use `create-branch-task` first when the user explicitly asks for a new branch.
676
+ 20. Run `scripts/context_guard.py validate-bad-cases` only after updating bad-case entries, changing bad-case schema/renderer/hook behavior, or intentionally auditing the register; do not run it on unrelated code turns. Historical resolved cases without the new fields may remain warnings until touched; use `--strict` only when intentionally migrating or auditing all resolved cases.
677
+ 21. Run `scripts/context_guard.py validate-roadmap-maintenance` only after adding route nodes, changing roadmap maintenance rules, or before showing the roadmap; if it reports too many hidden checkpoints after a route's latest visible node, promote or add a major node before finalizing.
540
678
 
541
679
  ## Completion Report
542
680
 
543
681
  At the end of every response, include a compact context summary when development work, context intake, task switching, bad-case intake, or a register was involved. Keep it one line for unrelated conversation.
544
682
 
545
683
  - Context folder used.
684
+ - User-message memory updated or skipped, with secret handling summarized only as redacted/local-only.
546
685
  - Current task index status.
547
686
  - Roadmap node updated, exported, or displayed.
548
687
  - If no roadmap node was created, the brief reason why it was not needed.