@michelj/context-guard 0.4.3 → 0.4.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (74) hide show
  1. package/README.md +103 -208
  2. package/README.zh-CN.md +103 -208
  3. package/SKILL.md +27 -678
  4. package/agents/openai.yaml +2 -2
  5. package/bin/context-guard-skill.js +159 -56
  6. package/bin/postinstall.js +1 -1
  7. package/hooks.json +11 -11
  8. package/package.json +9 -6
  9. package/prototype/workbench.html +4933 -0
  10. package/references/bug-record-template.md +37 -0
  11. package/references/context-template.md +14 -336
  12. package/scripts/context_guard.py +577 -7655
  13. package/scripts/context_guard_hook.py +179 -731
  14. package/scripts/map_owns.py +769 -0
  15. package/references/feature-chain-methodology.md +0 -228
  16. package/references/register-template.md +0 -85
  17. package/references/task-case-template.md +0 -63
  18. package/tests/BC-20260618-063.sh +0 -116
  19. package/tests/BC-20260618-065.sh +0 -66
  20. package/tests/BC-20260626-080.sh +0 -48
  21. package/tests/BC-20260626-081.sh +0 -40
  22. package/tests/BC-20260626-082.sh +0 -32
  23. package/tests/BC-20260626-083.sh +0 -66
  24. package/tests/BC-20260627-084.sh +0 -74
  25. package/tests/BC-20260630-086.sh +0 -50
  26. package/tests/BC-20260630-087.sh +0 -103
  27. package/tests/BC-20260630-088.sh +0 -32
  28. package/tests/BC-20260630-089.sh +0 -63
  29. package/tests/BC-20260701-090.sh +0 -84
  30. package/tests/BC-20260702-096.sh +0 -48
  31. package/tests/BC-20260706-098.sh +0 -66
  32. package/tests/BC-20260707-099.sh +0 -47
  33. package/tests/BC-20260707-100.sh +0 -46
  34. package/tests/BC-20260707-101.sh +0 -47
  35. package/tests/BC-20260707-102.sh +0 -68
  36. package/tests/BC-20260707-103.sh +0 -59
  37. package/tests/BC-20260707-104.sh +0 -103
  38. package/tests/BC-20260707-105.sh +0 -109
  39. package/tests/BC-20260707-106.sh +0 -80
  40. package/tests/BC-20260707-107.sh +0 -74
  41. package/tests/BC-20260707-108.sh +0 -48
  42. package/tests/BC-20260707-109.sh +0 -56
  43. package/tests/BC-20260707-110.sh +0 -71
  44. package/tests/BC-20260707-111.sh +0 -70
  45. package/tests/BC-20260707-112.sh +0 -45
  46. package/tests/BC-20260707-113.sh +0 -73
  47. package/tests/BC-20260707-115.sh +0 -77
  48. package/tests/BC-20260707-116.sh +0 -77
  49. package/tests/BC-20260707-118.sh +0 -115
  50. package/tests/BC-20260707-119.sh +0 -47
  51. package/tests/BC-20260707-120.sh +0 -60
  52. package/tests/BC-20260707-121.sh +0 -66
  53. package/tests/BC-20260707-122.sh +0 -48
  54. package/tests/BC-20260707-123.sh +0 -43
  55. package/tests/BC-20260707-124.sh +0 -56
  56. package/tests/BC-20260707-125.sh +0 -64
  57. package/tests/BC-20260707-126.sh +0 -80
  58. package/tests/BC-20260707-127.sh +0 -88
  59. package/tests/BC-20260707-129.sh +0 -59
  60. package/tests/BC-20260707-130.sh +0 -69
  61. package/tests/BC-20260707-131.sh +0 -140
  62. package/tests/BC-20260707-132.sh +0 -150
  63. package/tests/BC-20260707-133.sh +0 -70
  64. package/tests/BC-20260708-136.sh +0 -210
  65. package/tests/BC-20260708-137.sh +0 -106
  66. package/tests/BC-20260708-138.sh +0 -168
  67. package/tests/BC-20260708-139.sh +0 -79
  68. package/tests/BC-20260709-002.sh +0 -63
  69. package/tests/BC-20260709-003.sh +0 -239
  70. package/tests/BC-20260709-006.sh +0 -76
  71. package/tests/BC-20260709-008.sh +0 -168
  72. package/tests/BC-20260710-001.sh +0 -61
  73. package/tests/BC-20260710-002.sh +0 -111
  74. package/tests/npm-install-smoke.sh +0 -53
@@ -1,228 +0,0 @@
1
- # Feature Chain Methodology
2
-
3
- Use feature chains when the goal is to prevent fixed bad cases from reappearing without creating one durable test per bad case.
4
-
5
- ## Core Idea
6
-
7
- The durable test unit is a user-visible feature or workflow. A bad case is coverage attached to one checkpoint inside that workflow.
8
-
9
- ```text
10
- Feature chain
11
- Entry: the real trigger users or Codex will perform
12
- Checkpoint 1: expected intermediate state
13
- Covers: BC-...
14
- Checkpoint 2: expected transition or output
15
- Covers: BC-..., BC-...
16
- Exit check: strict final green condition
17
- ```
18
-
19
- ## Creation Rule
20
-
21
- When a bad case appears:
22
-
23
- 1. Identify the feature entry that can reproduce or guard the symptom.
24
- 2. Search existing feature chains for the same entry, workflow, component, route, or service. Use the read-only planning helper first:
25
-
26
- ```bash
27
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-plan --root <project> --query "<bad case or feature text>"
28
- ```
29
-
30
- The query can be natural language or a `BC-...` ID. The planner is read-only: it either says `action: review-existing-chain` with match evidence and an after-confirmation attach-command skeleton, or `action: propose-new-chain` with a compact user-confirmation prompt and a `feature-chain-propose` command skeleton. The proposed skeleton must contain a checkpoint and explicit coverage state: use `--bad-cases` for a real bad-case seed, or `--coverage-pending-reason` for a user-described test target that has no concrete bad case yet. If the input is only a `BC-...` ID, the prompt should describe the bad case by title or summary instead of asking the user to confirm an opaque ID. It must not create a feature chain, attach a bad case, or approve automation; run the skeleton only after user confirmation.
31
-
32
- 3. Use `feature-chain-suggest` when you only need raw candidate chains/checkpoints. If only a bad-case ID is available, the helper expands it from `bad-cases.md` first so matching uses the case title, summary, phenomenon, trigger, cause, and tags rather than the opaque ID alone.
33
- 4. Use `feature-chain-coverage` when you need the whole register view. It should show covered cases, unassigned candidates, and possible existing chains for visible candidates without changing the registry. Strong suggestions should include short match evidence terms; if the evidence is weak or unreadable, treat the suggestion as planning noise rather than coverage.
34
- 5. Use `feature-chain-candidates` when the unassigned list is too large. It groups unassigned bad cases by shared feature tags, prefers more specific tag combinations over broad single tags, suppresses repeated groups that add little new coverage, and proposes a small set of candidate feature chains. Treat `new coverage` as the key signal: a candidate with high total count but low new coverage may be a subcase of an earlier chain. It must not create, attach, or approve anything; it only helps choose which user-visible flow is worth designing.
35
- 6. Use `feature-chain-overlap` before approving automation, or whenever several proposed chains sound similar. It is read-only and flags pairs that likely describe the same workflow. If it reports overlap, merge the intent or extend one chain before creating another always-run test.
36
- 7. If a chain exists and the match is semantically correct, attach the bad case to the closest checkpoint and tighten that checkpoint.
37
- 8. If no chain exists, propose a new chain in one short business-facing sentence and wait for user confirmation before approving or automating it.
38
-
39
- For natural-language test requests, keep the user's business intent but remove the request wrapper before writing the confirmation prompt. For example, `写一个测试,检验每次开发完成后 Markdown 编辑器里的单行、多行和矩阵公式都能正常渲染` should become a compact subject such as `Markdown 编辑器里的单行、多行和矩阵公式能正常渲染`, not a verbatim copy of the whole chat sentence. This keeps the prompt useful for human confirmation while avoiding agent-invented workflow details.
40
-
41
- When the user already gives a workflow shape, preserve it. A request like `创建一个测试任务:从编辑器输入 Markdown 到预览正确渲染,主要验证公式渲染回归` should be confirmed as `从「编辑器输入 Markdown」到「预览正确渲染」,主要验证「公式渲染回归」`. Do not replace explicit entry/exit/risk wording with generic "相关入口到正确结果" language.
42
-
43
- For this explicit shape, `feature-chain-plan` may prefill the after-confirmation `feature-chain-propose` skeleton with the stated entry and exit check, and may print the stated risk as a suggested checkpoint. This is still only a confirmation aid: it must not create the chain, approve automation, or invent missing checkpoint details before the user confirms the business flow.
44
-
45
- CLI rule: `feature-chain-add` creates `status: proposed` by default. Do not treat this as an approved test. `feature-chain-add --test-status approved` is not allowed for `every-dev-completion` chains because it skips the user confirmation and approval dry-run gates. Use `feature-chain-approve` on the same proposed chain instead.
46
-
47
- After the user confirms a candidate flow, use `feature-chain-propose` when you need to record a safe draft with seed bad-case coverage:
48
-
49
- ```bash
50
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-propose \
51
- --root <project> \
52
- --title "<confirmed feature title>" \
53
- --entry "<confirmed user-visible entry>" \
54
- --exit-check "<confirmed strict final green condition>" \
55
- --node-title "<confirmed checkpoint>" \
56
- --bad-cases "BC-..., BC-..." \
57
- --check "<checkpoint recurrence check>"
58
- ```
59
-
60
- This command creates only a `proposed` chain. It records the user-confirmed shape and seed bad cases, but it does not add an executable command and must not enter `dev-complete` until `feature-chain-approve` is run with user-approved automation.
61
-
62
- If the user confirms the feature flow before any concrete bad case exists, keep the same command but replace `--bad-cases ...` with `--coverage-pending-reason "<why there is no linked bad case yet>"`. This records the draft so it is not lost, but it is not recurrence coverage and cannot be approved for `every-dev-completion` until a real bad case is attached to a checkpoint.
63
-
64
- When a later bad case appears, run `feature-chain-plan` first. If it points to a coverage-pending chain/checkpoint, attach the bad case there with `feature-chain-attach-bc` and tighten the checkpoint text. The attach step should clear the pending-coverage note, because the checkpoint now has real bad-case coverage.
65
-
66
- The expected lifecycle is:
67
-
68
- 1. `feature-chain-plan` turns user wording or a bad-case ID into a read-only confirmation prompt.
69
- 2. After user confirmation, `feature-chain-propose` records a non-executable draft with either linked bad cases or a coverage-pending reason.
70
- 3. Later bad cases are routed through `feature-chain-plan` and attached to the nearest existing checkpoint when semantically correct.
71
- 4. `feature-chain-summary` gives the fast coverage map; `feature-chain-overlap` checks duplicate workflow coverage before approval.
72
- 5. `feature-chain-approve` is the only path into the always-run set and must pass the checkpoint dry run.
73
- 6. `dev-complete` runs the approved chain with structured checkpoint markers and cleans success artifacts.
74
-
75
- This lifecycle is the core experiment: fewer feature chains should cover more bad-case recurrence checks without Codex inventing a broad test suite.
76
-
77
- Multi-project trials are the sanity check for this method. These trials must start from fresh sandbox projects instead of reusing an existing project, existing context folder, or previous test registry; otherwise the result may only prove that old context happened to work. The sandbox themes should also be genuinely different, such as a life utility, a creative tool, and a game or interaction, not merely three variants of the same engineering workflow. In small linear flows, one feature chain with two to four checkpoints can cover about three related bad cases and localize the failed phase. Treat planner checkpoint suggestions as hints, not business truth: Codex or the user must attach each bad case to the real phase where it can recur. When the workflow has queues, retries, multiple workers, recovery branches, or cross-process cleanup, upgrade the design to a task case instead of stretching a simple feature chain.
78
-
79
- A single-chain trial is not enough to validate Context Guard itself. A system-level regression should include at least two independent approved feature chains in one fresh project, then prove that `dev-complete` runs both, reports one chain failure without hiding the other chain's pass result, preserves the failing evidence, and returns to all-pass after the same chain is fixed. This checks the Test Hub orchestration layer rather than only the lifecycle of one feature chain.
80
-
81
- Before approving automation, use a dry run when the proposed command or checkpoint markers need validation:
82
-
83
- ```bash
84
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-dry-run \
85
- --root <project> \
86
- --chain-id FC-YYYYMMDD-001 \
87
- --command-text "<candidate command>"
88
- ```
89
-
90
- Dry run executes the candidate command against the proposed chain's registered checkpoints, reports missing/failed/unknown checkpoint markers, cleans success artifacts, and preserves failure evidence under `.codex/context/test-hub/dry-runs/`. It does not approve the chain, does not write the command into `feature-chains.json`, and does not add the chain to `dev-complete`.
91
-
92
- Dry-run evidence paths must be unique per run. Fast repeated or parallel dry runs must not reuse or overwrite a previous failure directory, because the preserved evidence is what lets Codex locate the failed checkpoint without reinterpreting the whole task.
93
-
94
- ## Approval Rule
95
-
96
- After the user confirms the feature flow and test design, promote the existing proposed chain with:
97
-
98
- ```bash
99
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-approve \
100
- --root <project> \
101
- --chain-id FC-YYYYMMDD-001 \
102
- --command-text "<approved command>"
103
- ```
104
-
105
- Approval is a safety gate and the only supported path from proposed feature-chain automation to `every-dev-completion`. It refuses a chain that has no checkpoint, no checkpoint check text, no linked bad-case coverage, or no automated command when the run policy is `every-dev-completion`. For `every-dev-completion` automation, approval must also run a dry-run preflight before mutating the registry. If the command misses a required checkpoint marker, emits an unknown marker, emits a `FAIL` marker, times out, or hits a blocker, approval fails and the chain stays `proposed`. Do not bypass this by hand-editing `feature-chains.json`, using `feature-chain-add --test-status approved`, or creating a second approved chain.
106
-
107
- ## Policy Rule
108
-
109
- After approval, the user's cadence still wins. If the user says a feature chain should not run after every development turn, update the existing chain instead of deleting or duplicating it:
110
-
111
- ```bash
112
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-set-policy \
113
- --root <project> \
114
- --chain-id FC-YYYYMMDD-001 \
115
- --run-policy relevant-only \
116
- --reason "User said this chain is only needed when touching GPU monitor flows."
117
- ```
118
-
119
- Keep the reason short and human-readable. Use `disabled-with-reason` only when the user asks to disable the chain or it cannot be run safely.
120
-
121
- ## Design Heuristic
122
-
123
- A good chain has:
124
-
125
- - one clear entry point
126
- - one realistic flow, not a pile of unrelated checks
127
- - two to five checkpoints
128
- - strict red and green conditions
129
- - failure localization, so the runner reports which checkpoint broke
130
- - cleanup-on-pass and preserve-on-fail behavior
131
-
132
- Avoid:
133
-
134
- - one script per bad case when one feature flow can cover them
135
- - broad suites that run unrelated product areas
136
- - checks that only prove code executed but cannot catch the old symptom
137
- - durable tests written from agent guesses without user confirmation
138
-
139
- ## Storage
140
-
141
- Store chain metadata in:
142
-
143
- ```text
144
- .codex/context/test-hub/feature-chains.json
145
- ```
146
-
147
- Store large scenario specs in `.codex/context/task-cases/` only when the workflow needs richer phases, logs, or human-readable execution notes.
148
-
149
- ## Execution
150
-
151
- Approved chains with `status: approved | active | stable` and `run_policy: every-dev-completion` are part of Test Hub. At development completion, run them through:
152
-
153
- ```bash
154
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py dev-complete --root <project>
155
- ```
156
-
157
- The runner should execute the approved command with minimal Codex reinterpretation, clean success artifacts, preserve failure evidence, and report the failed checkpoint or blocker.
158
-
159
- A feature-chain experiment is not complete just because one happy path passes. It should prove the closed loop: one chain covers multiple bad cases, a failing checkpoint preserves evidence with an actionable reason, and the fixed path passes while cleaning temporary artifacts.
160
-
161
- When an approved chain fails, do not design a new test to prove the same workflow. Treat the failed checkpoint as the recurrence signal, fix the cause, and rerun the same approved chain. A good runner makes this loop cheap by emitting readable checkpoint markers, preserving only the useful failure evidence, and cleaning success artifacts after the rerun passes.
162
-
163
- Feature-chain commands can report phase-level status with lightweight markers:
164
-
165
- ```text
166
- CG_CHECKPOINT:<checkpoint title or id>:PASS
167
- CG_CHECKPOINT:<checkpoint title or id>:FAIL:<short reason>
168
- ```
169
-
170
- Test Hub treats any `FAIL` marker as a failed feature chain, even if the command exits 0. Prefer these markers when one command covers several checkpoints, because the preserved result will point to the broken workflow step instead of only saying that the whole command failed.
171
-
172
- Marker names must match registered checkpoint titles or ids. Unknown markers fail the chain because they usually mean the script no longer matches the approved workflow. Keep non-English checkpoint titles readable and distinct; do not collapse them into generic ids.
173
-
174
- Approved feature-chain commands must report every registered checkpoint unless a checkpoint is explicitly optional (`optional: true` or `required: false`). Missing markers fail the chain, because an unreported checkpoint was not proven to run.
175
-
176
- If the user or business flow says one checkpoint should not be required every time, change that checkpoint explicitly instead of weakening the whole chain:
177
-
178
- ```bash
179
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-set-checkpoint \
180
- --root <project> \
181
- --chain-id FC-YYYYMMDD-001 \
182
- --node-title "前端打开监控页" \
183
- --required optional \
184
- --reason "Only runs in browser integration environment."
185
- ```
186
-
187
- Use `--required required` to restore the checkpoint to every-run coverage. This keeps the chain strict by default while allowing intentional, documented exceptions.
188
-
189
- Audit required/optional coverage without opening JSON:
190
-
191
- ```bash
192
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-list --root <project> --verbose
193
- ```
194
-
195
- Audit the compact coverage map before creating new coverage:
196
-
197
- ```bash
198
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-summary --root <project>
199
- ```
200
-
201
- This is the quickest way to see whether a small number of feature chains already covers several bad cases, and which checkpoints are still waiting for a real bad case before approval.
202
- Treat the `coverage density`, `reuse signal`, and `next:` lines as decision aids: they should push Codex toward reusing or extending an existing workflow when possible, not toward creating another standalone test.
203
-
204
- Audit possible duplicate feature chains before approval:
205
-
206
- ```bash
207
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-overlap --root <project>
208
- ```
209
-
210
- This is a route-choice guard, not a test runner. It compares existing chain wording and linked bad cases, then prints pairs that may be the same workflow. Use it to avoid turning one business flow into multiple always-run tests.
211
-
212
- Audit bad-case coverage across chains without mutating records:
213
-
214
- ```bash
215
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py feature-chain-coverage --root <project>
216
- ```
217
-
218
- Use this to decide whether a new bad case should attach to an existing chain or remain a standalone guard. Treat unassigned cases as candidates, not as required new tests.
219
-
220
- ## Quality Gate
221
-
222
- After editing feature chains, run:
223
-
224
- ```bash
225
- python3 ~/.agents/skills/context-guard/scripts/context_guard.py validate-feature-chains --root <project>
226
- ```
227
-
228
- This gate checks structure, not business judgment. It catches approved chains that are missing an entry point, exit check, automated command, checkpoint nodes, checkpoint check text, linked bad-case coverage, or a clear artifact policy. It does not create tests, approve tests, or decide whether the user's workflow deserves a durable chain.
@@ -1,85 +0,0 @@
1
- # Bad Case Register Template
2
-
3
- Use this format for `.codex/context/bad-cases.md`, task-local `.codex/context/tasks/<task-id>/bad-cases.md`, or the existing project context register.
4
-
5
- ```md
6
- # Bad Case Register
7
-
8
- This register tracks bad cases found during development and the guards that prevent them from recurring.
9
-
10
- Record only bad cases that are user-visible, recurring, risky, fixed, deferred, or needed to explain a guard. Do not turn the register into a defect diary.
11
-
12
- ## Active Cases
13
-
14
- ### BC-YYYYMMDD-001: Short descriptive title
15
-
16
- - Status: open | resolved | recurred | deferred | superseded-by-route-change
17
- - First observed: YYYY-MM-DD
18
- - Last checked: YYYY-MM-DD
19
- - Scope: feature, files, tests, route, UI flow, API, or subsystem
20
- - Context task: `CTX-...` folder or shared
21
- - Roadmap nodes: `NODE-...`
22
- - Tags: #hot | #flaky | #ui | #data-loss | #route-risk | custom tags
23
- - Frequency: first-seen | repeated-N | high-frequency
24
- - Phenomenon: one-line user-visible behavior or failing output
25
- - Trigger / reproduction: shortest command, step, input, environment, or precondition
26
- - Root cause: confirmed cause, suspected cause, or unknown, one line
27
- - Fix method: code/test/config/documentation change that fixed it, one line
28
- - Guard type: script | native-test | manual | browser-screenshot | browser-dom | curl | cli | prompt | log-invariant | fixture | unit | integration | e2e | custom
29
- - Guard / verification: native test, command, reusable script, manual check, screenshot, log, invariant, or reproduction note, one line
30
- - Run policy: every-dev-completion | relevant-only | manual | release-only | goal-final | disabled-with-reason | user-defined cadence
31
- - Artifact policy: cleanup-on-pass | preserve-on-fail | manual-preserve | none
32
- - Blocker handling: credentials | external-service | permissions | resource-limits | network | destructive-confirmation | user-judgment | none
33
- - Red condition: exact output, visual state, assertion, or symptom that means this bad case has recurred
34
- - Green condition: exact evidence that means this bad case is absent
35
- - Expected failure reason: why the guard should fail for the old symptom, not for a broken test or unrelated environment issue
36
- - Reusable guard path: project test file, `.codex/context/task-cases/...#phase-name`, `.codex/context/bad-case-tests/...`, or none
37
- - Covered by task case: TC-YYYYMMDD-short-slug phase/checkpoint, or none
38
- - Test-chain issue: false-positive | false-negative | wrong-granularity | missing-phase | wrong-assertion | unrealistic-setup | missing-cleanup | unclear-localization | none
39
- - Guard reuse rule: reuse this recorded guard before creating any new test or script for this case
40
- - Test chain: ordered checks only when multiple checks are genuinely needed
41
- - High-frequency note: warning text to show Codex when this pattern repeats often
42
- - Recurrence analysis: why it came back, if it ever did
43
- - Route-change note: only when an approved technical route change intentionally changes expected behavior
44
- - Evidence: links to tests, commands run, PRs, commits, screenshots, or logs
45
-
46
- ## Resolved History
47
-
48
- Move old resolved entries here only if the active section becomes noisy. Keep enough detail to replay the guard.
49
- ```
50
-
51
- Use the `### BC-...` section form as the canonical editable source. If a session accidentally records loose bullet blocks such as `- ID: BC-...`, `- Title: ...`, `- Status: ...`, or `- Nodes: ...`, the renderer should still project them, but future edits should normalize them back into formal case sections.
52
-
53
- ## Status Rules
54
-
55
- - `open`: bad case is known and not fixed.
56
- - `resolved`: fix is implemented and verification passed.
57
- - `recurred`: bad case came back after resolution; must be analyzed and fixed before completion.
58
- - `deferred`: intentionally not fixed in the current task; requires reason and owner/next step.
59
- - `superseded-by-route-change`: old behavior is no longer expected because an approved technical route changed it.
60
-
61
- ## Context Guard Rules
62
-
63
- - Use `.codex/context/` as the project folder for bad-case memory. Do not introduce a separate bad-case folder for new projects.
64
- - Use the configured `.codex/context/preferences.json` `record_language` for bad-case titles, phenomenon, root cause, fix method, guard summaries, and test-chain notes.
65
- - Preserve exact commands, paths, code identifiers, logs, API names, and error messages in their original language.
66
- - Prefer existing recorded context, user-approved commands, native tests, screenshots, logs, or manual checks over newly invented tests.
67
- - When the user creates or approves a test, default its `Run policy` to `every-dev-completion`; Codex must run it at every development completion unless the user sets another cadence.
68
- - Only the user can demote an approved test to `relevant-only`, `manual`, `release-only`, `goal-final`, `disabled-with-reason`, or a custom cadence. Record why.
69
- - After user approval, automate the check when feasible; successful automated checks should clean temporary files, while failed checks should preserve concise diagnostic evidence for Codex to analyze and rerun after fixing.
70
- - Store reusable user-approved automated tests in `.codex/context/test-hub/registry.json` or approved task-case files; do not auto-register ordinary bad-case guards or roadmap `Test chain:` notes as always-run tests.
71
- - At development completion, prefer `context_guard.py dev-complete --root <project>` so Test Hub runs the approved always-run set and writes `.codex/context/test-hub/last-run.json`.
72
- - If an approved automated check is blocked by credentials, external services, permissions, resource limits, network, destructive confirmation, or user-only judgment, record the blocker and ask or warn the user instead of looping.
73
- - For resolved or recurred cases, `Guard / verification`, `Guard type`, `Red condition`, `Green condition`, and `Expected failure reason` are required.
74
- - The guard must be red-capable: it should fail if the same user-visible symptom returns.
75
- - When the bad case is part of a longer workflow, attach it to a human-approved task-case checkpoint instead of creating a separate isolated script. The bad-case entry should say which task case phase covers it.
76
- - If the test chain itself is wrong, record that as a bad case and classify the test-chain issue. Fix the test-chain design before trusting its result.
77
- - Do not script every bad case. Store bad-case-specific scripts under `.codex/context/bad-case-tests/` only when the user approved the test design, the script is genuinely reusable, and it does not belong in the native test suite.
78
- - Name any guard script with the bad case ID so it is easy to find and reuse.
79
- - Update existing context when expected behavior changes; do not create parallel guards for the same case unless the old one is explicitly obsolete.
80
- - If a guard is manual-only, list the exact manual check and why that is acceptable for now.
81
- - Link bad cases to roadmap nodes so Codex can quickly see which mainline decisions created or fixed them.
82
- - Keep record/display linkage explicit: use `Roadmap nodes:` or `Nodes:` on the bad case, or `Linked bad cases:` on the roadmap node.
83
- - Add tags and frequency notes when a bad case repeats often; high-frequency cases should stand out during quick scanning.
84
- - Promote high-frequency cases into fixed pressure checks and rerun them whenever related code, UI, context, or hooks change.
85
- - Keep entries compact. If the same information appears in a roadmap node, link to it instead of duplicating it.
@@ -1,63 +0,0 @@
1
- # Task Case Template
2
-
3
- Use task cases for realistic multi-step verification. A task case should simulate a full user or agent workflow and log phase-level checkpoints so failures identify the broken step. Test design is human-owned: Codex can draft a proposal, but a durable task case stays `proposed` until the user confirms it.
4
-
5
- ```md
6
- # Task Case: short realistic workflow title
7
-
8
- - ID: TC-YYYYMMDD-short-slug
9
- - Status: proposed | approved | active | stable | deferred | obsolete
10
- - Route/task: `CTX-...` or branch name
11
- - Scope: feature, service, UI flow, agent workflow, or subsystem
12
- - Last checked: YYYY-MM-DD
13
- - Design confirmation: pending | user-approved YYYY-MM-DD
14
- - Run policy: every-dev-completion | relevant-only | manual | release-only | goal-final | disabled-with-reason | user-defined cadence
15
- - Automation entry: native command | script path | prompt/manual runner | none
16
- - Artifact policy: cleanup-on-pass | preserve-on-fail | manual-preserve
17
- - Linked roadmap nodes: NODE-...
18
- - Linked bad cases: BC-..., BC-...
19
- - Entry command/prompt: command, prompt, manual setup, or fixture
20
- - Not covered: explicit exclusions to avoid fake confidence
21
- - Stop condition: what means the workflow is complete
22
- - Cleanup: required cleanup or none
23
- - Blocker handling: credentials | external service | permissions | resource limits | network | destructive confirmation | user judgment | none
24
-
25
- ## Phases
26
-
27
- ### Phase 1: setup or trigger
28
-
29
- - Action: one realistic action
30
- - Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
31
- - Covers bad cases: BC-...
32
- - Failure localization: what this phase failure usually means
33
- - Log note: what the script/agent should record
34
-
35
- ### Phase 2: transition or recovery
36
-
37
- - Action: one realistic action
38
- - Expected checkpoint: invariant, log, UI state, file state, API result, or assertion
39
- - Covers bad cases: BC-...
40
- - Failure localization: what this phase failure usually means
41
- - Log note: what the script/agent should record
42
-
43
- ## Result Log
44
-
45
- - YYYY-MM-DD: pass/fail, failed phase/checkpoint if any, evidence path or command output summary
46
- ```
47
-
48
- Rules:
49
-
50
- - When the user explicitly asks to create, write, generate, design, or add a test/task case, begin the user-visible response with a compact intake line such as `测试创建识别:...` before proposing or implementing the case.
51
- - Prefer one task case with clear checkpoints over many disconnected bug-level scripts when the same workflow is being exercised.
52
- - Link bad cases to the checkpoint that catches them.
53
- - Keep the task case realistic enough to match actual product or agent usage.
54
- - Once a test is approved, automate it when it can be safely encapsulated. Future Codex turns should run the registered entry instead of reinterpreting the test design.
55
- - Automated approved tests should be registered in `.codex/context/test-hub/registry.json` or represented by this approved task-case file with `Run policy: every-dev-completion` and an executable `Entry command/prompt`.
56
- - At development completion, run the approved always-run set through `context_guard.py dev-complete --root <project>` instead of manually reconstructing each test.
57
- - Automated task cases should clean temporary files after full success and preserve the smallest useful evidence on failure.
58
- - If the automated case fails, Codex should analyze the failed phase/checkpoint, record or update the bad case, fix in scope, and rerun the same approved test until it passes or a non-actionable blocker is reached.
59
- - If blocked by credentials, unavailable services, permissions, hardware/resource limits, network, destructive-risk confirmation, or user-only judgment, stop and ask or warn the user with the blocker and evidence path.
60
- - Keep checkpoint logs concise and useful for localizing failure.
61
- - Keep the case in `proposed` state until the user confirms the design.
62
- - Once the user confirms the design, default `Run policy` to `every-dev-completion`; lower the cadence only when the user explicitly asks.
63
- - In goal mode, use human-approved task cases as phase gates and log phase progress; do not create broad new scripts or active task cases without confirmation unless the user explicitly asked for that exact test.
@@ -1,116 +0,0 @@
1
- #!/usr/bin/env bash
2
- set -euo pipefail
3
-
4
- SCRIPT="${1:-$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)/scripts/context_guard.py}"
5
-
6
- python3 - "$SCRIPT" <<'PY'
7
- from __future__ import annotations
8
-
9
- import json
10
- import shutil
11
- import subprocess
12
- import sys
13
- import tempfile
14
- from pathlib import Path
15
-
16
- script = Path(sys.argv[1])
17
- root = Path(tempfile.mkdtemp(prefix="context-guard-branch-test-"))
18
-
19
- try:
20
- subprocess.run(["python3", str(script), "init", "--root", str(root)], check=True, capture_output=True, text=True)
21
- subprocess.run(["python3", str(script), "set-language", "--language", "zh", "--root", str(root)], check=True, capture_output=True, text=True)
22
-
23
- ctx = root / ".codex" / "context"
24
- (ctx / "index.md").write_text(
25
- """# Context Index
26
-
27
- ## Quick Scan
28
-
29
- - Current: CTX-20260618-main
30
- - Latest roadmap node: NODE-20260618-001
31
- - Hot bad-case tags: none
32
- - Resume candidate: none
33
-
34
- ## Current
35
-
36
- - ID: CTX-20260618-main
37
- - Title: 主线任务
38
- - State: current
39
- - Folder: `.codex/context/tasks/CTX-20260618-main/`
40
- - Last updated: 2026-06-18
41
- - Summary: 正在推进主线任务。
42
- - Next step: 继续主线。
43
-
44
- ## Parked / Resume Candidates
45
-
46
- None.
47
-
48
- ## Archived
49
-
50
- None.
51
- """,
52
- encoding="utf-8",
53
- )
54
- (ctx / "roadmap.md").write_text(
55
- """# Context Roadmap
56
-
57
- ## Nodes
58
-
59
- ### NODE-20260618-001: 主线起点
60
-
61
- - Date: 2026-06-18
62
- - Status: done
63
- - Level: major
64
- - Branch: Main
65
- - Parent: none
66
- - Task: `CTX-20260618-main`
67
- - Outcome: 主线已经建立。
68
- - Decision / reason: 作为支线父节点。
69
- - Avoid going back: 不要把支线写回主线。
70
- - Next: 等待支线。
71
- - Linked bad cases: none
72
- - Test chain: none
73
- """,
74
- encoding="utf-8",
75
- )
76
-
77
- result = subprocess.run(
78
- [
79
- "python3",
80
- str(script),
81
- "create-branch-task",
82
- "--root",
83
- str(root),
84
- "--title",
85
- "后端状态机设计",
86
- "--branch",
87
- "后端状态机",
88
- "--parent-node",
89
- "NODE-20260618-001",
90
- ],
91
- check=True,
92
- capture_output=True,
93
- text=True,
94
- )
95
-
96
- output = result.stdout + result.stderr
97
- assert "CTX-" in output, output
98
- assert "NODE-" in output, output
99
-
100
- index = (ctx / "index.md").read_text(encoding="utf-8")
101
- roadmap = (ctx / "roadmap.md").read_text(encoding="utf-8")
102
- exports = json.loads((ctx / "roadmap" / "roadmap.json").read_text(encoding="utf-8"))
103
-
104
- assert "- Current: CTX-20260618-main" not in index, index
105
- assert "后端状态机设计" in index, index
106
- assert "主线任务" in index and "resume-candidate" in index, index
107
- assert "- Branch: 后端状态机" in roadmap, roadmap
108
- assert "- Parent: NODE-20260618-001" in roadmap, roadmap
109
- assert "- Task: `CTX-" in roadmap, roadmap
110
- assert any(route["branch"] == "后端状态机" for route in exports["routes"]), exports
111
-
112
- task_dirs = list((ctx / "tasks").glob("CTX-*"))
113
- assert any((path / "context.md").exists() and "后端状态机设计" in (path / "context.md").read_text(encoding="utf-8") for path in task_dirs), task_dirs
114
- finally:
115
- shutil.rmtree(root, ignore_errors=True)
116
- PY
@@ -1,66 +0,0 @@
1
- #!/usr/bin/env bash
2
- set -euo pipefail
3
-
4
- SCRIPT="${1:-$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)/scripts/context_guard.py}"
5
- ROOT="$(mktemp -d "${TMPDIR:-/tmp}/context-guard-marker-test-XXXXXX")"
6
- trap 'rm -rf "$ROOT"' EXIT
7
-
8
- python3 "$SCRIPT" init --root "$ROOT" >/dev/null
9
- python3 "$SCRIPT" set-language --language zh --root "$ROOT" >/dev/null
10
-
11
- CTX="$ROOT/.codex/context"
12
- cat > "$CTX/roadmap.md" <<'MD'
13
- # Context Roadmap
14
-
15
- ## Nodes
16
-
17
- ### NODE-20260618-001: 标记测试节点
18
-
19
- - Date: 2026-06-18
20
- - Status: done
21
- - Level: major
22
- - Branch: Main
23
- - Parent: none
24
- - Task: `CTX-20260618-marker`
25
- - Outcome: 检查问题案例标题行标记布局。
26
- - Decision / reason: 标记不能单独占一行。
27
- - Avoid going back: 不要把状态点放在标题前面。
28
- - Next: none
29
- - Linked bad cases: BC-20260618-999
30
- - Test chain: none
31
- MD
32
-
33
- cat > "$CTX/bad-cases.md" <<'MD'
34
- # Bad Case Register
35
-
36
- ## Active Cases
37
-
38
- ### BC-20260618-999: 这是一个很长的问题案例标题用于触发换行但标记不能单独成行
39
-
40
- - Status: resolved
41
- - First observed: 2026-06-18
42
- - Last checked: 2026-06-18
43
- - Scope: roadmap marker layout
44
- - Context task: `CTX-20260618-marker`
45
- - Roadmap nodes: NODE-20260618-001
46
- - Tags: #roadmap-ux
47
- - Frequency: repeated-2
48
- - Phenomenon: 状态点和频率点不应该单独一行。
49
- - Trigger / reproduction: 生成 roadmap 并检查 badcase-head。
50
- - Root cause: 标记位于标题前,窄列时容易变成单独一行。
51
- - Fix method: 将标记移到标题右侧的 inline marker 容器。
52
- - Guard / verification: 本脚本检查 HTML 结构。
53
- - Reusable guard path: none
54
- - Test chain: none
55
- MD
56
-
57
- python3 "$SCRIPT" show-roadmap --root "$ROOT" >/dev/null
58
- HTML="$CTX/roadmap/roadmap.html"
59
-
60
- grep -q 'class="badcase-markers"' "$HTML"
61
- grep -q '.badcase-head { display: grid; grid-template-columns: minmax(0, 1fr) auto;' "$HTML"
62
- if grep -q '<div class="badcase-head"><span class="status-dot' "$HTML"; then
63
- echo "badcase marker still starts a separate title row" >&2
64
- exit 1
65
- fi
66
- grep -q '<a class="detail-link" href="#case-1">' "$HTML"
@@ -1,48 +0,0 @@
1
- #!/usr/bin/env bash
2
- set -euo pipefail
3
-
4
- SCRIPT="${1:-$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)/scripts/context_guard.py}"
5
- ROOT="$(mktemp -d "${TMPDIR:-/tmp}/context-guard-readable-map-XXXXXX")"
6
- trap 'rm -rf "$ROOT"' EXIT
7
-
8
- python3 "$SCRIPT" init --root "$ROOT" >/dev/null
9
- python3 "$SCRIPT" set-language --language zh --root "$ROOT" >/dev/null
10
-
11
- CTX="$ROOT/.codex/context"
12
- cat > "$CTX/roadmap.md" <<'MD'
13
- # Context Roadmap
14
-
15
- ## Nodes
16
-
17
- ### NODE-20260626-001: 主线起点
18
-
19
- - Date: 2026-06-26
20
- - Status: done
21
- - Level: major
22
- - Branch: Main
23
- - Parent: none
24
- - Outcome: 这是一段很长的主线摘要,过去会直接显示在多路线概览卡片里,让路线图看起来像卡片墙而不是路线骨架。
25
- - Next: none
26
- - Linked bad cases: none
27
-
28
- ### NODE-20260626-002: 支线入口
29
-
30
- - Date: 2026-06-26
31
- - Status: done
32
- - Level: major
33
- - Branch: 可读性支线
34
- - Parent: NODE-20260626-001
35
- - Outcome: 这是一段很长的支线摘要,应该留在详情里,而不是占据概览卡片的主要视觉空间。
36
- - Next: none
37
- - Linked bad cases: none
38
- MD
39
-
40
- python3 "$SCRIPT" show-roadmap --root "$ROOT" >/dev/null
41
- HTML="$CTX/roadmap/roadmap.html"
42
-
43
- grep -q 'class="route-stack branch-map"' "$HTML"
44
- grep -q '.track-grid.route-only {' "$HTML"
45
- grep -q 'grid-auto-columns: minmax(180px, 230px);' "$HTML"
46
- grep -q '.branch-map .summary {' "$HTML"
47
- grep -q 'display: none;' "$HTML"
48
- grep -q -- '-webkit-line-clamp: 3;' "$HTML"
@@ -1,40 +0,0 @@
1
- #!/usr/bin/env bash
2
- set -euo pipefail
3
-
4
- SCRIPT="${1:-$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)/scripts/context_guard.py}"
5
- ROOT="$(mktemp -d "${TMPDIR:-/tmp}/context-guard-hidden-details-XXXXXX")"
6
- trap 'rm -rf "$ROOT"' EXIT
7
-
8
- python3 "$SCRIPT" init --root "$ROOT" >/dev/null
9
- python3 "$SCRIPT" set-language --language zh --root "$ROOT" >/dev/null
10
-
11
- CTX="$ROOT/.codex/context"
12
- cat > "$CTX/roadmap.md" <<'MD'
13
- # Context Roadmap
14
-
15
- ## Nodes
16
-
17
- ### NODE-20260626-001: 可点击路线节点
18
-
19
- - Date: 2026-06-26
20
- - Status: done
21
- - Level: major
22
- - Branch: Main
23
- - Parent: none
24
- - Outcome: 这个节点的详细概括只能在点击节点后显示,不能默认铺在 roadmap 下方。
25
- - Next: none
26
- - Linked bad cases: none
27
- MD
28
-
29
- python3 "$SCRIPT" show-roadmap --root "$ROOT" >/dev/null
30
- HTML="$CTX/roadmap/roadmap.html"
31
-
32
- grep -q '<a class="lane-link" href="#node-1">' "$HTML"
33
- grep -q '<section class="inline-details" hidden data-inline-details' "$HTML"
34
- grep -q 'function setupInlineDetails()' "$HTML"
35
- grep -q 'class="detail-card" id="node-1"' "$HTML"
36
-
37
- if grep -q '<section class="inline-details" aria-label="Roadmap details">' "$HTML"; then
38
- echo "inline details are visible by default" >&2
39
- exit 1
40
- fi