claude-dev-env 8.42.1 → 8.42.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +3 -1
- package/docs/rule-guides/anti-corollary-tests.md +51 -0
- package/docs/rule-guides/archiving-agent-config.md +34 -0
- package/docs/rule-guides/build-needs-a-user.md +27 -0
- package/docs/rule-guides/falsify-before-green.md +43 -0
- package/docs/rule-guides/flag-non-breaking-findings.md +59 -0
- package/docs/rule-guides/long-horizon-autonomy.md +40 -0
- package/docs/rule-guides/no-cross-skill-duplicate-helpers.md +24 -0
- package/docs/rule-guides/prefer-existing-tools.md +22 -0
- package/docs/rule-guides/retired-hook-prose.md +27 -0
- package/docs/rule-guides/testing.md +21 -0
- package/docs/rule-guides/workers-done-before-complete.md +30 -0
- package/docs/rule-guides/workflow-substitution-slots.md +13 -0
- package/package.json +1 -1
- package/rules/anti-corollary-tests.md +5 -57
- package/rules/archiving-agent-config.md +6 -74
- package/rules/build-needs-a-user.md +6 -50
- package/rules/falsify-before-green.md +5 -45
- package/rules/flag-non-breaking-findings.md +5 -119
- package/rules/long-horizon-autonomy.md +5 -48
- package/rules/no-cross-skill-duplicate-helpers.md +5 -21
- package/rules/prefer-existing-tools.md +6 -32
- package/rules/retired-hook-prose.md +5 -15
- package/rules/testing.md +5 -18
- package/rules/workers-done-before-complete.md +5 -5
- package/rules/workflow-substitution-slots.md +6 -4
- package/docs/worker-completion-gate.md +0 -32
package/AGENTS.md
CHANGED
|
@@ -74,6 +74,8 @@ Banned word: real
|
|
|
74
74
|
|
|
75
75
|
Never write real, really, or real-world. This ban has no exception. Emphasis and contrast are no exception. It covers chat, commits, pull requests, comments, documentation, headings, and variable names.
|
|
76
76
|
|
|
77
|
+
`One real failure` is `one failure`. `The real cause` is `the cause`. `Really fast` is `fast`, or the measured number. `Real users` is `users`. `A real bug, not a flake` is `a bug`, followed by the evidence that rules out a flake.
|
|
78
|
+
|
|
77
79
|
Delete the word, then read the sentence; when meaning thins, name the evidence. The failing check. The log line. The measured number. The file and the line.
|
|
78
80
|
|
|
79
|
-
Swapping in actual, actually, genuine, or true is the same move, and each is banned with it.
|
|
81
|
+
Swapping in actual, actually, genuine, or true is the same move, and each is banned with it. So is the invented contrast that invites the word back, such as `not a hypothetical problem but a problem`.
|
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
# Anti-corollary tests
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/anti-corollary-tests.md`](../../rules/anti-corollary-tests.md).
|
|
4
|
+
|
|
5
|
+
## The three questions
|
|
6
|
+
|
|
7
|
+
### 1. Is this a corollary?
|
|
8
|
+
|
|
9
|
+
When the code reduces each input to a canonical form and then compares the forms, every pairwise combination of input spellings follows once the reduction is proven canonical. Walking the full N-by-N matrix restates that fact. It adds cases and runtime, and it buries the few cases that carry information.
|
|
10
|
+
|
|
11
|
+
### 2. Could this test pass if the mechanism were dead?
|
|
12
|
+
|
|
13
|
+
The degenerate value may be an empty string, `None`, `False`, a blanket refusal, or an empty collection. A test expecting that value passes even when the mechanism stops.
|
|
14
|
+
|
|
15
|
+
### 3. What single change to the code would make this test fail?
|
|
16
|
+
|
|
17
|
+
The named mutation identifies the behavior a case observes. A case that survives it provides no evidence about that behavior.
|
|
18
|
+
|
|
19
|
+
**The policy-surface case.** A test that asserts a policy's wording appears in its declaring file detects edits to that sentence. A suite of such tests counts repeated text. The declaring file can supply the policy while assertions inspect the surfaces it governs.
|
|
20
|
+
|
|
21
|
+
## What a mechanism with a degenerate failure mode needs
|
|
22
|
+
|
|
23
|
+
The audit record pairs a specific code mutation with the number of tests it fails. A reviewer or audit skill checks that record; no hook computes it.
|
|
24
|
+
|
|
25
|
+
## Worked shape (sanitized)
|
|
26
|
+
|
|
27
|
+
A write guard decides whether a write is about to hit a production database. It reduces each database URL to a canonical endpoint identity, then compares identities.
|
|
28
|
+
|
|
29
|
+
Two independent mutations show the two halves of a useful suite:
|
|
30
|
+
|
|
31
|
+
- Gut the reduction so it always returns the empty string. The guard fails **closed** and refuses everything. The *allow*-expecting cases die.
|
|
32
|
+
- Abandon the reduction and compare raw hostnames. The guard fails **open** and allows a production write. The *refuse*-expecting cases die.
|
|
33
|
+
|
|
34
|
+
The two breaks exercise opposite outcomes. A suite with both outcomes detects each break.
|
|
35
|
+
|
|
36
|
+
## Why this requires judgment
|
|
37
|
+
|
|
38
|
+
A structural hook cannot infer whether a case follows from canonicalization or passes against a dead implementation. Counting `parametrize` cases or assertions would also flag useful suites. The review and audit lanes judge changed tests using their intended behavior.
|
|
39
|
+
|
|
40
|
+
## Sibling rules
|
|
41
|
+
|
|
42
|
+
| Rule | Role |
|
|
43
|
+
|---|---|
|
|
44
|
+
| [`code-standards.md`](../../rules/code-standards.md) | Points at CODE_RULES section 8 on a failing test before production code |
|
|
45
|
+
| [`testing.md`](../../rules/testing.md) | Mocks and test infrastructure standards |
|
|
46
|
+
| [`paired-test-coverage.md`](../../rules/paired-test-coverage.md) | Every public function in an established suite gets a behavioral test |
|
|
47
|
+
| [`anti-corollary-tests.md`](../../rules/anti-corollary-tests.md) | Each test carries information |
|
|
48
|
+
|
|
49
|
+
## Enforcement
|
|
50
|
+
|
|
51
|
+
The AI review lane and audit skills inspect the test lines a pull request changes. No blocking hook evaluates the intent behind a corollary or dead-default case.
|
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
# Archiving a rule, skill, agent, or command
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/archiving-agent-config.md`](../../rules/archiving-agent-config.md), which loads in sessions as the short form.
|
|
4
|
+
|
|
5
|
+
## Evidence of use
|
|
6
|
+
|
|
7
|
+
Enforcement in this package often has a different name from the file that describes it. The script, lint identifier, and running check show whether a capability remains in service. A rule can have no mentions of its slug and still have a checker under another name.
|
|
8
|
+
|
|
9
|
+
`CHANGELOG.md` records history. The `README.md` inventory table lists what ships. `bin/ever-shipped-skills.mjs` retains archived names on purpose. Mentions in these three files add no evidence of current use.
|
|
10
|
+
|
|
11
|
+
This package installs into other repositories. A path absent from this tree can resolve in the host repository where the rule or skill runs. The path's intended location determines whether it is stale.
|
|
12
|
+
|
|
13
|
+
## Archive layout
|
|
14
|
+
|
|
15
|
+
`rules-archived/` sits beside `rules/`, and `.agents/skills-archived/` sits beside `.agents/skills/`. The installer copies a content directory whole, so an archive inside a live directory would ship. `CONTENT_DIRECTORIES` in `bin/install.mjs` omits archive siblings, and skill installation enumerates `.agents/skills` alone.
|
|
16
|
+
|
|
17
|
+
The archive manifest records why each file left service, the `git mv` that restores it, and every companion edit a restore must undo. The sibling archive keeps the file available for a reverse move.
|
|
18
|
+
|
|
19
|
+
## Exempt files
|
|
20
|
+
|
|
21
|
+
The "Never archived" section of `rules-archived/ARCHIVE-MANIFEST.md` names files that stay in service. `rules/correction-lens.md` is listed there. Editing an exempt file to sharpen it remains ordinary work.
|
|
22
|
+
|
|
23
|
+
## Skill names and tests
|
|
24
|
+
|
|
25
|
+
The prune computes retired skill names as every name in `bin/ever-shipped-skills.mjs` minus the names installed now. The installer uses that set to move stale copies out of a host's agents home. `scripts/active_capability_references.py` uses the same registry to report leftover mentions of an archived name.
|
|
26
|
+
|
|
27
|
+
The node test command globs `.agents/skills/**/*.test.mjs`. Tests under a renamed archived skill directory leave that run without failing, so the test count shows whether coverage still runs.
|
|
28
|
+
|
|
29
|
+
## Sibling rules
|
|
30
|
+
|
|
31
|
+
| Rule | Role |
|
|
32
|
+
|---|---|
|
|
33
|
+
| [`retired-hook-prose.md`](../../rules/retired-hook-prose.md) | Hook prose names checks that still run |
|
|
34
|
+
| [`doc-inventory-integrity.md`](../../rules/doc-inventory-integrity.md) | Inventory tables reflect what ships |
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Build needs a user
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/build-needs-a-user.md`](../../rules/build-needs-a-user.md), which loads in sessions as the short form.
|
|
4
|
+
|
|
5
|
+
## First use and size
|
|
6
|
+
|
|
7
|
+
A caller can be a hook entry, workflow step, skill, command, or module in the same change. First use can be a date or a named event. Prior manual use and frequency show whether the job recurs. An unanswered point leaves the owner a one-line statement of the intended build and the missing answer while the rest of the task continues.
|
|
8
|
+
|
|
9
|
+
A correction starts as a row in a rule file. [`correction-lens.md`](../../rules/correction-lens.md) describes when repetition warrants a hook or lint. The first implementation needs only the behavior its first caller uses.
|
|
10
|
+
|
|
11
|
+
## Where the check runs
|
|
12
|
+
|
|
13
|
+
The staged policy lint's `uncalled-new-file` rule reads newly added code files under `scripts/`, `hooks/`, `bin/`, `ci/`, or `tools/`. It reports a file whose name appears only in its own tests, `CHANGELOG.md`, `README.md`, and `bin/ever-shipped-skills.mjs`. CI runs the lint against the merge base.
|
|
14
|
+
|
|
15
|
+
The code checks caller references. The other decision points, including first use and prior manual work, depend on the author.
|
|
16
|
+
|
|
17
|
+
## Why the caller matters
|
|
18
|
+
|
|
19
|
+
The Codex compatibility watcher shipped 502 lines and 599 lines of tests in July. PR 1488 deleted it in September. Its own test was the only file that named it.
|
|
20
|
+
|
|
21
|
+
## Sibling rules
|
|
22
|
+
|
|
23
|
+
| Rule | Role |
|
|
24
|
+
|---|---|
|
|
25
|
+
| [`prefer-existing-tools.md`](../../rules/prefer-existing-tools.md) | Search before building |
|
|
26
|
+
| [`correction-lens.md`](../../rules/correction-lens.md) | A repeated correction moves up a layer |
|
|
27
|
+
| [`archiving-agent-config.md`](../../rules/archiving-agent-config.md) | How an unused capability leaves service |
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# Falsify before green
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/falsify-before-green.md`](../../rules/falsify-before-green.md).
|
|
4
|
+
|
|
5
|
+
## The four shapes this stops
|
|
6
|
+
|
|
7
|
+
### 1. A probe whose trigger condition never fires
|
|
8
|
+
|
|
9
|
+
The probe reports zero while its counter stays at the start value. The trip input exposes whether it observes an event.
|
|
10
|
+
|
|
11
|
+
**Break to apply:** feed it one input that must trip it. A probe still at zero on that input measures nothing.
|
|
12
|
+
|
|
13
|
+
### 2. A sweep that reads a subset of the files it claims to cover
|
|
14
|
+
|
|
15
|
+
The sweep may compare against the wrong base or walk a slice of the tree while claiming the full set.
|
|
16
|
+
|
|
17
|
+
**Break to apply:** plant one violation in a file the sweep's coverage claim names. A sweep that misses the plant walks a smaller file set than the one it reports.
|
|
18
|
+
|
|
19
|
+
### 3. A mutation that survives
|
|
20
|
+
|
|
21
|
+
The test may miss the mutated code because a mock stands in for the call, a guard returns early, or the test drives a neighboring branch.
|
|
22
|
+
|
|
23
|
+
**Break to apply:** hold the mutation in place and run the test. A green test names a line nothing covers.
|
|
24
|
+
|
|
25
|
+
### 4. An assertion that counts an artifact the harness seeded
|
|
26
|
+
|
|
27
|
+
The harness writes the row, file, or event the assertion counts. Stubbing the production writer exposes that source.
|
|
28
|
+
|
|
29
|
+
**Break to apply:** stub the production writer to a no-op. A green assertion counts the seed.
|
|
30
|
+
|
|
31
|
+
## What a shown-red record holds
|
|
32
|
+
|
|
33
|
+
| Part | What it names |
|
|
34
|
+
|---|---|
|
|
35
|
+
| The break | The mutation, stub, or trip input, named by file and line or by the exact input text |
|
|
36
|
+
| The red | The failing output the check printed under that break |
|
|
37
|
+
| The control | The case that passes beside the red, run on the same command |
|
|
38
|
+
|
|
39
|
+
The red and control distinguish a selective check from one that always fails or always passes.
|
|
40
|
+
|
|
41
|
+
## Enforcement
|
|
42
|
+
|
|
43
|
+
A reviewer reads the shown-red record beside each new check a pull request adds. A regex cannot tell which code a green check reached at run time.
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Flag non-breaking findings
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/flag-non-breaking-findings.md`](../../rules/flag-non-breaking-findings.md), which loads in sessions as the short form.
|
|
4
|
+
|
|
5
|
+
## Why severity matters
|
|
6
|
+
|
|
7
|
+
A breaking finding means the change is wrong: a bug, a secret in the tree, a broken test, a syntax error, or an instruction file that fails to load. A smell means the change reads poorly: a length limit, a naming convention, a prose term, a comment rule, or a structural preference.
|
|
8
|
+
|
|
9
|
+
A smell that blocks a commit stops delivery for something the reader could fix in the next pass. The writer then reaches for a bypass flag, which disables the breaking checks beside it. Recording the smell keeps the blocking checks useful and gives later work a record to clear.
|
|
10
|
+
|
|
11
|
+
## The ledger
|
|
12
|
+
|
|
13
|
+
Recorded findings land in `.claude/followups/smells.jsonl` at the repository root, one JSON object per line. `hooks/followup_ledger.py` writes and reads it. Ledger writes are fail-safe, so a ledger failure leaves the gate's decision unchanged.
|
|
14
|
+
|
|
15
|
+
| Field | What it carries |
|
|
16
|
+
|---|---|
|
|
17
|
+
| `rule_id` | The rule that raised the finding |
|
|
18
|
+
| `check_id` | The single check behind it, which a severity table keys on |
|
|
19
|
+
| `file_path` | The repository-relative path the finding names |
|
|
20
|
+
| `message` | The text a reader acts on |
|
|
21
|
+
| `severity` | The class the gate put it in |
|
|
22
|
+
| `origin_commit` | The revision checked out when it was recorded |
|
|
23
|
+
|
|
24
|
+
The origin commit groups a follow-up pull request by the change that raised the findings. A smell seen again under a later revision keeps the revision that first raised it, so one smell stays one record. The ledger is per-checkout state and stays out of the repository.
|
|
25
|
+
|
|
26
|
+
## The check identifier
|
|
27
|
+
|
|
28
|
+
Most lint rules run one check, so their rule identifier already names it. The `code-rules` and `validators` rules bundle many checks behind one identifier. `scripts/policy_lint/check_catalog.py` resolves those to a `<rule>/<check>` identifier, which `cde lint --format json` emits as `check_id` on every diagnostic. Consumers partition findings by that identifier.
|
|
29
|
+
|
|
30
|
+
A message with no catalog entry resolves to `<rule>/unclassified`, which a partition treats as blocking. When check wording changes, the catalog synchronization tests expose a missing mapping.
|
|
31
|
+
|
|
32
|
+
## Reading and clearing the ledger
|
|
33
|
+
|
|
34
|
+
| Command | What it does |
|
|
35
|
+
|---|---|
|
|
36
|
+
| `cde followup list` | Names every recorded follow-up |
|
|
37
|
+
| `cde followup ingest REPORT` | Records the diagnostics a policy-lint JSON report carries |
|
|
38
|
+
| `cde followup brief` | Writes the task an agent fixes them from |
|
|
39
|
+
| `cde followup clear` | Empties the ledger |
|
|
40
|
+
| `cde followup count` | Reports the backlog against the threshold |
|
|
41
|
+
|
|
42
|
+
`count` exits non-zero once the backlog passes `FOLLOWUP_BACKLOG_THRESHOLD` in `scripts/dev_env_scripts_constants/followup_constants.py`, so a scheduled job escalates. The number is the repository's setting.
|
|
43
|
+
|
|
44
|
+
The `/fix-followups` command reads the brief, fixes each rule group, opens a pull request, and clears the ledger.
|
|
45
|
+
|
|
46
|
+
## Worked example
|
|
47
|
+
|
|
48
|
+
`scripts/validate_instruction_pairs.py` raises five findings. A missing governing `AGENTS.md`, mismatched import text, and an instruction path that is not a regular file each stop instructions loading. A non-canonical filename and a Git mode other than 100644 leave instructions loading, so the gate records them and passes.
|
|
49
|
+
|
|
50
|
+
`SEVERITY_BY_RULE_ID` in that module declares the split. `scripts/repository_policy.py` uses `SEVERITY_BY_CHECK_ID` in `repository_checks/config/constants.py`. Its `package-inventory` check is a smell because the production file still imports and runs when an inventory row is missing. The check prints an `advisory:` prefix and records the finding. A `CLAUDE.md` naming a missing file, an env-var row naming a file that never reads the variable, a test outside the testpaths allowlist, and a tracked secret each block.
|
|
51
|
+
|
|
52
|
+
`scripts/cde_lint.py` uses `SEVERITY_BY_CHECK_ID` in `scripts/policy_lint/config/check_catalog_constants.py`. `test-pairing` and the two paired-test coverage checks are smells there. The lint prints warnings, records them, and exits zero when no error remains.
|
|
53
|
+
|
|
54
|
+
## Sibling rules
|
|
55
|
+
|
|
56
|
+
| Rule | Role |
|
|
57
|
+
|---|---|
|
|
58
|
+
| [CI Owns the Gate](../../.agents/skills/pr-lifecycle/SKILL.md#ci-owns-the-gate) | The full check suite runs once, on CI |
|
|
59
|
+
| [Git workflow](../../.agents/skills/pr-lifecycle/SKILL.md#git-workflow) | A red required check blocks the branch |
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# Long-horizon autonomy
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/long-horizon-autonomy.md`](../../rules/long-horizon-autonomy.md).
|
|
4
|
+
|
|
5
|
+
Source: [Anthropic prompting guidance for Claude Fable 5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5).
|
|
6
|
+
|
|
7
|
+
## Act on what you have
|
|
8
|
+
|
|
9
|
+
Facts already settled in the conversation remain settled. A recommendation gives the user one course to evaluate. A list of discarded options adds reading without changing that course. This governs user-facing messages and leaves private reasoning free to examine alternatives.
|
|
10
|
+
|
|
11
|
+
Unclear intent calls for research and a recommendation before action. Clear intent and available evidence allow the run to proceed.
|
|
12
|
+
|
|
13
|
+
## Pause and resume
|
|
14
|
+
|
|
15
|
+
A pause asks for one missing input in the channel the user reads. It names what resumes when the answer arrives. `AskUserQuestion` is the Claude tool for that request. A status line, checklist, working document, or side thread leaves the request unseen by the user.
|
|
16
|
+
|
|
17
|
+
Independent work continues while an answer is outstanding. Authority granted earlier in the task remains in force. A later preference about tone or report length changes the message. The task's earlier authorization stays in force.
|
|
18
|
+
|
|
19
|
+
Before ending a turn, inspect its last paragraph. These endings leave work owed to the user:
|
|
20
|
+
|
|
21
|
+
- A summary that names the next step without starting it.
|
|
22
|
+
- An offer to continue unless the user objects.
|
|
23
|
+
- A list of decisions that does not block remaining work.
|
|
24
|
+
- A pause because the turn ran long or reached a milestone.
|
|
25
|
+
|
|
26
|
+
A harness reads a turn ending in text as a report. When checklist items remain and the report names no blocker, it sends a short continuation prompt. The harness stops after two or three continuations on one task. [Prompting Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#unattended-agentic-runs) describes this loop.
|
|
27
|
+
|
|
28
|
+
In an autonomous pipeline, the user cannot answer during the run. Reversible actions already authorized by the task continue, and follow-up offers wait until completion.
|
|
29
|
+
|
|
30
|
+
## Delegate and keep working
|
|
31
|
+
|
|
32
|
+
A large independent track can run in a subagent while the lead handles other work. A few reads, a handful of edits, or a simple check fit in the lead's turn. Reusing a long-lived subagent across related tasks preserves its context. The lead steps in when that agent drifts or lacks context.
|
|
33
|
+
|
|
34
|
+
For a task that spawned workers, the completion condition and checklist live in the [workers-done-before-complete guide](workers-done-before-complete.md).
|
|
35
|
+
|
|
36
|
+
## Ground progress and the closing report
|
|
37
|
+
|
|
38
|
+
Progress claims come from tool results in the current session. Failed tests carry their output. Skipped steps are named. The [ASD-STE100 language rule](../../rules/asd-ste100-language.md) governs the wording. The first progress update uses one sentence. Later updates cover important discoveries or a change in direction.
|
|
39
|
+
|
|
40
|
+
Terse notes between tool calls can support the run. The final message briefs a reader who saw none of them. It opens with the outcome, then explains any input needed using plain names for each file, commit, or flag. A remaining context or token count does not end an unfinished run. Content the user must see word for word goes through the channel the harness provides for it.
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
# Cross-skill duplicate helpers
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/no-cross-skill-duplicate-helpers.md`](../../rules/no-cross-skill-duplicate-helpers.md).
|
|
4
|
+
|
|
5
|
+
## The two duplication cases differ
|
|
6
|
+
|
|
7
|
+
CODE_RULES "Reuse before create" / DRY gives one helper one home within a skill. Two `.py` modules in the same skill's `scripts/` directory with the same top-level function body fail the `code_rules_duplicate_body` check. The staged policy lint runs it through `code_rules_enforcer.py`.
|
|
8
|
+
|
|
9
|
+
Each skill folder installs on its own. A shared module that lives in one skill would break the other skill after removal or reinstall. A small launch helper, such as one that reads the browser registry entry and starts `chrome.exe`, can live in each skill.
|
|
10
|
+
|
|
11
|
+
## Dependency choices
|
|
12
|
+
|
|
13
|
+
A large helper may hold business logic or drift in ways that change behavior. The user chooses between two arrangements:
|
|
14
|
+
|
|
15
|
+
- Copy the helper into each skill and accept the drift risk.
|
|
16
|
+
- Declare a shared dependency in both skills. A published package in both `requirements` files or a `_shared` module installed into each skill can survive independent installation.
|
|
17
|
+
|
|
18
|
+
## What the advisory tells you
|
|
19
|
+
|
|
20
|
+
The `advise_cross_skill_duplicate_helper` check in `code_rules_duplicate_body` prints to stderr when a changed top-level function matches a helper in another skill's `scripts/` directory. It leaves the lint passing. The message names the source skill and function for review. The advisory covers copies across skill folders; the blocking check covers copies within one skill.
|
|
21
|
+
|
|
22
|
+
## Why the checks have different scopes
|
|
23
|
+
|
|
24
|
+
The cross-skill advisory preserves the small-copy option that keeps skills independently installable. Its source-skill name gives the writer enough context to judge the copy. The blocking check handles duplication inside one skill.
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
# Prefer existing tools
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/prefer-existing-tools.md`](../../rules/prefer-existing-tools.md), which loads in sessions as the short form.
|
|
4
|
+
|
|
5
|
+
## What a candidate must cover
|
|
6
|
+
|
|
7
|
+
Internal code includes the repository and its shared packages. An established external option has a security process, a field reputation, broad use measured by stars, downloads, or dependents, and maintenance shown by a release or commit within the last year. An open critical advisory rules it out.
|
|
8
|
+
|
|
9
|
+
Custom rules belong in the chosen tool's configuration format. A wrapper adds code that this package must maintain.
|
|
10
|
+
|
|
11
|
+
The choice report identifies the tool by name and link, gives its license, and cites one adoption count. A custom build's report records the candidates and the missing capability that ruled out each one.
|
|
12
|
+
|
|
13
|
+
## Why this rule exists
|
|
14
|
+
|
|
15
|
+
Code written here needs local maintenance. A widely used tool draws fixes and new patterns from its maintainers and users.
|
|
16
|
+
|
|
17
|
+
## Sibling rules
|
|
18
|
+
|
|
19
|
+
| Rule | Role |
|
|
20
|
+
|---|---|
|
|
21
|
+
| [`explore-thoroughly.md`](../../rules/explore-thoroughly.md) | Read existing code before proposing a change |
|
|
22
|
+
| [`verify-before-asking.md`](../../rules/verify-before-asking.md) | Check a question with a tool before asking it |
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Rules prose names only hooks that run
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/retired-hook-prose.md`](../../rules/retired-hook-prose.md), which loads in sessions as the short form.
|
|
4
|
+
|
|
5
|
+
## How a hook reaches a tool call
|
|
6
|
+
|
|
7
|
+
A hook module runs when `hooks/hooks.json` registers it or a dispatcher roster hosts it. A module can survive on disk while appearing in neither place. The installer writes an inert stand-in at a retired hook path when stale `settings.json` still names it. A missing registered path makes the interpreter exit 2, which a PreToolUse harness reads as a block.
|
|
8
|
+
|
|
9
|
+
A retired hook's logic can move into a `*_parts/` package imported by a live module. The old hook name then stops while its check still fires. A present-tense claim about the old module sends the reader to a gate they will never encounter.
|
|
10
|
+
|
|
11
|
+
The staged policy lint registers `retired-hook-prose` in `scripts/policy_lint/registry.py`. `accepts_instruction_markdown` covers Markdown under `.agents/`, `commands/`, `docs/`, `output-styles/`, `rules/`, and `system-prompts/`. The detector reads backticked hook module names, checks registration and dispatcher rosters, and reports a present-tense action attributed to a module with no live registration. The lint does not find a forced detour that names only an agent, step, or token.
|
|
12
|
+
|
|
13
|
+
A module name counts as a hook when it ends in a hook family suffix such as `_blocker`, `_enforcer`, `_gate`, `_tracker`, or `_dispatcher`, and either sits under `hooks/` or appears in `RETIRED_HOOK_REGISTRATION_RELATIVE_PATHS` in `bin/install.mjs`. A support module without a hook suffix stays outside that check.
|
|
14
|
+
|
|
15
|
+
Past-tense history makes no live claim. A sentence naming the staged policy lint, `cde_lint`, or a repository check identifies where the work runs now.
|
|
16
|
+
|
|
17
|
+
## Why gate detours matter
|
|
18
|
+
|
|
19
|
+
A gate and the standing orders written to satisfy it form one unit. A required agent spawn, token, or extra step can remain after the gate ends because those orders often omit the module name the lint searches for. The stale order costs every later session an unnecessary action.
|
|
20
|
+
|
|
21
|
+
A withdrawn threshold can also remain in an audit rubric, an agent instruction table, or a review prompt. Removing only the mechanical check leaves a judgment lane enforcing the same threshold. The archive entry records each lane so a restore can recover the complete decision.
|
|
22
|
+
|
|
23
|
+
## Sibling rules
|
|
24
|
+
|
|
25
|
+
| Rule | Role |
|
|
26
|
+
|---|---|
|
|
27
|
+
| [`archiving-agent-config.md`](../../rules/archiving-agent-config.md) | Archive a retired rule or capability with a restore record |
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# Testing standards
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/testing.md`](../../rules/testing.md).
|
|
4
|
+
|
|
5
|
+
> **Reference:** [`TEST_QUALITY.md`](../TEST_QUALITY.md) for test review.
|
|
6
|
+
|
|
7
|
+
## Complete mocks for testability
|
|
8
|
+
|
|
9
|
+
If a component renders field X, a mock with a valid X value lets the test reach rendering behavior. An omitted field can make a failing render ambiguous.
|
|
10
|
+
|
|
11
|
+
## Tests exercise production behavior
|
|
12
|
+
|
|
13
|
+
A stand-in can satisfy an assertion while the production path stays untested. The behavior under review is the result of the production call.
|
|
14
|
+
|
|
15
|
+
## Test order
|
|
16
|
+
|
|
17
|
+
The TDD skill (`pstack:tdd`) carries the red, green, refactor procedure. No hook or lint checks the order. Review reads the reproducing test on the diff.
|
|
18
|
+
|
|
19
|
+
## Fix proof job
|
|
20
|
+
|
|
21
|
+
The `Fix test proof` job in `.github/workflows/pr-check.yml` runs `_shared/pr-loop/scripts/fix_pr_test_proof.py` on pull requests whose title starts with `fix`. Node proof tests end in `.test.mjs`, `.test.js`, or `.test.cjs`, and the job runs them with `node --test`. A docs-only or CI-only fix passes. A fix proven by a PowerShell test also needs Python or Node proof.
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
# Workers done before complete
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/workers-done-before-complete.md`](../../rules/workers-done-before-complete.md).
|
|
4
|
+
|
|
5
|
+
## The completion gate
|
|
6
|
+
|
|
7
|
+
The gate applies to subagents, workflow agents, and background shells spawned by a task. The run state can be `state.json`, `pr-converge-state.json`, a task list, or another record the task keeps. A finished worker with an unmerged result leaves the task open.
|
|
8
|
+
|
|
9
|
+
The gate controls task status. Work that can proceed while a worker runs continues. A worker that dies or hangs becomes a finding to record and report. A wakeup returns the run to the outstanding workers.
|
|
10
|
+
|
|
11
|
+
## Checklist before marking complete
|
|
12
|
+
|
|
13
|
+
| Check | Action |
|
|
14
|
+
|---|---|
|
|
15
|
+
| Are any spawned workers still running? | List them. If any run, stay `in_progress` and schedule a wakeup. |
|
|
16
|
+
| Did every finished worker return a result? | Read each result. Report a dead or hung worker as a finding. |
|
|
17
|
+
| Is each result merged into run state? | Write it to the run record before closing. |
|
|
18
|
+
| Does the task's goal hold? | Check the merged state and the repository or diff behind worker claims. |
|
|
19
|
+
|
|
20
|
+
The file list, count, description, and finding in each worker report need repository and diff evidence before they enter run state or a user report.
|
|
21
|
+
|
|
22
|
+
## Examples
|
|
23
|
+
|
|
24
|
+
An audit task has two workers still running. The lead lists them, keeps the task `in_progress`, and schedules a wakeup to collect their results.
|
|
25
|
+
|
|
26
|
+
A worker crashes while the others finish. The lead records the crash as a finding, reports it, and keeps the task open until the missing work is covered.
|
|
27
|
+
|
|
28
|
+
## Relationship to other rules
|
|
29
|
+
|
|
30
|
+
The [long-horizon-autonomy guide](long-horizon-autonomy.md) covers acting on available evidence and finishing owed work before ending a turn. This gate defines the completion condition for tasks with workers.
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# Workflow substitution slots
|
|
2
|
+
|
|
3
|
+
Full text behind [`rules/workflow-substitution-slots.md`](../../rules/workflow-substitution-slots.md).
|
|
4
|
+
|
|
5
|
+
## Values in agent-prompt templates
|
|
6
|
+
|
|
7
|
+
An agent fills angle-bracket slots such as `<plate.svg>`, `<object.svg>`, and `<glow_hex>` for each call. The same convention applies to an iteration index in a path or output key. `cand_<i>` names a changing path segment. `cand_i` names one literal directory, so repeated iterations can overwrite the same output and collapse an N-iteration gate into one run.
|
|
8
|
+
|
|
9
|
+
For a looped path or key, use `cand_<i>`. Step text can instead say `replace <i> with the iteration index 0, 1, 2`. Each call then gets its own value.
|
|
10
|
+
|
|
11
|
+
## Enforcement
|
|
12
|
+
|
|
13
|
+
The staged policy lint runs `workflow-substitution` on `.workflow.js` files. It reports a bare `<word>_<i|j|k>` token used as a path segment in looped content. CI runs the lint against the merge base. A write-time hook does not report this case, so the token remains on disk until lint runs.
|
package/package.json
CHANGED
|
@@ -8,64 +8,12 @@ paths:
|
|
|
8
8
|
- "**/tests/**"
|
|
9
9
|
---
|
|
10
10
|
|
|
11
|
-
# Anti-
|
|
11
|
+
# Anti-corollary tests
|
|
12
12
|
|
|
13
|
-
**When
|
|
13
|
+
**When:** Add or change a test.
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
Before keeping a test, name one code change it would catch; replace a case that catches none. Skip spelling cross products after proving canonicalization once; compare a few discriminating cases. Exercise the production path with a non-default expected result so a dead implementation cannot pass. Test policy compliance on governed surfaces; in an audit, name a mutation and record how many tests it fails.
|
|
16
16
|
|
|
17
|
-
|
|
17
|
+
**Enforcement:** none, the agent applies it.
|
|
18
18
|
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
### 1. Is this a corollary?
|
|
22
|
-
|
|
23
|
-
When the code reduces each input to a canonical form and then compares the forms, every pairwise combination of input spellings follows once the reduction is proven canonical. Walking the full N×N matrix restates that fact. It adds cases and runtime, and it buries the few cases that carry information.
|
|
24
|
-
|
|
25
|
-
**Do this:** test the reduction once. Test the comparison with a few discriminating cases. Do not walk the cross product of spellings.
|
|
26
|
-
|
|
27
|
-
### 2. Could this test pass if the mechanism were dead?
|
|
28
|
-
|
|
29
|
-
Name the degenerate value a dead implementation would return — empty string, `None`, `False`, a blanket refusal, an empty collection. When the test's expected value equals that default, the test passes whether the mechanism works or not. On its own it proves nothing.
|
|
30
|
-
|
|
31
|
-
**Do this:** keep at least one case that expects the **non-default** answer, and drive the production code path — not a mock that only records that a call happened.
|
|
32
|
-
|
|
33
|
-
### 3. What single change to the code would make this test fail?
|
|
34
|
-
|
|
35
|
-
When the honest answer is "none," or "only a change that also breaks everything else," the test is decoration. Drop it or rewrite it so one named mutation kills it.
|
|
36
|
-
|
|
37
|
-
**The policy-surface case.** A test that asserts a policy's own wording appears in the file that declares the policy has one mutation available: editing that sentence. It cannot report whether anything obeys the policy, so a suite full of such tests measures how many files repeat a rule, which is the duplication the reader was already complaining about. Assert instead that the surfaces the policy governs comply with it, and let the declaring file be the input rather than the subject.
|
|
38
|
-
|
|
39
|
-
## What a mechanism with a degenerate failure mode needs
|
|
40
|
-
|
|
41
|
-
1. **At least one non-default case** that exercises the live path and expects the non-default answer.
|
|
42
|
-
2. **A stated mutation (audit lane):** name one specific change to the code and record how many tests it kills. A mutation that kills zero tests means the suite proves nothing. A reviewer or an audit skill checks this; a hook does not compute it.
|
|
43
|
-
3. **A few discriminating cases** in place of a large matrix.
|
|
44
|
-
|
|
45
|
-
## Worked shape (sanitized)
|
|
46
|
-
|
|
47
|
-
A write guard decides whether a write is about to hit a production database. It reduces each database URL to a canonical endpoint identity, then compares identities.
|
|
48
|
-
|
|
49
|
-
Two independent mutations show the two halves of a useful suite:
|
|
50
|
-
|
|
51
|
-
- Gut the reduction so it always returns the empty string: the guard fails **closed** and refuses everything. Only the *allow*-expecting cases die.
|
|
52
|
-
- Abandon the reduction and compare raw hostnames: the guard fails **open** and allows a production write. Only the *refuse*-expecting cases die.
|
|
53
|
-
|
|
54
|
-
Opposite breaks kill opposite halves. A suite with only one half cannot see one of those breaks. Build both halves; skip the spelling matrix once the reduction is covered.
|
|
55
|
-
|
|
56
|
-
## What this is not
|
|
57
|
-
|
|
58
|
-
A structural hook is the wrong tool here. "Is this a corollary?" and "would this pass against a dead implementation?" need the intent of the code under test. A hook that pattern-matches `parametrize` breadth or counts assertions fires on correct suites and trains people to ignore it. A false-positive gate on a judgment call is worse than no gate. This rule stays judgment-only, with the stated-mutation check in the audit lane — the same pattern as other judgment rules in this package that have no Write/Edit blocker.
|
|
59
|
-
|
|
60
|
-
## Sibling rules
|
|
61
|
-
|
|
62
|
-
| Rule | Role |
|
|
63
|
-
|---|---|
|
|
64
|
-
| `code-standards.md` | Points at CODE_RULES §8: write a failing test before production code |
|
|
65
|
-
| `testing.md` | Mocks and test infrastructure standards |
|
|
66
|
-
| `paired-test-coverage.md` | Every public function in an established suite gets a behavioral test |
|
|
67
|
-
| `anti-corollary-tests.md` | Each test carries information; no corollary matrices; no suite that only matches the dead default |
|
|
68
|
-
|
|
69
|
-
## Enforcement
|
|
70
|
-
|
|
71
|
-
The AI review lane and audit skills carry this rule: an agent applies it to the test lines a PR changes. No blocking hook backs it, because corollary and dead-default judgments need meaning a regex cannot read.
|
|
19
|
+
**Full text:** [`docs/rule-guides/anti-corollary-tests.md`](../docs/rule-guides/anti-corollary-tests.md). Read it when choosing cases or auditing mutation evidence.
|
|
@@ -2,87 +2,19 @@
|
|
|
2
2
|
paths:
|
|
3
3
|
- "**/rules/**"
|
|
4
4
|
- "**/rules-archived/**"
|
|
5
|
-
- "
|
|
5
|
+
- "**/.agents/skills/**"
|
|
6
6
|
- "**/skills-archived/**"
|
|
7
7
|
- "**/agents/**"
|
|
8
8
|
- "**/commands/**"
|
|
9
9
|
- "**/ever-shipped-skills.mjs"
|
|
10
10
|
---
|
|
11
11
|
|
|
12
|
-
# Archiving a
|
|
12
|
+
# Archiving a rule, skill, agent, or command
|
|
13
13
|
|
|
14
|
-
**When
|
|
15
|
-
rule, skill, agent, or command out of service, and deciding whether one is
|
|
16
|
-
still used at all.
|
|
14
|
+
**When:** Before archiving configuration.
|
|
17
15
|
|
|
18
|
-
|
|
16
|
+
Trace checks and callers; `CHANGELOG.md`, `README.md`, shipped-skills registry, and host paths prove neither use nor disuse. Read `rules-archived/ARCHIVE-MANIFEST.md`; skip and report exemptions, then continue. Move unused files with `git mv` into sibling archives; record the reason, restore command, and companion edits. Retain skill names in `bin/ever-shipped-skills.mjs` and check tests lost when a skill directory moves.
|
|
19
17
|
|
|
20
|
-
|
|
21
|
-
confident wrong answers here.
|
|
18
|
+
**Enforcement:** none, the agent applies it.
|
|
22
19
|
|
|
23
|
-
**
|
|
24
|
-
this package is rarely named after the file that describes it. Resolve what the
|
|
25
|
-
file names instead: does the script exist, does the lint rule id appear in the
|
|
26
|
-
registry, does the check run. A rule with no mention of its slug anywhere can
|
|
27
|
-
still have a live checker behind it under a different name.
|
|
28
|
-
|
|
29
|
-
**Three files name every capability and prove nothing about any of them.**
|
|
30
|
-
`CHANGELOG.md` records history. The `README.md` inventory table lists what
|
|
31
|
-
ships. `bin/ever-shipped-skills.mjs` is a registry the prune reads, and it keeps
|
|
32
|
-
archived names on purpose. Subtract all three before you read a referrer count.
|
|
33
|
-
|
|
34
|
-
**A path this tree does not hold is not by itself a stale reference.** This
|
|
35
|
-
package installs into other repositories, so a rule or skill naming a file that
|
|
36
|
-
lives in the host repository is working as intended. Ask whether the specific
|
|
37
|
-
script or id it names resolves where it would run, not whether the path exists
|
|
38
|
-
here.
|
|
39
|
-
|
|
40
|
-
## Archive rather than delete
|
|
41
|
-
|
|
42
|
-
Move the file with `git mv` into the archive sibling: `rules-archived/` beside
|
|
43
|
-
`rules/`, `.agents/skills-archived/` beside `.agents/skills/`. The archive is a
|
|
44
|
-
sibling and never a child, because the installer copies a content directory
|
|
45
|
-
whole.
|
|
46
|
-
|
|
47
|
-
Archives ship nothing by omission. `CONTENT_DIRECTORIES` in `bin/install.mjs`
|
|
48
|
-
names the directories the installer writes, and no archive sibling is on that
|
|
49
|
-
list; skills are enumerated from `.agents/skills` alone. So an archived file
|
|
50
|
-
leaves every install without any further step, and it comes back with the
|
|
51
|
-
reverse `git mv`.
|
|
52
|
-
|
|
53
|
-
Record a row in the archive manifest for each file: why it went, the `git mv`
|
|
54
|
-
that restores it, and every other edit a restore has to undo. A restore that has
|
|
55
|
-
to be reconstructed from the diff is not recoverable in any useful sense.
|
|
56
|
-
|
|
57
|
-
## Exempt files
|
|
58
|
-
|
|
59
|
-
The "Never archived" section of `rules-archived/ARCHIVE-MANIFEST.md` names the
|
|
60
|
-
files this procedure leaves alone. Read that section before a prune moves
|
|
61
|
-
anything. A pass that reaches an exempt file stops there and reports which file
|
|
62
|
-
it was, so the prune's other moves still land.
|
|
63
|
-
|
|
64
|
-
`rules/correction-lens.md` is on that list today. Editing an exempt file to
|
|
65
|
-
sharpen it is ordinary work.
|
|
66
|
-
|
|
67
|
-
## Make the archive stick
|
|
68
|
-
|
|
69
|
-
Two things do not follow the file on their own.
|
|
70
|
-
|
|
71
|
-
**A skill's name stays in `bin/ever-shipped-skills.mjs`.** The prune computes
|
|
72
|
-
the retired set as every name ever shipped minus the names installed now, and
|
|
73
|
-
that is how a stale copy is moved out of a host's agents home. Delete the name
|
|
74
|
-
and the copy is stranded on every machine that has it. The reference checker in
|
|
75
|
-
`scripts/active_capability_references.py` derives its retired set the same way,
|
|
76
|
-
so the registry entry is also what makes a leftover mention of the archived
|
|
77
|
-
name report. No second list needs the name.
|
|
78
|
-
|
|
79
|
-
**A renamed skill directory takes its tests out of the run.** The node test
|
|
80
|
-
command globs `.agents/skills/**/*.test.mjs`, so a suite under an archived skill
|
|
81
|
-
stops running without failing. Check that the count you expect still runs.
|
|
82
|
-
|
|
83
|
-
## Sibling rules
|
|
84
|
-
|
|
85
|
-
| Rule | Role |
|
|
86
|
-
|---|---|
|
|
87
|
-
| [`retired-hook-prose.md`](retired-hook-prose.md) | Prose names only hooks that run, and retiring a gate drops the detours it required |
|
|
88
|
-
| [`doc-inventory-integrity.md`](doc-inventory-integrity.md) | Inventory tables stay in step with what ships |
|
|
20
|
+
**Full text:** [`docs/rule-guides/archiving-agent-config.md`](../docs/rule-guides/archiving-agent-config.md). Read it before moving or restoring.
|
|
@@ -5,60 +5,16 @@ paths:
|
|
|
5
5
|
- "**/bin/**"
|
|
6
6
|
- "**/ci/**"
|
|
7
7
|
- "**/tools/**"
|
|
8
|
-
- "**/skills/**"
|
|
8
|
+
- "**/skills/*/scripts/**"
|
|
9
9
|
- "**/commands/**"
|
|
10
10
|
---
|
|
11
11
|
|
|
12
|
-
# Build
|
|
12
|
+
# Build needs a user
|
|
13
13
|
|
|
14
|
-
**When
|
|
15
|
-
command, or workflow step. That includes a hook or check built to answer a
|
|
16
|
-
correction.
|
|
14
|
+
**When:** Before building a tool, script, hook, gate, skill, command, or workflow step.
|
|
17
15
|
|
|
18
|
-
|
|
16
|
+
Name its caller in this change, first-use date or event, and last manual use with frequency. Build only with all three answers; otherwise tell the owner in one line what the build is for and which answer is missing, then continue unaffected work. Search for an existing tool, size the build to first use, and turn a correction into a hook or lint only after it recurs.
|
|
19
17
|
|
|
20
|
-
|
|
18
|
+
**Enforcement:** `uncalled-new-file` checks new code under `scripts/`, `hooks/`, `bin/`, `ci/`, and `tools/` for a caller.
|
|
21
19
|
|
|
22
|
-
|
|
23
|
-
module that will run it in this same change.
|
|
24
|
-
2. **When does it first run?** Name the date or the event of first use.
|
|
25
|
-
3. **When was this job last done by hand?** Name the last time, and how often it
|
|
26
|
-
comes up.
|
|
27
|
-
|
|
28
|
-
Build when all three have a concrete answer. When the only answer is "it might
|
|
29
|
-
help" or "we could use it later", stop and send the owner one line: what the
|
|
30
|
-
build is for and which question has no answer. Keep the rest of the task moving
|
|
31
|
-
while that line waits.
|
|
32
|
-
|
|
33
|
-
Search first, per [`prefer-existing-tools.md`](prefer-existing-tools.md). A
|
|
34
|
-
tool that already exists answers the need, and the build is the glue.
|
|
35
|
-
|
|
36
|
-
## Size the first build to the first use
|
|
37
|
-
|
|
38
|
-
Build the smallest version that serves the first caller. A correction starts
|
|
39
|
-
as a row in a rule file. It becomes a hook or a lint only when the same
|
|
40
|
-
correction arrives a second time, per
|
|
41
|
-
[`correction-lens.md`](correction-lens.md).
|
|
42
|
-
|
|
43
|
-
## Where it is enforced
|
|
44
|
-
|
|
45
|
-
The staged policy lint's `uncalled-new-file` rule reads each code file a change
|
|
46
|
-
adds under `scripts/`, `hooks/`, `bin/`, `ci/`, or `tools/`. It reports the file when its name
|
|
47
|
-
appears only in its own tests, `CHANGELOG.md`, `README.md`, and
|
|
48
|
-
`bin/ever-shipped-skills.mjs`. CI runs that lint against the merge base, so a
|
|
49
|
-
file with no caller turns the pull request red.
|
|
50
|
-
|
|
51
|
-
## Why
|
|
52
|
-
|
|
53
|
-
The history of this package holds files that lived weeks and ran never. The
|
|
54
|
-
Codex compatibility watcher shipped 502 lines and 599 lines of tests in July,
|
|
55
|
-
and PR 1488 deleted it in September. Its own test was the only file that named
|
|
56
|
-
it.
|
|
57
|
-
|
|
58
|
-
## Sibling rules
|
|
59
|
-
|
|
60
|
-
| Rule | Role |
|
|
61
|
-
|---|---|
|
|
62
|
-
| [`prefer-existing-tools.md`](prefer-existing-tools.md) | Search before you build |
|
|
63
|
-
| [`correction-lens.md`](correction-lens.md) | A repeated correction moves up a layer |
|
|
64
|
-
| [`archiving-agent-config.md`](archiving-agent-config.md) | How an unused capability leaves service |
|
|
20
|
+
**Full text:** [`docs/rule-guides/build-needs-a-user.md`](../docs/rule-guides/build-needs-a-user.md). Read it when sizing a new check.
|
|
@@ -9,52 +9,12 @@ paths:
|
|
|
9
9
|
- "**/scripts/**"
|
|
10
10
|
---
|
|
11
11
|
|
|
12
|
-
# Falsify
|
|
12
|
+
# Falsify before green
|
|
13
13
|
|
|
14
|
-
**When
|
|
14
|
+
**When:** A new test, probe, sweep, mutation check, or measurement script first reports green.
|
|
15
15
|
|
|
16
|
-
|
|
16
|
+
Apply a named break that the same check must catch, and run a passing control beside it on the same command. Record the break, its failing output, and the control before counting the restored green as evidence. If the check stays green under its break, fix its reach; a probe needs a trip input, a sweep needs a planted violation in its claimed file set, and a writer assertion needs a disabled production writer.
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
**Enforcement:** none, the agent applies it.
|
|
19
19
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
## The four shapes this stops
|
|
23
|
-
|
|
24
|
-
### 1. A probe whose trigger condition never fires
|
|
25
|
-
|
|
26
|
-
The probe reports zero because it measured zero events, not because the code is clean. Its counter sits at its start value for the whole run.
|
|
27
|
-
|
|
28
|
-
**Break to apply:** feed it one input that must trip it. A probe still at zero on that input measures nothing.
|
|
29
|
-
|
|
30
|
-
### 2. A sweep that reads a subset of the files it claims to cover
|
|
31
|
-
|
|
32
|
-
The sweep compares against the wrong base, walks a slice of the tree, and reports its finding count over the full set.
|
|
33
|
-
|
|
34
|
-
**Break to apply:** plant one violation in a file the sweep's coverage claim names. A sweep that misses the plant walks a smaller file set than the one it reports.
|
|
35
|
-
|
|
36
|
-
### 3. A mutation that survives
|
|
37
|
-
|
|
38
|
-
The test meant to kill the mutation never reaches the mutated code — a mock stands in for the call, a guard returns early, or the test drives a neighboring branch.
|
|
39
|
-
|
|
40
|
-
**Break to apply:** hold the mutation in place and run the test. A green test names a line nothing covers.
|
|
41
|
-
|
|
42
|
-
### 4. An assertion that counts an artifact the harness seeded
|
|
43
|
-
|
|
44
|
-
The harness writes the row, file, or event the assertion counts, so the assertion tracks the harness rather than the code under test.
|
|
45
|
-
|
|
46
|
-
**Break to apply:** stub the production writer to a no-op. A green assertion counts the seed.
|
|
47
|
-
|
|
48
|
-
## What a shown-red record holds
|
|
49
|
-
|
|
50
|
-
| Part | What it names |
|
|
51
|
-
|---|---|
|
|
52
|
-
| The break | The mutation, stub, or trip input applied, named by file and line or by the exact input text |
|
|
53
|
-
| The red | The failing output the check printed under that break |
|
|
54
|
-
| The control | The case that passes beside the red, run on the same command |
|
|
55
|
-
|
|
56
|
-
All three land together. A record carrying the red alone shows a check that fails on everything; a record carrying the control alone shows a check that passes on everything.
|
|
57
|
-
|
|
58
|
-
## Enforcement
|
|
59
|
-
|
|
60
|
-
This rule binds as prose discipline: a reviewer reads the shown-red record beside each new check a PR adds. No hook backs it, because a green that measured the code and a green that measured nothing look the same to a regex — the difference sits in what the check reached at run time.
|
|
20
|
+
**Full text:** [`docs/rule-guides/falsify-before-green.md`](../docs/rule-guides/falsify-before-green.md). Read it when designing a break or reviewing a red record.
|
|
@@ -9,126 +9,12 @@ paths:
|
|
|
9
9
|
- "**/.github/workflows/**"
|
|
10
10
|
---
|
|
11
11
|
|
|
12
|
-
# Flag
|
|
12
|
+
# Flag non-breaking findings
|
|
13
13
|
|
|
14
|
-
**When
|
|
15
|
-
pre-commit entry, a Claude tool-use hook — and deciding what one of its checks
|
|
16
|
-
does when it finds something.
|
|
14
|
+
**When:** When writing a local gate or changing a check's severity.
|
|
17
15
|
|
|
18
|
-
|
|
16
|
+
Give every check ID a breaking or smell severity row; style and structure findings are smells. Fail the gate and stop for a bug, secret, broken test, syntax error, or instruction file that cannot load. Record a smell in `.claude/followups/smells.jsonl`, proceed, and fix it in a later pull request; ledger failure leaves the gate decision unchanged. Treat an unclassified check as breaking.
|
|
19
17
|
|
|
20
|
-
|
|
18
|
+
**Enforcement:** none, the agent applies it.
|
|
21
19
|
|
|
22
|
-
|
|
23
|
-
tree, a broken test, a syntax error, an instruction file that fails to
|
|
24
|
-
load. The gate fails and the work stops there.
|
|
25
|
-
|
|
26
|
-
A **smell** finding means the change reads poorly: a length limit, a naming
|
|
27
|
-
convention, a prose term, a comment rule, a structural preference. The gate
|
|
28
|
-
records it and lets the commit, the push, or the tool call proceed. A later
|
|
29
|
-
pull request carries the fix.
|
|
30
|
-
|
|
31
|
-
Every check a gate runs declares which of the two it is. A check that cannot
|
|
32
|
-
say which one it raises is a breaking check until someone decides otherwise.
|
|
33
|
-
|
|
34
|
-
## Why
|
|
35
|
-
|
|
36
|
-
A smell that blocks a commit stops delivery for something the reader would
|
|
37
|
-
have fixed in the next pass anyway. The writer then reaches for the bypass
|
|
38
|
-
flag, and the bypass turns off the breaking checks beside it. The gate that
|
|
39
|
-
blocks on everything ends up enforcing nothing.
|
|
40
|
-
|
|
41
|
-
Recording the smell keeps both halves working. The breaking checks stay
|
|
42
|
-
sharp because nobody routes around them, and the smells still get fixed
|
|
43
|
-
because the ledger holds them until a pull request does.
|
|
44
|
-
|
|
45
|
-
## The ledger
|
|
46
|
-
|
|
47
|
-
A recorded finding lands in `.claude/followups/smells.jsonl` at the
|
|
48
|
-
repository root, one JSON object per line. `hooks/followup_ledger.py` writes
|
|
49
|
-
and reads it. Every write there is fail-safe, so a ledger failure leaves the
|
|
50
|
-
gate's decision unchanged.
|
|
51
|
-
|
|
52
|
-
| Field | What it carries |
|
|
53
|
-
|---|---|
|
|
54
|
-
| `rule_id` | The rule that raised the finding |
|
|
55
|
-
| `check_id` | The single check behind it, which a severity table keys on |
|
|
56
|
-
| `file_path` | The repository-relative path the finding names |
|
|
57
|
-
| `message` | The text a reader acts on |
|
|
58
|
-
| `severity` | The class the gate put it in |
|
|
59
|
-
| `origin_commit` | The revision checked out when it was recorded |
|
|
60
|
-
|
|
61
|
-
The origin commit groups a follow-up pull request by the change that raised
|
|
62
|
-
the findings. A smell seen again under a later revision keeps the revision
|
|
63
|
-
that first raised it, so one smell stays one record.
|
|
64
|
-
|
|
65
|
-
The ledger is per-checkout state and stays out of the repository.
|
|
66
|
-
|
|
67
|
-
## The check identifier
|
|
68
|
-
|
|
69
|
-
Most lint rules run one check, so their rule identifier already names it. The
|
|
70
|
-
`code-rules` and `validators` rules bundle many checks behind one identifier.
|
|
71
|
-
`scripts/policy_lint/check_catalog.py` resolves those to a `<rule>/<check>`
|
|
72
|
-
identifier, which `cde lint --format json` emits as `check_id` on every
|
|
73
|
-
diagnostic. A consumer partitions findings by that identifier rather than by
|
|
74
|
-
message text.
|
|
75
|
-
|
|
76
|
-
A message no catalog entry names resolves to `<rule>/unclassified`, which a
|
|
77
|
-
partition treats as blocking. A check whose wording moves reports louder
|
|
78
|
-
rather than going quiet, and the catalog's synchronization tests fail on the
|
|
79
|
-
same change.
|
|
80
|
-
|
|
81
|
-
## Reading and clearing it
|
|
82
|
-
|
|
83
|
-
| Command | What it does |
|
|
84
|
-
|---|---|
|
|
85
|
-
| `cde followup list` | Names every recorded follow-up |
|
|
86
|
-
| `cde followup ingest REPORT` | Records the diagnostics a policy-lint JSON report carries |
|
|
87
|
-
| `cde followup brief` | Writes the task an agent fixes them from |
|
|
88
|
-
| `cde followup clear` | Empties the ledger |
|
|
89
|
-
| `cde followup count` | Reports the backlog against the threshold |
|
|
90
|
-
|
|
91
|
-
`count` exits non-zero once the backlog passes
|
|
92
|
-
`FOLLOWUP_BACKLOG_THRESHOLD` in
|
|
93
|
-
`scripts/dev_env_scripts_constants/followup_constants.py`, so a scheduled job
|
|
94
|
-
escalates rather than opening one more quiet pull request. The number is the
|
|
95
|
-
repository's setting, and raising or lowering it is one edit there.
|
|
96
|
-
|
|
97
|
-
The `/fix-followups` command drives the whole pass: it reads the brief, fixes
|
|
98
|
-
each rule group, opens a pull request, and clears the ledger.
|
|
99
|
-
|
|
100
|
-
## Worked example
|
|
101
|
-
|
|
102
|
-
`scripts/validate_instruction_pairs.py` raises five findings and splits them
|
|
103
|
-
on this line. A missing governing `AGENTS.md`, an import text that differs,
|
|
104
|
-
and an instruction path that is not a regular file each stop the instructions
|
|
105
|
-
loading, so the gate fails. A non-canonical filename and a Git mode other than
|
|
106
|
-
100644 leave the instructions loading, so the gate records them and passes.
|
|
107
|
-
|
|
108
|
-
`SEVERITY_BY_RULE_ID` in that module is the whole declaration. Add a check,
|
|
109
|
-
add its row.
|
|
110
|
-
|
|
111
|
-
`scripts/repository_policy.py` splits the committed-tree checks the same way.
|
|
112
|
-
`SEVERITY_BY_CHECK_ID` in `repository_checks/config/constants.py` carries one
|
|
113
|
-
row per check id, and a check id with no row reads as breaking. The
|
|
114
|
-
`package-inventory` check is the one smell there. A production file whose
|
|
115
|
-
package inventory omits its row still imports and still runs, so the check
|
|
116
|
-
prints its finding with an `advisory:` prefix, records it in the ledger, and
|
|
117
|
-
leaves the tree passing. Every other committed-tree check blocks: a `CLAUDE.md`
|
|
118
|
-
naming a file that does not exist, an env-var row naming a file that never
|
|
119
|
-
reads the variable, a test outside the testpaths allowlist, and a tracked
|
|
120
|
-
secret.
|
|
121
|
-
|
|
122
|
-
`scripts/cde_lint.py` splits the policy-lint findings the same way.
|
|
123
|
-
`SEVERITY_BY_CHECK_ID` in `scripts/policy_lint/config/check_catalog_constants.py`
|
|
124
|
-
carries one row per smell, and a check id with no row reads as breaking. Three
|
|
125
|
-
checks are smells there: `test-pairing` and the two paired-test coverage
|
|
126
|
-
checks. The lint prints each as a warning, records it in the ledger, and exits
|
|
127
|
-
zero when no error remains.
|
|
128
|
-
|
|
129
|
-
## Sibling rules
|
|
130
|
-
|
|
131
|
-
| Rule | Role |
|
|
132
|
-
|---|---|
|
|
133
|
-
| [CI Owns the Gate](../.agents/skills/pr-lifecycle/SKILL.md#ci-owns-the-gate) | The full check suite runs once, on CI |
|
|
134
|
-
| [Git workflow](../.agents/skills/pr-lifecycle/SKILL.md#git-workflow) | A red required check blocks the branch |
|
|
20
|
+
**Full text:** [`docs/rule-guides/flag-non-breaking-findings.md`](../docs/rule-guides/flag-non-breaking-findings.md). Read it when adding a check.
|
|
@@ -4,55 +4,12 @@ paths:
|
|
|
4
4
|
- "**/skills/orchestrator-refresh/**"
|
|
5
5
|
---
|
|
6
6
|
|
|
7
|
-
# Long-
|
|
7
|
+
# Long-horizon autonomy
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
**When:** Run a long or unattended task.
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
Act on settled facts and authority, and recommend after researching unclear intent. When protected or irreversible actions, scope changes, or user-only input block the task, ask for one answer in the user-facing channel, name what resumes, and keep other work moving. Finish owed work without ending a turn on a promise, milestone, or context limit. Delegate independent tracks while working, ground claims in tool output, name skipped checks, make the first update one sentence, reserve later updates for discoveries, and close with the outcome and input.
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
**Enforcement:** none, the agent applies it.
|
|
14
14
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
This is the autonomous-run partner to the ambiguous case: research and recommend first when intent is unclear; this rule covers the clear case, where the evidence is already in hand.
|
|
18
|
-
|
|
19
|
-
## Do not end a turn on a promise
|
|
20
|
-
|
|
21
|
-
Pause for the user only when the work needs them: a destructive or irreversible action, a scope change, or input only they can give. When you hit one, ask through `AskUserQuestion` and end the turn. Do not end on a promise about work you have not done.
|
|
22
|
-
|
|
23
|
-
A pause is a request you deliver, not a state you enter. Put it in the channel the user reads, name the one input you need and what resumes the moment it arrives, and shape it so a single word answers it. A decision recorded in a status line, a checklist, a working document, or a side thread is a note; the user never learns they are the bottleneck, and the work sits. Keep everything that does not depend on the answer moving while it is outstanding.
|
|
24
|
-
|
|
25
|
-
Before you end any turn, read your last paragraph. If it is a plan, an analysis, a list of next steps, or a statement of intent ("I'll run the tests", "next I'll wire it up"), do that work with tool calls before you stop. End the turn only when the task is done or you are blocked on input only the user can give.
|
|
26
|
-
|
|
27
|
-
Four turn endings leave owed work unstarted. Each one is a stop to remove:
|
|
28
|
-
|
|
29
|
-
- A summary of the work so far that closes by naming the next step, with no tool call to start it.
|
|
30
|
-
- An offer to continue unless the user objects.
|
|
31
|
-
- A list of decisions for the user when none of them blocks the remaining work.
|
|
32
|
-
- A pause chosen because the turn ran long or a milestone finished.
|
|
33
|
-
|
|
34
|
-
A status note or a recommendation on an open decision rides in the same message as the next tool call. The stops that stay are the two where nothing can move: the next step needs the user, or the blocker is a control that is protected from you on purpose. Confirmation before a risky or destructive action still applies.
|
|
35
|
-
|
|
36
|
-
A harness or routine that drives an unattended run reads a turn that ends in text as a report. When open checklist items remain and the turn names no blocker, it sends one short message naming those items. It stops after two or three such continuations on one task, so a stuck run ends where a reader can review it. [Prompting Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#unattended-agentic-runs) is the source for this list and this loop.
|
|
37
|
-
|
|
38
|
-
In an autonomous pipeline the user cannot answer mid-task. For reversible actions that follow from the original request, act without asking; save any follow-up offers for after the task is done.
|
|
39
|
-
|
|
40
|
-
Authority the task already granted stays granted. A later preference about tone, brevity, or reply format changes how you report, never what you are allowed to do. Re-asking for permission the task already gave hands the work back to the user, who then does it by hand.
|
|
41
|
-
|
|
42
|
-
## Delegate and keep working
|
|
43
|
-
|
|
44
|
-
Delegate a large, independent track to a subagent, and keep working while it runs in the background. Do work you can finish in a few tool calls yourself: a few reads, a handful of edits, or a simple check. Reuse a long-lived subagent across related subtasks so its context carries forward and saves repeated reads. Step in when a subagent drifts off track or is missing context.
|
|
45
|
-
|
|
46
|
-
## Ground every progress claim
|
|
47
|
-
|
|
48
|
-
Before you report progress, check each claim against a tool result from this session. State only what the evidence backs; name anything unverified as unverified. If tests fail, say so with the output; if a step was skipped, say that.
|
|
49
|
-
|
|
50
|
-
Visible progress follows the [ASD-STE100 language policy](asd-ste100-language.md) and this rule's run-completion contract: the **first progress update is one sentence**; later updates mark only **important discoveries or a direction change**.
|
|
51
|
-
|
|
52
|
-
## Re-ground the final message
|
|
53
|
-
|
|
54
|
-
Terse shorthand between tool calls is fine — that is you thinking. The final message is for a reader who saw none of it. After a long or unwatched run, write it as a fresh briefing: the **outcome in one sentence first**, then the one or two things you need from the reader, each explained as if new. Drop the working vocabulary, arrow chains, and stacked-hyphen compounds; give each file, commit, or flag its own plain clause. When short and clear pull apart, choose clear.
|
|
55
|
-
|
|
56
|
-
## Keep going on context
|
|
57
|
-
|
|
58
|
-
A remaining-context or token count is not a reason to stop. Do not pause, summarize, or float a fresh session on account of context limits; keep working. When the user must see content word-for-word (a partial deliverable, a direct answer to a mid-run question), surface it through the channel the harness gives for that, not by ending the turn.
|
|
15
|
+
**Full text:** [`docs/rule-guides/long-horizon-autonomy.md`](../docs/rule-guides/long-horizon-autonomy.md). Read it when preparing the run.
|
|
@@ -2,28 +2,12 @@
|
|
|
2
2
|
paths: "**/skills/*/scripts/**/*.py"
|
|
3
3
|
---
|
|
4
4
|
|
|
5
|
-
# Cross-
|
|
5
|
+
# Cross-skill duplicate helpers
|
|
6
6
|
|
|
7
|
-
**When
|
|
7
|
+
**When:** Write a top-level Python helper in a skill `scripts/` directory.
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
Within one skill, extract a shared module and import it at both call sites. Across skill folders, copy a small self-contained helper when independent installation needs it, then confirm the advisory names the source. For a large or behavior-bearing body, ask the user to choose an intentional copy with drift risk or a dependency that survives independent install.
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
**Enforcement:** `code_rules_duplicate_body.py` blocks sibling-module copies and emits cross-skill advisories through `code_rules_enforcer.py`; the agent judges copy size.
|
|
12
12
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
## Decision
|
|
16
|
-
|
|
17
|
-
Before you copy a top-level helper from one skill's `scripts/` directory into another:
|
|
18
|
-
|
|
19
|
-
- **Same skill, two modules.** Extract one shared module in that skill and import it from both. The `code_rules_duplicate_body` check reports the copy.
|
|
20
|
-
- **Two skill folders, a small self-contained helper.** Copy it, and the skill-isolation tradeoff stands. An advisory `[CODE_RULES advisory]` line names the source skill when the lint runs, so the copy is a deliberate choice on record, not an oversight.
|
|
21
|
-
- **Two skill folders, a large or behavior-bearing body** — when the copied body is large, holds business logic, or would drift in a way that changes behavior, raise the choice through `AskUserQuestion`: copy and accept drift, or stand up a shared dependency both skills declare (for example a published package both `requirements` files name, or a `_shared` module the install step writes into each skill). A shared dependency that survives independent install is the only shared-home path that does not break the install model.
|
|
22
|
-
|
|
23
|
-
## What the advisory tells you
|
|
24
|
-
|
|
25
|
-
The `advise_cross_skill_duplicate_helper` check in `code_rules_duplicate_body` prints to stderr (never fails the lint) when a top-level function in a changed file has the same normalized body as a top-level function in another skill's `scripts/` directory. The message names the source skill and function so a reviewer can confirm the copy was intentional. It fires only across skill folders. Within one skill the reporting check already covers the copy.
|
|
26
|
-
|
|
27
|
-
## Why this is a rule, not a wider check
|
|
28
|
-
|
|
29
|
-
Extending the duplicate-body check to span skill folders would report the exact skill-isolation copy that keeps skills independently installable, a false positive on a sanctioned pattern. The boundary between "same skill, report" and "two skills, signal" is a judgment the writer makes with the source skill named in front of them. The rule states the judgment, and the `[CODE_RULES advisory]` line surfaces the signal. Neither one fails the defensible copy.
|
|
13
|
+
**Full text:** [`docs/rule-guides/no-cross-skill-duplicate-helpers.md`](../docs/rule-guides/no-cross-skill-duplicate-helpers.md). Read it when judging copy size or dependency ownership.
|
|
@@ -5,44 +5,18 @@ paths:
|
|
|
5
5
|
- "**/bin/**"
|
|
6
6
|
- "**/ci/**"
|
|
7
7
|
- "**/tools/**"
|
|
8
|
-
- "**/skills/**"
|
|
8
|
+
- "**/skills/*/scripts/**"
|
|
9
9
|
- "**/package.json"
|
|
10
10
|
- "**/pyproject.toml"
|
|
11
11
|
- "**/requirements*.txt"
|
|
12
12
|
---
|
|
13
13
|
|
|
14
|
-
# Prefer
|
|
14
|
+
# Prefer existing tools
|
|
15
15
|
|
|
16
|
-
**When
|
|
16
|
+
**When:** Before building a tool, check, script, or library.
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
Search local and shared code, then open-source options. Stop at first fit; build if neither fits. Choose an external option with an active security process, no open critical advisory, field reputation, broad use, and a release or commit within a year; add missing rules in its configuration without a wrapper. Report its name, link, license, and one use count; for custom code, name rejected candidates and reasons.
|
|
19
19
|
|
|
20
|
-
|
|
20
|
+
**Enforcement:** none, the agent applies it.
|
|
21
21
|
|
|
22
|
-
|
|
23
|
-
2. **An established open-source option.** Search for a plugin, package, or tool that already does the job.
|
|
24
|
-
3. **Your own code.** Build it only when steps 1 and 2 find nothing that fits.
|
|
25
|
-
|
|
26
|
-
Use an open-source option only when it meets all four conditions:
|
|
27
|
-
|
|
28
|
-
- It is secure: it has an active security process and no open critical advisory.
|
|
29
|
-
- It is well known in its field.
|
|
30
|
-
- It is widely used: many stars, downloads, or dependents.
|
|
31
|
-
- It is maintained: it had a release or commit in the last year.
|
|
32
|
-
|
|
33
|
-
Add only what the option lacks. Put custom rules in its own configuration format. Do not wrap it in a new layer of your own code.
|
|
34
|
-
|
|
35
|
-
## Report the choice
|
|
36
|
-
|
|
37
|
-
Tell the user which tool you chose before the work is marked ready. Give its name, link, license, and one number that shows how widely it is used. When you build your own code, name what you searched and why each candidate did not fit.
|
|
38
|
-
|
|
39
|
-
## Why
|
|
40
|
-
|
|
41
|
-
Code of your own is code you maintain alone. A trusted tool carries fixes and new patterns from its community, and its users find its defects first.
|
|
42
|
-
|
|
43
|
-
## Sibling rules
|
|
44
|
-
|
|
45
|
-
| Rule | Role |
|
|
46
|
-
|---|---|
|
|
47
|
-
| [`explore-thoroughly.md`](explore-thoroughly.md) | Read what exists before you propose a change |
|
|
48
|
-
| [`verify-before-asking.md`](verify-before-asking.md) | Answer a question with a tool before you ask it |
|
|
22
|
+
**Full text:** [`docs/rule-guides/prefer-existing-tools.md`](../docs/rule-guides/prefer-existing-tools.md). Read it when comparing candidates.
|
|
@@ -3,22 +3,12 @@ paths:
|
|
|
3
3
|
- "**/rules/*.md"
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Rules
|
|
6
|
+
# Rules prose names only hooks that run
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
**When:** When describing a hook or retiring a gate.
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
Verify the module and its hook registration or dispatcher roster before claiming it runs; describe a retired hook in past tense or name the check that carries its work now. When retiring a gate, remove the forced agent calls, tokens, extra steps, thresholds, audit rubrics, and review prompts it required across every instruction lane. Add a retired hook path to the installer roster when deleting its module.
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
**Enforcement:** `retired-hook-prose` in `scripts/policy_lint/registry.py` checks instruction Markdown, including `rules/` and `docs/`; the detour sweep is manual.
|
|
13
13
|
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
## Retiring a gate drops the detours it required
|
|
17
|
-
|
|
18
|
-
A gate and the standing orders written to satisfy it are one unit. An order that says to spawn a named agent first, to carry a magic token into a spawn, or to route through an extra step exists because some check demanded it. When the check goes, that order is a toll every later session pays for nothing, and it costs more than a stale sentence does, because a reader obeys it.
|
|
19
|
-
|
|
20
|
-
So retiring a hook is finished only when the detours it required are gone in the same change. Search for the order, not for the module: the module name is what the lint looks for, and a detour usually names the agent, the step, or the token instead, which is how these survive a lint that reads the same file. Drop the requirement and keep the agent, the skill, or the command itself; a capability nobody is forced to call costs nothing, while a forced call costs every run.
|
|
21
|
-
|
|
22
|
-
The same duty covers every lane a withdrawn standard rode on. A threshold enforced by a hook is often also stated in an audit rubric, an agent's instruction table, or a review prompt. Withdraw all of them together, and name the lanes in the archive entry. Removing the fast mechanical lane and leaving the slow judgment lane keeps the standard in force while deleting the text that a reader could argue with.
|
|
23
|
-
|
|
24
|
-
When the rule reports a line, say what the reader faces now. Name the harness permission prompt, the staged policy lint, or the repository check that carries the work, or drop the claim.
|
|
14
|
+
**Full text:** [`docs/rule-guides/retired-hook-prose.md`](../docs/rule-guides/retired-hook-prose.md). Read it when tracing a hook's current path or removing its gate.
|
package/rules/testing.md
CHANGED
|
@@ -8,25 +8,12 @@ paths:
|
|
|
8
8
|
- "**/tests/**"
|
|
9
9
|
---
|
|
10
10
|
|
|
11
|
-
# Testing
|
|
11
|
+
# Testing standards
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
**When:** Write or review behavior tests.
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
Read [`TEST_QUALITY.md`](../docs/TEST_QUALITY.md) and give mocks every field the component reads, with valid values. Assert results from production data and code paths. Use red, green, refactor for fixes and new behavior; add prototype tests before a pull request goes ready, and ship a reproducing test with each fix. A production fix titled `fix` needs a changed Python or Node test that fails on base and passes on head.
|
|
16
16
|
|
|
17
|
-
**
|
|
17
|
+
**Enforcement:** `useless_test_checks.py` and `code_rules_test_assertions.py` catch weak assertions; `Fix test proof` checks fix tests.
|
|
18
18
|
|
|
19
|
-
|
|
20
|
-
Incomplete mocks make it impossible to distinguish "broken code" from "missing data".
|
|
21
|
-
|
|
22
|
-
## Tests Exercise Production Behavior
|
|
23
|
-
|
|
24
|
-
Tests exercise production behavior, production data, and production code paths. A test that asserts on a stand-in for the production path proves the stand-in works.
|
|
25
|
-
|
|
26
|
-
## No Gate Holds the Test-First Order
|
|
27
|
-
|
|
28
|
-
Red, green, refactor is the default loop for a bug fix and for new behavior. The TDD skill (`pstack:tdd`) carries the procedure, and no hook or lint checks the order. A prototype may run ahead of its tests and adds them before the pull request goes ready. A bug fix ships with a test that reproduces the bug. Review reads the tests on the diff.
|
|
29
|
-
|
|
30
|
-
## A Fix Carries Its Proof Test
|
|
31
|
-
|
|
32
|
-
The `Fix test proof` job in `.github/workflows/pr-check.yml` runs `_shared/pr-loop/scripts/fix_pr_test_proof.py` on every pull request whose title starts with `fix`. A fix that changes production code must change at least one Python or Node test that fails on the base and passes on the head. Node tests are files ending in `.test.mjs`, `.test.js`, or `.test.cjs`, and the job runs them with `node --test`. A fix that changes only docs or CI config passes. A fix proven by a PowerShell test needs a Python or Node test beside it.
|
|
19
|
+
**Full text:** [`docs/rule-guides/testing.md`](../docs/rule-guides/testing.md). Read it for mock fields, test order, and fix proof details.
|
|
@@ -4,12 +4,12 @@ paths:
|
|
|
4
4
|
- "**/skills/orchestrator-refresh/**"
|
|
5
5
|
---
|
|
6
6
|
|
|
7
|
-
# Workers
|
|
7
|
+
# Workers done before complete
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
**When:** Mark a task `completed` after spawning subagents, workflow agents, or background shells.
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
List every worker and confirm that each has finished and its result is merged into run state. Verify worker file lists, counts, descriptions, and findings against the repository and diff before repeating them. While a worker runs or output is missing, keep the task `in_progress`, report dead or hung workers, schedule a wakeup, and keep other work moving. Check the task goal against merged state before closing.
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
**Enforcement:** none, the agent applies it.
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
**Full text:** [`docs/rule-guides/workers-done-before-complete.md`](../docs/rule-guides/workers-done-before-complete.md). Read it before closing a task that spawned workers.
|
|
@@ -3,10 +3,12 @@ paths:
|
|
|
3
3
|
- "**/*.workflow.js"
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Workflow
|
|
6
|
+
# Workflow substitution slots
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
**When:** Write a `.workflow.js` agent-prompt template with values filled per call or iteration.
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
Mark each changing value with angle brackets, including loop indices in paths or output keys such as `cand_<i>`.
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
**Enforcement:** `workflow-substitution` in the staged policy lint.
|
|
13
|
+
|
|
14
|
+
**Full text:** [`docs/rule-guides/workflow-substitution-slots.md`](../docs/rule-guides/workflow-substitution-slots.md). Read it when a loop builds a path or output key.
|
|
@@ -1,32 +0,0 @@
|
|
|
1
|
-
# Worker Completion Gate
|
|
2
|
-
|
|
3
|
-
Full detail behind the always-on `rules/workers-done-before-complete.md` kernel. It applies before marking any task `completed` when the task spawned workers — subagents, workflow agents, or background shells.
|
|
4
|
-
|
|
5
|
-
## The gate
|
|
6
|
-
|
|
7
|
-
A task reaches `completed` only when two things hold: every worker it spawned has finished, and each worker's result is merged into run state — `state.json`, `pr-converge-state.json`, the task list, or whatever record the task keeps. A worker that still runs, or one that finished but whose output never landed in run state, leaves the task open.
|
|
8
|
-
|
|
9
|
-
List the live workers before you mark the task complete. When a worker is dead or hung, that is a finding to record and report, not a result to drop in silence. A step that waits on workers ends its turn `in_progress` with a wakeup scheduled, so the run picks the workers back up rather than closing the task without them.
|
|
10
|
-
|
|
11
|
-
## Checklist before marking complete
|
|
12
|
-
|
|
13
|
-
| Check | Action |
|
|
14
|
-
|---|---|
|
|
15
|
-
| Are any spawned workers still running? | List them; when yes, stay `in_progress` and schedule a wakeup. |
|
|
16
|
-
| Did every finished worker return a result? | Read each result; a dead or hung worker is a finding to report. |
|
|
17
|
-
| Is each result merged into run state? | Write it to `state.json` or the task list before closing. |
|
|
18
|
-
| Does the task's own goal now hold? | Confirm against the merged state, not a worker's self-report. |
|
|
19
|
-
|
|
20
|
-
Mark `completed` only when every row passes.
|
|
21
|
-
|
|
22
|
-
## Examples
|
|
23
|
-
|
|
24
|
-
**Wrong:** Marking the audit task complete while two bugteam workers still run in the background.
|
|
25
|
-
**Right:** List the workers, see two still running, keep the task `in_progress`, and schedule a wakeup to collect them.
|
|
26
|
-
|
|
27
|
-
**Wrong:** A worker crashes; the task closes as complete because the other workers finished.
|
|
28
|
-
**Right:** Record the crashed worker as a finding, report it, and hold the task open until its work is covered.
|
|
29
|
-
|
|
30
|
-
## Relationship to other rules
|
|
31
|
-
|
|
32
|
-
- `long-horizon-autonomy` covers acting on what you have and not ending a turn on a promise. This gate names the specific completion condition: workers finished and their results merged.
|