universal-dev-standards 6.4.0 → 6.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled/ai/standards/agent-dispatch.ai.yaml +162 -0
- package/bundled/ai/standards/ai-friendly-architecture.ai.yaml +1 -1
- package/bundled/ai/standards/ai-instruction-standards.ai.yaml +190 -15
- package/bundled/ai/standards/ai-response-navigation.ai.yaml +43 -3
- package/bundled/ai/standards/class-level-fix.ai.yaml +38 -3
- package/bundled/ai/standards/commit-message.ai.yaml +2 -0
- package/bundled/ai/standards/model-selection.ai.yaml +370 -72
- package/bundled/ai/standards/mutation-testing.ai.yaml +105 -2
- package/bundled/ai/standards/project-structure.ai.yaml +1 -1
- package/bundled/ai/standards/security-standards.ai.yaml +22 -1
- package/bundled/ai/standards/spec-driven-development.ai.yaml +86 -2
- package/bundled/ai/standards/test-governance.ai.yaml +49 -2
- package/bundled/ai/standards/testing.ai.yaml +49 -3
- package/bundled/ai/standards/translation-lifecycle-standards.ai.yaml +4 -4
- package/bundled/ai/standards/verification-evidence.ai.yaml +48 -4
- package/bundled/core/ai-response-navigation.md +75 -2
- package/bundled/core/class-level-fix.md +26 -3
- package/bundled/core/model-selection.md +383 -125
- package/bundled/core/mutation-testing.md +41 -2
- package/bundled/core/spec-driven-development.md +57 -2
- package/bundled/core/test-governance.md +22 -2
- package/bundled/core/translation-lifecycle-standards.md +6 -6
- package/bundled/core/verification-evidence.md +42 -3
- package/bundled/locales/zh-CN/CHANGELOG.md +24 -3
- package/bundled/locales/zh-CN/README.md +1 -1
- package/bundled/locales/zh-CN/SECURITY.md +1 -1
- package/bundled/locales/zh-CN/core/ai-response-navigation.md +69 -5
- package/bundled/locales/zh-CN/core/model-selection.md +375 -60
- package/bundled/locales/zh-CN/core/mutation-testing.md +1 -1
- package/bundled/locales/zh-CN/core/spec-driven-development.md +1 -1
- package/bundled/locales/zh-CN/core/test-governance.md +1 -1
- package/bundled/locales/zh-CN/core/translation-lifecycle-standards.md +1 -1
- package/bundled/locales/zh-CN/core/verification-evidence.md +1 -1
- package/bundled/locales/zh-CN/docs/CHEATSHEET.md +7 -12
- package/bundled/locales/zh-CN/docs/FEATURE-REFERENCE.md +10 -15
- package/bundled/locales/zh-TW/CHANGELOG.md +49 -3
- package/bundled/locales/zh-TW/README.md +1 -1
- package/bundled/locales/zh-TW/SECURITY.md +1 -1
- package/bundled/locales/zh-TW/core/ai-response-navigation.md +69 -5
- package/bundled/locales/zh-TW/core/class-level-fix.md +22 -7
- package/bundled/locales/zh-TW/core/model-selection.md +385 -47
- package/bundled/locales/zh-TW/core/mutation-testing.md +45 -6
- package/bundled/locales/zh-TW/core/spec-driven-development.md +1 -1
- package/bundled/locales/zh-TW/core/test-governance.md +22 -3
- package/bundled/locales/zh-TW/core/translation-lifecycle-standards.md +1 -1
- package/bundled/locales/zh-TW/core/verification-evidence.md +33 -6
- package/bundled/locales/zh-TW/docs/CHEATSHEET.md +7 -12
- package/bundled/locales/zh-TW/docs/FEATURE-REFERENCE.md +10 -15
- package/bundled/locales/zh-TW/integrations/claude-code/README.md +31 -5
- package/package.json +1 -1
- package/src/utils/reference-sync.js +83 -16
- package/standards-registry.json +20 -8
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Mutation Testing Standards
|
|
2
2
|
|
|
3
|
-
**Version**: 1.
|
|
4
|
-
**Last Updated**: 2026-
|
|
3
|
+
**Version**: 1.1.0
|
|
4
|
+
**Last Updated**: 2026-08-14
|
|
5
5
|
**Applicability**: All software projects with unit/integration tests
|
|
6
6
|
**Scope**: universal
|
|
7
7
|
**Industry Standards**: ISTQB Foundation Syllabus (test effectiveness metrics)
|
|
@@ -30,6 +30,42 @@ A test with `expect(x).toBeDefined()` can achieve 100% line coverage but survive
|
|
|
30
30
|
|
|
31
31
|
---
|
|
32
32
|
|
|
33
|
+
## Attribution: Kill Is Credited to the First Failing Test
|
|
34
|
+
|
|
35
|
+
`Killed` means *some* test in the run failed against the mutant — most tools do not record *which* test killed it, and none records *which test level*. A 7/7 kill score therefore verifies the **whole suite that ran against the mutant**, not any single test, and not any single test level (unit vs. integration vs. property).
|
|
36
|
+
|
|
37
|
+
**Consequence**: "the property suite verifies X" is not a claim an aggregate mutation run supports. To support it, re-run mutation testing with *only* the property suite active:
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
npx stryker run --mutate 'src/module/**' -- --project=property
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
If the isolated run kills fewer mutants than the aggregate run, the gap is exactly what the unit/integration tests were quietly covering.
|
|
44
|
+
|
|
45
|
+
**Rule**: high-risk modules (the same set the 80% threshold applies to — auth/license/payment/security) must re-run mutants against the property suite in isolation before any claim of "property-verified" is made.
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
49
|
+
## One-Sided Invariants Miss Fail-Closed Defects
|
|
50
|
+
|
|
51
|
+
A property like "output never exceeds the limit" is one-sided: it cannot detect a mutant that makes the code **fail closed** — reject everything, including valid input — because a fail-closed mutant never produces an over-limit output. The mutation score looks unaffected while an entire class of defect (denial of service, wrongly rejected requests) is invisible to the suite.
|
|
52
|
+
|
|
53
|
+
**Fix**: pair every one-sided invariant with its opposite boundary. "Never exceeds the limit" needs a companion property — e.g. "accepts everything at or below the limit" — so that both over-permissive and over-restrictive mutants have a path to detection.
|
|
54
|
+
|
|
55
|
+
---
|
|
56
|
+
|
|
57
|
+
## Equivalent Mutants Are Not Survivors to Chase
|
|
58
|
+
|
|
59
|
+
A surviving mutant is not automatically a test gap. Some mutants are **semantically equivalent** to the original — no input can distinguish their behavior — and no test can kill them, regardless of how it's written. Forcing a kill with an assertion that exists only to move the score (`expect(x).toBeDefined()` on an incidental value) produces a hollow test without closing any real gap.
|
|
60
|
+
|
|
61
|
+
**Required classification for every reviewed survivor**:
|
|
62
|
+
- **Genuine gap** → write a test that exercises the distinguishing behavior.
|
|
63
|
+
- **`equivalent, because <reason>`** → recorded with the input(s) checked to reach that conclusion.
|
|
64
|
+
|
|
65
|
+
An unclassified survivor is not the same as a classified-equivalent one. Only a classified-equivalent mutant may be excluded from the score's denominator.
|
|
66
|
+
|
|
67
|
+
---
|
|
68
|
+
|
|
33
69
|
## Tools
|
|
34
70
|
|
|
35
71
|
| Language | Tool | Command |
|
|
@@ -87,6 +123,9 @@ npm install --save-dev @stryker-mutator/core @stryker-mutator/vitest-runner
|
|
|
87
123
|
- Adding mutation testing to CI for every PR (too slow)
|
|
88
124
|
- Accepting AI-generated tests without mutation score validation
|
|
89
125
|
- Killing mutations by adding `toBeDefined()` assertions
|
|
126
|
+
- Claiming "the property suite verifies X" from an aggregate (not isolated) mutation run
|
|
127
|
+
- Using only a one-sided invariant for a property that has a fail-closed failure mode
|
|
128
|
+
- Forcing a kill on a semantically equivalent mutant instead of classifying it `equivalent, because <reason>`
|
|
90
129
|
|
|
91
130
|
---
|
|
92
131
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Spec-Driven Development (SDD) Standards
|
|
2
2
|
|
|
3
|
-
**Version**: 2.
|
|
4
|
-
**Last Updated**: 2026-
|
|
3
|
+
**Version**: 2.4.0
|
|
4
|
+
**Last Updated**: 2026-08-17
|
|
5
5
|
**Applicability**: All projects adopting Spec-Driven Development
|
|
6
6
|
**Scope**: universal
|
|
7
7
|
**Industry Standards**: None (Emerging 2025+ methodology)
|
|
@@ -50,6 +50,61 @@ UDS supports two AC notations. **GWT is the default and preferred** (Forward Der
|
|
|
50
50
|
|
|
51
51
|
Provide **GWT or EARS** per AC (`.ac.yaml`: `given/when/then` **or** `ears`). Prefer GWT for BDD-derivable behaviour; reach for EARS when GWT feels forced. Do not require both; do not remove GWT.
|
|
52
52
|
|
|
53
|
+
## An AC With No Verification Item Is Not an AC | 沒有驗證項的 AC 不是 AC
|
|
54
|
+
|
|
55
|
+
Every acceptance criterion must have a **verification item that points at it** — a test, a
|
|
56
|
+
check, a gate, or an explicitly recorded manual step. An AC that no verification item
|
|
57
|
+
references is a **promise nobody kept**, and it does not fail loudly: it simply stops being
|
|
58
|
+
true while the spec continues to assert it.
|
|
59
|
+
|
|
60
|
+
每一條驗收標準都必須有一個**指向它的驗證項**——測試、檢查、閘門,或一則明確記錄的
|
|
61
|
+
手動步驟。沒有任何驗證項引用的 AC 是**一張沒有人兌現的支票**,而且它不會大聲失敗:
|
|
62
|
+
它只是安靜地停止成立,而規格繼續宣稱它為真。
|
|
63
|
+
|
|
64
|
+
**Rule**: an AC without a verification item must be **demoted to a design intent**, not
|
|
65
|
+
carried as an AC. Demotion is honest; an unverified AC is not.
|
|
66
|
+
**規則**:沒有驗證項的 AC 必須**降級為設計意圖**,不得繼續掛在 AC 欄。降級是誠實的,
|
|
67
|
+
未經驗證的 AC 不是。
|
|
68
|
+
|
|
69
|
+
### The measured instance | 實測案例
|
|
70
|
+
|
|
71
|
+
A 2026-05-14 spec carried `AC-7: the legacy system-report is fully preserved, no
|
|
72
|
+
regression`. Its Test Plan had seven items and **none of them pointed at AC-7**. The
|
|
73
|
+
report's timer was disabled and its deploy function was never called — **from the same day
|
|
74
|
+
the AC was written**. It was found three months later, by accident, while verifying an
|
|
75
|
+
unrelated install.
|
|
76
|
+
|
|
77
|
+
**It was never wired to a check and later came loose. It was never wired at all.**
|
|
78
|
+
|
|
79
|
+
一份 2026-05-14 的規格寫著 `AC-7:舊版系統報告完整保留(無回歸)`。它的 Test Plan
|
|
80
|
+
有七項,**沒有一項指向 AC-7**。該報告的 timer 是 disabled、部署函式從未被呼叫——
|
|
81
|
+
**從那條 AC 被寫下的同一天起**。三個月後在驗證另一件無關的安裝時偶然發現。
|
|
82
|
+
|
|
83
|
+
**它不是後來斷線的。它從來沒有被接上過。**
|
|
84
|
+
|
|
85
|
+
### Why "someone will check it" is not a verification item
|
|
86
|
+
|
|
87
|
+
A verification item must be **executable or recorded**, not implied. "The reviewer will
|
|
88
|
+
notice" is not one, because a reviewer reads the spec — and the spec says the AC holds.
|
|
89
|
+
An AC is a claim **about the world**, and only something that touches the world can
|
|
90
|
+
falsify it.
|
|
91
|
+
|
|
92
|
+
「有人會檢查」不是驗證項。它必須**可執行或有紀錄**,不能是隱含的。「審查者會注意到」
|
|
93
|
+
不算——因為審查者讀的是規格,而規格說那條 AC 成立。**AC 是一個關於世界的宣稱,
|
|
94
|
+
只有碰得到世界的東西才能證偽它。**
|
|
95
|
+
|
|
96
|
+
### Related | 關聯
|
|
97
|
+
|
|
98
|
+
- [verification-evidence](verification-evidence.md) — VE-011 requires evidence to come from a
|
|
99
|
+
fresh run **after the last edit**; this standard is the upstream question of whether any
|
|
100
|
+
run was ever pointed at the claim in the first place.
|
|
101
|
+
- [class-level-fix](class-level-fix.md) — the same discipline applied to *scope*: traverse the
|
|
102
|
+
set rather than enumerate it.
|
|
103
|
+
|
|
104
|
+
## What's New in v2.4.0
|
|
105
|
+
|
|
106
|
+
- **An AC with no verification item is not an AC** (XSPEC-380 R5). Every acceptance criterion must have a verification item pointing at it; one that has none is demoted to a design intent rather than carried as an AC. Measured instance: a spec's `AC-7` had no matching Test Plan item, and the thing it protected stopped running **on the day the AC was written** — found three months later by accident. An AC is a claim about the world, and only something that touches the world can falsify it.
|
|
107
|
+
|
|
53
108
|
## What's New in v2.3.0
|
|
54
109
|
|
|
55
110
|
- **EARS notation** as an optional AC format (XSPEC-263): 5 EARS templates + `.ac.yaml` `ears` field. GWT remains default & preferred; `given/when/then` relaxed from `required` (backward compatible).
|
|
@@ -2,8 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
> **Language**: English | [繁體中文](../locales/zh-TW/core/test-governance.md)
|
|
4
4
|
|
|
5
|
-
**Version**: 1.
|
|
6
|
-
**Last Updated**: 2026-
|
|
5
|
+
**Version**: 1.2.0
|
|
6
|
+
**Last Updated**: 2026-08-14
|
|
7
7
|
**Applicability**: All software projects
|
|
8
8
|
**Scope**: universal
|
|
9
9
|
|
|
@@ -62,6 +62,23 @@ Run static analysis tools before test execution:
|
|
|
62
62
|
- `eslint .` (TypeScript/JavaScript)
|
|
63
63
|
- `mypy .` (Python type checking)
|
|
64
64
|
|
|
65
|
+
### Threshold Gates Must Fail Closed
|
|
66
|
+
|
|
67
|
+
A measurement layer that prints a percentage but exits `0` regardless of whether that number cleared its threshold is a report, not a gate. It stays green while the number it prints drifts downward across commits, and nothing stops the next merge.
|
|
68
|
+
|
|
69
|
+
Every check with a pass/fail threshold — coverage, lint, mutation score, or any other bounded metric — must translate "below threshold" into a non-zero exit code, using the tool's own enforcement flag rather than a wrapper script that re-parses printed output after the fact:
|
|
70
|
+
|
|
71
|
+
| Tool | Fail-closed flag |
|
|
72
|
+
|------|-------------------|
|
|
73
|
+
| pytest-cov | `--cov-fail-under=<N>` |
|
|
74
|
+
| coverage.py | `coverage report --fail-under=<N>` |
|
|
75
|
+
| diff-cover | `diff-cover coverage.xml --fail-under=<N>` |
|
|
76
|
+
| nyc / Istanbul | `--check-coverage --lines <N>` |
|
|
77
|
+
| Stryker Mutator | `thresholds.break` in `stryker.config.json` |
|
|
78
|
+
| ESLint | `--max-warnings 0` |
|
|
79
|
+
|
|
80
|
+
A wrapper script that computes the number, prints it, and always `exit 0` satisfies neither this rule nor `verification-evidence`'s Evidence Validity rule 1 — the tool's exit code no longer carries any information about the artefact it measured.
|
|
81
|
+
|
|
65
82
|
---
|
|
66
83
|
|
|
67
84
|
## Completion Criteria
|
|
@@ -124,6 +141,7 @@ Release completion criteria (checked before release):
|
|
|
124
141
|
| pyramid-compliance | Planning test strategy | Follow the 70/20/7/3 pyramid ratio as a guideline. Deviation is acceptable with documented justification | Required |
|
|
125
142
|
| sit-isolation | Running system tests | System tests should stub external dependencies but use real internal services. Use SIT environment for system-level validation | Recommended |
|
|
126
143
|
| test-execution-continuity | Adding or completing a test case | A test case must be wired to an automated execution trigger (CI gate, build hook, or scheduled run). A test that exists but is never run provides false confidence and is worse than no test at all. Verify that execution history exists before marking test coverage as complete. | Required |
|
|
144
|
+
| fail-closed-threshold-gate | Configuring or reviewing any coverage/lint/mutation/other threshold check | The check must use the tool's fail-under (or equivalent) enforcement flag so it exits non-zero when the threshold is not met. A script that only prints the number and always exits 0 is a report, not a gate, and does not satisfy this rule | Required |
|
|
127
145
|
|
|
128
146
|
---
|
|
129
147
|
|
|
@@ -139,6 +157,7 @@ Release completion criteria (checked before release):
|
|
|
139
157
|
- [Testing Standards](testing-standards.md) — Testing pyramid and coverage ratios
|
|
140
158
|
- [Test Completeness Dimensions](test-completeness-dimensions.md) — 8-dimension test coverage
|
|
141
159
|
- [Check-in Standards](checkin-standards.md) — Pre-commit quality gates
|
|
160
|
+
- [Verification Evidence](verification-evidence.md) — Evidence Validity governs how to *read* an exit code once produced; `fail-closed-threshold-gate` governs how the gate must be *built* so that exit code carries real information in the first place
|
|
142
161
|
|
|
143
162
|
---
|
|
144
163
|
|
|
@@ -147,3 +166,4 @@ Release completion criteria (checked before release):
|
|
|
147
166
|
| Version | Date | Changes |
|
|
148
167
|
|---------|------|---------|
|
|
149
168
|
| 1.0.0 | 2026-03-11 | Initial version — test policy, completion criteria, environment management |
|
|
169
|
+
| 1.2.0 | 2026-08-14 | Added `fail-closed-threshold-gate` — coverage/lint/mutation threshold checks must exit non-zero below threshold, not just print the number |
|
|
@@ -2,8 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
> **Language**: English | [繁體中文](../locales/zh-TW/core/translation-lifecycle-standards.md)
|
|
4
4
|
|
|
5
|
-
**Version**: 1.0.
|
|
6
|
-
**Last Updated**: 2026-
|
|
5
|
+
**Version**: 1.0.1
|
|
6
|
+
**Last Updated**: 2026-08-12
|
|
7
7
|
**Status**: Trial (expires 2026-10-20)
|
|
8
8
|
**Applicability**: All projects with multi-language documentation
|
|
9
9
|
**Scope**: universal
|
|
@@ -100,7 +100,7 @@ When updating a translation after a source change:
|
|
|
100
100
|
|
|
101
101
|
When `core/*.md` files are staged, the pre-commit hook runs `check-translation-sync.sh` and shows OUTDATED warnings. The hook **never blocks** the commit (blocking at commit time is too disruptive) — it is a reminder only.
|
|
102
102
|
|
|
103
|
-
Setup:
|
|
103
|
+
Setup: `node scripts/install-hooks.mjs` (one-time, after clone; `./scripts/install-hooks.sh` is a POSIX wrapper around the same logic)
|
|
104
104
|
|
|
105
105
|
### Release Gate (`check-translation-sync.sh`)
|
|
106
106
|
|
|
@@ -112,9 +112,9 @@ bash scripts/check-translation-sync.sh
|
|
|
112
112
|
# exit 0 if only MINOR/PATCH gaps (with advisory output)
|
|
113
113
|
```
|
|
114
114
|
|
|
115
|
-
### Version Bump Integration (`bump-version.
|
|
115
|
+
### Version Bump Integration (`bump-version.mjs`)
|
|
116
116
|
|
|
117
|
-
`bump-version.
|
|
117
|
+
`bump-version.mjs` automatically runs `check-translation-sync.sh` after version files are updated, showing the translation health snapshot at the moment of bump — giving the author immediate feedback on what needs updating before publish. `scripts/bump-version.sh` is a thin POSIX wrapper that execs `bump-version.mjs` — the bump logic itself lives only in the `.mjs` file.
|
|
118
118
|
|
|
119
119
|
---
|
|
120
120
|
|
|
@@ -159,4 +159,4 @@ bash scripts/check-translation-sync.sh
|
|
|
159
159
|
- [standard-admission-criteria.md](standard-admission-criteria.md) — How new standards are admitted (triggers TRANS-001)
|
|
160
160
|
- `scripts/check-translation-sync.sh` — Implementation of this standard's automation rules
|
|
161
161
|
- `.githooks/pre-commit` — Pre-commit integration
|
|
162
|
-
- `scripts/install-hooks.
|
|
162
|
+
- `scripts/install-hooks.mjs` — Hook installation (`scripts/install-hooks.sh` is a POSIX wrapper)
|
|
@@ -2,8 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
> **Language**: English | [繁體中文](../locales/zh-TW/core/verification-evidence.md)
|
|
4
4
|
|
|
5
|
-
**Version**: 1.
|
|
6
|
-
**Last Updated**: 2026-
|
|
5
|
+
**Version**: 1.3.0
|
|
6
|
+
**Last Updated**: 2026-08-14
|
|
7
7
|
**Applicability**: All AI-assisted development workflows
|
|
8
8
|
**Scope**: universal
|
|
9
9
|
**Inspired by**: [Superpowers](https://github.com/obra/superpowers) — verification-before-completion (MIT)
|
|
@@ -39,6 +39,7 @@ The first is a form of hallucination. **The second is not** — nothing was inve
|
|
|
39
39
|
| Environment Layer | Which environment the evidence was collected from (`local` / `uat` / `prd`) |
|
|
40
40
|
| Evidence Validity | Whether the evidence *itself* is trustworthy — i.e. whether the verification command actually ran and actually measured what it claims |
|
|
41
41
|
| Silent Tool Failure | A verification command that fails to run, or runs without measuring anything, yet produces output indistinguishable from a genuine result |
|
|
42
|
+
| Evidence Freshness | Whether evidence was produced by a run executed *after* the most recent edit to the code, prompt, or configuration under test — not carried over from a run against an earlier state |
|
|
42
43
|
|
|
43
44
|
---
|
|
44
45
|
|
|
@@ -178,7 +179,7 @@ The Iron Law stops an agent from claiming success **without** evidence. It does
|
|
|
178
179
|
|
|
179
180
|
This is not hallucination. Hallucination is inventing what you did not check. This is the reverse: **you checked, and the checking tool lied to you.** The `anti-hallucination` standard does not cover it — every prohibition there is a form of "don't make things up", and here nothing was made up.
|
|
180
181
|
|
|
181
|
-
### The
|
|
182
|
+
### The five validity rules
|
|
182
183
|
|
|
183
184
|
**1. `exit_code = 0` means success only for tools that return 0 on success.**
|
|
184
185
|
|
|
@@ -196,6 +197,20 @@ Suppressing stderr (`2>/dev/null` and equivalents) silences precisely the channe
|
|
|
196
197
|
|
|
197
198
|
With `set -o pipefail`, `producer | grep -q pattern` inherits a non-zero from `producer` regardless of whether `grep` matched. When the decision depends on content, **capture the output first, then evaluate it** — do not let a pipeline collapse two questions into one number.
|
|
198
199
|
|
|
200
|
+
**5. Evidence must postdate the last edit to what it verifies.**
|
|
201
|
+
|
|
202
|
+
A verification run captured *before* the most recent change to the code, prompt, or configuration under test proves something — just not the thing being claimed. The trap is silent: the full suite ran green, something was edited afterward (a prompt, a config value, a threshold), and only a narrower check (lint, a partial suite) ran after that edit — yet the earlier green run gets cited as evidence for the current state. **Evidence is a single fresh execution captured after the last edit, not the most recent execution that happened to pass.**
|
|
203
|
+
|
|
204
|
+
### Stale Evidence — a failure shape
|
|
205
|
+
|
|
206
|
+
Rules 1–4 assume the evidence was collected in the right place and read correctly. They say nothing about *when*. A command that ran, measured correctly, and returned a true result can still fail to support a completion claim, if something changed after it ran.
|
|
207
|
+
|
|
208
|
+
**Shape of the failure**: the full test suite runs and passes. Afterward, a prompt or a configuration value is edited — no test exercises that edit directly. Before committing, only `lint` is run, and it also passes. The commit is made citing "tests pass" as evidence. CI then fails, red on a test that pins the exact content that was just changed — the suite that "passed" ran against a version of the artefact that no longer exists by the time the claim was made.
|
|
209
|
+
|
|
210
|
+
規則 1–4 假設證據**在正確的地方**被蒐集且被正確解讀。它們沒有回答「**什麼時候**」這個問題。一個執行了、量測正確、也回傳真結果的指令,依然可能撐不起一個完成聲明——如果它跑完之後有東西又變了。
|
|
211
|
+
|
|
212
|
+
**失效樣態**:完整測試套件跑過且通過。之後,某個提示詞或設定值被編輯——沒有任何測試直接涵蓋那次編輯。提交前只跑了 `lint`,也通過。commit 引用「測試通過」作為證據。接著 CI 紅在一個釘住剛剛被改掉那份內容的測試上——那個「通過」的套件,跑的是一個在主張被提出時已經不存在的版本。
|
|
213
|
+
|
|
199
214
|
### Evidence of the failure mode
|
|
200
215
|
|
|
201
216
|
The following are real verification commands run by an AI agent on 2026-07-17, each of which produced a confident, wrong conclusion that was then reported as fact:
|
|
@@ -225,11 +240,32 @@ The following are real verification commands run by an AI agent on 2026-07-17, e
|
|
|
225
240
|
| `verification_evidence` exists but `exit_code ≠ 0` | Mark as **verification failed** — **unless** the tool is known to exit non-zero in the state under test (Evidence Validity rule 1), in which case judge by output |
|
|
226
241
|
| `exit_code = 0` but the command could not have measured the claim | Mark as **unverified** — a passing command that ran in the wrong place, or never ran at all, is not evidence |
|
|
227
242
|
| Evidence asserts absence (`0`, empty, "not found") | Mark as **unverified** until the query tool is shown to have executed successfully |
|
|
243
|
+
| Evidence's timestamp precedes the most recent edit to what it verifies | Mark as **stale** — re-run after the edit; do not cite the earlier pass |
|
|
228
244
|
| Multiple verification steps | **All** steps must pass |
|
|
229
245
|
| Agent provides evidence for wrong command | Mark as **unverified** |
|
|
230
246
|
|
|
231
247
|
---
|
|
232
248
|
|
|
249
|
+
## Narrow Coverage Must Be Registered, Not Just Disclosed
|
|
250
|
+
|
|
251
|
+
Evidence Validity and the Environment Stratification Matrix both allow a documented gap — an environment that cannot exercise a dimension, a check whose measurement is narrower than the claim it is cited for. A written disclosure of that gap is not, by itself, enough: prose is cheaper than closing the gap, and a disclosure that is never revisited quietly becomes a permanent excuse for evidence that no longer supports what it is used to support.
|
|
252
|
+
|
|
253
|
+
**Requirement**: any such documented gap must also be registered in a dated exception inventory — external to the standard or the matrix entry itself — naming the gap, its reason, and a review or expiry date. A footnote with no corresponding inventory entry does not satisfy this rule.
|
|
254
|
+
|
|
255
|
+
**Falsifiable condition**: an inventory entry left unchanged across two consecutive review cycles means the disclosure has become an escape hatch, and the rule is violated for that entry — not "partially satisfied", violated. The inventory's location, format, and cadence are left to the adopting project; this standard requires only that one exists and that entries move.
|
|
256
|
+
|
|
257
|
+
This is the same requirement as [class-level-fix](class-level-fix.md)'s narrow-coverage rule, applied to verification evidence rather than to class-level checks: in both cases, a disclosure that never has to change is not a disclosure — it is a permanent exemption wearing a disclosure's clothes.
|
|
258
|
+
|
|
259
|
+
Evidence Validity 與 Environment Stratification Matrix 都允許一個被記錄下來的缺口——一個無法驗證某個維度的環境層次、一項量測範圍窄於它所支撐主張的檢查。單純寫下這個缺口本身並不夠:散文比真正補上缺口便宜,而一個從未被重新檢視的揭露,會悄悄變成「這份證據其實撐不起它被用來支撐的主張」的永久藉口。
|
|
260
|
+
|
|
261
|
+
**要求**:任何這一類已記錄的缺口,必須同時登記到一份**帶到期日的例外清冊**——獨立於標準本文或 matrix 條目本身之外——指名缺口、其成因、並附上覆核或到期日期。只有一則註腳而沒有對應清冊條目,不滿足本條。
|
|
262
|
+
|
|
263
|
+
**可證偽條件**:一則清冊條目連續兩期審查都未變動,代表該揭露已經變成逃生口,本條對該條目失效——不是「部分滿足」,是失效。清冊放在哪裡、什麼格式、多久審一次,留給採用專案自行決定;本標準只要求「有這麼一份東西存在,且條目會動」。
|
|
264
|
+
|
|
265
|
+
這與 [class-level-fix](class-level-fix.md) 的窄涵蓋規則是同一條要求,套用在驗證證據而非類別層檢查上:兩種情況下,一個永遠不必改變的揭露都不是揭露,是一份穿著揭露外衣的永久豁免。
|
|
266
|
+
|
|
267
|
+
---
|
|
268
|
+
|
|
233
269
|
## Rules
|
|
234
270
|
|
|
235
271
|
| ID | Trigger | Action | Priority |
|
|
@@ -244,6 +280,8 @@ The following are real verification commands run by an AI agent on 2026-07-17, e
|
|
|
244
280
|
| VE-008 | Evidence claims absence (`0`, empty output, "not found") | Invalid until the query tool is shown to have run successfully. Re-run without stderr suppression | High |
|
|
245
281
|
| VE-009 | An existence/absence check suppresses stderr (`2>/dev/null` or equivalent) | Evidence does not stand. Re-run with stderr visible | High |
|
|
246
282
|
| VE-010 | Evidence's `exit_code` comes from a pipeline (esp. under `pipefail`) | The code attributes to no single stage. Capture the output and evaluate content instead | Medium |
|
|
283
|
+
| VE-011 | Evidence's execution timestamp precedes the most recent edit to the artefact it claims to verify | Mark **stale** — not valid evidence for the current state; require a fresh single run captured after the edit | High |
|
|
284
|
+
| VE-012 | A documented evidence/coverage gap (environment-stratification ⚠️/❌, or any check narrower than the claim it supports) has no entry in a dated exception inventory | Disclosure alone does not satisfy this rule — require a registered, dated inventory entry naming the gap and its review/expiry | Required |
|
|
247
285
|
|
|
248
286
|
---
|
|
249
287
|
|
|
@@ -323,6 +361,7 @@ verification_evidence:
|
|
|
323
361
|
- [Testing Standards](testing-standards.md) — how the tests that produce most evidence are written
|
|
324
362
|
- [Agent Dispatch](agent-dispatch.md) — delegated work returns claims that need this standard applied to them
|
|
325
363
|
- [Deployment Standards](deployment-standards.md) — defines the Environment Stratification Responsibility Matrix referenced by VE-005 / VE-006
|
|
364
|
+
- [Class-Level Fix](class-level-fix.md) — the same narrow-coverage-must-be-registered requirement (VE-012), applied to class-level checks rather than to verification evidence
|
|
326
365
|
|
|
327
366
|
---
|
|
328
367
|
|
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
---
|
|
2
2
|
source: ../../CHANGELOG.md
|
|
3
|
-
source_version: 6.
|
|
4
|
-
translation_version: 6.
|
|
5
|
-
last_synced: 2026-08-
|
|
3
|
+
source_version: 6.6.0
|
|
4
|
+
translation_version: 6.6.0
|
|
5
|
+
last_synced: 2026-08-17
|
|
6
6
|
status: current
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -17,6 +17,27 @@ status: current
|
|
|
17
17
|
|
|
18
18
|
## [Unreleased]
|
|
19
19
|
|
|
20
|
+
## [6.6.0] - 2026-08-17
|
|
21
|
+
|
|
22
|
+
### 新增
|
|
23
|
+
|
|
24
|
+
- **`ai-response-navigation` 1.1.0 → 1.2.0 —— 可选规则 R7–R9,管答案本身**(XSPEC 借鉴 B-10,来源 [`ayghri/i-have-adhd`](https://github.com/ayghri/i-have-adhd),MIT)。规则 1–6 管的是答案**之后**要附什么:导航区块、标记过的推荐、匹配响应类型的模板。**答案本身没有任何规则在管。** 于是一个响应可以把结论埋在一整面证据底下,只要结尾附上正确的导航区块,它仍然满足**本标准的每一条**——而找不到答案的读者,不会因为被告知下一步而得到帮助。
|
|
25
|
+
- **R7 —— 先讲发现,不要先讲过程。** *触发*:回答问题、汇报调查结果、或提出决策的响应。第一行写**查到了什么**或**该做什么**——不是方法、不是把问题复述一遍、不是回答的计划。证据(`file:line`、命令输出、表格、测量数字)是**佐证**,应放在它所支持的论断之后;以证据开场会迫使读者自行重建结论,而那正是他请你做的工作。本条规范的是**顺序**,**不**代表可以省略证据。
|
|
26
|
+
- **R8 —— 每一轮重述进度。** *触发*:跨 3 轮以上的对话,或含 3 个以上步骤的任务。用一行说明工作进行到哪里。不能假设读者能在消息之间记住「我们在 5 步中的第 3 步」,而重述它的成本是一个句子。与模板 4(进行中)互补:**R8 管开头,模板管结尾。**
|
|
27
|
+
- **R9 —— 不要开场白。** *触发*:任何实质性响应。本条把一项既有禁令一般化:[`anti-sycophancy-prompting`](../../core/anti-sycophancy-prompting.md) 已经禁止「以正面肯定开场批评」,但**仅限批评情境**。R9 把同一项禁令扩及每一个实质响应,理由不同——不是为了防拍马屁,而是为了消除它在读者与答案之间制造的延迟。**R9 不适用于结语**;R1 的导航区块要求依然成立。
|
|
28
|
+
- **来源十条只取三条,其余七条的淘汰理由写进标准本文**,不是只写在待办清单里。两条与 R1–R2 重复。三条与本标准或其他标准冲突:它的「不要 recap/不要结语」**与 R1 的导航区块直接矛盾**;它的「列表上限 5 项」会截断证据表格与遍历分母;它的「具体时间估计」已由 [`estimation-standards`](../../core/estimation-standards.md) 涵盖。
|
|
29
|
+
- **可选的语义同 R6**(模型级别标注):采用者不必启用、既有 skill 不需回头补,项目**可以**在自己的配置中把任一条提升为必须。**不可选的是每一条都带有精确的触发条件**——一条松到永远不会启动的规则,与没有这条规则无从分辨,那正是 XSPEC-378 记录的失效模式。
|
|
30
|
+
- **扩充既有标准而非新建一支**:再开一支管「AI 怎么对人类写回答」的标准,会让同一条轴出现两个实现。
|
|
31
|
+
|
|
32
|
+
## [6.5.0] - 2026-08-14
|
|
33
|
+
|
|
34
|
+
> ⚠️ **本节简体译文待补。** 本次发布包含两批内容:XSPEC 借鉴 B-01 的五条标准补强
|
|
35
|
+
> (`verification-evidence` VE-011/VE-012、`test-governance` 门槛闸须 fail-closed、
|
|
36
|
+
> `mutation-testing` kill 归因、`class-level-fix` 负向控制极限声明),
|
|
37
|
+
> 以及 XSPEC-362 的 `model-selection` 2.1.0 与 `agent-dispatch` 复位。
|
|
38
|
+
> 完整内容见 [English](../../CHANGELOG.md) 或 [繁體中文](../zh-TW/CHANGELOG.md)。
|
|
39
|
+
> **此处明示未翻译,而非以繁体内容充当简体译文。**
|
|
40
|
+
|
|
20
41
|
## [6.4.0] - 2026-08-10
|
|
21
42
|
|
|
22
43
|
### 新增
|
|
@@ -15,7 +15,7 @@ status: current
|
|
|
15
15
|
|
|
16
16
|
> **语言**: [English](../../README.md) | [繁體中文](../zh-TW/README.md) | 简体中文
|
|
17
17
|
|
|
18
|
-
**版本**: 6.
|
|
18
|
+
**版本**: 6.6.0 | **发布日期**: 2026-08-17 | **授权**: [双重授权](../../LICENSE) (CC BY 4.0 + MIT)
|
|
19
19
|
|
|
20
20
|
语言无关、框架无关的软件项目文档标准。通过 AI 原生工作流,确保不同技术栈之间的一致性、质量和可维护性。
|
|
21
21
|
|
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
---
|
|
2
2
|
source: ../../../core/ai-response-navigation.md
|
|
3
|
-
source_version: 1.
|
|
4
|
-
translation_version: 1.
|
|
5
|
-
last_synced: 2026-
|
|
3
|
+
source_version: 1.2.0
|
|
4
|
+
translation_version: 1.2.0
|
|
5
|
+
last_synced: 2026-08-17
|
|
6
6
|
status: current
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -10,8 +10,8 @@ status: current
|
|
|
10
10
|
|
|
11
11
|
> **语言**: [English](../../../core/ai-response-navigation.md) | [繁體中文](../../zh-TW/core/ai-response-navigation.md) | 简体中文
|
|
12
12
|
|
|
13
|
-
**版本**: 1.
|
|
14
|
-
**最后更新**: 2026-
|
|
13
|
+
**版本**: 1.2.0
|
|
14
|
+
**最后更新**: 2026-08-17
|
|
15
15
|
**适用范围**: 所有使用 AI 辅助开发的项目
|
|
16
16
|
**范围**: universal
|
|
17
17
|
**行业标准**: 无(新兴 AI 工具实践)
|
|
@@ -27,6 +27,11 @@ status: current
|
|
|
27
27
|
|
|
28
28
|
**解决方案**:在每个实质性 AI 响应结尾附加标准化的「导航区块」,包含情境模板、推荐标记和弹性选项数量。
|
|
29
29
|
|
|
30
|
+
**范围注记(v1.2.0)**:规则 1–6 管的是答案**之后**要附什么。规则 7–9 于 1.2.0 新增、属**可选**,
|
|
31
|
+
管的是答案本身:先讲发现、每轮重述进度、不要开场白。新增的理由是——
|
|
32
|
+
一个响应可以满足规则 1–6 的每一条,同时把结论埋起来;
|
|
33
|
+
**而找不到答案的读者,不会因为结尾有一个正确的导航区块被告知下一步而得到帮助。**
|
|
34
|
+
|
|
30
35
|
---
|
|
31
36
|
|
|
32
37
|
## 核心规则
|
|
@@ -103,6 +108,61 @@ status: current
|
|
|
103
108
|
|
|
104
109
|
---
|
|
105
110
|
|
|
111
|
+
## 导航之前的那个答案(规则 7–9,可选)
|
|
112
|
+
|
|
113
|
+
> **借鉴自**:[`ayghri/i-have-adhd`](https://github.com/ayghri/i-have-adhd)(MIT),十条中取三条。
|
|
114
|
+
> 其余七条删去:两条已被上方规则 1–2 涵盖,五条与本标准冲突
|
|
115
|
+
> (它的「不要 recap/不要结语」与规则 1 的导航区块直接矛盾;它的「列表上限 5 项」
|
|
116
|
+
> 会截断证据表格与遍历分母)或与 [estimation-standards](estimation-standards.md) 重复。
|
|
117
|
+
|
|
118
|
+
**这一节为何存在**:规则 1–6 管的是答案**之后**要附什么,而答案本身没有任何规则在管——
|
|
119
|
+
一个响应可以把结论埋在一整面证据底下,只要结尾附上正确的导航区块,
|
|
120
|
+
它仍然满足本标准的每一条。**找不到答案的读者,不会因为被告知下一步而得到帮助。**
|
|
121
|
+
|
|
122
|
+
**这三条是可选的**,语义同规则 6:采用项目不必启用,既有 skill 也不需回头补。
|
|
123
|
+
项目**可以**在自己的配置中把任一条提升为必须。**不可选的是它们必须有精确的触发条件**——
|
|
124
|
+
一条松到永远不会启动的规则,与没有这条规则无从分辨。
|
|
125
|
+
|
|
126
|
+
### 规则 7:先讲发现,不要先讲过程(可选)
|
|
127
|
+
|
|
128
|
+
**触发条件**:回答问题、汇报调查结果、或提出决策的响应。
|
|
129
|
+
|
|
130
|
+
第一行写**查到了什么**或**该做什么**。不是方法、不是把问题复述一遍、不是回答的计划。
|
|
131
|
+
|
|
132
|
+
证据——`file:line`、命令输出、表格、测量数字——是**佐证**,应放在它所支持的论断**之后**。
|
|
133
|
+
以证据开场会迫使读者自行重建结论,而那正是他请你做的工作。
|
|
134
|
+
|
|
135
|
+
| 不要写 | 改写成 |
|
|
136
|
+
|---|---|
|
|
137
|
+
| 「我检查了 44 天、63 个域名的数据,发现……」 | 「那三组查询删掉。它返回的东西有 46% 是下载页。」 |
|
|
138
|
+
| 「让我看看这是怎么配置的。」 | 「配置在 `x.yaml:12`,值是错的,因为……」 |
|
|
139
|
+
|
|
140
|
+
**这不代表可以省略证据**,它规范的是顺序。
|
|
141
|
+
|
|
142
|
+
### 规则 8:多轮工作中每一轮重述进度(可选)
|
|
143
|
+
|
|
144
|
+
**触发条件**:跨 3 轮以上的对话,或含 3 个以上步骤的任务。
|
|
145
|
+
|
|
146
|
+
每条响应用一行说明工作进行到哪里。不能假设读者能在消息之间记住
|
|
147
|
+
「我们在 5 步中的第 3 步」,而重述它的成本是一个句子。
|
|
148
|
+
|
|
149
|
+
本条与下方模板 4(进行中)互补:**规则 8 管开头,模板 4 管结尾。**
|
|
150
|
+
|
|
151
|
+
### 规则 9:不要开场白(可选)
|
|
152
|
+
|
|
153
|
+
**触发条件**:任何实质性响应。
|
|
154
|
+
|
|
155
|
+
从答案开始。不要以「我接下来要做什么」的预告、对请求的确认、或对问题本身的评价开场。
|
|
156
|
+
|
|
157
|
+
本条把一项既有禁令一般化:[anti-sycophancy-prompting](anti-sycophancy-prompting.md)
|
|
158
|
+
已经禁止「以正面肯定开场批评」,但**仅限批评情境**。规则 9 把同一项禁令扩及每一个实质响应,
|
|
159
|
+
理由不同:不是为了防拍马屁,而是为了消除它在读者与答案之间制造的延迟。
|
|
160
|
+
|
|
161
|
+
**规则 9 不适用于结语。** 规则 1 要求的导航区块依然成立——
|
|
162
|
+
响应的结尾正是本标准安放「读者下一步」的位置。
|
|
163
|
+
|
|
164
|
+
---
|
|
165
|
+
|
|
106
166
|
## 情境模板
|
|
107
167
|
|
|
108
168
|
### 模板 1:任务完成
|
|
@@ -290,6 +350,9 @@ AI 需要用户做出选择或提供信息时使用。
|
|
|
290
350
|
| R4 | 1–5 个选项,依情境调整 |
|
|
291
351
|
| R5 | 适用时使用 `/command` 格式 |
|
|
292
352
|
| R6 | *(可选)* 级别明确时附加 `〔模型:Fast|Standard|Capable〕` |
|
|
353
|
+
| R7 | *(可选)* 先讲发现;证据放在它所支持的论断之后 |
|
|
354
|
+
| R8 | *(可选)* 跨 3 轮或 3 步以上 → 用一行重述进度 |
|
|
355
|
+
| R9 | *(可选)* 不要开场白。结语仍为必须——见 R1 |
|
|
293
356
|
|
|
294
357
|
| 豁免 | 不豁免 |
|
|
295
358
|
|------|--------|
|
|
@@ -313,6 +376,7 @@ AI 需要用户做出选择或提供信息时使用。
|
|
|
313
376
|
|
|
314
377
|
| 版本 | 日期 | 变更 |
|
|
315
378
|
|------|------|------|
|
|
379
|
+
| 1.2.0 | 2026-08-17 | 新增可选规则 R7–R9,管答案本身(先讲发现、重述进度、不要开场白)。借鉴自 `ayghri/i-have-adhd`(MIT),十条取三;其余七条因已被 R1–R2 涵盖、与 R1 冲突、或与 estimation-standards 重复而删去。规则 1–6 全部可以被一个把结论埋起来的响应满足——R7–R9 补上这个缺口 |
|
|
316
380
|
| 1.1.0 | 2026-06-10 | 新增规则 R6 可选模型级别标注(`〔模型:Fast|Standard|Capable〕`);与厂商无关;不强制既有技能回改 |
|
|
317
381
|
| 1.0.0 | 2026-03-25 | 初始版本 |
|
|
318
382
|
|