universal-dev-standards 6.4.0 → 6.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/bundled/ai/standards/agent-dispatch.ai.yaml +162 -0
  2. package/bundled/ai/standards/ai-friendly-architecture.ai.yaml +1 -1
  3. package/bundled/ai/standards/ai-instruction-standards.ai.yaml +190 -15
  4. package/bundled/ai/standards/ai-response-navigation.ai.yaml +43 -3
  5. package/bundled/ai/standards/class-level-fix.ai.yaml +38 -3
  6. package/bundled/ai/standards/commit-message.ai.yaml +2 -0
  7. package/bundled/ai/standards/model-selection.ai.yaml +370 -72
  8. package/bundled/ai/standards/mutation-testing.ai.yaml +105 -2
  9. package/bundled/ai/standards/project-structure.ai.yaml +1 -1
  10. package/bundled/ai/standards/security-standards.ai.yaml +22 -1
  11. package/bundled/ai/standards/spec-driven-development.ai.yaml +86 -2
  12. package/bundled/ai/standards/test-governance.ai.yaml +49 -2
  13. package/bundled/ai/standards/testing.ai.yaml +49 -3
  14. package/bundled/ai/standards/translation-lifecycle-standards.ai.yaml +4 -4
  15. package/bundled/ai/standards/verification-evidence.ai.yaml +48 -4
  16. package/bundled/core/ai-response-navigation.md +75 -2
  17. package/bundled/core/class-level-fix.md +26 -3
  18. package/bundled/core/model-selection.md +383 -125
  19. package/bundled/core/mutation-testing.md +41 -2
  20. package/bundled/core/spec-driven-development.md +57 -2
  21. package/bundled/core/test-governance.md +22 -2
  22. package/bundled/core/translation-lifecycle-standards.md +6 -6
  23. package/bundled/core/verification-evidence.md +42 -3
  24. package/bundled/locales/zh-CN/CHANGELOG.md +24 -3
  25. package/bundled/locales/zh-CN/README.md +1 -1
  26. package/bundled/locales/zh-CN/SECURITY.md +1 -1
  27. package/bundled/locales/zh-CN/core/ai-response-navigation.md +69 -5
  28. package/bundled/locales/zh-CN/core/model-selection.md +375 -60
  29. package/bundled/locales/zh-CN/core/mutation-testing.md +1 -1
  30. package/bundled/locales/zh-CN/core/spec-driven-development.md +1 -1
  31. package/bundled/locales/zh-CN/core/test-governance.md +1 -1
  32. package/bundled/locales/zh-CN/core/translation-lifecycle-standards.md +1 -1
  33. package/bundled/locales/zh-CN/core/verification-evidence.md +1 -1
  34. package/bundled/locales/zh-CN/docs/CHEATSHEET.md +7 -12
  35. package/bundled/locales/zh-CN/docs/FEATURE-REFERENCE.md +10 -15
  36. package/bundled/locales/zh-TW/CHANGELOG.md +49 -3
  37. package/bundled/locales/zh-TW/README.md +1 -1
  38. package/bundled/locales/zh-TW/SECURITY.md +1 -1
  39. package/bundled/locales/zh-TW/core/ai-response-navigation.md +69 -5
  40. package/bundled/locales/zh-TW/core/class-level-fix.md +22 -7
  41. package/bundled/locales/zh-TW/core/model-selection.md +385 -47
  42. package/bundled/locales/zh-TW/core/mutation-testing.md +45 -6
  43. package/bundled/locales/zh-TW/core/spec-driven-development.md +1 -1
  44. package/bundled/locales/zh-TW/core/test-governance.md +22 -3
  45. package/bundled/locales/zh-TW/core/translation-lifecycle-standards.md +1 -1
  46. package/bundled/locales/zh-TW/core/verification-evidence.md +33 -6
  47. package/bundled/locales/zh-TW/docs/CHEATSHEET.md +7 -12
  48. package/bundled/locales/zh-TW/docs/FEATURE-REFERENCE.md +10 -15
  49. package/bundled/locales/zh-TW/integrations/claude-code/README.md +31 -5
  50. package/package.json +1 -1
  51. package/src/utils/reference-sync.js +83 -16
  52. package/standards-registry.json +20 -8
@@ -2,8 +2,8 @@
2
2
 
3
3
  > **Language**: English | [繁體中文](../locales/zh-TW/core/model-selection.md)
4
4
 
5
- **Version**: 1.0.1
6
- **Last Updated**: 2026-06-10
5
+ **Version**: 2.1.0
6
+ **Last Updated**: 2026-08-11
7
7
  **Applicability**: AI-assisted development with multiple model tiers
8
8
  **Scope**: universal
9
9
  **Inspired by**: [Superpowers](https://github.com/obra/superpowers) — subagent-driven-development (MIT)
@@ -12,9 +12,9 @@
12
12
 
13
13
  ## Purpose
14
14
 
15
- Define a cost-effective strategy for selecting AI model tiers based on task complexity signals. Use the cheapest model that can handle the job, and escalate only when necessary.
15
+ Define how to choose **which model** and **how deeply it should think** — two independent decisions that must not be collapsed into one. Use the cheapest combination that can handle the job, and escalate along the correct axis when it cannot.
16
16
 
17
- 定義基於任務複雜度的 AI 模型分級選擇策略。使用能勝任的最便宜模型,僅在必要時升級。
17
+ 定義兩個獨立決策:**選哪個模型**與**要它想多深**。兩者不可壓成同一軸。使用能勝任的最便宜組合,失敗時沿**正確的軸**升級。
18
18
 
19
19
  ---
20
20
 
@@ -22,98 +22,251 @@ Define a cost-effective strategy for selecting AI model tiers based on task comp
22
22
 
23
23
  | Term | Definition |
24
24
  |------|-----------|
25
- | Model Tier | A classification level representing model capability and cost |
26
- | Complexity Signal | Observable characteristics of a task that indicate required capability |
27
- | Escalation | Upgrading to a higher-tier model after a lower tier fails |
25
+ | Model Tier | A classification level representing a model's reasoning ceiling and cost |
26
+ | Effort Level | How much reasoning depth is requested **for one dispatch**; a request parameter, not a model property |
27
+ | Reasoning Ceiling | The limit beyond which more thinking time yields no better answer |
28
+ | Specification Definiteness | Whether the steps are already defined, or only the goal and constraints are known |
29
+ | Hard Boundary | A capability the model does not have at all (not "worse at") — e.g. context capacity, effort-parameter support |
30
+ | Refusal Marker | A field in an otherwise-normal response indicating the request was declined |
31
+ | Escalation | Moving up the model axis after the effort axis is exhausted |
28
32
 
29
33
  ---
30
34
 
31
- ## Core Principle — Cost Efficiency
35
+ ## Version Field Semantics
32
36
 
33
- > **Always start with the cheapest model tier that matches the task's complexity signals.**
37
+ > **This document carries exactly one version.** The `Version` field in the header covers the
38
+ > whole file, including every section. **Sections MUST NOT carry their own version markers.**
39
+ > Section-level change history belongs in the [Version History](#version-history) table.
40
+ >
41
+ > The `standard.meta.version` field in `ai/standards/model-selection.ai.yaml` corresponds to
42
+ > **this whole-file version**, never to a section version.
34
43
 
35
- 始終使用與任務複雜度匹配的最便宜模型。
44
+ 本檔只有一個版本欄,位於檔頭,涵蓋全檔。**章節不得自帶版本標記**;章節的變更歷程一律記於 Version History 表。`.ai.yaml` 的 `standard.meta.version` 對應**全檔版本**。
36
45
 
37
46
  ---
38
47
 
39
- ## Three-Tier Model Classification
48
+ ## Core Principle — Two Orthogonal Axes
40
49
 
41
- ### Tier 1: Fast(快速層)
50
+ > **The model axis buys a ceiling and a temperament. The effort axis buys depth on this one dispatch. Choosing on one axis tells you nothing about the other.**
42
51
 
43
- **Purpose**: Mechanical implementation — single file, clear spec, no judgment needed.
52
+ 模型軸買的是「天花板與性格」,effort 軸買的是「這一次要它想多深」。在一軸上的選擇,不決定另一軸。
44
53
 
45
- 機械性實作 — 單一檔案、明確規格、無需判斷。
54
+ ```
55
+ effort → low medium high very-high max
56
+ model tier ↓
57
+ fast · · · ? ?
58
+ standard · · · ? ?
59
+ capable · · · ? ?
60
+ ```
61
+
62
+ Every cell is in principle a legal dispatch: the two axes are chosen independently.
63
+
64
+ `?` marks a cell whose availability **this standard cannot state**. Whether a given model accepts a given effort level is a property of that model, and tier names here are vendor-neutral labels — so the answer lives in each platform's tier→model mapping, not in this table. An unsupported level is a **hard boundary**, not a poor choice: see [R3a](#r3a-hard-boundaries). Adopters MUST record, per model, which levels it accepts.
65
+
66
+ `?` 標示的格子,**本標準無法斷言其可用性**——某個模型接受哪些 effort 級距是該模型的屬性,而此處的層級名稱是與廠商無關的標籤。答案在各平台自己的「層級 → 模型」映射裡。
67
+
68
+ ---
69
+
70
+ ## Axis 1 — Model Tier
71
+
72
+ ### Selection Criteria
73
+
74
+ Two criteria decide the tier. **Neither is "how many files does this change".**
75
+
76
+ 兩個判準決定分層。**兩者都不是「改幾個檔案」**。
77
+
78
+ #### Criterion 1 — Reasoning Ceiling Requirement(推理天花板需求)
79
+
80
+ Does the task contain a component that **more thinking time cannot solve**?
81
+
82
+ - **No** → a lower tier suffices; buy depth with effort instead of buying a ceiling with money.
83
+ - **Yes** → a higher ceiling is required; no amount of effort on a lower tier reaches it.
84
+
85
+ 這個任務是否存在「想更久也想不出來」的成分?若否,較低層加 effort 即可;若是,才需要更高天花板。
86
+
87
+ #### Criterion 2 — Specification Definiteness(規格明確度)
88
+
89
+ **This criterion is bidirectional. It is not monotonically increasing.**
90
+
91
+ | Specification state | Prefer | Failure mode if inverted |
92
+ |---|---|---|
93
+ | Steps already defined, path known | Literal-following tier (lower) | Feeding a written-out step list to a high-ceiling model **lowers** output quality — it reinterprets what was already decided |
94
+ | Only goal and constraints known, path unknown | Ambiguity-navigating tier (higher) | Handing an ambiguous task to a literal-following model yields **"precisely executed the wrong sentence"** |
95
+
96
+ 規格明確度是**雙向**判準:把模糊任務給字面遵循型,會得到「精確執行了錯誤的那句話」;把已寫死的步驟餵給高天花板型,反而降低品質。
97
+
98
+ > **Why file count was removed.** A 3-file redesign of module boundaries is harder than an 8-file
99
+ > mechanical rename. File count mis-sorts "deep and narrow" work **in a fixed direction** — it is a
100
+ > bias, not noise. It survived as a signal because it is cheap to measure, not because it predicts
101
+ > anything.
102
+
103
+ ### The Three Tiers
104
+
105
+ Tier ids (`fast` / `standard` / `capable`) are unchanged. Only the criteria changed.
106
+
107
+ #### Tier 1: Fast(快速層)
108
+
109
+ **Purpose**: Work with no reasoning ceiling requirement and a fully definite specification.
110
+
111
+ 無推理天花板需求、規格完全明確的工作。
46
112
 
47
- **Complexity Signals**:
48
- - Modifies a single file
49
- - Specification is completely unambiguous
113
+ **Signals**:
114
+ - No component that "more thinking cannot solve"
115
+ - Steps are written out; no path-finding needed
50
116
  - No design judgment required
51
- - Repetitive, pattern-based work
117
+ - Literal following is exactly what is wanted
52
118
 
53
119
  **Examples**:
54
- - Update `package.json` version number
55
- - Add a new export statement
120
+ - Apply a defined edit across a set of files (mechanical rename, version bump)
56
121
  - Fix a typo
57
- - Rename a variable across a file
122
+ - Add an export statement named in the spec
58
123
 
59
- ### Tier 2: Standard(標準層)
124
+ #### Tier 2: Standard(標準層)
60
125
 
61
- **Purpose**: Integration work — multiple files, requires judgment.
126
+ **Purpose**: A modest reasoning ceiling, with a specification that is mostly defined but leaves local decisions open.
62
127
 
63
- 整合性實作 — 多檔案、需要判斷力。
128
+ 需要一定推理天花板,規格大致明確但留有局部決策空間。
64
129
 
65
- **Complexity Signals**:
66
- - Modifies 2–5 files
130
+ **Signals**:
131
+ - Some components need reasoning, but none are "unreachable by thinking longer"
132
+ - Goal and most steps defined; a bounded set of local choices remains
67
133
  - Requires understanding inter-module relationships
68
- - Needs some design judgment
69
- - Cross-cutting but within a bounded context
70
134
 
71
135
  **Examples**:
72
- - Add an API endpoint (route + handler + test)
73
- - Refactor a module's internal structure
74
- - Implement a feature with database migration
75
- - Write integration tests for a subsystem
136
+ - Implement a defined feature whose integration points are known
137
+ - Refactor a module against a stated target structure
138
+ - Write integration tests for a specified subsystem
76
139
 
77
- ### Tier 3: Capable(能力層)
140
+ #### Tier 3: Capable(能力層)
78
141
 
79
- **Purpose**: Architectural work — design, review, complex debugging.
142
+ **Purpose**: A high reasoning ceiling, or a specification that gives only goal and constraints.
80
143
 
81
- 架構性工作 — 設計、審查、複雜除錯。
144
+ 需要高推理天花板,或規格僅有目標與約束、路徑未知。
82
145
 
83
- **Complexity Signals**:
84
- - Modifies 5+ files
85
- - Requires architectural decisions
86
- - Cross-module coordination
87
- - Complex debugging or performance analysis
88
- - Ambiguous requirements needing interpretation
146
+ **Signals**:
147
+ - Contains a component that more thinking time alone cannot resolve
148
+ - Path is unknown; the answer must be found, not applied
149
+ - Requires navigating ambiguity rather than following text
89
150
 
90
151
  **Examples**:
91
- - Design a new subsystem architecture
92
- - Review a large pull request
93
- - Diagnose cross-service performance issues
94
- - Refactor a major component with many dependents
152
+ - Design a subsystem architecture (path unknown)
153
+ - Review a large pull request (what matters is not stated in advance)
154
+ - Diagnose a failure whose cause is not yet hypothesised
95
155
 
96
- ---
156
+ > **Note on "large refactor"**: a major refactor with a *defined* target structure is **not**
157
+ > automatically Tier 3 — under Criterion 2 it is a definite specification, and a high-ceiling model
158
+ > may produce worse results than a literal-following one. What raises it to Tier 3 is an *undecided*
159
+ > target structure.
97
160
 
98
- ## Selection Decision Flow
161
+ ### Selection Decision Flow
99
162
 
100
163
  ```
101
- Analyze task complexity signals
102
- ├── Single file, clear spec, no judgment? → Tier 1 (Fast)
103
- ├── 2-5 files, some judgment needed? → Tier 2 (Standard)
104
- └── 5+ files, architectural decisions? → Tier 3 (Capable)
164
+ Does the task contain something more thinking cannot solve?
165
+ ├── Yes → high ceiling needed → Tier 3 (Capable)
166
+ └── No → how definite is the specification?
167
+ ├── Fully defined steps, path known → Tier 1 (Fast)
168
+ ├── Goal + most steps, local choices → Tier 2 (Standard)
169
+ └── Goal + constraints only → Tier 3 (Capable)
170
+ Then, independently: choose an effort level (Axis 2).
105
171
  ```
106
172
 
107
173
  ---
108
174
 
175
+ ## Axis 2 — Effort (Reasoning Depth)
176
+
177
+ Effort is **a parameter of one dispatch**, not a property of a model. The same model at `low` and at `max` is not the same worker.
178
+
179
+ effort 是**單次派工的參數**,不是模型屬性。
180
+
181
+ ### Vendor-Neutral Effort Levels
182
+
183
+ Each platform maps these labels to its own parameter. The labels are the contract; the mapping is local.
184
+
185
+ | Level | Semantics | Typical use |
186
+ |---|---|---|
187
+ | `low`(低) | Minimal deliberation; answer largely on pattern | Mechanical edits, lookups, formatting |
188
+ | `medium`(中) | Ordinary deliberation; the default | Most defined implementation work |
189
+ | `high`(高) | Extended deliberation; considers alternatives before answering | Integration work with local decisions |
190
+ | `very-high`(極高) | Explores and discards candidate approaches | Design, review, non-obvious debugging |
191
+ | `max`(極限) | Maximum available depth for this model | Last attempt before escalating the model axis |
192
+
193
+ **Not every tier supports every level.** Support for the effort parameter is a **hard boundary**, not a quality gradient — see [R3a](#r3a-hard-boundaries). A platform's tier→model mapping MUST record which levels each model accepts.
194
+
195
+ ### Failure Diagnosis — Depth or Ceiling?
196
+
197
+ > **When a dispatch fails, first decide which axis was insufficient. The two remedies are not interchangeable.**
198
+
199
+ 任務失敗時,**先判斷是「深度不足」還是「天花板不足」**——兩者的補救方式不可互換。
200
+
201
+ | Observation | Diagnosis | Remedy | Wrong remedy costs |
202
+ |---|---|---|---|
203
+ | Output is shallow, skipped considerations, stopped early, but the reasoning it *did* do was sound | **Depth insufficient** | Raise effort on the **same** model | Escalating the tier pays more for a ceiling that was never the constraint |
204
+ | Output is wrong in kind — misunderstood the problem, produced an approach that cannot work — **and this persists at `max` effort** | **Ceiling insufficient** | Escalate the **model** tier | Raising effort again buys nothing; the level is already exhausted |
205
+
206
+ **Ordering rule**: raise effort before escalating the tier. Escalate the tier only after the current tier's **highest supported** effort level has failed. A tier escalation performed without exhausting effort is an untested assumption about which axis was short.
207
+
208
+ **Exception**: if [Criterion 1](#criterion-1--reasoning-ceiling-requirement推理天花板需求) already identified a component that more thinking cannot solve, start at the higher tier. The ordering rule is about *diagnosing failures*, not about ignoring the criteria up front.
209
+
210
+ ---
211
+
212
+ ## Reverse Exclusion Rules
213
+
214
+ The tier criteria say what work should be **raised** to a tier. This section says what work must **not** be sent to one. These are separate questions, and the second one has no answer derivable from the first.
215
+
216
+ ### R3a Hard Boundaries
217
+
218
+ > **A hard boundary is the absence of a capability, not a lower degree of it.** A model that cannot hold the input does not "do the task worse" — it does not do the task.
219
+
220
+ 硬邊界是**能力有無**,不是程度差。
221
+
222
+ Adopters MUST record hard boundaries in their tier→model mapping, **alongside and separately from** capability scores:
223
+
224
+ | Hard boundary | Question it answers | Consequence when violated |
225
+ |---|---|---|
226
+ | **Context capacity** | Does the input fit? | Truncation or error — the model never saw part of the task |
227
+ | **Effort parameter support** | Does this model accept an effort level at all? | Request rejected, or silently executed at the model's fixed depth |
228
+ | **Modality support** | Can it accept this input type at all (image, audio)? | Input dropped or rejected |
229
+
230
+ **Rule**: a model failing any required hard boundary is **excluded before cost comparison**, not ranked lower within it. Excluding it afterwards allows a cheaper, incapable model to win on price.
231
+
232
+ **Recording format**: hard boundaries map to `declared` in the [capability registry](#capability_dimensions--capability-dimensions) — a boolean, obtainable at connection time, at near-zero cost. A score of `1` is **not** a hard boundary; it is a measured low quality (see [R7a state table](#routing_rules--four-state-routing)).
233
+
234
+ ### R3b Reverse Risks
235
+
236
+ Higher capability is not uniformly better. Two risks run in the opposite direction from the tier criteria:
237
+
238
+ #### 1. Safety-classifier refusal on specification-sensitive work
239
+
240
+ Higher-capability tiers may carry stricter safety classification. Work that is legitimate but resembles a prohibited category — security hardening, exploit-mitigation implementation, credential-handling code, red-team tooling — may be **declined**.
241
+
242
+ > **The refusal is not an error.** The call returns a normal response with a refusal marker.
243
+ > For a caller that checks only the exit code, or only whether an exception was raised, **this is a
244
+ > silent failure**: the pipeline records success, and the work was never done.
245
+
246
+ 拒絕**不是錯誤**——回傳的是正常回應加上拒絕標記。對只看 exit code/有無例外的呼叫端,這是**靜默失效**。
247
+
248
+ **Requirements**:
249
+
250
+ 1. Callers **MUST inspect the response for a refusal marker**. Absence of an error is not evidence of completion. (See the `verification-evidence` standard: "the checking tool ran and returned nothing useful" is a distinct failure class from "the check was not run".)
251
+ 2. Specification-sensitive work **MUST be feasibility-checked before batch dispatch** to a high-capability tier — a cheap probe request, not the full task.
252
+ 3. On refusal, **re-dispatch to a different tier or provider**; do not retry the same request at higher effort. Effort does not move a classifier decision.
253
+
254
+ #### 2. Over-specified prompts degrade high-ceiling models
255
+
256
+ Writing out every step is the correct practice for a literal-following tier and the **wrong** practice for a high-ceiling one: the latter reinterprets decisions that were already made, and the output gets worse.
257
+
258
+ **Rule**: prompt granularity follows the tier. If a fully written-out step list is the only prompt available, [Criterion 2](#criterion-2--specification-definiteness規格明確度) already says to send it to a lower tier — sending it to a high tier is a double mistake, paying more for worse output.
259
+
260
+ ---
261
+
109
262
  ## Escalation Rules
110
263
 
111
- When a model tier fails (returns `BLOCKED`), escalate to the next tier:
264
+ Escalation now has two axes; the table applies **after** effort has been exhausted at the current tier.
112
265
 
113
- | Current Tier | On BLOCKED | Action |
266
+ | Current Tier | On failure at max supported effort | Action |
114
267
  |-------------|-----------|--------|
115
- | Fast | → Standard | Re-dispatch with Standard tier |
116
- | Standard | → Capable | Re-dispatch with Capable tier |
268
+ | Fast | → Standard | Re-dispatch at Standard tier |
269
+ | Standard | → Capable | Re-dispatch at Capable tier |
117
270
  | Capable | → Human | Flag for human intervention |
118
271
 
119
272
  ### Escalation is Not Retry
@@ -121,9 +274,10 @@ When a model tier fails (returns `BLOCKED`), escalate to the next tier:
121
274
  Escalation means using a more capable model, not repeating the same action. The higher-tier model receives:
122
275
  - The original task
123
276
  - The lower-tier model's output and failure reason
277
+ - The effort level already attempted
124
278
  - Additional context if available
125
279
 
126
- 升級不是重試。更高層級的模型會收到原始任務、低層級的輸出與失敗原因。
280
+ 升級不是重試。更高層級的模型會收到原始任務、低層級的輸出與失敗原因,以及**已嘗試過的 effort 等級**。
127
281
 
128
282
  ---
129
283
 
@@ -131,128 +285,232 @@ Escalation means using a more capable model, not repeating the same action. The
131
285
 
132
286
  | ID | Trigger | Action | Priority |
133
287
  |----|---------|--------|----------|
134
- | MS-001 | BLOCKED at Fast tier | Escalate to Standard | High |
135
- | MS-002 | BLOCKED at Standard tier | Escalate to Capable | High |
136
- | MS-003 | BLOCKED at Capable tier | Flag for human intervention | Critical |
137
- | MS-004 | Task has ambiguous requirements | Start at Standard or higher | Medium |
288
+ | MS-001 | BLOCKED at Fast tier, effort exhausted | Escalate to Standard | High |
289
+ | MS-002 | BLOCKED at Standard tier, effort exhausted | Escalate to Capable | High |
290
+ | MS-003 | BLOCKED at Capable tier, effort exhausted | Flag for human intervention | Critical |
291
+ | MS-004 | Task specification gives only goal and constraints | Start at Standard or higher | Medium |
292
+ | MS-005 | Output shallow but sound; effort not yet at max | Raise effort on the same model — do **not** escalate tier | High |
293
+ | MS-006 | Output wrong in kind and effort already at max | Escalate model tier — raising effort buys nothing | High |
294
+ | MS-007 | Candidate model fails a required hard boundary (context, effort support, modality) | Exclude **before** cost comparison | Critical |
295
+ | MS-008 | Dispatching specification-sensitive work to a high-capability tier | Probe feasibility first; inspect every response for a refusal marker | Critical |
296
+ | MS-009 | Prompt is a fully written-out step list | Prefer a literal-following tier; do not escalate to a high-ceiling tier | Medium |
297
+ | MS-010 | `pin_date` or `measured.at` older than 90 days | WARN and queue for re-measurement (does not block) | Medium |
138
298
 
139
299
  ---
140
300
 
141
301
  ## Cost Optimization Tips
142
302
 
143
- 1. **Batch simple tasks** — send multiple Fast-tier tasks in one session
144
- 2. **Pre-classify tasks** — use task metadata to auto-select tiers
145
- 3. **Track escalation rates** — high escalation from Fast → Standard may indicate poor task decomposition
146
- 4. **Review Capable usage** — ensure Capable-tier tasks genuinely need that level
303
+ 1. **Try the effort axis first** — it is the cheaper of the two axes to move
304
+ 2. **Batch simple tasks** — send multiple Fast-tier tasks in one session
305
+ 3. **Track escalation rates by axis** — frequent effort escalation and frequent tier escalation have different root causes and different fixes
306
+ 4. **Review Capable usage** — a Capable dispatch on a fully defined specification is both more expensive and likely worse
147
307
 
148
308
  ---
149
309
 
150
- ## References
310
+ ## LLM Capability Management (XSPEC-027)
151
311
 
152
- - **Superpowers**: [subagent-driven-development](https://github.com/obra/superpowers) (MIT)
153
- - **Cost-Effective AI**: Principle of using the minimum capability needed
312
+ The two axes above answer "which model, how deep". This section answers a third, independent question: **can this model do this kind of thing at all, and how well** — needed in multi-model-pool environments.
154
313
 
155
- ---
314
+ ### capability_dimensions — Capability Dimensions
156
315
 
157
- ## LLM 能力管理(XSPEC-027)
316
+ Capabilities are grouped in four categories, ten sub-dimensions:
158
317
 
159
- > **Version**: 2.0.0 — 本節於 2026-04-13 新增,對應 XSPEC-027 Phase 1。
318
+ | Category | Sub-dimension | Description | Benchmark |
319
+ |------|--------|------|---------|
320
+ | modality | vision | Image / screenshot understanding | internal-vision-bench |
321
+ | modality | audio | Speech understanding | future-audio-bench |
322
+ | modality | image_generation | Image generation | provider-specific |
323
+ | reasoning | code_reasoning | Code understanding and generation quality | humaneval-plus |
324
+ | reasoning | math_reasoning | Mathematical reasoning accuracy | gsm8k |
325
+ | reasoning | instruction_following | Multi-step instruction adherence | internal-instruction-bench |
326
+ | reasoning | long_context_quality | Mid-document information access | needle-in-haystack |
327
+ | output | structured_output | JSON / schema output success rate | internal-json-bench |
328
+ | output | tool_use | Function-calling correctness | internal-tool-bench |
329
+ | language | multilingual_zh_tw | Traditional Chinese quality | internal-zh-tw-bench |
160
330
 
161
- 三層模型選擇策略處理「任務複雜度」維度;本節新增「模型能力維度」管理,用於多模型池環境。
331
+ #### Each sub-dimension has TWO independent fields
162
332
 
163
- ### capability_dimensions — 能力維度分類
333
+ > **A single 1–5 score cannot express two different things.** "Does it support this" is binary,
334
+ > vendor-declared, and obtainable at connection time. "How good is it" is continuous and requires a
335
+ > benchmark run. Encoding both in one number means the cheap fact and the expensive fact become
336
+ > indistinguishable — and the expensive one silently wins.
164
337
 
165
- 能力維度分為四大類,共 9 個子維度:
338
+ 每個子維度拆為兩個獨立欄位:
166
339
 
167
- | 大類 | 子維度 | 說明 | 基準測試 |
168
- |------|--------|------|---------|
169
- | modality | vision | 圖片/截圖理解(UI 分析、圖表解讀) | internal-vision-bench |
170
- | modality | audio | 語音理解能力 | future-audio-bench |
171
- | modality | image_generation | 圖片生成能力 | provider-specific |
172
- | reasoning | code_reasoning | 程式碼理解與生成品質 | humaneval-plus |
173
- | reasoning | math_reasoning | 數學推理準確率 | gsm8k |
174
- | reasoning | instruction_following | 複雜多步驟指令遵循率 | internal-instruction-bench |
175
- | reasoning | long_context_quality | 長文件中間段資訊存取 | needle-in-haystack |
176
- | output | structured_output | JSON/Schema 格式輸出成功率 | internal-json-bench |
177
- | output | tool_use | Function Calling 正確率 | internal-tool-bench |
178
- | language | multilingual_zh_tw | 繁體中文品質(本系統優先語言) | internal-zh-tw-bench |
179
-
180
- **評分量表(1–5)**:
181
-
182
- | 分數 | 意義 |
340
+ | Field | Type | Source | Cost | Meaning |
341
+ |---|---|---|---|---|
342
+ | `declared` | boolean | Provider model-description endpoint or configuration declaration | ~0, at connection time | **Can it do this at all** |
343
+ | `measured` | `{ score: 1–5, at: date, version_identifier: string }` | Benchmark run | High, offline | **How well does it do it** |
344
+
345
+ **Normative rules**:
346
+
347
+ 1. **`declared: false` is a hard boundary.** The model cannot do this, regardless of any score. It is excluded, and it is **not** queued for calibration — measuring it is meaningless.
348
+ 2. **`declared: true` with `measured` absent = `UNKNOWN`.** It can do this; we do not know how well. This is an information gap, not a conclusion.
349
+ 3. **`supported` MUST NOT be derived from `score`.** Any implementation computing `supported = score > 0` (or similar) has merged the two fields back into one axis and reintroduced the defect this rule exists to prevent.
350
+ 4. **A failed measurement MUST NOT be recorded as a score.** Record it as absent (`UNKNOWN`). Filling a failed measurement with a default value — including `0`, which is not on the 1–5 scale — makes "we could not measure it" indistinguishable from "we measured it and it was bad".
351
+
352
+ **Scoring scale (1–5), applies to `measured.score` only**:
353
+
354
+ | Score | Meaning |
183
355
  |------|------|
184
- | 5 | 生產就緒 — 高準確率,可直接使用 |
185
- | 4 | 良好 — 偶有遺漏,可接受 |
186
- | 3 | 基本可用 — 需人工補充 |
187
- | 2 | 部分可用 — 僅供參考 |
188
- | 1 | 不可靠 — 不建議使用 |
356
+ | 5 | Production-ready — high accuracy, usable directly |
357
+ | 4 | Good — occasional gaps, acceptable |
358
+ | 3 | Basically usable — needs human supplementation |
359
+ | 2 | Partially usable — reference only |
360
+ | 1 | Unreliable — not recommended |
361
+
362
+ There is no `0`. Absence of a measurement is expressed by the **absence of the `measured` object**, never by a score.
363
+
364
+ ### capability_registry — Model Capability Registry
189
365
 
190
- ### capability_registry — 模型能力登記表
366
+ Each project maintains scores for its own models in its own `capability_registry`, based on its own measurements.
191
367
 
192
- 各子專案依實際測試,在自己的 `capability_registry` 中維護每個模型的能力分數。
368
+ > **This standard registers no concrete vendor model IDs, by rule.** A model ID written into a
369
+ > standard is a citation with an expiry date and no clock: it goes stale, and a stale entry is
370
+ > indistinguishable on the page from a current one. Examples below use placeholders.
193
371
 
194
- **格式**:
372
+ **本標準的 examples 不得登記任何具體廠商模型 ID**——那是等著過期的引用端。具體登記由採用者在自己的 registry 維護。
373
+
374
+ **Format**:
195
375
  ```yaml
196
- - model_id: "provider/model-name"
197
- version_pinned: "版本鎖定識別(SHA、日期戳或 model_version)"
198
- pin_date: "YYYY-MM-DD"
199
- eol_date: "YYYY-MM-DD" # 可選
376
+ - model_id: "<provider>/<model-name>" # placeholder — adopters fill in
377
+ version_pinned: "<version-identifier>" # SHA, date stamp, or model_version
378
+ pin_date: "<YYYY-MM-DD>"
379
+ eol_date: "<YYYY-MM-DD>" # optional
200
380
  capabilities:
201
- "modality.vision": 4
202
- "reasoning.instruction_following": 5
203
- "output.structured_output": 5
381
+ "modality.vision":
382
+ declared: true
383
+ measured:
384
+ score: 4
385
+ at: "<YYYY-MM-DD>"
386
+ version_identifier: "<version-identifier measured against>"
387
+ "modality.audio":
388
+ declared: false # hard boundary — no measured field, none needed
389
+ "output.tool_use":
390
+ declared: true # measured absent → UNKNOWN → calibration queue
204
391
  ```
205
392
 
206
- **版本鎖定(DEC-031 D1)**:必須記錄 `version_pinned` 與 `pin_date`,避免模型靜默升級造成能力變化。
393
+ **Version pinning (DEC-031 D1)**: `version_pinned` and `pin_date` are REQUIRED, to prevent silent model upgrades from changing capability without notice.
394
+
395
+ **Staleness check (WARN, not BLOCK)**: a `pin_date` or `measured.at` older than **90 days** MUST raise a WARN naming the affected `model_id` and sub-dimension. It MUST NOT block a release — the false-positive rate of purely in-file invariants is too high to gate on.
396
+
397
+ ### routing_rules — Four-State Routing
398
+
399
+ > **"Never measured" and "measured and unreliable" are not the same state.** Collapsing them means a
400
+ > newly detected model is excluded on its first evaluation and never re-enters the pool — the exact
401
+ > opposite of supporting more models.
402
+
403
+ 「沒測過」與「測過且不可靠」不是同一件事。壓成同一態的後果是**新模型永遠進不了候選池**。
404
+
405
+ | State | Condition | Action |
406
+ |---|---|---|
407
+ | `SUPPORTED` | All required capabilities `measured.score` ≥ `min_score` | Execute normally |
408
+ | `DEGRADED` | `measured.score` present, ≥ 2, but below `min_score` | Execute degraded; mark output `[DEGRADED]` |
409
+ | `UNSUPPORTED` | **`measured` present** and `score` ≤ 1 | Exclude; do **not** queue for calibration (a conclusion exists) |
410
+ | `UNKNOWN` | **`measured` absent** — never registered, measurement failed, or data expired — while `declared: true` | **Queue for calibration**; MUST NOT be silently excluded before calibration completes |
411
+
412
+ **`declared: false`** is handled before this table: it is a hard boundary ([R3a](#r3a-hard-boundaries)), excluded and **not** queued.
207
413
 
208
- ### routing_rules — 路由策略決策樹
414
+ **Observability requirement**: `UNKNOWN` and `UNSUPPORTED` MUST be distinguishable **in the return structure**, not merely in logs. A caller has to be able to tell "this model cannot do it" from "I do not yet know whether this model can do it" — they lead to different next actions. A routing API returning a single boolean, or returning three states, cannot express this.
209
415
 
210
- 根據模型能力分數,路由規則分三態:
416
+ **Decision tree**:
211
417
 
212
418
  ```
213
- 任務需要 capability X
214
- ├── 模型得分 ≥ min_score → SUPPORTED — 正常執行
215
- ├── 模型得分 2–3(低於 min_score) → DEGRADED — 降級執行,產出標記 [DEGRADED]
216
- └── 模型得分 ≤ 1 或未登記 → UNSUPPORTED — 替代流程或提示用戶
419
+ Task requires capability X
420
+ ├── declared == false → HARD BOUNDARY — exclude, no calibration
421
+ ├── measured absent / failed / expired → UNKNOWN — queue for calibration
422
+ ├── measured.score ≤ 1 → UNSUPPORTED — exclude, no calibration
423
+ ├── measured.score ≥ 2 and < min_score → DEGRADED — run, mark [DEGRADED]
424
+ └── measured.score ≥ min_score → SUPPORTED — run
217
425
  ```
218
426
 
219
- **降智偵測(DEC-033)額外規則**:
427
+ ### Re-measurement Triggers — Three Independent Paths
220
428
 
221
- | 規則 | 條件 | 動作 |
222
- |------|------|------|
223
- | CAP-004 | 降智偵測觸發 moderate 信號 | 啟動金絲雀測試,記錄降智警告 |
224
- | CAP-005 | 降智偵測觸發 critical 信號 | 自動切換備用模型,上報 P1 Issue |
429
+ > **Version change is a sufficient condition, not a necessary one.** Degradation detection
430
+ > (DEC-033) exists precisely because behaviour changes while the model ID and version string
431
+ > stay the same. An implementation triggered only by version change misses the entire scenario
432
+ > DEC-033 was built to catch.
225
433
 
226
- **選擇策略**:`pareto_weighted` — 優先選擇在所需維度得分最高、成本最低的模型。
434
+ **版本變更是充分條件,不是必要條件。**
227
435
 
228
- ### 與三層分級的關係
436
+ | # | Trigger | Effect |
437
+ |---|---|---|
438
+ | 1 | `version_identifier` differs from the one recorded in `measured` | Existing measurement is **invalidated immediately** → state becomes `UNKNOWN` |
439
+ | 2 | `measured.at` older than 90 days | Queue for re-measurement; WARN (see MS-010) |
440
+ | 3 | Degradation-detection alert (DEC-033, CAP-004 / CAP-005) | Queue for re-measurement **even though the version string did not change** |
229
441
 
230
- - **三層分級(Tier 1/2/3)** — 基於任務複雜度(How hard is the task?)
231
- - **能力維度(XSPEC-027)** — 基於模型能力(Can this model do it?)
442
+ These three are **independent**: each fires on its own, and none is a precondition for another.
232
443
 
233
- 兩者並用:先依複雜度選擇 Tier,再依能力維度確認選中的模型支援任務所需的 capability。
444
+ ### Capability Rules
445
+
446
+ | ID | Condition | Action | Priority |
447
+ |------|------|------|------|
448
+ | CAP-001 | Required capability's `measured.score` ≥ `min_score` | SUPPORTED — execute normally | High |
449
+ | CAP-002 | `measured.score` ≥ 2 but below `min_score` | DEGRADED — degraded flow, mark output `[DEGRADED]` | Medium |
450
+ | CAP-003 | **`measured` present** and `score` ≤ 1 | UNSUPPORTED — alternative flow or prompt user; do not calibrate | High |
451
+ | CAP-004 | Degradation detection (DEC-033) raises a moderate signal | Start canary testing, record degradation warning, **queue affected capabilities for re-measurement** | High |
452
+ | CAP-005 | Degradation detection raises a critical signal | Switch to fallback model, file a P1 issue, **queue affected capabilities for re-measurement** | Critical |
453
+ | CAP-006 | `declared: true` and `measured` absent, failed, or expired | UNKNOWN — queue for calibration; MUST NOT be silently excluded | High |
454
+ | CAP-007 | `declared: false` for a required capability | Hard boundary — exclude before cost comparison; do **not** queue for calibration | Critical |
455
+ | CAP-008 | `version_identifier` changed since `measured.version_identifier` | Invalidate the measurement immediately → `UNKNOWN` | High |
456
+
457
+ **Selection strategy**: `pareto_weighted` — prefer the model with the highest scores on the required dimensions at the lowest cost, **among candidates that passed hard-boundary exclusion**.
458
+
459
+ ### Relationship to the Two Axes
460
+
461
+ - **Model axis** — reasoning ceiling and temperament (which model)
462
+ - **Effort axis** — reasoning depth for this dispatch (how deep)
463
+ - **Capability dimensions** — whether the chosen model can do this *kind* of thing at all, and how well
464
+
465
+ Order of application: exclude on hard boundaries → select tier by the two criteria → confirm the required capabilities are `SUPPORTED` or acceptably `DEGRADED` → choose an effort level.
234
466
 
235
467
  ---
236
468
 
237
469
  ## Navigation Integration | 與下一步建議整合
238
470
 
239
- The [ai-response-navigation](ai-response-navigation.md) standard (Rule R6, optional) allows each next-step suggestion option to carry a Tier annotation. The complexity signals in this standard serve as the judgment basis.
471
+ The [ai-response-navigation](ai-response-navigation.md) standard (Rule R6, optional) allows each next-step suggestion option to carry a Tier annotation. The criteria in this standard serve as the judgment basis.
472
+
473
+ > **Changed in 2.1.0**: annotations previously derived from file count. Existing annotations made
474
+ > under the old criteria may no longer be accurate. Tier ids are unchanged, so no data migration is
475
+ > required; re-evaluate annotations when the surrounding text is next edited.
240
476
 
241
- **Vendor-neutral principle**: Tier names (Fast / Standard / Capable) are tool-agnostic labels. They do not correspond to any specific vendor's model identifiers. Each platform or tool independently maintains its own Tier → concrete model mapping.
477
+ **Vendor-neutral principle**: Tier names (Fast / Standard / Capable) and effort labels (`low` … `max`) are tool-agnostic. They do not correspond to any specific vendor's model identifiers or parameter names. Each platform or tool independently maintains its own tier → model mapping, effort-label → parameter mapping, and hard-boundary register.
242
478
 
243
479
  Example mapping (illustrative only, not normative):
244
480
 
245
481
  | Tier | Typical capability level |
246
482
  |------|--------------------------|
247
- | Fast | Lightweight / instruction-following |
483
+ | Fast | Lightweight / literal instruction-following |
248
484
  | Standard | Balanced reasoning + code generation |
249
- | Capable | Advanced reasoning, architectural analysis |
485
+ | Capable | High reasoning ceiling, ambiguity navigation |
486
+
487
+ ---
488
+
489
+ ## Related Standards
490
+
491
+ - [agent-dispatch](agent-dispatch.md) — **how to dispatch**: parallel safety, independent-domain criteria, status protocol, prompt design. This standard covers **who to dispatch to and how deeply**; the two are complementary and **must not duplicate each other**. Parallel-safety rules belong there, not here.
492
+ - [verification-evidence](verification-evidence.md) — why "no error was raised" is not evidence of success; the basis for the refusal-marker requirement in [R3b](#r3b-reverse-risks)
493
+ - [systematic-debugging](systematic-debugging.md) — diagnosing before changing, applied here to the depth-vs-ceiling decision
494
+
495
+ ---
496
+
497
+ ## References
498
+
499
+ - **Superpowers**: [subagent-driven-development](https://github.com/obra/superpowers) (MIT)
500
+ - **Cost-Effective AI**: Principle of using the minimum capability needed
250
501
 
251
502
  ---
252
503
 
253
504
  ## Version History
254
505
 
506
+ > **Note on ordering**: rows are in version order, not date order. Version `2.0.0` (2026-04-13)
507
+ > predates `1.0.1` (2026-06-10) because the capability-management section carried its own version
508
+ > sequence at the time. Version 2.1.0 removes section-level versions; see
509
+ > [Version Field Semantics](#version-field-semantics).
510
+
255
511
  | Version | Date | Changes |
256
512
  |---------|------|---------|
513
+ | 2.1.0 | 2026-08-11 | **Two-axis restructure (XSPEC-362)**. Model-axis criteria changed from file count to reasoning-ceiling requirement × specification definiteness (R1). Added the orthogonal effort axis with vendor-neutral levels and the depth-vs-ceiling failure diagnosis (R2). Added reverse-exclusion rules: hard boundaries and safety-classifier refusal as silent failure (R3). `capability_dimensions` sub-dimensions split into `declared` / `measured` (R7b); `routing_rules` extended to four states separating `UNKNOWN` from `UNSUPPORTED` (R7a); three independent re-measurement triggers (R7c). `capability_registry` examples replaced with placeholders and a 90-day staleness WARN added (R4). Section versions removed; `.ai.yaml` `meta.version` defined as the whole-file version (R6a). New rules MS-005–MS-010, CAP-006–CAP-008. |
514
+ | 2.0.0 | 2026-04-13 | Add LLM Capability Management section (XSPEC-027 Phase 1): `capability_dimensions`, `capability_registry`, `routing_rules`. *Recorded retroactively in 2.1.0 — this change was never entered in the version history when it was made.* |
257
515
  | 1.0.1 | 2026-06-10 | Add Navigation Integration section (R6 cross-reference; vendor-neutral principle) |
258
516
  | 1.0.0 | 2026-03-20 | Initial release |