universal-dev-standards 6.3.10 → 6.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/uds.js +2 -2
- package/bundled/ai/standards/agent-dispatch.ai.yaml +162 -0
- package/bundled/ai/standards/ai-friendly-architecture.ai.yaml +1 -1
- package/bundled/ai/standards/ai-instruction-standards.ai.yaml +190 -15
- package/bundled/ai/standards/class-level-fix.ai.yaml +178 -0
- package/bundled/ai/standards/commit-message.ai.yaml +2 -0
- package/bundled/ai/standards/model-selection.ai.yaml +370 -72
- package/bundled/ai/standards/mutation-testing.ai.yaml +105 -2
- package/bundled/ai/standards/project-structure.ai.yaml +1 -1
- package/bundled/ai/standards/security-standards.ai.yaml +22 -1
- package/bundled/ai/standards/spec-driven-development.ai.yaml +59 -2
- package/bundled/ai/standards/test-governance.ai.yaml +49 -2
- package/bundled/ai/standards/testing.ai.yaml +49 -3
- package/bundled/ai/standards/translation-lifecycle-standards.ai.yaml +4 -4
- package/bundled/ai/standards/verification-evidence.ai.yaml +48 -4
- package/bundled/core/class-level-fix.md +184 -0
- package/bundled/core/model-selection.md +383 -125
- package/bundled/core/mutation-testing.md +41 -2
- package/bundled/core/test-governance.md +22 -2
- package/bundled/core/translation-lifecycle-standards.md +6 -6
- package/bundled/core/verification-evidence.md +42 -3
- package/bundled/locales/zh-CN/CHANGELOG.md +26 -3
- package/bundled/locales/zh-CN/CLAUDE.md +1 -1
- package/bundled/locales/zh-CN/README.md +3 -3
- package/bundled/locales/zh-CN/SECURITY.md +1 -1
- package/bundled/locales/zh-CN/core/model-selection.md +375 -60
- package/bundled/locales/zh-CN/core/mutation-testing.md +1 -1
- package/bundled/locales/zh-CN/core/test-governance.md +1 -1
- package/bundled/locales/zh-CN/core/translation-lifecycle-standards.md +1 -1
- package/bundled/locales/zh-CN/core/verification-evidence.md +1 -1
- package/bundled/locales/zh-CN/docs/CHEATSHEET.md +9 -12
- package/bundled/locales/zh-CN/docs/FEATURE-REFERENCE.md +16 -19
- package/bundled/locales/zh-TW/CHANGELOG.md +51 -3
- package/bundled/locales/zh-TW/CLAUDE.md +1 -1
- package/bundled/locales/zh-TW/README.md +3 -3
- package/bundled/locales/zh-TW/SECURITY.md +1 -1
- package/bundled/locales/zh-TW/core/class-level-fix.md +163 -0
- package/bundled/locales/zh-TW/core/model-selection.md +385 -47
- package/bundled/locales/zh-TW/core/mutation-testing.md +45 -6
- package/bundled/locales/zh-TW/core/test-governance.md +22 -3
- package/bundled/locales/zh-TW/core/translation-lifecycle-standards.md +1 -1
- package/bundled/locales/zh-TW/core/verification-evidence.md +33 -6
- package/bundled/locales/zh-TW/docs/CHEATSHEET.md +9 -12
- package/bundled/locales/zh-TW/docs/FEATURE-REFERENCE.md +16 -19
- package/bundled/locales/zh-TW/integrations/claude-code/README.md +31 -5
- package/package.json +1 -1
- package/src/utils/reference-sync.js +83 -16
- package/standards-registry.json +32 -8
|
@@ -2,8 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
> **Language**: English | [繁體中文](../locales/zh-TW/core/model-selection.md)
|
|
4
4
|
|
|
5
|
-
**Version**: 1.0
|
|
6
|
-
**Last Updated**: 2026-
|
|
5
|
+
**Version**: 2.1.0
|
|
6
|
+
**Last Updated**: 2026-08-11
|
|
7
7
|
**Applicability**: AI-assisted development with multiple model tiers
|
|
8
8
|
**Scope**: universal
|
|
9
9
|
**Inspired by**: [Superpowers](https://github.com/obra/superpowers) — subagent-driven-development (MIT)
|
|
@@ -12,9 +12,9 @@
|
|
|
12
12
|
|
|
13
13
|
## Purpose
|
|
14
14
|
|
|
15
|
-
Define
|
|
15
|
+
Define how to choose **which model** and **how deeply it should think** — two independent decisions that must not be collapsed into one. Use the cheapest combination that can handle the job, and escalate along the correct axis when it cannot.
|
|
16
16
|
|
|
17
|
-
|
|
17
|
+
定義兩個獨立決策:**選哪個模型**與**要它想多深**。兩者不可壓成同一軸。使用能勝任的最便宜組合,失敗時沿**正確的軸**升級。
|
|
18
18
|
|
|
19
19
|
---
|
|
20
20
|
|
|
@@ -22,98 +22,251 @@ Define a cost-effective strategy for selecting AI model tiers based on task comp
|
|
|
22
22
|
|
|
23
23
|
| Term | Definition |
|
|
24
24
|
|------|-----------|
|
|
25
|
-
| Model Tier | A classification level representing model
|
|
26
|
-
|
|
|
27
|
-
|
|
|
25
|
+
| Model Tier | A classification level representing a model's reasoning ceiling and cost |
|
|
26
|
+
| Effort Level | How much reasoning depth is requested **for one dispatch**; a request parameter, not a model property |
|
|
27
|
+
| Reasoning Ceiling | The limit beyond which more thinking time yields no better answer |
|
|
28
|
+
| Specification Definiteness | Whether the steps are already defined, or only the goal and constraints are known |
|
|
29
|
+
| Hard Boundary | A capability the model does not have at all (not "worse at") — e.g. context capacity, effort-parameter support |
|
|
30
|
+
| Refusal Marker | A field in an otherwise-normal response indicating the request was declined |
|
|
31
|
+
| Escalation | Moving up the model axis after the effort axis is exhausted |
|
|
28
32
|
|
|
29
33
|
---
|
|
30
34
|
|
|
31
|
-
##
|
|
35
|
+
## Version Field Semantics
|
|
32
36
|
|
|
33
|
-
> **
|
|
37
|
+
> **This document carries exactly one version.** The `Version` field in the header covers the
|
|
38
|
+
> whole file, including every section. **Sections MUST NOT carry their own version markers.**
|
|
39
|
+
> Section-level change history belongs in the [Version History](#version-history) table.
|
|
40
|
+
>
|
|
41
|
+
> The `standard.meta.version` field in `ai/standards/model-selection.ai.yaml` corresponds to
|
|
42
|
+
> **this whole-file version**, never to a section version.
|
|
34
43
|
|
|
35
|
-
|
|
44
|
+
本檔只有一個版本欄,位於檔頭,涵蓋全檔。**章節不得自帶版本標記**;章節的變更歷程一律記於 Version History 表。`.ai.yaml` 的 `standard.meta.version` 對應**全檔版本**。
|
|
36
45
|
|
|
37
46
|
---
|
|
38
47
|
|
|
39
|
-
##
|
|
48
|
+
## Core Principle — Two Orthogonal Axes
|
|
40
49
|
|
|
41
|
-
|
|
50
|
+
> **The model axis buys a ceiling and a temperament. The effort axis buys depth on this one dispatch. Choosing on one axis tells you nothing about the other.**
|
|
42
51
|
|
|
43
|
-
|
|
52
|
+
模型軸買的是「天花板與性格」,effort 軸買的是「這一次要它想多深」。在一軸上的選擇,不決定另一軸。
|
|
44
53
|
|
|
45
|
-
|
|
54
|
+
```
|
|
55
|
+
effort → low medium high very-high max
|
|
56
|
+
model tier ↓
|
|
57
|
+
fast · · · ? ?
|
|
58
|
+
standard · · · ? ?
|
|
59
|
+
capable · · · ? ?
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Every cell is in principle a legal dispatch: the two axes are chosen independently.
|
|
63
|
+
|
|
64
|
+
`?` marks a cell whose availability **this standard cannot state**. Whether a given model accepts a given effort level is a property of that model, and tier names here are vendor-neutral labels — so the answer lives in each platform's tier→model mapping, not in this table. An unsupported level is a **hard boundary**, not a poor choice: see [R3a](#r3a-hard-boundaries). Adopters MUST record, per model, which levels it accepts.
|
|
65
|
+
|
|
66
|
+
`?` 標示的格子,**本標準無法斷言其可用性**——某個模型接受哪些 effort 級距是該模型的屬性,而此處的層級名稱是與廠商無關的標籤。答案在各平台自己的「層級 → 模型」映射裡。
|
|
67
|
+
|
|
68
|
+
---
|
|
69
|
+
|
|
70
|
+
## Axis 1 — Model Tier
|
|
71
|
+
|
|
72
|
+
### Selection Criteria
|
|
73
|
+
|
|
74
|
+
Two criteria decide the tier. **Neither is "how many files does this change".**
|
|
75
|
+
|
|
76
|
+
兩個判準決定分層。**兩者都不是「改幾個檔案」**。
|
|
77
|
+
|
|
78
|
+
#### Criterion 1 — Reasoning Ceiling Requirement(推理天花板需求)
|
|
79
|
+
|
|
80
|
+
Does the task contain a component that **more thinking time cannot solve**?
|
|
81
|
+
|
|
82
|
+
- **No** → a lower tier suffices; buy depth with effort instead of buying a ceiling with money.
|
|
83
|
+
- **Yes** → a higher ceiling is required; no amount of effort on a lower tier reaches it.
|
|
84
|
+
|
|
85
|
+
這個任務是否存在「想更久也想不出來」的成分?若否,較低層加 effort 即可;若是,才需要更高天花板。
|
|
86
|
+
|
|
87
|
+
#### Criterion 2 — Specification Definiteness(規格明確度)
|
|
88
|
+
|
|
89
|
+
**This criterion is bidirectional. It is not monotonically increasing.**
|
|
90
|
+
|
|
91
|
+
| Specification state | Prefer | Failure mode if inverted |
|
|
92
|
+
|---|---|---|
|
|
93
|
+
| Steps already defined, path known | Literal-following tier (lower) | Feeding a written-out step list to a high-ceiling model **lowers** output quality — it reinterprets what was already decided |
|
|
94
|
+
| Only goal and constraints known, path unknown | Ambiguity-navigating tier (higher) | Handing an ambiguous task to a literal-following model yields **"precisely executed the wrong sentence"** |
|
|
95
|
+
|
|
96
|
+
規格明確度是**雙向**判準:把模糊任務給字面遵循型,會得到「精確執行了錯誤的那句話」;把已寫死的步驟餵給高天花板型,反而降低品質。
|
|
97
|
+
|
|
98
|
+
> **Why file count was removed.** A 3-file redesign of module boundaries is harder than an 8-file
|
|
99
|
+
> mechanical rename. File count mis-sorts "deep and narrow" work **in a fixed direction** — it is a
|
|
100
|
+
> bias, not noise. It survived as a signal because it is cheap to measure, not because it predicts
|
|
101
|
+
> anything.
|
|
102
|
+
|
|
103
|
+
### The Three Tiers
|
|
104
|
+
|
|
105
|
+
Tier ids (`fast` / `standard` / `capable`) are unchanged. Only the criteria changed.
|
|
106
|
+
|
|
107
|
+
#### Tier 1: Fast(快速層)
|
|
108
|
+
|
|
109
|
+
**Purpose**: Work with no reasoning ceiling requirement and a fully definite specification.
|
|
110
|
+
|
|
111
|
+
無推理天花板需求、規格完全明確的工作。
|
|
46
112
|
|
|
47
|
-
**
|
|
48
|
-
-
|
|
49
|
-
-
|
|
113
|
+
**Signals**:
|
|
114
|
+
- No component that "more thinking cannot solve"
|
|
115
|
+
- Steps are written out; no path-finding needed
|
|
50
116
|
- No design judgment required
|
|
51
|
-
-
|
|
117
|
+
- Literal following is exactly what is wanted
|
|
52
118
|
|
|
53
119
|
**Examples**:
|
|
54
|
-
-
|
|
55
|
-
- Add a new export statement
|
|
120
|
+
- Apply a defined edit across a set of files (mechanical rename, version bump)
|
|
56
121
|
- Fix a typo
|
|
57
|
-
-
|
|
122
|
+
- Add an export statement named in the spec
|
|
58
123
|
|
|
59
|
-
|
|
124
|
+
#### Tier 2: Standard(標準層)
|
|
60
125
|
|
|
61
|
-
**Purpose**:
|
|
126
|
+
**Purpose**: A modest reasoning ceiling, with a specification that is mostly defined but leaves local decisions open.
|
|
62
127
|
|
|
63
|
-
|
|
128
|
+
需要一定推理天花板,規格大致明確但留有局部決策空間。
|
|
64
129
|
|
|
65
|
-
**
|
|
66
|
-
-
|
|
130
|
+
**Signals**:
|
|
131
|
+
- Some components need reasoning, but none are "unreachable by thinking longer"
|
|
132
|
+
- Goal and most steps defined; a bounded set of local choices remains
|
|
67
133
|
- Requires understanding inter-module relationships
|
|
68
|
-
- Needs some design judgment
|
|
69
|
-
- Cross-cutting but within a bounded context
|
|
70
134
|
|
|
71
135
|
**Examples**:
|
|
72
|
-
-
|
|
73
|
-
- Refactor a module
|
|
74
|
-
-
|
|
75
|
-
- Write integration tests for a subsystem
|
|
136
|
+
- Implement a defined feature whose integration points are known
|
|
137
|
+
- Refactor a module against a stated target structure
|
|
138
|
+
- Write integration tests for a specified subsystem
|
|
76
139
|
|
|
77
|
-
|
|
140
|
+
#### Tier 3: Capable(能力層)
|
|
78
141
|
|
|
79
|
-
**Purpose**:
|
|
142
|
+
**Purpose**: A high reasoning ceiling, or a specification that gives only goal and constraints.
|
|
80
143
|
|
|
81
|
-
|
|
144
|
+
需要高推理天花板,或規格僅有目標與約束、路徑未知。
|
|
82
145
|
|
|
83
|
-
**
|
|
84
|
-
-
|
|
85
|
-
-
|
|
86
|
-
-
|
|
87
|
-
- Complex debugging or performance analysis
|
|
88
|
-
- Ambiguous requirements needing interpretation
|
|
146
|
+
**Signals**:
|
|
147
|
+
- Contains a component that more thinking time alone cannot resolve
|
|
148
|
+
- Path is unknown; the answer must be found, not applied
|
|
149
|
+
- Requires navigating ambiguity rather than following text
|
|
89
150
|
|
|
90
151
|
**Examples**:
|
|
91
|
-
- Design a
|
|
92
|
-
- Review a large pull request
|
|
93
|
-
- Diagnose
|
|
94
|
-
- Refactor a major component with many dependents
|
|
152
|
+
- Design a subsystem architecture (path unknown)
|
|
153
|
+
- Review a large pull request (what matters is not stated in advance)
|
|
154
|
+
- Diagnose a failure whose cause is not yet hypothesised
|
|
95
155
|
|
|
96
|
-
|
|
156
|
+
> **Note on "large refactor"**: a major refactor with a *defined* target structure is **not**
|
|
157
|
+
> automatically Tier 3 — under Criterion 2 it is a definite specification, and a high-ceiling model
|
|
158
|
+
> may produce worse results than a literal-following one. What raises it to Tier 3 is an *undecided*
|
|
159
|
+
> target structure.
|
|
97
160
|
|
|
98
|
-
|
|
161
|
+
### Selection Decision Flow
|
|
99
162
|
|
|
100
163
|
```
|
|
101
|
-
|
|
102
|
-
├──
|
|
103
|
-
|
|
104
|
-
|
|
164
|
+
Does the task contain something more thinking cannot solve?
|
|
165
|
+
├── Yes → high ceiling needed → Tier 3 (Capable)
|
|
166
|
+
└── No → how definite is the specification?
|
|
167
|
+
├── Fully defined steps, path known → Tier 1 (Fast)
|
|
168
|
+
├── Goal + most steps, local choices → Tier 2 (Standard)
|
|
169
|
+
└── Goal + constraints only → Tier 3 (Capable)
|
|
170
|
+
Then, independently: choose an effort level (Axis 2).
|
|
105
171
|
```
|
|
106
172
|
|
|
107
173
|
---
|
|
108
174
|
|
|
175
|
+
## Axis 2 — Effort (Reasoning Depth)
|
|
176
|
+
|
|
177
|
+
Effort is **a parameter of one dispatch**, not a property of a model. The same model at `low` and at `max` is not the same worker.
|
|
178
|
+
|
|
179
|
+
effort 是**單次派工的參數**,不是模型屬性。
|
|
180
|
+
|
|
181
|
+
### Vendor-Neutral Effort Levels
|
|
182
|
+
|
|
183
|
+
Each platform maps these labels to its own parameter. The labels are the contract; the mapping is local.
|
|
184
|
+
|
|
185
|
+
| Level | Semantics | Typical use |
|
|
186
|
+
|---|---|---|
|
|
187
|
+
| `low`(低) | Minimal deliberation; answer largely on pattern | Mechanical edits, lookups, formatting |
|
|
188
|
+
| `medium`(中) | Ordinary deliberation; the default | Most defined implementation work |
|
|
189
|
+
| `high`(高) | Extended deliberation; considers alternatives before answering | Integration work with local decisions |
|
|
190
|
+
| `very-high`(極高) | Explores and discards candidate approaches | Design, review, non-obvious debugging |
|
|
191
|
+
| `max`(極限) | Maximum available depth for this model | Last attempt before escalating the model axis |
|
|
192
|
+
|
|
193
|
+
**Not every tier supports every level.** Support for the effort parameter is a **hard boundary**, not a quality gradient — see [R3a](#r3a-hard-boundaries). A platform's tier→model mapping MUST record which levels each model accepts.
|
|
194
|
+
|
|
195
|
+
### Failure Diagnosis — Depth or Ceiling?
|
|
196
|
+
|
|
197
|
+
> **When a dispatch fails, first decide which axis was insufficient. The two remedies are not interchangeable.**
|
|
198
|
+
|
|
199
|
+
任務失敗時,**先判斷是「深度不足」還是「天花板不足」**——兩者的補救方式不可互換。
|
|
200
|
+
|
|
201
|
+
| Observation | Diagnosis | Remedy | Wrong remedy costs |
|
|
202
|
+
|---|---|---|---|
|
|
203
|
+
| Output is shallow, skipped considerations, stopped early, but the reasoning it *did* do was sound | **Depth insufficient** | Raise effort on the **same** model | Escalating the tier pays more for a ceiling that was never the constraint |
|
|
204
|
+
| Output is wrong in kind — misunderstood the problem, produced an approach that cannot work — **and this persists at `max` effort** | **Ceiling insufficient** | Escalate the **model** tier | Raising effort again buys nothing; the level is already exhausted |
|
|
205
|
+
|
|
206
|
+
**Ordering rule**: raise effort before escalating the tier. Escalate the tier only after the current tier's **highest supported** effort level has failed. A tier escalation performed without exhausting effort is an untested assumption about which axis was short.
|
|
207
|
+
|
|
208
|
+
**Exception**: if [Criterion 1](#criterion-1--reasoning-ceiling-requirement推理天花板需求) already identified a component that more thinking cannot solve, start at the higher tier. The ordering rule is about *diagnosing failures*, not about ignoring the criteria up front.
|
|
209
|
+
|
|
210
|
+
---
|
|
211
|
+
|
|
212
|
+
## Reverse Exclusion Rules
|
|
213
|
+
|
|
214
|
+
The tier criteria say what work should be **raised** to a tier. This section says what work must **not** be sent to one. These are separate questions, and the second one has no answer derivable from the first.
|
|
215
|
+
|
|
216
|
+
### R3a Hard Boundaries
|
|
217
|
+
|
|
218
|
+
> **A hard boundary is the absence of a capability, not a lower degree of it.** A model that cannot hold the input does not "do the task worse" — it does not do the task.
|
|
219
|
+
|
|
220
|
+
硬邊界是**能力有無**,不是程度差。
|
|
221
|
+
|
|
222
|
+
Adopters MUST record hard boundaries in their tier→model mapping, **alongside and separately from** capability scores:
|
|
223
|
+
|
|
224
|
+
| Hard boundary | Question it answers | Consequence when violated |
|
|
225
|
+
|---|---|---|
|
|
226
|
+
| **Context capacity** | Does the input fit? | Truncation or error — the model never saw part of the task |
|
|
227
|
+
| **Effort parameter support** | Does this model accept an effort level at all? | Request rejected, or silently executed at the model's fixed depth |
|
|
228
|
+
| **Modality support** | Can it accept this input type at all (image, audio)? | Input dropped or rejected |
|
|
229
|
+
|
|
230
|
+
**Rule**: a model failing any required hard boundary is **excluded before cost comparison**, not ranked lower within it. Excluding it afterwards allows a cheaper, incapable model to win on price.
|
|
231
|
+
|
|
232
|
+
**Recording format**: hard boundaries map to `declared` in the [capability registry](#capability_dimensions--capability-dimensions) — a boolean, obtainable at connection time, at near-zero cost. A score of `1` is **not** a hard boundary; it is a measured low quality (see [R7a state table](#routing_rules--four-state-routing)).
|
|
233
|
+
|
|
234
|
+
### R3b Reverse Risks
|
|
235
|
+
|
|
236
|
+
Higher capability is not uniformly better. Two risks run in the opposite direction from the tier criteria:
|
|
237
|
+
|
|
238
|
+
#### 1. Safety-classifier refusal on specification-sensitive work
|
|
239
|
+
|
|
240
|
+
Higher-capability tiers may carry stricter safety classification. Work that is legitimate but resembles a prohibited category — security hardening, exploit-mitigation implementation, credential-handling code, red-team tooling — may be **declined**.
|
|
241
|
+
|
|
242
|
+
> **The refusal is not an error.** The call returns a normal response with a refusal marker.
|
|
243
|
+
> For a caller that checks only the exit code, or only whether an exception was raised, **this is a
|
|
244
|
+
> silent failure**: the pipeline records success, and the work was never done.
|
|
245
|
+
|
|
246
|
+
拒絕**不是錯誤**——回傳的是正常回應加上拒絕標記。對只看 exit code/有無例外的呼叫端,這是**靜默失效**。
|
|
247
|
+
|
|
248
|
+
**Requirements**:
|
|
249
|
+
|
|
250
|
+
1. Callers **MUST inspect the response for a refusal marker**. Absence of an error is not evidence of completion. (See the `verification-evidence` standard: "the checking tool ran and returned nothing useful" is a distinct failure class from "the check was not run".)
|
|
251
|
+
2. Specification-sensitive work **MUST be feasibility-checked before batch dispatch** to a high-capability tier — a cheap probe request, not the full task.
|
|
252
|
+
3. On refusal, **re-dispatch to a different tier or provider**; do not retry the same request at higher effort. Effort does not move a classifier decision.
|
|
253
|
+
|
|
254
|
+
#### 2. Over-specified prompts degrade high-ceiling models
|
|
255
|
+
|
|
256
|
+
Writing out every step is the correct practice for a literal-following tier and the **wrong** practice for a high-ceiling one: the latter reinterprets decisions that were already made, and the output gets worse.
|
|
257
|
+
|
|
258
|
+
**Rule**: prompt granularity follows the tier. If a fully written-out step list is the only prompt available, [Criterion 2](#criterion-2--specification-definiteness規格明確度) already says to send it to a lower tier — sending it to a high tier is a double mistake, paying more for worse output.
|
|
259
|
+
|
|
260
|
+
---
|
|
261
|
+
|
|
109
262
|
## Escalation Rules
|
|
110
263
|
|
|
111
|
-
|
|
264
|
+
Escalation now has two axes; the table applies **after** effort has been exhausted at the current tier.
|
|
112
265
|
|
|
113
|
-
| Current Tier | On
|
|
266
|
+
| Current Tier | On failure at max supported effort | Action |
|
|
114
267
|
|-------------|-----------|--------|
|
|
115
|
-
| Fast | → Standard | Re-dispatch
|
|
116
|
-
| Standard | → Capable | Re-dispatch
|
|
268
|
+
| Fast | → Standard | Re-dispatch at Standard tier |
|
|
269
|
+
| Standard | → Capable | Re-dispatch at Capable tier |
|
|
117
270
|
| Capable | → Human | Flag for human intervention |
|
|
118
271
|
|
|
119
272
|
### Escalation is Not Retry
|
|
@@ -121,9 +274,10 @@ When a model tier fails (returns `BLOCKED`), escalate to the next tier:
|
|
|
121
274
|
Escalation means using a more capable model, not repeating the same action. The higher-tier model receives:
|
|
122
275
|
- The original task
|
|
123
276
|
- The lower-tier model's output and failure reason
|
|
277
|
+
- The effort level already attempted
|
|
124
278
|
- Additional context if available
|
|
125
279
|
|
|
126
|
-
|
|
280
|
+
升級不是重試。更高層級的模型會收到原始任務、低層級的輸出與失敗原因,以及**已嘗試過的 effort 等級**。
|
|
127
281
|
|
|
128
282
|
---
|
|
129
283
|
|
|
@@ -131,128 +285,232 @@ Escalation means using a more capable model, not repeating the same action. The
|
|
|
131
285
|
|
|
132
286
|
| ID | Trigger | Action | Priority |
|
|
133
287
|
|----|---------|--------|----------|
|
|
134
|
-
| MS-001 | BLOCKED at Fast tier | Escalate to Standard | High |
|
|
135
|
-
| MS-002 | BLOCKED at Standard tier | Escalate to Capable | High |
|
|
136
|
-
| MS-003 | BLOCKED at Capable tier | Flag for human intervention | Critical |
|
|
137
|
-
| MS-004 | Task
|
|
288
|
+
| MS-001 | BLOCKED at Fast tier, effort exhausted | Escalate to Standard | High |
|
|
289
|
+
| MS-002 | BLOCKED at Standard tier, effort exhausted | Escalate to Capable | High |
|
|
290
|
+
| MS-003 | BLOCKED at Capable tier, effort exhausted | Flag for human intervention | Critical |
|
|
291
|
+
| MS-004 | Task specification gives only goal and constraints | Start at Standard or higher | Medium |
|
|
292
|
+
| MS-005 | Output shallow but sound; effort not yet at max | Raise effort on the same model — do **not** escalate tier | High |
|
|
293
|
+
| MS-006 | Output wrong in kind and effort already at max | Escalate model tier — raising effort buys nothing | High |
|
|
294
|
+
| MS-007 | Candidate model fails a required hard boundary (context, effort support, modality) | Exclude **before** cost comparison | Critical |
|
|
295
|
+
| MS-008 | Dispatching specification-sensitive work to a high-capability tier | Probe feasibility first; inspect every response for a refusal marker | Critical |
|
|
296
|
+
| MS-009 | Prompt is a fully written-out step list | Prefer a literal-following tier; do not escalate to a high-ceiling tier | Medium |
|
|
297
|
+
| MS-010 | `pin_date` or `measured.at` older than 90 days | WARN and queue for re-measurement (does not block) | Medium |
|
|
138
298
|
|
|
139
299
|
---
|
|
140
300
|
|
|
141
301
|
## Cost Optimization Tips
|
|
142
302
|
|
|
143
|
-
1. **
|
|
144
|
-
2. **
|
|
145
|
-
3. **Track escalation rates** —
|
|
146
|
-
4. **Review Capable usage** —
|
|
303
|
+
1. **Try the effort axis first** — it is the cheaper of the two axes to move
|
|
304
|
+
2. **Batch simple tasks** — send multiple Fast-tier tasks in one session
|
|
305
|
+
3. **Track escalation rates by axis** — frequent effort escalation and frequent tier escalation have different root causes and different fixes
|
|
306
|
+
4. **Review Capable usage** — a Capable dispatch on a fully defined specification is both more expensive and likely worse
|
|
147
307
|
|
|
148
308
|
---
|
|
149
309
|
|
|
150
|
-
##
|
|
310
|
+
## LLM Capability Management (XSPEC-027)
|
|
151
311
|
|
|
152
|
-
|
|
153
|
-
- **Cost-Effective AI**: Principle of using the minimum capability needed
|
|
312
|
+
The two axes above answer "which model, how deep". This section answers a third, independent question: **can this model do this kind of thing at all, and how well** — needed in multi-model-pool environments.
|
|
154
313
|
|
|
155
|
-
|
|
314
|
+
### capability_dimensions — Capability Dimensions
|
|
156
315
|
|
|
157
|
-
|
|
316
|
+
Capabilities are grouped in four categories, ten sub-dimensions:
|
|
158
317
|
|
|
159
|
-
|
|
318
|
+
| Category | Sub-dimension | Description | Benchmark |
|
|
319
|
+
|------|--------|------|---------|
|
|
320
|
+
| modality | vision | Image / screenshot understanding | internal-vision-bench |
|
|
321
|
+
| modality | audio | Speech understanding | future-audio-bench |
|
|
322
|
+
| modality | image_generation | Image generation | provider-specific |
|
|
323
|
+
| reasoning | code_reasoning | Code understanding and generation quality | humaneval-plus |
|
|
324
|
+
| reasoning | math_reasoning | Mathematical reasoning accuracy | gsm8k |
|
|
325
|
+
| reasoning | instruction_following | Multi-step instruction adherence | internal-instruction-bench |
|
|
326
|
+
| reasoning | long_context_quality | Mid-document information access | needle-in-haystack |
|
|
327
|
+
| output | structured_output | JSON / schema output success rate | internal-json-bench |
|
|
328
|
+
| output | tool_use | Function-calling correctness | internal-tool-bench |
|
|
329
|
+
| language | multilingual_zh_tw | Traditional Chinese quality | internal-zh-tw-bench |
|
|
160
330
|
|
|
161
|
-
|
|
331
|
+
#### Each sub-dimension has TWO independent fields
|
|
162
332
|
|
|
163
|
-
|
|
333
|
+
> **A single 1–5 score cannot express two different things.** "Does it support this" is binary,
|
|
334
|
+
> vendor-declared, and obtainable at connection time. "How good is it" is continuous and requires a
|
|
335
|
+
> benchmark run. Encoding both in one number means the cheap fact and the expensive fact become
|
|
336
|
+
> indistinguishable — and the expensive one silently wins.
|
|
164
337
|
|
|
165
|
-
|
|
338
|
+
每個子維度拆為兩個獨立欄位:
|
|
166
339
|
|
|
167
|
-
|
|
|
168
|
-
|
|
169
|
-
|
|
|
170
|
-
|
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
| 分數 | 意義 |
|
|
340
|
+
| Field | Type | Source | Cost | Meaning |
|
|
341
|
+
|---|---|---|---|---|
|
|
342
|
+
| `declared` | boolean | Provider model-description endpoint or configuration declaration | ~0, at connection time | **Can it do this at all** |
|
|
343
|
+
| `measured` | `{ score: 1–5, at: date, version_identifier: string }` | Benchmark run | High, offline | **How well does it do it** |
|
|
344
|
+
|
|
345
|
+
**Normative rules**:
|
|
346
|
+
|
|
347
|
+
1. **`declared: false` is a hard boundary.** The model cannot do this, regardless of any score. It is excluded, and it is **not** queued for calibration — measuring it is meaningless.
|
|
348
|
+
2. **`declared: true` with `measured` absent = `UNKNOWN`.** It can do this; we do not know how well. This is an information gap, not a conclusion.
|
|
349
|
+
3. **`supported` MUST NOT be derived from `score`.** Any implementation computing `supported = score > 0` (or similar) has merged the two fields back into one axis and reintroduced the defect this rule exists to prevent.
|
|
350
|
+
4. **A failed measurement MUST NOT be recorded as a score.** Record it as absent (`UNKNOWN`). Filling a failed measurement with a default value — including `0`, which is not on the 1–5 scale — makes "we could not measure it" indistinguishable from "we measured it and it was bad".
|
|
351
|
+
|
|
352
|
+
**Scoring scale (1–5), applies to `measured.score` only**:
|
|
353
|
+
|
|
354
|
+
| Score | Meaning |
|
|
183
355
|
|------|------|
|
|
184
|
-
| 5 |
|
|
185
|
-
| 4 |
|
|
186
|
-
| 3 |
|
|
187
|
-
| 2 |
|
|
188
|
-
| 1 |
|
|
356
|
+
| 5 | Production-ready — high accuracy, usable directly |
|
|
357
|
+
| 4 | Good — occasional gaps, acceptable |
|
|
358
|
+
| 3 | Basically usable — needs human supplementation |
|
|
359
|
+
| 2 | Partially usable — reference only |
|
|
360
|
+
| 1 | Unreliable — not recommended |
|
|
361
|
+
|
|
362
|
+
There is no `0`. Absence of a measurement is expressed by the **absence of the `measured` object**, never by a score.
|
|
363
|
+
|
|
364
|
+
### capability_registry — Model Capability Registry
|
|
189
365
|
|
|
190
|
-
|
|
366
|
+
Each project maintains scores for its own models in its own `capability_registry`, based on its own measurements.
|
|
191
367
|
|
|
192
|
-
|
|
368
|
+
> **This standard registers no concrete vendor model IDs, by rule.** A model ID written into a
|
|
369
|
+
> standard is a citation with an expiry date and no clock: it goes stale, and a stale entry is
|
|
370
|
+
> indistinguishable on the page from a current one. Examples below use placeholders.
|
|
193
371
|
|
|
194
|
-
|
|
372
|
+
**本標準的 examples 不得登記任何具體廠商模型 ID**——那是等著過期的引用端。具體登記由採用者在自己的 registry 維護。
|
|
373
|
+
|
|
374
|
+
**Format**:
|
|
195
375
|
```yaml
|
|
196
|
-
- model_id: "provider
|
|
197
|
-
version_pinned: "
|
|
198
|
-
pin_date: "YYYY-MM-DD"
|
|
199
|
-
eol_date: "YYYY-MM-DD"
|
|
376
|
+
- model_id: "<provider>/<model-name>" # placeholder — adopters fill in
|
|
377
|
+
version_pinned: "<version-identifier>" # SHA, date stamp, or model_version
|
|
378
|
+
pin_date: "<YYYY-MM-DD>"
|
|
379
|
+
eol_date: "<YYYY-MM-DD>" # optional
|
|
200
380
|
capabilities:
|
|
201
|
-
"modality.vision":
|
|
202
|
-
|
|
203
|
-
|
|
381
|
+
"modality.vision":
|
|
382
|
+
declared: true
|
|
383
|
+
measured:
|
|
384
|
+
score: 4
|
|
385
|
+
at: "<YYYY-MM-DD>"
|
|
386
|
+
version_identifier: "<version-identifier measured against>"
|
|
387
|
+
"modality.audio":
|
|
388
|
+
declared: false # hard boundary — no measured field, none needed
|
|
389
|
+
"output.tool_use":
|
|
390
|
+
declared: true # measured absent → UNKNOWN → calibration queue
|
|
204
391
|
```
|
|
205
392
|
|
|
206
|
-
|
|
393
|
+
**Version pinning (DEC-031 D1)**: `version_pinned` and `pin_date` are REQUIRED, to prevent silent model upgrades from changing capability without notice.
|
|
394
|
+
|
|
395
|
+
**Staleness check (WARN, not BLOCK)**: a `pin_date` or `measured.at` older than **90 days** MUST raise a WARN naming the affected `model_id` and sub-dimension. It MUST NOT block a release — the false-positive rate of purely in-file invariants is too high to gate on.
|
|
396
|
+
|
|
397
|
+
### routing_rules — Four-State Routing
|
|
398
|
+
|
|
399
|
+
> **"Never measured" and "measured and unreliable" are not the same state.** Collapsing them means a
|
|
400
|
+
> newly detected model is excluded on its first evaluation and never re-enters the pool — the exact
|
|
401
|
+
> opposite of supporting more models.
|
|
402
|
+
|
|
403
|
+
「沒測過」與「測過且不可靠」不是同一件事。壓成同一態的後果是**新模型永遠進不了候選池**。
|
|
404
|
+
|
|
405
|
+
| State | Condition | Action |
|
|
406
|
+
|---|---|---|
|
|
407
|
+
| `SUPPORTED` | All required capabilities `measured.score` ≥ `min_score` | Execute normally |
|
|
408
|
+
| `DEGRADED` | `measured.score` present, ≥ 2, but below `min_score` | Execute degraded; mark output `[DEGRADED]` |
|
|
409
|
+
| `UNSUPPORTED` | **`measured` present** and `score` ≤ 1 | Exclude; do **not** queue for calibration (a conclusion exists) |
|
|
410
|
+
| `UNKNOWN` | **`measured` absent** — never registered, measurement failed, or data expired — while `declared: true` | **Queue for calibration**; MUST NOT be silently excluded before calibration completes |
|
|
411
|
+
|
|
412
|
+
**`declared: false`** is handled before this table: it is a hard boundary ([R3a](#r3a-hard-boundaries)), excluded and **not** queued.
|
|
207
413
|
|
|
208
|
-
|
|
414
|
+
**Observability requirement**: `UNKNOWN` and `UNSUPPORTED` MUST be distinguishable **in the return structure**, not merely in logs. A caller has to be able to tell "this model cannot do it" from "I do not yet know whether this model can do it" — they lead to different next actions. A routing API returning a single boolean, or returning three states, cannot express this.
|
|
209
415
|
|
|
210
|
-
|
|
416
|
+
**Decision tree**:
|
|
211
417
|
|
|
212
418
|
```
|
|
213
|
-
|
|
214
|
-
├──
|
|
215
|
-
├──
|
|
216
|
-
|
|
419
|
+
Task requires capability X
|
|
420
|
+
├── declared == false → HARD BOUNDARY — exclude, no calibration
|
|
421
|
+
├── measured absent / failed / expired → UNKNOWN — queue for calibration
|
|
422
|
+
├── measured.score ≤ 1 → UNSUPPORTED — exclude, no calibration
|
|
423
|
+
├── measured.score ≥ 2 and < min_score → DEGRADED — run, mark [DEGRADED]
|
|
424
|
+
└── measured.score ≥ min_score → SUPPORTED — run
|
|
217
425
|
```
|
|
218
426
|
|
|
219
|
-
|
|
427
|
+
### Re-measurement Triggers — Three Independent Paths
|
|
220
428
|
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
|
|
429
|
+
> **Version change is a sufficient condition, not a necessary one.** Degradation detection
|
|
430
|
+
> (DEC-033) exists precisely because behaviour changes while the model ID and version string
|
|
431
|
+
> stay the same. An implementation triggered only by version change misses the entire scenario
|
|
432
|
+
> DEC-033 was built to catch.
|
|
225
433
|
|
|
226
|
-
|
|
434
|
+
**版本變更是充分條件,不是必要條件。**
|
|
227
435
|
|
|
228
|
-
|
|
436
|
+
| # | Trigger | Effect |
|
|
437
|
+
|---|---|---|
|
|
438
|
+
| 1 | `version_identifier` differs from the one recorded in `measured` | Existing measurement is **invalidated immediately** → state becomes `UNKNOWN` |
|
|
439
|
+
| 2 | `measured.at` older than 90 days | Queue for re-measurement; WARN (see MS-010) |
|
|
440
|
+
| 3 | Degradation-detection alert (DEC-033, CAP-004 / CAP-005) | Queue for re-measurement **even though the version string did not change** |
|
|
229
441
|
|
|
230
|
-
|
|
231
|
-
- **能力維度(XSPEC-027)** — 基於模型能力(Can this model do it?)
|
|
442
|
+
These three are **independent**: each fires on its own, and none is a precondition for another.
|
|
232
443
|
|
|
233
|
-
|
|
444
|
+
### Capability Rules
|
|
445
|
+
|
|
446
|
+
| ID | Condition | Action | Priority |
|
|
447
|
+
|------|------|------|------|
|
|
448
|
+
| CAP-001 | Required capability's `measured.score` ≥ `min_score` | SUPPORTED — execute normally | High |
|
|
449
|
+
| CAP-002 | `measured.score` ≥ 2 but below `min_score` | DEGRADED — degraded flow, mark output `[DEGRADED]` | Medium |
|
|
450
|
+
| CAP-003 | **`measured` present** and `score` ≤ 1 | UNSUPPORTED — alternative flow or prompt user; do not calibrate | High |
|
|
451
|
+
| CAP-004 | Degradation detection (DEC-033) raises a moderate signal | Start canary testing, record degradation warning, **queue affected capabilities for re-measurement** | High |
|
|
452
|
+
| CAP-005 | Degradation detection raises a critical signal | Switch to fallback model, file a P1 issue, **queue affected capabilities for re-measurement** | Critical |
|
|
453
|
+
| CAP-006 | `declared: true` and `measured` absent, failed, or expired | UNKNOWN — queue for calibration; MUST NOT be silently excluded | High |
|
|
454
|
+
| CAP-007 | `declared: false` for a required capability | Hard boundary — exclude before cost comparison; do **not** queue for calibration | Critical |
|
|
455
|
+
| CAP-008 | `version_identifier` changed since `measured.version_identifier` | Invalidate the measurement immediately → `UNKNOWN` | High |
|
|
456
|
+
|
|
457
|
+
**Selection strategy**: `pareto_weighted` — prefer the model with the highest scores on the required dimensions at the lowest cost, **among candidates that passed hard-boundary exclusion**.
|
|
458
|
+
|
|
459
|
+
### Relationship to the Two Axes
|
|
460
|
+
|
|
461
|
+
- **Model axis** — reasoning ceiling and temperament (which model)
|
|
462
|
+
- **Effort axis** — reasoning depth for this dispatch (how deep)
|
|
463
|
+
- **Capability dimensions** — whether the chosen model can do this *kind* of thing at all, and how well
|
|
464
|
+
|
|
465
|
+
Order of application: exclude on hard boundaries → select tier by the two criteria → confirm the required capabilities are `SUPPORTED` or acceptably `DEGRADED` → choose an effort level.
|
|
234
466
|
|
|
235
467
|
---
|
|
236
468
|
|
|
237
469
|
## Navigation Integration | 與下一步建議整合
|
|
238
470
|
|
|
239
|
-
The [ai-response-navigation](ai-response-navigation.md) standard (Rule R6, optional) allows each next-step suggestion option to carry a Tier annotation. The
|
|
471
|
+
The [ai-response-navigation](ai-response-navigation.md) standard (Rule R6, optional) allows each next-step suggestion option to carry a Tier annotation. The criteria in this standard serve as the judgment basis.
|
|
472
|
+
|
|
473
|
+
> **Changed in 2.1.0**: annotations previously derived from file count. Existing annotations made
|
|
474
|
+
> under the old criteria may no longer be accurate. Tier ids are unchanged, so no data migration is
|
|
475
|
+
> required; re-evaluate annotations when the surrounding text is next edited.
|
|
240
476
|
|
|
241
|
-
**Vendor-neutral principle**: Tier names (Fast / Standard / Capable) are tool-agnostic
|
|
477
|
+
**Vendor-neutral principle**: Tier names (Fast / Standard / Capable) and effort labels (`low` … `max`) are tool-agnostic. They do not correspond to any specific vendor's model identifiers or parameter names. Each platform or tool independently maintains its own tier → model mapping, effort-label → parameter mapping, and hard-boundary register.
|
|
242
478
|
|
|
243
479
|
Example mapping (illustrative only, not normative):
|
|
244
480
|
|
|
245
481
|
| Tier | Typical capability level |
|
|
246
482
|
|------|--------------------------|
|
|
247
|
-
| Fast | Lightweight / instruction-following |
|
|
483
|
+
| Fast | Lightweight / literal instruction-following |
|
|
248
484
|
| Standard | Balanced reasoning + code generation |
|
|
249
|
-
| Capable |
|
|
485
|
+
| Capable | High reasoning ceiling, ambiguity navigation |
|
|
486
|
+
|
|
487
|
+
---
|
|
488
|
+
|
|
489
|
+
## Related Standards
|
|
490
|
+
|
|
491
|
+
- [agent-dispatch](agent-dispatch.md) — **how to dispatch**: parallel safety, independent-domain criteria, status protocol, prompt design. This standard covers **who to dispatch to and how deeply**; the two are complementary and **must not duplicate each other**. Parallel-safety rules belong there, not here.
|
|
492
|
+
- [verification-evidence](verification-evidence.md) — why "no error was raised" is not evidence of success; the basis for the refusal-marker requirement in [R3b](#r3b-reverse-risks)
|
|
493
|
+
- [systematic-debugging](systematic-debugging.md) — diagnosing before changing, applied here to the depth-vs-ceiling decision
|
|
494
|
+
|
|
495
|
+
---
|
|
496
|
+
|
|
497
|
+
## References
|
|
498
|
+
|
|
499
|
+
- **Superpowers**: [subagent-driven-development](https://github.com/obra/superpowers) (MIT)
|
|
500
|
+
- **Cost-Effective AI**: Principle of using the minimum capability needed
|
|
250
501
|
|
|
251
502
|
---
|
|
252
503
|
|
|
253
504
|
## Version History
|
|
254
505
|
|
|
506
|
+
> **Note on ordering**: rows are in version order, not date order. Version `2.0.0` (2026-04-13)
|
|
507
|
+
> predates `1.0.1` (2026-06-10) because the capability-management section carried its own version
|
|
508
|
+
> sequence at the time. Version 2.1.0 removes section-level versions; see
|
|
509
|
+
> [Version Field Semantics](#version-field-semantics).
|
|
510
|
+
|
|
255
511
|
| Version | Date | Changes |
|
|
256
512
|
|---------|------|---------|
|
|
513
|
+
| 2.1.0 | 2026-08-11 | **Two-axis restructure (XSPEC-362)**. Model-axis criteria changed from file count to reasoning-ceiling requirement × specification definiteness (R1). Added the orthogonal effort axis with vendor-neutral levels and the depth-vs-ceiling failure diagnosis (R2). Added reverse-exclusion rules: hard boundaries and safety-classifier refusal as silent failure (R3). `capability_dimensions` sub-dimensions split into `declared` / `measured` (R7b); `routing_rules` extended to four states separating `UNKNOWN` from `UNSUPPORTED` (R7a); three independent re-measurement triggers (R7c). `capability_registry` examples replaced with placeholders and a 90-day staleness WARN added (R4). Section versions removed; `.ai.yaml` `meta.version` defined as the whole-file version (R6a). New rules MS-005–MS-010, CAP-006–CAP-008. |
|
|
514
|
+
| 2.0.0 | 2026-04-13 | Add LLM Capability Management section (XSPEC-027 Phase 1): `capability_dimensions`, `capability_registry`, `routing_rules`. *Recorded retroactively in 2.1.0 — this change was never entered in the version history when it was made.* |
|
|
257
515
|
| 1.0.1 | 2026-06-10 | Add Navigation Integration section (R6 cross-reference; vendor-neutral principle) |
|
|
258
516
|
| 1.0.0 | 2026-03-20 | Initial release |
|