universal-dev-standards 6.4.0 → 6.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/bundled/ai/standards/agent-dispatch.ai.yaml +162 -0
  2. package/bundled/ai/standards/ai-friendly-architecture.ai.yaml +1 -1
  3. package/bundled/ai/standards/ai-instruction-standards.ai.yaml +190 -15
  4. package/bundled/ai/standards/class-level-fix.ai.yaml +38 -3
  5. package/bundled/ai/standards/commit-message.ai.yaml +2 -0
  6. package/bundled/ai/standards/model-selection.ai.yaml +370 -72
  7. package/bundled/ai/standards/mutation-testing.ai.yaml +105 -2
  8. package/bundled/ai/standards/project-structure.ai.yaml +1 -1
  9. package/bundled/ai/standards/security-standards.ai.yaml +22 -1
  10. package/bundled/ai/standards/spec-driven-development.ai.yaml +59 -2
  11. package/bundled/ai/standards/test-governance.ai.yaml +49 -2
  12. package/bundled/ai/standards/testing.ai.yaml +49 -3
  13. package/bundled/ai/standards/translation-lifecycle-standards.ai.yaml +4 -4
  14. package/bundled/ai/standards/verification-evidence.ai.yaml +48 -4
  15. package/bundled/core/class-level-fix.md +26 -3
  16. package/bundled/core/model-selection.md +383 -125
  17. package/bundled/core/mutation-testing.md +41 -2
  18. package/bundled/core/test-governance.md +22 -2
  19. package/bundled/core/translation-lifecycle-standards.md +6 -6
  20. package/bundled/core/verification-evidence.md +42 -3
  21. package/bundled/locales/zh-CN/CHANGELOG.md +12 -3
  22. package/bundled/locales/zh-CN/README.md +1 -1
  23. package/bundled/locales/zh-CN/SECURITY.md +1 -1
  24. package/bundled/locales/zh-CN/core/model-selection.md +375 -60
  25. package/bundled/locales/zh-CN/core/mutation-testing.md +1 -1
  26. package/bundled/locales/zh-CN/core/test-governance.md +1 -1
  27. package/bundled/locales/zh-CN/core/translation-lifecycle-standards.md +1 -1
  28. package/bundled/locales/zh-CN/core/verification-evidence.md +1 -1
  29. package/bundled/locales/zh-CN/docs/CHEATSHEET.md +7 -12
  30. package/bundled/locales/zh-CN/docs/FEATURE-REFERENCE.md +10 -15
  31. package/bundled/locales/zh-TW/CHANGELOG.md +37 -3
  32. package/bundled/locales/zh-TW/README.md +1 -1
  33. package/bundled/locales/zh-TW/SECURITY.md +1 -1
  34. package/bundled/locales/zh-TW/core/class-level-fix.md +22 -7
  35. package/bundled/locales/zh-TW/core/model-selection.md +385 -47
  36. package/bundled/locales/zh-TW/core/mutation-testing.md +45 -6
  37. package/bundled/locales/zh-TW/core/test-governance.md +22 -3
  38. package/bundled/locales/zh-TW/core/translation-lifecycle-standards.md +1 -1
  39. package/bundled/locales/zh-TW/core/verification-evidence.md +33 -6
  40. package/bundled/locales/zh-TW/docs/CHEATSHEET.md +7 -12
  41. package/bundled/locales/zh-TW/docs/FEATURE-REFERENCE.md +10 -15
  42. package/bundled/locales/zh-TW/integrations/claude-code/README.md +31 -5
  43. package/package.json +1 -1
  44. package/src/utils/reference-sync.js +83 -16
  45. package/standards-registry.json +20 -8
@@ -5,74 +5,259 @@
5
5
  standard:
6
6
  id: model-selection
7
7
  name: AI Model Selection Strategy
8
- description: AI 模型分級選擇策略
8
+ description: AI 模型選擇策略(模型軸 × effort 軸)
9
9
 
10
10
  meta:
11
- version: "2.0.1"
12
- updated: "2026-06-10"
11
+ version: "2.1.0"
12
+ updated: "2026-08-11"
13
13
  source: core/model-selection.md
14
- description: "三層模型分級選擇策略 + 多模型池能力管理(XSPEC-027)"
14
+ description: "模型軸(天花板 × 規格明確度)× effort 軸(思考深度)+ 反向排除 + 多模型池能力管理(XSPEC-027, XSPEC-362)"
15
15
  inspired_by: superpowers/subagent-driven-development
16
+ version_semantics: >
17
+ This version field corresponds to the WHOLE FILE version in core/model-selection.md's header,
18
+ never to a section version. Sections do not carry their own version markers (XSPEC-362 R6a).
16
19
 
17
20
  guidelines:
18
- - "使用最便宜能勝任的模型(成本效益原則)"
19
- - "根據任務複雜度信號選擇模型等級"
20
- - "BLOCKED 狀態時考慮升級模型再試"
21
+ - "模型軸與 effort 軸正交:模型軸選天花板與性格,effort 軸選這次要它想多深"
22
+ - "使用能勝任的最便宜組合(成本效益原則)"
23
+ - "模型軸判準是「推理天花板需求 × 規格明確度」,不是修改檔案數"
24
+ - "任務失敗時先判斷是深度不足(調 effort)還是天花板不足(換模型),兩者不可互換"
25
+ - "硬邊界(context 上限、effort 參數支援、modality)在成本比較之前排除,不是排名較低"
26
+ - "拒絕不是錯誤:呼叫端必須檢查拒絕標記,而非只檢查有無例外"
27
+
28
+ version_field_semantics:
29
+ rule: "core/model-selection.md 只有一個版本欄(檔頭),涵蓋全檔;章節不得自帶版本標記"
30
+ ai_yaml_mapping: "standard.meta.version 對應全檔版本"
31
+ section_history_location: "章節變更歷程記於 core/model-selection.md 的 Version History 表"
32
+
33
+ axes:
34
+ model:
35
+ purpose: "選天花板與性格(which model)"
36
+ criteria:
37
+ - id: reasoning_ceiling
38
+ name: "推理天花板需求"
39
+ question: "這個任務是否存在「想更久也想不出來」的成分?"
40
+ if_no: "較低層即可,用 effort 買深度而非用錢買天花板"
41
+ if_yes: "需要更高天花板;在較低層加多少 effort 都到不了"
42
+ - id: specification_definiteness
43
+ name: "規格明確度"
44
+ bidirectional: true
45
+ note: "非單調遞增判準——兩個方向都會出錯"
46
+ definite_spec:
47
+ prefer: "字面遵循型(較低層)"
48
+ inverted_failure: "把已寫死的步驟餵給高天花板型反而降低品質——它會重新詮釋已經決定過的事"
49
+ ambiguous_spec:
50
+ prefer: "模糊性導航型(較高層)"
51
+ inverted_failure: "把模糊任務給字面遵循型會得到「精確執行了錯誤的那句話」"
52
+ removed_signal:
53
+ signal: "修改檔案數"
54
+ removed_in: "2.1.0"
55
+ reason: >
56
+ 3 檔的深度重構可能遠難於 8 檔的機械改名。檔案數會系統性把「深而窄」的任務誤判到低層,
57
+ 且誤判方向固定——是偏差不是噪音。它之所以存活是因為便宜好量,不是因為它預測得準。
58
+ effort:
59
+ purpose: "選這次要它想多深(how deep)"
60
+ nature: "單次派工的請求參數,不是模型屬性"
61
+ vendor_neutral: true
62
+ levels:
63
+ - id: low
64
+ label_zh: "低"
65
+ semantics: "最少斟酌,主要依模式作答"
66
+ typical_use: "機械性編輯、查詢、格式化"
67
+ - id: medium
68
+ label_zh: "中"
69
+ semantics: "一般斟酌;預設值"
70
+ typical_use: "多數已定義的實作工作"
71
+ - id: high
72
+ label_zh: "高"
73
+ semantics: "延伸斟酌,作答前考慮替代方案"
74
+ typical_use: "含局部決策的整合工作"
75
+ - id: very-high
76
+ label_zh: "極高"
77
+ semantics: "探索並淘汰候選方法"
78
+ typical_use: "設計、審查、非顯而易見的除錯"
79
+ - id: max
80
+ label_zh: "極限"
81
+ semantics: "該模型可用的最大深度"
82
+ typical_use: "升級模型軸前的最後一次嘗試"
83
+ support_is_a_hard_boundary: >
84
+ 並非每一層都支援每一個 effort 級距。是否支援 effort 參數是硬邊界(能力有無),
85
+ 不是品質梯度。平台的 tier→模型映射必須登記每個模型接受哪些級距。
86
+
87
+ failure_diagnosis:
88
+ rule: "任務失敗時,先判斷哪一軸不足。兩者的補救方式不可互換。"
89
+ ordering: "先調 effort,再升級模型層;只有在當前層的最高可用 effort 也失敗後才升級模型軸"
90
+ exception: >
91
+ 若判準 1(推理天花板需求)事前已識別出「想更久也想不出來」的成分,直接從較高層開始。
92
+ 排序規則是關於「診斷失敗」,不是關於忽略事前判準。
93
+ cases:
94
+ - id: depth_insufficient
95
+ observation: "產出淺、略過考量、提早收手,但它做過的推理本身是對的"
96
+ diagnosis: "深度不足"
97
+ remedy: "在同一個模型上提高 effort"
98
+ wrong_remedy_cost: "升級模型層=為一個從來不是瓶頸的天花板多付錢"
99
+ - id: ceiling_insufficient
100
+ observation: "產出在種類上就錯了——誤解問題、提出不可行的方法——且在 max effort 下仍如此"
101
+ diagnosis: "天花板不足"
102
+ remedy: "升級模型層"
103
+ wrong_remedy_cost: "再提高 effort 買不到任何東西;級距已經用盡"
104
+
105
+ reverse_exclusion:
106
+ description: "什麼工作不該給哪一層(XSPEC-362 R3)。此問題無法由「該升到哪層」的判準推導出來。"
107
+ hard_boundaries:
108
+ id: R3a
109
+ principle: "硬邊界是能力有無,不是程度差。裝不下的模型不是「做得比較差」,是做不了。"
110
+ exclusion_timing: "在成本比較「之前」排除,不是在其中排名較低"
111
+ exclusion_timing_reason: "事後排除會讓比較便宜但做不到的模型在價格上勝出"
112
+ boundaries:
113
+ - id: context_capacity
114
+ question: "輸入裝得下嗎?"
115
+ violation: "截斷或報錯——模型從未看過任務的一部分"
116
+ - id: effort_parameter_support
117
+ question: "這個模型接受 effort 參數嗎?"
118
+ violation: "請求被拒,或靜默以模型固定深度執行"
119
+ - id: modality_support
120
+ question: "它接受這種輸入型別嗎(影像、語音)?"
121
+ violation: "輸入被丟棄或拒絕"
122
+ registration: "硬邊界對應 capability registry 的 declared(布林,連線時可得,成本趨近零)"
123
+ score_1_is_not_a_hard_boundary: "分數 1 不是硬邊界,是量到的低品質;兩者的動作不同"
124
+ reverse_risks:
125
+ id: R3b
126
+ risks:
127
+ - id: safety_classifier_refusal
128
+ description: >
129
+ 高能力層可能帶更嚴格的安全分類器。合法但形似禁止類別的工作(資安加固、
130
+ 緩解措施實作、憑證處理程式碼、red-team 工具)可能被拒絕。
131
+ silent_failure_mechanism: >
132
+ 拒絕不是錯誤。呼叫回傳的是正常回應加上拒絕標記。
133
+ 對只看 exit code、或只看有無擲出例外的呼叫端而言,這是靜默失效:
134
+ 管線記錄成功,而工作根本沒做。
135
+ requirements:
136
+ - "呼叫端必須檢查回應中的拒絕標記;沒有錯誤不等於完成"
137
+ - "規格敏感類工作在批次派往高能力層前,須先以廉價探測請求驗證可行性"
138
+ - "遇到拒絕時改派其他層或其他 provider;不要以更高 effort 重試同一請求——effort 移動不了分類器的判定"
139
+ related_standard: verification-evidence
140
+ - id: over_specified_prompt_degrades_high_tier
141
+ description: >
142
+ 寫出每一個步驟對字面遵循型是正確做法,對高天花板型是錯誤做法——
143
+ 後者會重新詮釋已經決定過的事,產出反而變差。
144
+ rule: "prompt 顆粒度隨層級走。若唯一可用的 prompt 是完整步驟清單,判準 2 本來就要求送到較低層"
21
145
 
22
146
  tiers:
23
147
  - id: fast
24
- description: "機械性實作 — 單一檔案、明確 spec、無需判斷"
148
+ description: "無推理天花板需求、規格完全明確"
25
149
  signals:
26
- - "修改單一檔案"
27
- - "spec 完全明確,無歧義"
150
+ - "不存在「想更久也想不出來」的成分"
151
+ - "步驟已寫出,無需尋路"
28
152
  - "不需要設計判斷"
29
- - "模式化的重複工作"
153
+ - "字面遵循正是想要的行為"
30
154
  examples:
31
- - "更新 package.json 版本號"
32
- - "新增一個 export"
155
+ - "在一組檔案上套用已定義的編輯(機械改名、版本號提升)"
33
156
  - "修正 typo"
157
+ - "新增規格中已指名的 export"
34
158
  - id: standard
35
- description: "整合性實作 — 多檔案、需要判斷力"
159
+ description: "需要一定推理天花板,規格大致明確但留有局部決策空間"
36
160
  signals:
37
- - "修改 2-5 個檔案"
161
+ - "部分成分需要推理,但沒有一項是「想更久也想不出來」"
162
+ - "目標與多數步驟已定義,剩下有限的局部選擇"
38
163
  - "需要理解模組間關係"
39
- - "需要一定設計判斷"
40
164
  examples:
41
- - "新增一個 API endpoint(含 route + handler + test)"
42
- - "重構一個模組的內部結構"
165
+ - "實作整合點已知的既定功能"
166
+ - "依既定目標結構重構模組"
167
+ - "為指定子系統撰寫整合測試"
43
168
  - id: capable
44
- description: "架構性工作 — 設計、審查、除錯"
169
+ description: "需要高推理天花板,或規格僅有目標與約束、路徑未知"
45
170
  signals:
46
- - "修改 5+ 個檔案"
47
- - "需要架構決策"
48
- - "跨模組協調"
49
- - "除錯複雜問題"
171
+ - "存在單靠更多思考時間解決不了的成分"
172
+ - "路徑未知;答案要找出來,不是套用"
173
+ - "需要導航模糊性而非遵循文字"
50
174
  examples:
51
- - "設計新的子系統架構"
52
- - "審查大型 PR"
53
- - "診斷跨服務的效能問題"
175
+ - "設計子系統架構(路徑未知)"
176
+ - "審查大型 PR(重點事前未載明)"
177
+ - "診斷尚未形成假設的失效"
178
+ note: >
179
+ 目標結構「已定」的大型重構不自動屬於 capable——依判準 2 它是明確規格,
180
+ 高天花板模型的產出可能比字面遵循型更差。使它升到 capable 的是「目標結構未定」。
54
181
 
55
182
  rules:
56
183
  - id: MS-001
57
- trigger: "BLOCKED 狀態且當前為 fast"
58
- action: "升級至 standard 再試"
184
+ trigger: "BLOCKED 且當前為 fast,effort 已用盡"
185
+ action: "升級至 standard"
59
186
  priority: high
60
187
  - id: MS-002
61
- trigger: "BLOCKED 狀態且當前為 standard"
62
- action: "升級至 capable 再試"
188
+ trigger: "BLOCKED 且當前為 standard,effort 已用盡"
189
+ action: "升級至 capable"
63
190
  priority: high
64
191
  - id: MS-003
65
- trigger: "BLOCKED 狀態且當前為 capable"
192
+ trigger: "BLOCKED 且當前為 capable,effort 已用盡"
66
193
  action: "標記為需要人工介入"
67
194
  priority: critical
68
195
  - id: MS-004
69
- trigger: "任務需求不明確"
196
+ trigger: "任務規格僅有目標與約束"
70
197
  action: "從 standard 或更高層級開始"
71
198
  priority: medium
199
+ - id: MS-005
200
+ trigger: "產出淺但推理健全,且 effort 尚未達 max"
201
+ action: "在同一模型上提高 effort — 不要升級模型層"
202
+ priority: high
203
+ - id: MS-006
204
+ trigger: "產出在種類上就錯了,且 effort 已達 max"
205
+ action: "升級模型層 — 再提高 effort 買不到任何東西"
206
+ priority: high
207
+ - id: MS-007
208
+ trigger: "候選模型不滿足某項必要硬邊界(context、effort 支援、modality)"
209
+ action: "在成本比較「之前」排除"
210
+ priority: critical
211
+ - id: MS-008
212
+ trigger: "將規格敏感類工作派往高能力層"
213
+ action: "先探測可行性;每一則回應都要檢查拒絕標記"
214
+ priority: critical
215
+ - id: MS-009
216
+ trigger: "prompt 是完整寫出的步驟清單"
217
+ action: "偏好字面遵循型;不要升到高天花板層"
218
+ priority: medium
219
+ - id: MS-010
220
+ trigger: "pin_date 或 measured.at 距今超過 90 天"
221
+ action: "發出 WARN 並排入重測佇列(不阻擋發版)"
222
+ priority: medium
72
223
 
73
224
  capability_dimensions:
74
- description: "能力維度分類與評分說明(XSPEC-027, DEC-031, DEC-032)"
225
+ description: "能力維度分類與評分說明(XSPEC-027, DEC-031, DEC-032, XSPEC-362 R7b)"
226
+ two_fields_per_subdimension:
227
+ rationale: >
228
+ 單一 1–5 分無法同時表達兩件事:「支不支援」是二元、廠商宣告、連線時可低成本查得;
229
+ 「有多好」是連續、需實測。壓成同一個數字,會讓便宜的事實與昂貴的事實無從分辨,
230
+ 而昂貴的那個會靜默勝出。
231
+ declared:
232
+ type: boolean
233
+ source: "provider 的模型描述端點或設定宣告"
234
+ cost: "趨近零,連線建立時可得"
235
+ meaning: "做不做得到這件事"
236
+ measured:
237
+ type: object
238
+ fields:
239
+ score: "1–5 整數"
240
+ at: "量測日期 YYYY-MM-DD"
241
+ version_identifier: "量測時所用的版本識別"
242
+ source: "benchmark"
243
+ cost: "高,離線"
244
+ meaning: "做得多好"
245
+ normative_rules:
246
+ - id: declared_false_is_hard_boundary
247
+ rule: "declared: false 是硬邊界——無論任何分數皆排除,且「不」排入校準佇列(校準無意義)"
248
+ - id: declared_true_measured_absent_is_unknown
249
+ rule: "declared: true 且 measured 缺失 = UNKNOWN(做得到,但不知道做得多好);這是資訊缺口不是結論"
250
+ - id: supported_must_not_derive_from_score
251
+ rule: >
252
+ 禁止由分數推導支援與否。任何以 supported = score > 0(或類似)計算的實作,
253
+ 都已把兩個欄位合併回單軸,重新製造了本規則要防止的缺陷。
254
+ - id: failed_measurement_must_not_be_scored
255
+ rule: >
256
+ 量測失敗不得記為任何分數,必須記為缺失(UNKNOWN)。
257
+ 以預設值填補(包括 0,而 0 不在 1–5 量表內)會讓「量不到」與「量到了且很差」無從分辨。
75
258
  scoring_scale:
259
+ applies_to: "measured.score only"
260
+ zero_does_not_exist: "沒有 0 分。量測缺失以「measured 物件不存在」表達,絕不以分數表達。"
76
261
  5: "生產就緒 — 高準確率,可直接使用"
77
262
  4: "良好 — 偶有遺漏,可接受"
78
263
  3: "基本可用 — 需人工補充"
@@ -122,77 +307,190 @@ standard:
122
307
 
123
308
  capability_registry:
124
309
  description: "模型能力登記表(DEC-031 D1)。各子專案依實際測試維護分數。"
310
+ no_concrete_model_ids:
311
+ rule: "本標準的 examples 不得登記任何具體廠商模型 ID"
312
+ reason: >
313
+ 寫進標準的模型 ID 是「有到期日卻沒有時鐘」的引用端:它會過期,
314
+ 而過期的樣子與正常的樣子在檔案上無從分辨。具體登記由採用者在自己的 registry 維護。
315
+ enforced_since: "2.1.0(XSPEC-362 R4)"
125
316
  format:
126
- model_id: "provider/model-name"
127
- version_pinned: "版本鎖定識別(SHA、日期戳或 model_version)"
128
- pin_date: "鎖定日期(YYYY-MM-DD)"
129
- eol_date: "預期淘汰日期(可選,YYYY-MM-DD)"
317
+ model_id: "<provider>/<model-name>"
318
+ version_pinned: "<version-identifier>(SHA、日期戳或 model_version)"
319
+ pin_date: "<YYYY-MM-DD>"
320
+ eol_date: "<YYYY-MM-DD>(可選)"
130
321
  capabilities:
131
- "<dimension>.<subdimension>": "<1-5 整數分>"
322
+ "<dimension>.<subdimension>":
323
+ declared: "<boolean>"
324
+ measured:
325
+ score: "<1-5 整數>"
326
+ at: "<YYYY-MM-DD>"
327
+ version_identifier: "<量測時所用的版本識別>"
132
328
  examples:
133
- - model_id: "anthropic/claude-sonnet-4-6"
134
- version_pinned: "claude-sonnet-4-6"
135
- pin_date: "2026-04-13"
136
- capabilities:
137
- "modality.vision": 4
138
- "reasoning.instruction_following": 5
139
- "reasoning.code_reasoning": 5
140
- "output.structured_output": 5
141
- "output.tool_use": 5
142
- "language.multilingual_zh_tw": 5
143
- - model_id: "anthropic/claude-haiku-4-5"
144
- version_pinned: "claude-haiku-4-5-20251001"
145
- pin_date: "2026-04-13"
329
+ - model_id: "<provider>/<model-name-a>"
330
+ version_pinned: "<version-identifier-a>"
331
+ pin_date: "<YYYY-MM-DD>"
146
332
  capabilities:
147
- "modality.vision": 3
148
- "reasoning.instruction_following": 4
149
- "reasoning.code_reasoning": 4
150
- "output.structured_output": 4
151
- "output.tool_use": 4
152
- "language.multilingual_zh_tw": 4
333
+ "modality.vision":
334
+ declared: true
335
+ measured:
336
+ score: 4
337
+ at: "<YYYY-MM-DD>"
338
+ version_identifier: "<version-identifier-a>"
339
+ "modality.audio":
340
+ declared: false
341
+ "output.tool_use":
342
+ declared: true
343
+ example_notes:
344
+ - "modality.audio 的 declared: false 是硬邊界——無 measured 欄位,也不需要"
345
+ - "output.tool_use 的 measured 缺失 → UNKNOWN → 排入校準佇列(不是 UNSUPPORTED)"
346
+ staleness_check:
347
+ threshold_days: 90
348
+ applies_to: ["pin_date", "measured.at"]
349
+ severity: WARN
350
+ must_not_block: true
351
+ rationale: "依 XSPEC-361 R8 量測結論,純檔內不變量的誤報率偏高,不適合作為 BLOCK 閘門"
352
+ output_requirement: "WARN 須指名受影響的 model_id 與子維度"
153
353
 
154
354
  routing_rules:
155
- description: "能力路由決策規則(DEC-031 D2, DEC-032 D3)"
355
+ description: "能力路由決策規則(DEC-031 D2, DEC-032 D3, XSPEC-362 R7a)"
156
356
  selection_strategy: "pareto_weighted"
357
+ selection_strategy_scope: "僅在通過硬邊界排除的候選之中比較"
358
+ states:
359
+ - id: SUPPORTED
360
+ condition: "所需能力的 measured.score 皆 ≥ min_score"
361
+ action: "正常執行"
362
+ calibration_queue: false
363
+ - id: DEGRADED
364
+ condition: "measured.score 存在、≥ 2、但低於 min_score"
365
+ action: "降級執行,產出標記 [DEGRADED]"
366
+ calibration_queue: false
367
+ - id: UNSUPPORTED
368
+ condition: "measured 存在且 score ≤ 1"
369
+ action: "排除,不再嘗試"
370
+ calibration_queue: false
371
+ calibration_queue_reason: "已有結論"
372
+ - id: UNKNOWN
373
+ condition: "declared: true 但 measured 缺失——未登記、量測失敗或資料逾期"
374
+ action: "排入校準佇列;在校準完成前不得靜默排除"
375
+ calibration_queue: true
376
+ hard_boundary_precedes_states: "declared: false 在本表之前處理:硬邊界排除,且不排入校準佇列"
377
+ observability_requirement:
378
+ rule: "UNKNOWN 與 UNSUPPORTED 必須在「回傳結構」上可分辨,不能只在 log 上分辨"
379
+ reason: >
380
+ 呼叫端必須能分辨「這個模型不行」與「我還不知道這個模型行不行」——兩者導向不同的下一步動作。
381
+ 回傳單一布林、或回傳三態的路由 API,表達不了這件事。
382
+ anti_pattern:
383
+ description: "把「查無」與「低分」放進同一個判斷式,並以預設值填補缺資料"
384
+ example: "if (!cap || !cap.supported || cap.score < 2) — 三件事一個判斷式;以 ?? 0 填補缺資料"
385
+ consequence: "任何新偵測到的模型必定被排除,新模型永遠進不了候選池——與「支援更多模型」的目標直接相反"
386
+ decision_tree: |
387
+ 任務需要 capability X
388
+ ├── declared == false → HARD BOUNDARY — 排除,不校準
389
+ ├── measured 缺失/失敗/逾期 → UNKNOWN — 排入校準佇列
390
+ ├── measured.score ≤ 1 → UNSUPPORTED — 排除,不校準
391
+ ├── measured.score ≥ 2 且 < min_score → DEGRADED — 執行並標記 [DEGRADED]
392
+ └── measured.score ≥ min_score → SUPPORTED — 執行
157
393
  rules:
158
394
  - id: CAP-001
159
- condition: "任務需要 capability X,模型得分 ≥ min_score"
395
+ condition: "任務需要 capability X,measured.score ≥ min_score"
160
396
  action: "SUPPORTED — 正常執行"
161
397
  priority: high
162
398
  - id: CAP-002
163
- condition: "任務需要 capability X,模型得分 2-3(min_score 未達但 ≥ 2)"
399
+ condition: "measured.score ≥ 2 但低於 min_score"
164
400
  action: "DEGRADED — 降級流程執行,產出標記 [DEGRADED]"
165
401
  priority: medium
166
402
  - id: CAP-003
167
- condition: "任務需要 capability X,模型得分 ≤ 1 或未登記"
168
- action: "UNSUPPORTED — 使用替代流程或提示用戶"
403
+ condition: "measured 存在且 score ≤ 1"
404
+ action: "UNSUPPORTED — 使用替代流程或提示用戶;不排入校準佇列"
169
405
  priority: high
170
406
  - id: CAP-004
171
407
  condition: "模型降智偵測(DEC-033)觸發 moderate 信號"
172
- action: "啟動金絲雀測試,同時記錄降智警告"
408
+ action: "啟動金絲雀測試,記錄降智警告,並將受影響能力排入重測"
173
409
  priority: high
174
410
  - id: CAP-005
175
411
  condition: "模型降智偵測觸發 critical 信號"
176
- action: "自動切換備用模型,上報 P1 Issue"
412
+ action: "自動切換備用模型,上報 P1 Issue,並將受影響能力排入重測"
413
+ priority: critical
414
+ - id: CAP-006
415
+ condition: "declared: true 且 measured 缺失、失敗或逾期"
416
+ action: "UNKNOWN — 排入校準佇列;不得靜默排除"
417
+ priority: high
418
+ - id: CAP-007
419
+ condition: "所需 capability 的 declared: false"
420
+ action: "硬邊界 — 在成本比較之前排除;不排入校準佇列"
177
421
  priority: critical
422
+ - id: CAP-008
423
+ condition: "version_identifier 與 measured.version_identifier 不同"
424
+ action: "立即失效該量測 → UNKNOWN"
425
+ priority: high
426
+
427
+ remeasurement_triggers:
428
+ description: "重測觸發(XSPEC-362 R7c)"
429
+ principle: "版本變更是充分條件,不是必要條件"
430
+ principle_reason: >
431
+ DEC-033 降智偵測的存在前提,正是模型 ID 與版本字串不變而行為改變。
432
+ 僅依賴版本變更的實作,會漏掉 DEC-033 整個要防的情境。
433
+ independence: "三條路徑彼此獨立:各自單獨觸發,且沒有任何一條是另一條的前提"
434
+ triggers:
435
+ - id: version_identifier_changed
436
+ trigger: "version_identifier 與 measured 中記錄的不同"
437
+ effect: "既有量測立即失效 → 狀態變為 UNKNOWN"
438
+ - id: measurement_expired
439
+ trigger: "measured.at 距今超過 90 天"
440
+ effect: "排入重測;發出 WARN(見 MS-010)"
441
+ - id: degradation_alert
442
+ trigger: "降智偵測告警(DEC-033 CAP-004/CAP-005)"
443
+ effect: "排入重測——即使版本字串未變"
444
+
445
+ relationship_to_axes:
446
+ model_axis: "推理天花板與性格(選哪個模型)"
447
+ effort_axis: "這次派工的思考深度(想多深)"
448
+ capability_dimensions: "選中的模型做不做得到這「種」事,以及做得多好"
449
+ order_of_application: >
450
+ 硬邊界排除 → 依兩個判準選 tier → 確認所需能力為 SUPPORTED 或可接受的 DEGRADED → 選 effort 級距
178
451
 
179
452
  navigation_integration:
180
453
  description: >
181
454
  The ai-response-navigation standard (Rule R6, optional) allows next-step suggestion
182
455
  options to carry a Tier annotation (〔model: Fast〕 / 〔model: Standard〕 / 〔model: Capable〕).
183
- The complexity signals in this standard serve as the judgment basis for those annotations.
456
+ The criteria in this standard serve as the judgment basis for those annotations.
457
+ changed_in_2_1_0: >
458
+ Annotations previously derived from file count. Existing annotations made under the old
459
+ criteria may no longer be accurate. Tier ids are unchanged, so no data migration is required.
184
460
  vendor_neutral_principle: >
185
- Tier names (Fast / Standard / Capable) are vendor-agnostic labels.
186
- Each platform independently maps these tiers to its available models.
461
+ Tier names (Fast / Standard / Capable) and effort labels (low … max) are vendor-agnostic.
462
+ They do not correspond to any specific vendor's model identifiers or parameter names.
463
+ Each platform independently maps these tiers to its available models, maps effort labels to
464
+ its own parameter, and maintains its own hard-boundary register.
465
+
466
+ related_standards:
467
+ - id: agent-dispatch
468
+ relationship: >
469
+ agent-dispatch 管「怎麼派」(並行安全、獨立域判準、狀態協定、prompt 設計);
470
+ model-selection 管「派給誰、想多深」(模型軸 × effort 軸)。
471
+ 兩者互補,不得互相重寫——並行安全規則屬於 agent-dispatch,不屬於這裡。
472
+ - id: verification-evidence
473
+ relationship: "「沒有擲出錯誤」不是完成的證據;R3b 拒絕標記要求的依據"
474
+ - id: systematic-debugging
475
+ relationship: "先診斷再修改,在此適用於「深度不足 vs 天花板不足」的判斷"
187
476
 
188
477
  physical_spec:
189
478
  type: checklist
190
479
  validator:
191
480
  type: ai_review
192
- rule: "檢查模型選擇是否基於任務複雜度信號,而非固定使用最高等級"
481
+ rule: "檢查模型選擇是否依「推理天花板需求 × 規格明確度」判準,且 effort 與模型軸分開決定"
193
482
  checks:
194
- - "是否根據複雜度信號選擇模型等級"
195
- - "BLOCKED 時是否嘗試升級而非直接失敗"
196
- - "簡單任務是否避免使用過高等級模型"
483
+ - "模型層級是否依推理天花板需求與規格明確度選擇,而非依修改檔案數"
484
+ - "effort 級距是否與模型層級分開決定"
485
+ - "失敗時是否先判斷「深度不足」或「天花板不足」,再選擇對應補救"
486
+ - "升級模型層之前,當前層的最高可用 effort 是否已嘗試過"
487
+ - "硬邊界(context、effort 支援、modality)是否在成本比較之前排除"
488
+ - "規格敏感類工作的呼叫端是否檢查拒絕標記,而非只檢查有無例外"
197
489
  - "能力登記表是否包含 version_pinned 與 pin_date"
198
- - "路由決策是否依 SUPPORTED/DEGRADED/UNSUPPORTED 三態處理"
490
+ - "能力登記表的 examples 是否不含具體廠商模型 ID"
491
+ - "每個子維度是否同時具備 declared(布林)與 measured(分數+時間+版本識別)兩個獨立欄位"
492
+ - "是否不存在任何以分數推導支援與否的實作(如 supported = score > 0)"
493
+ - "量測失敗是否記為缺失,而非以預設值(如 0)填補"
494
+ - "路由決策是否依 SUPPORTED/DEGRADED/UNSUPPORTED/UNKNOWN 四態處理"
495
+ - "UNKNOWN 與 UNSUPPORTED 是否在回傳結構上可分辨"
496
+ - "重測觸發是否包含版本變更、時間逾期、降智告警三條獨立路徑"