universal-dev-standards 6.4.0 → 6.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled/ai/standards/agent-dispatch.ai.yaml +162 -0
- package/bundled/ai/standards/ai-friendly-architecture.ai.yaml +1 -1
- package/bundled/ai/standards/ai-instruction-standards.ai.yaml +190 -15
- package/bundled/ai/standards/class-level-fix.ai.yaml +38 -3
- package/bundled/ai/standards/commit-message.ai.yaml +2 -0
- package/bundled/ai/standards/model-selection.ai.yaml +370 -72
- package/bundled/ai/standards/mutation-testing.ai.yaml +105 -2
- package/bundled/ai/standards/project-structure.ai.yaml +1 -1
- package/bundled/ai/standards/security-standards.ai.yaml +22 -1
- package/bundled/ai/standards/spec-driven-development.ai.yaml +59 -2
- package/bundled/ai/standards/test-governance.ai.yaml +49 -2
- package/bundled/ai/standards/testing.ai.yaml +49 -3
- package/bundled/ai/standards/translation-lifecycle-standards.ai.yaml +4 -4
- package/bundled/ai/standards/verification-evidence.ai.yaml +48 -4
- package/bundled/core/class-level-fix.md +26 -3
- package/bundled/core/model-selection.md +383 -125
- package/bundled/core/mutation-testing.md +41 -2
- package/bundled/core/test-governance.md +22 -2
- package/bundled/core/translation-lifecycle-standards.md +6 -6
- package/bundled/core/verification-evidence.md +42 -3
- package/bundled/locales/zh-CN/CHANGELOG.md +12 -3
- package/bundled/locales/zh-CN/README.md +1 -1
- package/bundled/locales/zh-CN/SECURITY.md +1 -1
- package/bundled/locales/zh-CN/core/model-selection.md +375 -60
- package/bundled/locales/zh-CN/core/mutation-testing.md +1 -1
- package/bundled/locales/zh-CN/core/test-governance.md +1 -1
- package/bundled/locales/zh-CN/core/translation-lifecycle-standards.md +1 -1
- package/bundled/locales/zh-CN/core/verification-evidence.md +1 -1
- package/bundled/locales/zh-CN/docs/CHEATSHEET.md +7 -12
- package/bundled/locales/zh-CN/docs/FEATURE-REFERENCE.md +10 -15
- package/bundled/locales/zh-TW/CHANGELOG.md +37 -3
- package/bundled/locales/zh-TW/README.md +1 -1
- package/bundled/locales/zh-TW/SECURITY.md +1 -1
- package/bundled/locales/zh-TW/core/class-level-fix.md +22 -7
- package/bundled/locales/zh-TW/core/model-selection.md +385 -47
- package/bundled/locales/zh-TW/core/mutation-testing.md +45 -6
- package/bundled/locales/zh-TW/core/test-governance.md +22 -3
- package/bundled/locales/zh-TW/core/translation-lifecycle-standards.md +1 -1
- package/bundled/locales/zh-TW/core/verification-evidence.md +33 -6
- package/bundled/locales/zh-TW/docs/CHEATSHEET.md +7 -12
- package/bundled/locales/zh-TW/docs/FEATURE-REFERENCE.md +10 -15
- package/bundled/locales/zh-TW/integrations/claude-code/README.md +31 -5
- package/package.json +1 -1
- package/src/utils/reference-sync.js +83 -16
- package/standards-registry.json +20 -8
|
@@ -5,74 +5,259 @@
|
|
|
5
5
|
standard:
|
|
6
6
|
id: model-selection
|
|
7
7
|
name: AI Model Selection Strategy
|
|
8
|
-
description: AI
|
|
8
|
+
description: AI 模型選擇策略(模型軸 × effort 軸)
|
|
9
9
|
|
|
10
10
|
meta:
|
|
11
|
-
version: "2.0
|
|
12
|
-
updated: "2026-
|
|
11
|
+
version: "2.1.0"
|
|
12
|
+
updated: "2026-08-11"
|
|
13
13
|
source: core/model-selection.md
|
|
14
|
-
description: "
|
|
14
|
+
description: "模型軸(天花板 × 規格明確度)× effort 軸(思考深度)+ 反向排除 + 多模型池能力管理(XSPEC-027, XSPEC-362)"
|
|
15
15
|
inspired_by: superpowers/subagent-driven-development
|
|
16
|
+
version_semantics: >
|
|
17
|
+
This version field corresponds to the WHOLE FILE version in core/model-selection.md's header,
|
|
18
|
+
never to a section version. Sections do not carry their own version markers (XSPEC-362 R6a).
|
|
16
19
|
|
|
17
20
|
guidelines:
|
|
18
|
-
- "
|
|
19
|
-
- "
|
|
20
|
-
- "
|
|
21
|
+
- "模型軸與 effort 軸正交:模型軸選天花板與性格,effort 軸選這次要它想多深"
|
|
22
|
+
- "使用能勝任的最便宜組合(成本效益原則)"
|
|
23
|
+
- "模型軸判準是「推理天花板需求 × 規格明確度」,不是修改檔案數"
|
|
24
|
+
- "任務失敗時先判斷是深度不足(調 effort)還是天花板不足(換模型),兩者不可互換"
|
|
25
|
+
- "硬邊界(context 上限、effort 參數支援、modality)在成本比較之前排除,不是排名較低"
|
|
26
|
+
- "拒絕不是錯誤:呼叫端必須檢查拒絕標記,而非只檢查有無例外"
|
|
27
|
+
|
|
28
|
+
version_field_semantics:
|
|
29
|
+
rule: "core/model-selection.md 只有一個版本欄(檔頭),涵蓋全檔;章節不得自帶版本標記"
|
|
30
|
+
ai_yaml_mapping: "standard.meta.version 對應全檔版本"
|
|
31
|
+
section_history_location: "章節變更歷程記於 core/model-selection.md 的 Version History 表"
|
|
32
|
+
|
|
33
|
+
axes:
|
|
34
|
+
model:
|
|
35
|
+
purpose: "選天花板與性格(which model)"
|
|
36
|
+
criteria:
|
|
37
|
+
- id: reasoning_ceiling
|
|
38
|
+
name: "推理天花板需求"
|
|
39
|
+
question: "這個任務是否存在「想更久也想不出來」的成分?"
|
|
40
|
+
if_no: "較低層即可,用 effort 買深度而非用錢買天花板"
|
|
41
|
+
if_yes: "需要更高天花板;在較低層加多少 effort 都到不了"
|
|
42
|
+
- id: specification_definiteness
|
|
43
|
+
name: "規格明確度"
|
|
44
|
+
bidirectional: true
|
|
45
|
+
note: "非單調遞增判準——兩個方向都會出錯"
|
|
46
|
+
definite_spec:
|
|
47
|
+
prefer: "字面遵循型(較低層)"
|
|
48
|
+
inverted_failure: "把已寫死的步驟餵給高天花板型反而降低品質——它會重新詮釋已經決定過的事"
|
|
49
|
+
ambiguous_spec:
|
|
50
|
+
prefer: "模糊性導航型(較高層)"
|
|
51
|
+
inverted_failure: "把模糊任務給字面遵循型會得到「精確執行了錯誤的那句話」"
|
|
52
|
+
removed_signal:
|
|
53
|
+
signal: "修改檔案數"
|
|
54
|
+
removed_in: "2.1.0"
|
|
55
|
+
reason: >
|
|
56
|
+
3 檔的深度重構可能遠難於 8 檔的機械改名。檔案數會系統性把「深而窄」的任務誤判到低層,
|
|
57
|
+
且誤判方向固定——是偏差不是噪音。它之所以存活是因為便宜好量,不是因為它預測得準。
|
|
58
|
+
effort:
|
|
59
|
+
purpose: "選這次要它想多深(how deep)"
|
|
60
|
+
nature: "單次派工的請求參數,不是模型屬性"
|
|
61
|
+
vendor_neutral: true
|
|
62
|
+
levels:
|
|
63
|
+
- id: low
|
|
64
|
+
label_zh: "低"
|
|
65
|
+
semantics: "最少斟酌,主要依模式作答"
|
|
66
|
+
typical_use: "機械性編輯、查詢、格式化"
|
|
67
|
+
- id: medium
|
|
68
|
+
label_zh: "中"
|
|
69
|
+
semantics: "一般斟酌;預設值"
|
|
70
|
+
typical_use: "多數已定義的實作工作"
|
|
71
|
+
- id: high
|
|
72
|
+
label_zh: "高"
|
|
73
|
+
semantics: "延伸斟酌,作答前考慮替代方案"
|
|
74
|
+
typical_use: "含局部決策的整合工作"
|
|
75
|
+
- id: very-high
|
|
76
|
+
label_zh: "極高"
|
|
77
|
+
semantics: "探索並淘汰候選方法"
|
|
78
|
+
typical_use: "設計、審查、非顯而易見的除錯"
|
|
79
|
+
- id: max
|
|
80
|
+
label_zh: "極限"
|
|
81
|
+
semantics: "該模型可用的最大深度"
|
|
82
|
+
typical_use: "升級模型軸前的最後一次嘗試"
|
|
83
|
+
support_is_a_hard_boundary: >
|
|
84
|
+
並非每一層都支援每一個 effort 級距。是否支援 effort 參數是硬邊界(能力有無),
|
|
85
|
+
不是品質梯度。平台的 tier→模型映射必須登記每個模型接受哪些級距。
|
|
86
|
+
|
|
87
|
+
failure_diagnosis:
|
|
88
|
+
rule: "任務失敗時,先判斷哪一軸不足。兩者的補救方式不可互換。"
|
|
89
|
+
ordering: "先調 effort,再升級模型層;只有在當前層的最高可用 effort 也失敗後才升級模型軸"
|
|
90
|
+
exception: >
|
|
91
|
+
若判準 1(推理天花板需求)事前已識別出「想更久也想不出來」的成分,直接從較高層開始。
|
|
92
|
+
排序規則是關於「診斷失敗」,不是關於忽略事前判準。
|
|
93
|
+
cases:
|
|
94
|
+
- id: depth_insufficient
|
|
95
|
+
observation: "產出淺、略過考量、提早收手,但它做過的推理本身是對的"
|
|
96
|
+
diagnosis: "深度不足"
|
|
97
|
+
remedy: "在同一個模型上提高 effort"
|
|
98
|
+
wrong_remedy_cost: "升級模型層=為一個從來不是瓶頸的天花板多付錢"
|
|
99
|
+
- id: ceiling_insufficient
|
|
100
|
+
observation: "產出在種類上就錯了——誤解問題、提出不可行的方法——且在 max effort 下仍如此"
|
|
101
|
+
diagnosis: "天花板不足"
|
|
102
|
+
remedy: "升級模型層"
|
|
103
|
+
wrong_remedy_cost: "再提高 effort 買不到任何東西;級距已經用盡"
|
|
104
|
+
|
|
105
|
+
reverse_exclusion:
|
|
106
|
+
description: "什麼工作不該給哪一層(XSPEC-362 R3)。此問題無法由「該升到哪層」的判準推導出來。"
|
|
107
|
+
hard_boundaries:
|
|
108
|
+
id: R3a
|
|
109
|
+
principle: "硬邊界是能力有無,不是程度差。裝不下的模型不是「做得比較差」,是做不了。"
|
|
110
|
+
exclusion_timing: "在成本比較「之前」排除,不是在其中排名較低"
|
|
111
|
+
exclusion_timing_reason: "事後排除會讓比較便宜但做不到的模型在價格上勝出"
|
|
112
|
+
boundaries:
|
|
113
|
+
- id: context_capacity
|
|
114
|
+
question: "輸入裝得下嗎?"
|
|
115
|
+
violation: "截斷或報錯——模型從未看過任務的一部分"
|
|
116
|
+
- id: effort_parameter_support
|
|
117
|
+
question: "這個模型接受 effort 參數嗎?"
|
|
118
|
+
violation: "請求被拒,或靜默以模型固定深度執行"
|
|
119
|
+
- id: modality_support
|
|
120
|
+
question: "它接受這種輸入型別嗎(影像、語音)?"
|
|
121
|
+
violation: "輸入被丟棄或拒絕"
|
|
122
|
+
registration: "硬邊界對應 capability registry 的 declared(布林,連線時可得,成本趨近零)"
|
|
123
|
+
score_1_is_not_a_hard_boundary: "分數 1 不是硬邊界,是量到的低品質;兩者的動作不同"
|
|
124
|
+
reverse_risks:
|
|
125
|
+
id: R3b
|
|
126
|
+
risks:
|
|
127
|
+
- id: safety_classifier_refusal
|
|
128
|
+
description: >
|
|
129
|
+
高能力層可能帶更嚴格的安全分類器。合法但形似禁止類別的工作(資安加固、
|
|
130
|
+
緩解措施實作、憑證處理程式碼、red-team 工具)可能被拒絕。
|
|
131
|
+
silent_failure_mechanism: >
|
|
132
|
+
拒絕不是錯誤。呼叫回傳的是正常回應加上拒絕標記。
|
|
133
|
+
對只看 exit code、或只看有無擲出例外的呼叫端而言,這是靜默失效:
|
|
134
|
+
管線記錄成功,而工作根本沒做。
|
|
135
|
+
requirements:
|
|
136
|
+
- "呼叫端必須檢查回應中的拒絕標記;沒有錯誤不等於完成"
|
|
137
|
+
- "規格敏感類工作在批次派往高能力層前,須先以廉價探測請求驗證可行性"
|
|
138
|
+
- "遇到拒絕時改派其他層或其他 provider;不要以更高 effort 重試同一請求——effort 移動不了分類器的判定"
|
|
139
|
+
related_standard: verification-evidence
|
|
140
|
+
- id: over_specified_prompt_degrades_high_tier
|
|
141
|
+
description: >
|
|
142
|
+
寫出每一個步驟對字面遵循型是正確做法,對高天花板型是錯誤做法——
|
|
143
|
+
後者會重新詮釋已經決定過的事,產出反而變差。
|
|
144
|
+
rule: "prompt 顆粒度隨層級走。若唯一可用的 prompt 是完整步驟清單,判準 2 本來就要求送到較低層"
|
|
21
145
|
|
|
22
146
|
tiers:
|
|
23
147
|
- id: fast
|
|
24
|
-
description: "
|
|
148
|
+
description: "無推理天花板需求、規格完全明確"
|
|
25
149
|
signals:
|
|
26
|
-
- "
|
|
27
|
-
- "
|
|
150
|
+
- "不存在「想更久也想不出來」的成分"
|
|
151
|
+
- "步驟已寫出,無需尋路"
|
|
28
152
|
- "不需要設計判斷"
|
|
29
|
-
- "
|
|
153
|
+
- "字面遵循正是想要的行為"
|
|
30
154
|
examples:
|
|
31
|
-
- "
|
|
32
|
-
- "新增一個 export"
|
|
155
|
+
- "在一組檔案上套用已定義的編輯(機械改名、版本號提升)"
|
|
33
156
|
- "修正 typo"
|
|
157
|
+
- "新增規格中已指名的 export"
|
|
34
158
|
- id: standard
|
|
35
|
-
description: "
|
|
159
|
+
description: "需要一定推理天花板,規格大致明確但留有局部決策空間"
|
|
36
160
|
signals:
|
|
37
|
-
- "
|
|
161
|
+
- "部分成分需要推理,但沒有一項是「想更久也想不出來」"
|
|
162
|
+
- "目標與多數步驟已定義,剩下有限的局部選擇"
|
|
38
163
|
- "需要理解模組間關係"
|
|
39
|
-
- "需要一定設計判斷"
|
|
40
164
|
examples:
|
|
41
|
-
- "
|
|
42
|
-
- "
|
|
165
|
+
- "實作整合點已知的既定功能"
|
|
166
|
+
- "依既定目標結構重構模組"
|
|
167
|
+
- "為指定子系統撰寫整合測試"
|
|
43
168
|
- id: capable
|
|
44
|
-
description: "
|
|
169
|
+
description: "需要高推理天花板,或規格僅有目標與約束、路徑未知"
|
|
45
170
|
signals:
|
|
46
|
-
- "
|
|
47
|
-
- "
|
|
48
|
-
- "
|
|
49
|
-
- "除錯複雜問題"
|
|
171
|
+
- "存在單靠更多思考時間解決不了的成分"
|
|
172
|
+
- "路徑未知;答案要找出來,不是套用"
|
|
173
|
+
- "需要導航模糊性而非遵循文字"
|
|
50
174
|
examples:
|
|
51
|
-
- "
|
|
52
|
-
- "審查大型 PR"
|
|
53
|
-
- "
|
|
175
|
+
- "設計子系統架構(路徑未知)"
|
|
176
|
+
- "審查大型 PR(重點事前未載明)"
|
|
177
|
+
- "診斷尚未形成假設的失效"
|
|
178
|
+
note: >
|
|
179
|
+
目標結構「已定」的大型重構不自動屬於 capable——依判準 2 它是明確規格,
|
|
180
|
+
高天花板模型的產出可能比字面遵循型更差。使它升到 capable 的是「目標結構未定」。
|
|
54
181
|
|
|
55
182
|
rules:
|
|
56
183
|
- id: MS-001
|
|
57
|
-
trigger: "BLOCKED
|
|
58
|
-
action: "升級至 standard
|
|
184
|
+
trigger: "BLOCKED 且當前為 fast,effort 已用盡"
|
|
185
|
+
action: "升級至 standard"
|
|
59
186
|
priority: high
|
|
60
187
|
- id: MS-002
|
|
61
|
-
trigger: "BLOCKED
|
|
62
|
-
action: "升級至 capable
|
|
188
|
+
trigger: "BLOCKED 且當前為 standard,effort 已用盡"
|
|
189
|
+
action: "升級至 capable"
|
|
63
190
|
priority: high
|
|
64
191
|
- id: MS-003
|
|
65
|
-
trigger: "BLOCKED
|
|
192
|
+
trigger: "BLOCKED 且當前為 capable,effort 已用盡"
|
|
66
193
|
action: "標記為需要人工介入"
|
|
67
194
|
priority: critical
|
|
68
195
|
- id: MS-004
|
|
69
|
-
trigger: "
|
|
196
|
+
trigger: "任務規格僅有目標與約束"
|
|
70
197
|
action: "從 standard 或更高層級開始"
|
|
71
198
|
priority: medium
|
|
199
|
+
- id: MS-005
|
|
200
|
+
trigger: "產出淺但推理健全,且 effort 尚未達 max"
|
|
201
|
+
action: "在同一模型上提高 effort — 不要升級模型層"
|
|
202
|
+
priority: high
|
|
203
|
+
- id: MS-006
|
|
204
|
+
trigger: "產出在種類上就錯了,且 effort 已達 max"
|
|
205
|
+
action: "升級模型層 — 再提高 effort 買不到任何東西"
|
|
206
|
+
priority: high
|
|
207
|
+
- id: MS-007
|
|
208
|
+
trigger: "候選模型不滿足某項必要硬邊界(context、effort 支援、modality)"
|
|
209
|
+
action: "在成本比較「之前」排除"
|
|
210
|
+
priority: critical
|
|
211
|
+
- id: MS-008
|
|
212
|
+
trigger: "將規格敏感類工作派往高能力層"
|
|
213
|
+
action: "先探測可行性;每一則回應都要檢查拒絕標記"
|
|
214
|
+
priority: critical
|
|
215
|
+
- id: MS-009
|
|
216
|
+
trigger: "prompt 是完整寫出的步驟清單"
|
|
217
|
+
action: "偏好字面遵循型;不要升到高天花板層"
|
|
218
|
+
priority: medium
|
|
219
|
+
- id: MS-010
|
|
220
|
+
trigger: "pin_date 或 measured.at 距今超過 90 天"
|
|
221
|
+
action: "發出 WARN 並排入重測佇列(不阻擋發版)"
|
|
222
|
+
priority: medium
|
|
72
223
|
|
|
73
224
|
capability_dimensions:
|
|
74
|
-
description: "能力維度分類與評分說明(XSPEC-027, DEC-031, DEC-032)"
|
|
225
|
+
description: "能力維度分類與評分說明(XSPEC-027, DEC-031, DEC-032, XSPEC-362 R7b)"
|
|
226
|
+
two_fields_per_subdimension:
|
|
227
|
+
rationale: >
|
|
228
|
+
單一 1–5 分無法同時表達兩件事:「支不支援」是二元、廠商宣告、連線時可低成本查得;
|
|
229
|
+
「有多好」是連續、需實測。壓成同一個數字,會讓便宜的事實與昂貴的事實無從分辨,
|
|
230
|
+
而昂貴的那個會靜默勝出。
|
|
231
|
+
declared:
|
|
232
|
+
type: boolean
|
|
233
|
+
source: "provider 的模型描述端點或設定宣告"
|
|
234
|
+
cost: "趨近零,連線建立時可得"
|
|
235
|
+
meaning: "做不做得到這件事"
|
|
236
|
+
measured:
|
|
237
|
+
type: object
|
|
238
|
+
fields:
|
|
239
|
+
score: "1–5 整數"
|
|
240
|
+
at: "量測日期 YYYY-MM-DD"
|
|
241
|
+
version_identifier: "量測時所用的版本識別"
|
|
242
|
+
source: "benchmark"
|
|
243
|
+
cost: "高,離線"
|
|
244
|
+
meaning: "做得多好"
|
|
245
|
+
normative_rules:
|
|
246
|
+
- id: declared_false_is_hard_boundary
|
|
247
|
+
rule: "declared: false 是硬邊界——無論任何分數皆排除,且「不」排入校準佇列(校準無意義)"
|
|
248
|
+
- id: declared_true_measured_absent_is_unknown
|
|
249
|
+
rule: "declared: true 且 measured 缺失 = UNKNOWN(做得到,但不知道做得多好);這是資訊缺口不是結論"
|
|
250
|
+
- id: supported_must_not_derive_from_score
|
|
251
|
+
rule: >
|
|
252
|
+
禁止由分數推導支援與否。任何以 supported = score > 0(或類似)計算的實作,
|
|
253
|
+
都已把兩個欄位合併回單軸,重新製造了本規則要防止的缺陷。
|
|
254
|
+
- id: failed_measurement_must_not_be_scored
|
|
255
|
+
rule: >
|
|
256
|
+
量測失敗不得記為任何分數,必須記為缺失(UNKNOWN)。
|
|
257
|
+
以預設值填補(包括 0,而 0 不在 1–5 量表內)會讓「量不到」與「量到了且很差」無從分辨。
|
|
75
258
|
scoring_scale:
|
|
259
|
+
applies_to: "measured.score only"
|
|
260
|
+
zero_does_not_exist: "沒有 0 分。量測缺失以「measured 物件不存在」表達,絕不以分數表達。"
|
|
76
261
|
5: "生產就緒 — 高準確率,可直接使用"
|
|
77
262
|
4: "良好 — 偶有遺漏,可接受"
|
|
78
263
|
3: "基本可用 — 需人工補充"
|
|
@@ -122,77 +307,190 @@ standard:
|
|
|
122
307
|
|
|
123
308
|
capability_registry:
|
|
124
309
|
description: "模型能力登記表(DEC-031 D1)。各子專案依實際測試維護分數。"
|
|
310
|
+
no_concrete_model_ids:
|
|
311
|
+
rule: "本標準的 examples 不得登記任何具體廠商模型 ID"
|
|
312
|
+
reason: >
|
|
313
|
+
寫進標準的模型 ID 是「有到期日卻沒有時鐘」的引用端:它會過期,
|
|
314
|
+
而過期的樣子與正常的樣子在檔案上無從分辨。具體登記由採用者在自己的 registry 維護。
|
|
315
|
+
enforced_since: "2.1.0(XSPEC-362 R4)"
|
|
125
316
|
format:
|
|
126
|
-
model_id: "provider
|
|
127
|
-
version_pinned: "
|
|
128
|
-
pin_date: "
|
|
129
|
-
eol_date: "
|
|
317
|
+
model_id: "<provider>/<model-name>"
|
|
318
|
+
version_pinned: "<version-identifier>(SHA、日期戳或 model_version)"
|
|
319
|
+
pin_date: "<YYYY-MM-DD>"
|
|
320
|
+
eol_date: "<YYYY-MM-DD>(可選)"
|
|
130
321
|
capabilities:
|
|
131
|
-
"<dimension>.<subdimension>":
|
|
322
|
+
"<dimension>.<subdimension>":
|
|
323
|
+
declared: "<boolean>"
|
|
324
|
+
measured:
|
|
325
|
+
score: "<1-5 整數>"
|
|
326
|
+
at: "<YYYY-MM-DD>"
|
|
327
|
+
version_identifier: "<量測時所用的版本識別>"
|
|
132
328
|
examples:
|
|
133
|
-
- model_id: "
|
|
134
|
-
version_pinned: "
|
|
135
|
-
pin_date: "
|
|
136
|
-
capabilities:
|
|
137
|
-
"modality.vision": 4
|
|
138
|
-
"reasoning.instruction_following": 5
|
|
139
|
-
"reasoning.code_reasoning": 5
|
|
140
|
-
"output.structured_output": 5
|
|
141
|
-
"output.tool_use": 5
|
|
142
|
-
"language.multilingual_zh_tw": 5
|
|
143
|
-
- model_id: "anthropic/claude-haiku-4-5"
|
|
144
|
-
version_pinned: "claude-haiku-4-5-20251001"
|
|
145
|
-
pin_date: "2026-04-13"
|
|
329
|
+
- model_id: "<provider>/<model-name-a>"
|
|
330
|
+
version_pinned: "<version-identifier-a>"
|
|
331
|
+
pin_date: "<YYYY-MM-DD>"
|
|
146
332
|
capabilities:
|
|
147
|
-
"modality.vision":
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
333
|
+
"modality.vision":
|
|
334
|
+
declared: true
|
|
335
|
+
measured:
|
|
336
|
+
score: 4
|
|
337
|
+
at: "<YYYY-MM-DD>"
|
|
338
|
+
version_identifier: "<version-identifier-a>"
|
|
339
|
+
"modality.audio":
|
|
340
|
+
declared: false
|
|
341
|
+
"output.tool_use":
|
|
342
|
+
declared: true
|
|
343
|
+
example_notes:
|
|
344
|
+
- "modality.audio 的 declared: false 是硬邊界——無 measured 欄位,也不需要"
|
|
345
|
+
- "output.tool_use 的 measured 缺失 → UNKNOWN → 排入校準佇列(不是 UNSUPPORTED)"
|
|
346
|
+
staleness_check:
|
|
347
|
+
threshold_days: 90
|
|
348
|
+
applies_to: ["pin_date", "measured.at"]
|
|
349
|
+
severity: WARN
|
|
350
|
+
must_not_block: true
|
|
351
|
+
rationale: "依 XSPEC-361 R8 量測結論,純檔內不變量的誤報率偏高,不適合作為 BLOCK 閘門"
|
|
352
|
+
output_requirement: "WARN 須指名受影響的 model_id 與子維度"
|
|
153
353
|
|
|
154
354
|
routing_rules:
|
|
155
|
-
description: "能力路由決策規則(DEC-031 D2, DEC-032 D3)"
|
|
355
|
+
description: "能力路由決策規則(DEC-031 D2, DEC-032 D3, XSPEC-362 R7a)"
|
|
156
356
|
selection_strategy: "pareto_weighted"
|
|
357
|
+
selection_strategy_scope: "僅在通過硬邊界排除的候選之中比較"
|
|
358
|
+
states:
|
|
359
|
+
- id: SUPPORTED
|
|
360
|
+
condition: "所需能力的 measured.score 皆 ≥ min_score"
|
|
361
|
+
action: "正常執行"
|
|
362
|
+
calibration_queue: false
|
|
363
|
+
- id: DEGRADED
|
|
364
|
+
condition: "measured.score 存在、≥ 2、但低於 min_score"
|
|
365
|
+
action: "降級執行,產出標記 [DEGRADED]"
|
|
366
|
+
calibration_queue: false
|
|
367
|
+
- id: UNSUPPORTED
|
|
368
|
+
condition: "measured 存在且 score ≤ 1"
|
|
369
|
+
action: "排除,不再嘗試"
|
|
370
|
+
calibration_queue: false
|
|
371
|
+
calibration_queue_reason: "已有結論"
|
|
372
|
+
- id: UNKNOWN
|
|
373
|
+
condition: "declared: true 但 measured 缺失——未登記、量測失敗或資料逾期"
|
|
374
|
+
action: "排入校準佇列;在校準完成前不得靜默排除"
|
|
375
|
+
calibration_queue: true
|
|
376
|
+
hard_boundary_precedes_states: "declared: false 在本表之前處理:硬邊界排除,且不排入校準佇列"
|
|
377
|
+
observability_requirement:
|
|
378
|
+
rule: "UNKNOWN 與 UNSUPPORTED 必須在「回傳結構」上可分辨,不能只在 log 上分辨"
|
|
379
|
+
reason: >
|
|
380
|
+
呼叫端必須能分辨「這個模型不行」與「我還不知道這個模型行不行」——兩者導向不同的下一步動作。
|
|
381
|
+
回傳單一布林、或回傳三態的路由 API,表達不了這件事。
|
|
382
|
+
anti_pattern:
|
|
383
|
+
description: "把「查無」與「低分」放進同一個判斷式,並以預設值填補缺資料"
|
|
384
|
+
example: "if (!cap || !cap.supported || cap.score < 2) — 三件事一個判斷式;以 ?? 0 填補缺資料"
|
|
385
|
+
consequence: "任何新偵測到的模型必定被排除,新模型永遠進不了候選池——與「支援更多模型」的目標直接相反"
|
|
386
|
+
decision_tree: |
|
|
387
|
+
任務需要 capability X
|
|
388
|
+
├── declared == false → HARD BOUNDARY — 排除,不校準
|
|
389
|
+
├── measured 缺失/失敗/逾期 → UNKNOWN — 排入校準佇列
|
|
390
|
+
├── measured.score ≤ 1 → UNSUPPORTED — 排除,不校準
|
|
391
|
+
├── measured.score ≥ 2 且 < min_score → DEGRADED — 執行並標記 [DEGRADED]
|
|
392
|
+
└── measured.score ≥ min_score → SUPPORTED — 執行
|
|
157
393
|
rules:
|
|
158
394
|
- id: CAP-001
|
|
159
|
-
condition: "任務需要 capability X
|
|
395
|
+
condition: "任務需要 capability X,measured.score ≥ min_score"
|
|
160
396
|
action: "SUPPORTED — 正常執行"
|
|
161
397
|
priority: high
|
|
162
398
|
- id: CAP-002
|
|
163
|
-
condition: "
|
|
399
|
+
condition: "measured.score ≥ 2 但低於 min_score"
|
|
164
400
|
action: "DEGRADED — 降級流程執行,產出標記 [DEGRADED]"
|
|
165
401
|
priority: medium
|
|
166
402
|
- id: CAP-003
|
|
167
|
-
condition: "
|
|
168
|
-
action: "UNSUPPORTED —
|
|
403
|
+
condition: "measured 存在且 score ≤ 1"
|
|
404
|
+
action: "UNSUPPORTED — 使用替代流程或提示用戶;不排入校準佇列"
|
|
169
405
|
priority: high
|
|
170
406
|
- id: CAP-004
|
|
171
407
|
condition: "模型降智偵測(DEC-033)觸發 moderate 信號"
|
|
172
|
-
action: "
|
|
408
|
+
action: "啟動金絲雀測試,記錄降智警告,並將受影響能力排入重測"
|
|
173
409
|
priority: high
|
|
174
410
|
- id: CAP-005
|
|
175
411
|
condition: "模型降智偵測觸發 critical 信號"
|
|
176
|
-
action: "自動切換備用模型,上報 P1 Issue"
|
|
412
|
+
action: "自動切換備用模型,上報 P1 Issue,並將受影響能力排入重測"
|
|
413
|
+
priority: critical
|
|
414
|
+
- id: CAP-006
|
|
415
|
+
condition: "declared: true 且 measured 缺失、失敗或逾期"
|
|
416
|
+
action: "UNKNOWN — 排入校準佇列;不得靜默排除"
|
|
417
|
+
priority: high
|
|
418
|
+
- id: CAP-007
|
|
419
|
+
condition: "所需 capability 的 declared: false"
|
|
420
|
+
action: "硬邊界 — 在成本比較之前排除;不排入校準佇列"
|
|
177
421
|
priority: critical
|
|
422
|
+
- id: CAP-008
|
|
423
|
+
condition: "version_identifier 與 measured.version_identifier 不同"
|
|
424
|
+
action: "立即失效該量測 → UNKNOWN"
|
|
425
|
+
priority: high
|
|
426
|
+
|
|
427
|
+
remeasurement_triggers:
|
|
428
|
+
description: "重測觸發(XSPEC-362 R7c)"
|
|
429
|
+
principle: "版本變更是充分條件,不是必要條件"
|
|
430
|
+
principle_reason: >
|
|
431
|
+
DEC-033 降智偵測的存在前提,正是模型 ID 與版本字串不變而行為改變。
|
|
432
|
+
僅依賴版本變更的實作,會漏掉 DEC-033 整個要防的情境。
|
|
433
|
+
independence: "三條路徑彼此獨立:各自單獨觸發,且沒有任何一條是另一條的前提"
|
|
434
|
+
triggers:
|
|
435
|
+
- id: version_identifier_changed
|
|
436
|
+
trigger: "version_identifier 與 measured 中記錄的不同"
|
|
437
|
+
effect: "既有量測立即失效 → 狀態變為 UNKNOWN"
|
|
438
|
+
- id: measurement_expired
|
|
439
|
+
trigger: "measured.at 距今超過 90 天"
|
|
440
|
+
effect: "排入重測;發出 WARN(見 MS-010)"
|
|
441
|
+
- id: degradation_alert
|
|
442
|
+
trigger: "降智偵測告警(DEC-033 CAP-004/CAP-005)"
|
|
443
|
+
effect: "排入重測——即使版本字串未變"
|
|
444
|
+
|
|
445
|
+
relationship_to_axes:
|
|
446
|
+
model_axis: "推理天花板與性格(選哪個模型)"
|
|
447
|
+
effort_axis: "這次派工的思考深度(想多深)"
|
|
448
|
+
capability_dimensions: "選中的模型做不做得到這「種」事,以及做得多好"
|
|
449
|
+
order_of_application: >
|
|
450
|
+
硬邊界排除 → 依兩個判準選 tier → 確認所需能力為 SUPPORTED 或可接受的 DEGRADED → 選 effort 級距
|
|
178
451
|
|
|
179
452
|
navigation_integration:
|
|
180
453
|
description: >
|
|
181
454
|
The ai-response-navigation standard (Rule R6, optional) allows next-step suggestion
|
|
182
455
|
options to carry a Tier annotation (〔model: Fast〕 / 〔model: Standard〕 / 〔model: Capable〕).
|
|
183
|
-
The
|
|
456
|
+
The criteria in this standard serve as the judgment basis for those annotations.
|
|
457
|
+
changed_in_2_1_0: >
|
|
458
|
+
Annotations previously derived from file count. Existing annotations made under the old
|
|
459
|
+
criteria may no longer be accurate. Tier ids are unchanged, so no data migration is required.
|
|
184
460
|
vendor_neutral_principle: >
|
|
185
|
-
Tier names (Fast / Standard / Capable) are vendor-agnostic
|
|
186
|
-
|
|
461
|
+
Tier names (Fast / Standard / Capable) and effort labels (low … max) are vendor-agnostic.
|
|
462
|
+
They do not correspond to any specific vendor's model identifiers or parameter names.
|
|
463
|
+
Each platform independently maps these tiers to its available models, maps effort labels to
|
|
464
|
+
its own parameter, and maintains its own hard-boundary register.
|
|
465
|
+
|
|
466
|
+
related_standards:
|
|
467
|
+
- id: agent-dispatch
|
|
468
|
+
relationship: >
|
|
469
|
+
agent-dispatch 管「怎麼派」(並行安全、獨立域判準、狀態協定、prompt 設計);
|
|
470
|
+
model-selection 管「派給誰、想多深」(模型軸 × effort 軸)。
|
|
471
|
+
兩者互補,不得互相重寫——並行安全規則屬於 agent-dispatch,不屬於這裡。
|
|
472
|
+
- id: verification-evidence
|
|
473
|
+
relationship: "「沒有擲出錯誤」不是完成的證據;R3b 拒絕標記要求的依據"
|
|
474
|
+
- id: systematic-debugging
|
|
475
|
+
relationship: "先診斷再修改,在此適用於「深度不足 vs 天花板不足」的判斷"
|
|
187
476
|
|
|
188
477
|
physical_spec:
|
|
189
478
|
type: checklist
|
|
190
479
|
validator:
|
|
191
480
|
type: ai_review
|
|
192
|
-
rule: "
|
|
481
|
+
rule: "檢查模型選擇是否依「推理天花板需求 × 規格明確度」判準,且 effort 與模型軸分開決定"
|
|
193
482
|
checks:
|
|
194
|
-
- "
|
|
195
|
-
- "
|
|
196
|
-
- "
|
|
483
|
+
- "模型層級是否依推理天花板需求與規格明確度選擇,而非依修改檔案數"
|
|
484
|
+
- "effort 級距是否與模型層級分開決定"
|
|
485
|
+
- "失敗時是否先判斷「深度不足」或「天花板不足」,再選擇對應補救"
|
|
486
|
+
- "升級模型層之前,當前層的最高可用 effort 是否已嘗試過"
|
|
487
|
+
- "硬邊界(context、effort 支援、modality)是否在成本比較之前排除"
|
|
488
|
+
- "規格敏感類工作的呼叫端是否檢查拒絕標記,而非只檢查有無例外"
|
|
197
489
|
- "能力登記表是否包含 version_pinned 與 pin_date"
|
|
198
|
-
- "
|
|
490
|
+
- "能力登記表的 examples 是否不含具體廠商模型 ID"
|
|
491
|
+
- "每個子維度是否同時具備 declared(布林)與 measured(分數+時間+版本識別)兩個獨立欄位"
|
|
492
|
+
- "是否不存在任何以分數推導支援與否的實作(如 supported = score > 0)"
|
|
493
|
+
- "量測失敗是否記為缺失,而非以預設值(如 0)填補"
|
|
494
|
+
- "路由決策是否依 SUPPORTED/DEGRADED/UNSUPPORTED/UNKNOWN 四態處理"
|
|
495
|
+
- "UNKNOWN 與 UNSUPPORTED 是否在回傳結構上可分辨"
|
|
496
|
+
- "重測觸發是否包含版本變更、時間逾期、降智告警三條獨立路徑"
|