universal-dev-standards 6.11.0 → 6.13.0-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/bundled/ai/standards/developer-memory.ai.yaml +24 -2
  2. package/bundled/ai/standards/open-work-tracking.ai.yaml +216 -0
  3. package/bundled/ai/standards/turn-completion-integrity.ai.yaml +34 -2
  4. package/bundled/core/developer-memory.md +58 -2
  5. package/bundled/core/open-work-tracking.md +333 -0
  6. package/bundled/core/turn-completion-integrity.md +36 -2
  7. package/bundled/hooks/check-turn-completion-codex.mjs +116 -0
  8. package/bundled/hooks/check-turn-completion-gemini.mjs +81 -0
  9. package/bundled/hooks/check-turn-completion.mjs +24 -149
  10. package/bundled/hooks/turn-completion/engine.mjs +206 -0
  11. package/bundled/hooks/turn-completion/locales/en.mjs +29 -1
  12. package/bundled/hooks/turn-completion/locales/zh-TW.mjs +29 -1
  13. package/bundled/locales/zh-CN/CHANGELOG.md +23 -3
  14. package/bundled/locales/zh-CN/CLAUDE.md +1 -1
  15. package/bundled/locales/zh-CN/README.md +2 -2
  16. package/bundled/locales/zh-CN/SECURITY.md +2 -1
  17. package/bundled/locales/zh-CN/core/turn-completion-integrity.md +36 -6
  18. package/bundled/locales/zh-CN/docs/CHEATSHEET.md +2 -1
  19. package/bundled/locales/zh-CN/docs/CLI-INIT-OPTIONS.md +24 -5
  20. package/bundled/locales/zh-CN/docs/FEATURE-REFERENCE.md +4 -3
  21. package/bundled/locales/zh-TW/CHANGELOG.md +23 -3
  22. package/bundled/locales/zh-TW/CLAUDE.md +1 -1
  23. package/bundled/locales/zh-TW/README.md +2 -2
  24. package/bundled/locales/zh-TW/SECURITY.md +2 -1
  25. package/bundled/locales/zh-TW/core/open-work-tracking.md +255 -0
  26. package/bundled/locales/zh-TW/core/turn-completion-integrity.md +36 -6
  27. package/bundled/locales/zh-TW/docs/CHEATSHEET.md +2 -1
  28. package/bundled/locales/zh-TW/docs/CLI-INIT-OPTIONS.md +24 -5
  29. package/bundled/locales/zh-TW/docs/FEATURE-REFERENCE.md +4 -3
  30. package/package.json +1 -1
  31. package/src/commands/init.js +26 -0
  32. package/src/installers/hooks-installer.js +104 -0
  33. package/standards-registry.json +19 -7
@@ -0,0 +1,333 @@
1
+ # Open Work Tracking Standard
2
+
3
+ > **Language**: English | [繁體中文](../locales/zh-TW/core/open-work-tracking.md)
4
+
5
+ **Version**: 1.0.0
6
+ **Last Updated**: 2026-09-23
7
+ **Applicability**: Any project that carries work across more than one working session and risks losing an item between them
8
+ **Scope**: universal
9
+
10
+ ---
11
+
12
+ ## Purpose
13
+
14
+ The deferred-item-exit standard requires that a deferred item leave its document for a traceable exit. It deliberately does not say what that exit is made of, or what keeps it from rotting once the item has arrived there. **This standard is that downstream half**: given that a carrier for open work exists, what must be true of the carrier itself so that it stays trustworthy. See [deferred-item-exit](deferred-item-exit.md).
15
+
16
+ 延後項目出口標準(deferred-item-exit)要求延後項目離開文件、抵達一個可追蹤的出口,
17
+ 但刻意不規定那個出口長什麼樣、也不規定項目抵達之後什麼東西防止它腐壞。
18
+ **本標準是那個下游的一半**:假設一個承載開放工作的地方已經存在,
19
+ 它自己必須具備什麼性質才不會慢慢變得不可信。見 [deferred-item-exit](deferred-item-exit.md)。
20
+
21
+ Three distinct ways work goes missing between sessions are routinely folded into one undifferentiated list, and the fold is itself part of the failure — a list built to catch all three catches none of them well:
22
+
23
+ 三種不同的「工作不見了」的方式,常被塞進同一份沒有分別的清單,而**這個合併本身就是失敗的一部分**——
24
+ 一份想同時接住三者的清單,通常一個都接不好:
25
+
26
+ | Symptom | Underlying problem | Mechanism needed |
27
+ |---|---|---|
28
+ | A new item surfaces mid-task and there is no low-friction way to record it | The item is time-sensitive; by the time recording it is convenient, it is forgotten | A capture point cheap enough to use without breaking the current task |
29
+ | An item is paused pending some other event | "Waiting" with no recorded release condition is indistinguishable from "forgotten" | A release condition recorded alongside the wait |
30
+ | A planned item has not been started and time passes | An item with no clock decays silently — nothing ever points back at it | A threshold, or a forced periodic look, that surfaces it again |
31
+
32
+ | 症狀 | 背後的問題 | 需要的機制 |
33
+ |---|---|---|
34
+ | 工作進行中冒出新項目,沒有低摩擦的地方可以記下它 | 項目有時效性,等到方便記錄時已經忘了 | 一個便宜到不會打斷當前工作的收件點 |
35
+ | 某項目因等待別的事件而暫停 | 「等待中」若沒有記錄解除條件,與「被忘記」無法分辨 | 與等待一起記錄的解除條件 |
36
+ | 已規劃的項目還沒動工,時間過去 | 沒有時鐘的項目會無聲腐爛——沒有東西會再指向它 | 一個門檻,或一次被迫的定期檢視,讓它重新浮現 |
37
+
38
+ Each requirement below traces to one row of that table, or to one of four failures observed the day this standard's design was drafted (see [Evidence and calibration](#evidence-and-calibration)). **None of the mechanisms is prescribed** — per the same constraint [deferred-item-exit](deferred-item-exit.md) states for its own exits, and for the same reason (DEC-049: UDS defines relations that must hold, adoption layers choose what maintains them).
39
+
40
+ 下面每一條要求都對應這張表的一列,或對應本標準設計當天觀察到的四個失效之一
41
+ (見〈[證據與校準](#evidence-and-calibration)〉)。**沒有任何一個機制被規定**——
42
+ 理由與 [deferred-item-exit](deferred-item-exit.md) 對自己出口的約束相同(DEC-049:
43
+ UDS 定義必須成立的關係,維持它的機制由採用層選擇)。
44
+
45
+ ---
46
+
47
+ ## How this standard is written — and why it is written that way
48
+
49
+ **Read this section before reading any requirement below. It governs all of them.**
50
+
51
+ UDS defines **activities**; adoption layers **orchestrate** them (DEC-049). A standard written as a workflow protocol, a file format, or a specific tool's configuration belongs to the adoption layer, not here — this is the same boundary [deferred-item-exit](deferred-item-exit.md) is held to, applied one layer downstream.
52
+
53
+ UDS 定義**活動**,採用層負責**編排**(DEC-049)。一份寫成工作流協定、檔案格式、
54
+ 或特定工具設定的標準屬於採用層,不屬於這裡——這與 [deferred-item-exit](deferred-item-exit.md)
55
+ 受的約束相同,只是套用在下游一層。
56
+
57
+ | Admissible here — **what** | Not admissible here — **how** |
58
+ |---|---|
59
+ | A property a capture point's field count must have | Which app, file, or ticket system is the capture point |
60
+ | A relation between a waiting item and its release condition | A scheduler or bot that polls for that condition |
61
+ | A property a "this is current" claim must be provable from | A specific hash function, diff tool, or CI provider |
62
+ | A relation between a reported count and what it could not see | A dashboard layout or report template |
63
+ | That a checkpoint exists at the point control returns to a human, and never blocks | Which hook system, shell, or cron implements it |
64
+
65
+ | 這裡容許——**what** | 這裡不容許——**how** |
66
+ |---|---|
67
+ | 收件點欄位數必須具備的性質 | 收件點是哪個 app、檔案或工單系統 |
68
+ | 等待項目與解除條件之間必須存在的關係 | 輪詢那個條件的排程器或機器人 |
69
+ | 「這是最新的」這句宣稱必須從什麼可被證明 | 具體用哪個雜湊函式、diff 工具或 CI 供應商 |
70
+ | 一個回報數字與它看不到的部分之間的關係 | 儀表板版面或報告範本 |
71
+ | 控制權交回人的那一點存在一個確認點,且它永不阻斷 | 用什麼 hook 系統、shell 或 cron 實作它 |
72
+
73
+ **Consequence, stated plainly**: this standard ships **no gate**. It says what must be true of a carrier for open work. Whether anything checks it is the adopting project's decision — [OWT-014](#requirements) and [OWT-015](#requirements) exist so that decision cannot be made silently.
74
+
75
+ **直說它的後果**:本標準**不附帶任何閘門**。它只說一個承載開放工作的地方必須具備什麼性質;
76
+ 有沒有東西在檢查,是採用專案的決定——[OWT-014](#requirements) 與 [OWT-015](#requirements)
77
+ 存在的目的,是讓那個決定沒辦法被默默做掉。
78
+
79
+ ---
80
+
81
+ ## The invariant
82
+
83
+ **A carrier of open work must (1) accept a new item without demanding classification, (2) record a release condition for every item it marks waiting, (3) generate any field a reliable source already determines, (4) disclose what it cannot see whenever it reports what remains, and (5) be checked at the moment control returns from agent to human — by something that cannot fail the turn.**
84
+
85
+ **一個承載開放工作的地方,必須:(1)不要求分類就能收下新項目、(2)為每一個標為等待中的項目記下解除條件、
86
+ (3)對任何有可靠來源可推導的欄位改用生成、(4)回報還剩什麼時同時揭露看不到什麼、
87
+ (5)在控制權從 agent 交回人的那一刻被檢視——而且那個檢視不能讓回合失敗。**
88
+
89
+ ---
90
+
91
+ ## Requirements
92
+
93
+ | ID | Requirement | Severity |
94
+ |---|---|---|
95
+ | **OWT-001** | A capture point for a new item requires no more than two fields to record it. Classification, priority, and ownership are triage-time actions, never entry-time gates | error |
96
+ | **OWT-002** | An item recorded as waiting states both what it is waiting for and what event counts as its release | error |
97
+ | **OWT-003** | A field whose value is fully derivable from version control, a spec marker, or a CI result is generated, not hand-written | error |
98
+ | **OWT-004** | A generated section's claim to be current is provable from the content it was generated from (e.g. a hash of the source), not asserted by a bare human-editable date | error |
99
+ | **OWT-005** | "Content-proven current" and "a date claims current, content unverified" are reported as two distinct states. A report that merges them into one pass state does not satisfy OWT-004 | error |
100
+ | **OWT-006** | Any "N items remain" figure states, beside it, how many sources it could see and how many it could not. The unseen portion is never read as zero | error |
101
+ | **OWT-007** | A summary of open work occurs at the point control returns from agent to human — not only at session start, in CI, or when a tracking document happens to be edited | error |
102
+ | **OWT-008** | The open-work summary's own exit path never changes the turn's outcome, on any input, including "many items remain" | error |
103
+ | **OWT-009** | Where the summary and a blocking check share an end-of-turn event, the summary's output precedes the blocking check's verdict | warning |
104
+ | **OWT-010** | Whether an item is waiting, unclassified, or dropped is determined from a structural field the carrier defines, not from scanning free-text wording alone | error |
105
+ | **OWT-011** | Where a wording heuristic supplements the structural field, its coverage is declared unknown, and a clean pass over it is not reported as "nothing was missed" | warning |
106
+ | **OWT-012** | An item left unclassified past the declared threshold is named individually in the open-work summary, not folded into an aggregate count | error |
107
+ | **OWT-013** | An item removed from the carrier without becoming a spec, a tracked item, or any other named destination carries a one-line reason. A drop with no reason is indistinguishable from silent deletion | error |
108
+ | **OWT-014** | Every requirement of this standard is expressible as a decidable relation over named artefacts. A requirement that cannot be so expressed does not belong in this standard | error |
109
+ | **OWT-015** | A check offered as evidence for any requirement here has been observed to report failure against a sample built to violate it. A check never observed red is not admissible evidence | error |
110
+ | **OWT-016** | Any window or threshold this standard's requirements reference carries its provenance, or is marked uncalibrated | warning |
111
+
112
+ ---
113
+
114
+ ## Capture must cost almost nothing
115
+
116
+ **OWT-001** exists because a capture point with a third required field measurably stops being used. The failure is not hypothetical friction — it is the specific, observed shape of "a new idea surfaces mid-task, and recording it competes with the task that produced it." A field for classification, priority, or ownership asked for **at entry time** is a bet that the person interrupting their own work will pay that cost; the bet is lost more often than it is won, and a capture point nobody uses is not a capture point, it is a form.
117
+
118
+ **OWT-001** 之所以存在,是因為多一個必填欄位的收件點,量測到的結果是**不被使用**。
119
+ 這不是假想的摩擦——它是「工作進行中冒出新想法,記下它要跟正在做的事搶時間」這個情境的具體形狀。
120
+ 在**輸入當下**就要求分類、優先級或負責人,是在賭「正在被打斷的人願意付那個成本」,
121
+ 而這個賭注輸的次數比贏的多;一個沒有人用的收件點不是收件點,是一張表單。
122
+
123
+ Triage — deciding where an item belongs — is a separate, later act. **OWT-012** and **OWT-013** govern what happens if triage never comes: the item is not allowed to sit invisible forever, and it is not allowed to disappear without a reason either.
124
+
125
+ 分類(決定項目屬於哪裡)是另一個、之後才做的動作。**OWT-012** 與 **OWT-013**
126
+ 規範分類永遠不來時會發生什麼:項目不准永遠隱形地待著,也不准無理由地消失。
127
+
128
+ ---
129
+
130
+ ## A waiting item without a release condition is a forgotten item wearing a status label
131
+
132
+ **OWT-002** names the difference between "paused, and something will bring it back" and "paused, forever, with a word attached that makes that not look like what it is." A release condition should be **machine-observable** where that is possible — a date, an identifier appearing somewhere, a file existing — so the item has a chance of surfacing itself instead of depending on a person remembering it exists. Where a machine-observable condition genuinely does not exist, a human-readable one is still required; **OWT-002 does not require automation, it requires that the condition be recorded at all.**
133
+
134
+ **OWT-002** 指出「暫停中、有東西會讓它回來」與「暫停中、永遠、只是貼了一個讓它看起來不像永遠的標籤」
135
+ 之間的差別。解除條件應盡可能是**機器看得見的**——一個日期、一個會出現的識別字、一個檔案存在——
136
+ 讓項目有機會自己跳出來,而不是依賴某個人記得它存在。真的找不到機器看得見的條件時,
137
+ 仍然要求一個人看得懂的條件;**OWT-002 不要求自動化,只要求那個條件被記下來這件事本身**。
138
+
139
+ ---
140
+
141
+ ## A stamp is cheaper to write than the truth, and a check that reads only the stamp cannot tell the difference
142
+
143
+ This is the same failure DEX-006 names for a different artefact, one layer removed. There, an identifier being present was mistaken for the exit it points to being correct. Here, **a generated section's timestamp being recent is mistaken for its content being current** — and the two diverge in a way that is invisible to any check comparing only dates:
144
+
145
+ 這是 DEX-006 在另一個 artefact 上點名的同一種失敗,只是換了一層。DEX-006 那邊,
146
+ 識別字的存在被誤讀成它指向的出口是對的;這裡,**生成區段的時間戳很新,被誤讀成內容是新的**——
147
+ 而這兩者分歧的方式,對任何只比較日期的檢查是隱形的:
148
+
149
+ - A stamp older than the content: caught trivially, by comparing the stamp to the file's own modification history.
150
+ - A stamp *newer* than the content, where the content itself went stale: **invisible**, because "the stamp is recent" is exactly what a correct reconciliation also looks like.
151
+
152
+ - 戳比內容舊:拿戳跟檔案自己的修改紀錄一比就抓到,微不足道。
153
+ - 戳**比內容新**,而內容本身已經過期:**隱形**,因為「戳是新的」正是一次正確對帳看起來的樣子。
154
+
155
+ **OWT-004** requires the currency claim to be provable from the content itself — for example, a hash of the source the section was generated from, stored beside the section, so a mismatch is detectable without trusting that whoever last touched the date also actually reconciled the content. **OWT-005** requires that "provably current" and "a date says so, unverified" never share one pass/fail bit, for the same reason DEX-005/DEX-006 require it of a deferred item's exit: an unknown reported as a pass is worse than an unknown reported as unknown, because the second one is still findable.
156
+
157
+ **OWT-004** 要求「這是最新的」這句宣稱可以從內容本身被證明——例如儲存一份該區段
158
+ 是從哪個來源生成的雜湊、放在區段旁邊,這樣不比對內容也能偵測到不一致,
159
+ 不必信任「最後動手改日期的人也真的對過帳」。**OWT-005** 要求「內容可證明是最新的」
160
+ 與「日期這麼說、內容未驗證」永遠不共用同一個通過/失敗位元,理由與 DEX-005/DEX-006
161
+ 要求延後項目出口做同一件事相同:一個被回報成通過的未知,比一個被回報成未知的未知更糟,
162
+ 因為後者還找得到。
163
+
164
+ ---
165
+
166
+ ## Coverage must state its own blindness
167
+
168
+ **OWT-006** requires that any "N items remain" figure be printed beside the count of sources it could see and the count it could not — not because the unseen count is expected to be zero, but because a reader cannot tell the difference between "11.7% coverage, and 386 items are invisible to this figure" and "11.7% coverage is complete" unless the denominator is printed next to it. A coverage figure with no stated blindness reads as complete by default, and that default is the failure this requirement exists to prevent.
169
+
170
+ **OWT-006** 要求任何「還有 N 項」的數字旁邊,同時印出它看得見多少來源、看不見多少來源——
171
+ 不是因為預期看不見的數字會是零,而是因為讀者分不出「涵蓋率 11.7%,而且有 386 項
172
+ 對這個數字完全隱形」與「涵蓋率 11.7% 就是全貌」,除非分母被印在旁邊。
173
+ 一個沒有聲明盲區的涵蓋率數字,預設會被讀成完整——而那個預設正是這條要求要防的失效。
174
+
175
+ ---
176
+
177
+ ## The checkpoint is a report, not a gate
178
+
179
+ **turn-completion-integrity** ([TCI](turn-completion-integrity.md)) and this standard's OWT-007–OWT-009 both attach to the same event — the moment an agent's turn ends and control returns to a human — and they are built to behave in opposite ways on purpose. Wiring both to one event without understanding why they differ produces either a checkpoint that blocks on something that is nearly always true, or a report mistaken for a gate:
180
+
181
+ **turn-completion-integrity**([TCI](turn-completion-integrity.md))與本標準的 OWT-007–OWT-009
182
+ 都掛在同一個事件——agent 的回合結束、控制權交回人類的那一刻——而它們被刻意設計成
183
+ **行為相反**。把兩者接到同一個事件卻不理解為什麼不同,會產出「擋在一件幾乎永遠為真的事情上的
184
+ 確認點」,或是「被誤認成閘門的報告」,兩者都不對:
185
+
186
+ | | [turn-completion-integrity](turn-completion-integrity.md) | This standard (OWT-007–009) |
187
+ |---|---|---|
188
+ | What it watches | The agent's own last message, for a first-person commitment that was stated and then abandoned | Whatever the carrier holds, for items with no release condition, no exit, or left unclassified past threshold |
189
+ | Default state | Rare — it fires only when a specific commitment was made in that message and then dropped | Common — "some work is still open" is close to always true |
190
+ | What a violation does | Blocks the turn from ending until the commitment is resolved or its blocker is named | Never blocks. It can only report (OWT-008) |
191
+ | Why the difference | The event it watches for is rare enough that blocking on it does not wear out its welcome | TCI's own rule names the reason this one cannot be a gate: *"A gate that is true on every turn is turned off, and then it protects nothing"* (TCI R4). Open work being non-empty is close to always true, so this checkpoint is built to never withhold control |
192
+ | Ordering when both are wired to the same event | — | Reports first (OWT-009), so its output is visible even on a turn TCI then blocks |
193
+
194
+ | | [turn-completion-integrity](turn-completion-integrity.md) | 本標準(OWT-007–009) |
195
+ |---|---|---|
196
+ | 它在看什麼 | agent 自己最後一則訊息,看有沒有一個第一人稱承諾被說出口又被放棄 | 承載庫裡的任何項目,看有沒有沒解除條件的、沒出口的、或過門檻還沒分類的 |
197
+ | 預設狀態 | 罕見——只在那則訊息裡明確做了承諾又被丟下時才觸發 | 常見——「還有工作沒做完」幾乎永遠為真 |
198
+ | 違反時會怎樣 | 擋住回合結束,直到承諾被解決或說明卡在誰身上 | 永不阻斷。只能回報(OWT-008) |
199
+ | 為什麼行為相反 | 它在看的事件本身夠稀少,擋在它上面不會把耐性用完 | TCI 自己的規則已經寫出這裡不能做成閘門的理由:**「一個在每個回合都為真的閘門會被關掉,關掉之後它什麼都不保護」**(TCI R4)。開放工作非空幾乎永遠為真,所以這個確認點被設計成永不保留控制權 |
200
+ | 兩者掛同一事件時的順序 | — | 先回報(OWT-009),所以即使那個回合隨後被 TCI 擋下,它的輸出仍然可見 |
201
+
202
+ ---
203
+
204
+ ## Anchors: structure, not wording
205
+
206
+ To decide whether an item is waiting, unclassified, or dropped, **read the carrier's own structural field for that state** — a status column, a typed marker, a section heading — the same way [deferred-item-exit](deferred-item-exit.md)'s DEX-007 requires walking a document's structure rather than its wording to find deferred items. **OWT-010** requires the structural field to exist and be the primary source of truth.
207
+
208
+ 判定一個項目是等待中、未分類、還是已丟棄,要**讀承載庫自己描述那個狀態的結構欄位**——
209
+ 一個狀態欄、一個型別化標記、一個小節標題——與 [deferred-item-exit](deferred-item-exit.md)
210
+ 的 DEX-007 要求走訪文件結構而非措辭來找延後項目是同一個道理。**OWT-010** 要求那個結構欄位
211
+ 存在,並且是真相的主要來源。
212
+
213
+ A free-text wording scan ("contains the phrase 'waiting on'") can legitimately supplement the structural field — it catches items dropped into prose that never made it into the structured field. But it inherits the same limit [class-level-fix](class-level-fix.md) names for any enumerated list: **it is correct until the next member arrives, phrased a way the list did not anticipate.** **OWT-011** requires its coverage be declared unknown, and forbids a clean pass over it from being reported as "nothing was missed."
214
+
215
+ 一次自由文字措辭掃描(「含有『等待』這個詞」)可以正當地補充結構欄位——
216
+ 它能抓到那些寫進散文、從沒真的填進結構欄位的項目。但它繼承了 [class-level-fix](class-level-fix.md)
217
+ 對任何列舉清單指出的同一個限制:**它正確到下一個成員用清單沒預料到的寫法出現為止。**
218
+ **OWT-011** 要求它的涵蓋率明示為未知,且禁止它跑出乾淨結果就被回報成「沒有漏掉」。
219
+
220
+ ---
221
+
222
+ ## A requirement that cannot be checked is not a requirement here
223
+
224
+ **OWT-014** is a constraint on this standard's own contents, the same role [deferred-item-exit](deferred-item-exit.md)'s DEX-003 plays for that standard. Every requirement above names artefacts and a relation between them that a reader — or something a project builds — can decide. A property this standard cared about but could not phrase this way was left out of the table rather than included as an unenforceable aspiration. One example: "the capture point actually gets used" is exactly the outcome OWT-001 exists to protect, but it is a claim about human behavior over time, not a decidable relation over an artefact at a point in time — so it is stated here, in prose, as the *reason* for OWT-001, and is not itself a numbered requirement.
225
+
226
+ **OWT-014** 是對本標準自身內容的約束,與 [deferred-item-exit](deferred-item-exit.md) 的
227
+ DEX-003 扮演的角色相同。上面每一條都指名了 artefact 與它們之間可被判定的關係。
228
+ 一個本標準在意、卻無法這樣措辭的性質,會被排除在表格之外,而不是被寫成一條無法執行的期望。
229
+ 舉一例:「收件點真的有被使用」正是 OWT-001 存在要保護的結果,但那是一句關於人類長期行為的宣稱,
230
+ 不是某個時間點上 artefact 之間可判定的關係——所以它以散文形式出現在這裡,
231
+ 作為 OWT-001 存在的**理由**,而不是一條有編號的要求。
232
+
233
+ ### A check that has never been red
234
+
235
+ **OWT-015** carries [deferred-item-exit](deferred-item-exit.md)'s DEX-004 forward unchanged in substance: **a check that has never failed and a check that cannot fail produce identical output.** Until a check claimed as evidence for any requirement above has been observed reporting failure against a sample deliberately built to violate that requirement, its passing is evidence that something ran, not evidence that the requirement holds. The procedure for producing that evidence, and why it must be done per sub-requirement rather than in aggregate, is not restated here — see [class-level-fix](class-level-fix.md) and [verification-evidence](verification-evidence.md).
236
+
237
+ **OWT-015** 原封不動地延續 [deferred-item-exit](deferred-item-exit.md) 的 DEX-004:
238
+ **一支從未失敗過的檢查,與一支不可能失敗的檢查,輸出一模一樣。** 在一支被宣稱為上面
239
+ 任一要求之證據的檢查,被觀察到「對一個刻意違反該要求的樣本回報失敗」之前,
240
+ 它的通過只是「有東西跑過」的證據,不是「要求成立」的證據。產生這份證據的程序、
241
+ 以及為何必須逐條而非整體進行,此處不複述——見 [class-level-fix](class-level-fix.md)
242
+ 與 [verification-evidence](verification-evidence.md)。
243
+
244
+ ### Thresholds carry their provenance
245
+
246
+ **OWT-016** carries DEX-009 forward: a threshold with no recorded origin is a threshold nobody can evaluate changing. This standard's own two numeric thresholds are marked accordingly in [Evidence and calibration](#evidence-and-calibration) below, rather than being asserted as settled.
247
+
248
+ **OWT-016** 延續 DEX-009:一個沒有來歷的閾值,是一個沒有人能評估要不要改的閾值。
249
+ 本標準自己的兩個數字閾值在下方〈[證據與校準](#evidence-and-calibration)〉裡照此標示,而非被斷言為已定案。
250
+
251
+ ---
252
+
253
+ ## Anti-patterns
254
+
255
+ | Anti-pattern | Why it fails |
256
+ |---|---|
257
+ | A capture form with three or more required fields | Measurably stops being used; the friction it adds is paid by whoever is interrupting their own work |
258
+ | "We'll revisit this" with no release condition | Indistinguishable from forgotten; nothing brings it back |
259
+ | A hand-typed status that duplicates what git or CI already know | Two owners, one of which is never updated |
260
+ | A "last reconciled" date with no content-derived proof | Looks identical whether the content was actually re-checked or the date was just typed |
261
+ | "47 items remain" with no stated denominator | Reads as complete by default; the invisible majority is mistaken for "done" |
262
+ | A checkpoint at shell startup instead of at turn end | Drifts for exactly as long as nobody happens to open a new shell |
263
+ | A checkpoint that blocks the turn on "some work remains" | Fires on every turn; a gate that is always true gets disabled, and then protects nothing |
264
+ | Triage status read only from prose wording | Correct until an item is phrased a way the wording list did not anticipate |
265
+ | An item that silently vanishes from the carrier | Indistinguishable from a bug that lost it |
266
+
267
+ | 反模式 | 為什麼會失敗 |
268
+ |---|---|
269
+ | 三個以上必填欄位的收件表單 | 可量測地不再被使用;那份摩擦由正在打斷自己工作的人承擔 |
270
+ | 「之後再看」而沒有解除條件 | 與被忘記無法分辨;沒有東西會讓它回來 |
271
+ | 手動輸入、重複 git 或 CI 已知資訊的狀態 | 兩個擁有者,其中一個永遠不會被更新 |
272
+ | 沒有內容證明的「最後對過帳」日期 | 內容真的被重新核對過,跟日期只是被打上去,看起來一模一樣 |
273
+ | 「還有 47 項」而不寫分母 | 預設被讀成完整;看不見的大多數被誤讀成「都做完了」 |
274
+ | 確認點掛在 shell 啟動而不是回合結束 | 只要沒人剛好開新 shell,它就持續漂移 |
275
+ | 確認點擋住回合結束、理由是「還有工作沒做完」 | 每個回合都會觸發;永遠為真的閘門會被關掉,關掉之後什麼都不保護 |
276
+ | 分類狀態只靠散文措辭判讀 | 正確到某個項目用清單沒預料到的方式寫出來為止 |
277
+ | 項目從承載庫裡無聲消失 | 與一個弄丟它的 bug 無從分辨 |
278
+
279
+ ---
280
+
281
+ ## What enforces this standard
282
+
283
+ **Nothing in UDS does, and that is recorded rather than implied.** UDS states the relations a carrier of open work must satisfy; whether anything decides them is the adopting project's call, per the [writing constraint](#how-this-standard-is-written--and-why-it-is-written-that-way) above — the same boundary [deferred-item-exit](deferred-item-exit.md) draws for its own exits.
284
+
285
+ **本標準沒有任何 UDS 側的閘門,而這件事是被記錄的,不是被暗示的。** UDS 陳述一個承載開放工作的地方
286
+ 必須滿足的關係;有沒有東西去判定它,依上面的[寫法約束](#how-this-standard-is-written--and-why-it-is-written-that-way),
287
+ 是採用專案的決定——與 [deferred-item-exit](deferred-item-exit.md) 對自己出口劃的界線相同。
288
+
289
+ What this standard does do is make that call visible: OWT-014 guarantees every requirement here **can** be decided, OWT-015 fixes what it takes for a decision to count, and OWT-005/OWT-011 fix what a partial decision is allowed to print.
290
+
291
+ 本標準做的事,是讓那個決定顯形:OWT-014 保證這裡每一條**能**被判定,OWT-015 固定
292
+ 「一次判定要算數需要什麼」,OWT-005/OWT-011 固定「一次不完整的判定容許印出什麼」。
293
+
294
+ ---
295
+
296
+ ## Evidence and calibration
297
+
298
+ This standard's shape comes from one adopting project's observations made and acted on the same day the standard was drafted (XSPEC-427, 2026-09-23): a capture point, a summary script with self-test arms, and an end-of-turn hook, built and run for the first time that day. **That reference implementation is hours old at the time of writing, has one user, and has run in one repository.** It is cited here only as the origin of the requirements' shape, never as validation of the specific thresholds below.
299
+
300
+ 本標準的形狀來自一個採用專案在標準草擬**同一天**做出並實跑的觀察(XSPEC-427,2026-09-23):
301
+ 一個收件點、一支帶自測臂的摘要腳本、一個掛在回合結束的 hook,當天第一次建立並執行。
302
+ **寫下這段文字時,那個參考實作只有幾小時大、只有一個使用者、只在一個 repo 跑過。**
303
+ 它在此被引用,僅作為要求形狀的出處,**絕不作為下面具體閾值的驗證**。
304
+
305
+ - **OWT-001's "no more than two fields"** and **OWT-012's "past the declared threshold"** (illustrated at two weeks in the originating observation) are **initial judgments, not measurements** — per OWT-016. No controlled comparison exists yet between two fields and three, or between a two-week and a four-week unclassified threshold.
306
+ - Recalibrating either number against real usage, or downgrading either into project-specific guidance, is the adopting project's decision to make and to date — this standard does not carry that commitment, the same way it carries no gate.
307
+
308
+ - **OWT-001 的「不超過兩個欄位」**與**OWT-012 的「過了宣告的門檻」**(在原始觀察中以兩週為例)
309
+ 依 OWT-016 是**初始判斷,不是量測結果**——兩個欄位跟三個欄位、兩週跟四週的未分類門檻,
310
+ 目前都沒有對照比較過。
311
+ - 依實際使用情況重新校準這兩個數字、或將其中任一個降級為專案特定指引,是採用專案自己的決定
312
+ 與自己的時程——本標準不承諾這件事,如同它不附帶閘門一樣。
313
+
314
+ ---
315
+
316
+ ## Relationship to other standards
317
+
318
+ - [deferred-item-exit](deferred-item-exit.md) — the upstream half of the same shape: DEX requires that a deferred item leave its document for a traceable exit and deliberately leaves the exit's carrier unspecified. This standard picks up **after** the exit exists, requiring the carrier itself not to become the next document things get lost in.
319
+ - [turn-completion-integrity](turn-completion-integrity.md) — attaches to the same event (turn end) and is built to behave oppositely: TCI blocks on a rare, specific abandoned commitment; this standard's checkpoint (OWT-007–OWT-009) never blocks, because the condition it watches for is close to always true. See [the comparison table](#the-checkpoint-is-a-report-not-a-gate).
320
+ - [class-level-fix](class-level-fix.md) — the general form of the wording-list limit OWT-011 discloses, and the source of the non-vacuous-evidence procedure OWT-015 requires.
321
+ - [verification-evidence](verification-evidence.md) — the source of the exit-code and evidence-validity reasoning OWT-015 depends on; also where a partial-coverage exception (OWT-006, OWT-011) is registered rather than merely disclosed once.
322
+
323
+ - [deferred-item-exit](deferred-item-exit.md) — 同一個形狀的上游一半:DEX 要求延後項目離開文件、
324
+ 抵達可追蹤的出口,並刻意不規定出口的載體。本標準接手**出口存在之後**的事,
325
+ 要求那個載體自己不要變成下一份東西會不見的文件。
326
+ - [turn-completion-integrity](turn-completion-integrity.md) — 掛在同一個事件(回合結束)
327
+ 上,且被設計成行為相反:TCI 擋在一個罕見、明確的被放棄承諾上;本標準的確認點
328
+ (OWT-007–OWT-009)永不阻斷,因為它在看的條件幾乎永遠為真。見〈[對照表](#the-checkpoint-is-a-report-not-a-gate)〉。
329
+ - [class-level-fix](class-level-fix.md) — OWT-011 揭露的措辭清單限制的通則形式,
330
+ 也是 OWT-015 所要求「非空跑證據」程序的來源。
331
+ - [verification-evidence](verification-evidence.md) — OWT-015 所依賴的 exit code
332
+ 與證據有效性推理的來源;也是 OWT-006/OWT-011 的部分涵蓋例外該被登記的地方,
333
+ 而不是揭露一次就放著。
@@ -2,8 +2,8 @@
2
2
 
3
3
  > **Language**: English | [繁體中文](../locales/zh-TW/core/turn-completion-integrity.md)
4
4
 
5
- **Version**: 1.3.0
6
- **Last Updated**: 2026-09-08
5
+ **Version**: 1.4.0
6
+ **Last Updated**: 2026-09-25
7
7
  **Applicability**: Any harness where an agent ends a turn and hands control back to a human
8
8
  **Scope**: universal
9
9
  **Industry Standards**: none claimed — derived from observed failures, see Evidence
@@ -146,6 +146,38 @@ prevent, one level up.
146
146
 
147
147
  ---
148
148
 
149
+ ## Supported harnesses
150
+
151
+ The check is enforced only where a harness adapter exists and a hook is
152
+ actually wired into that harness's own config. As of v1.4.0:
153
+
154
+ | Harness | Event | Config file | Block contract |
155
+ |---|---|---|---|
156
+ | Claude Code | Stop | `.claude/settings.json` | stdout `{"decision":"block","reason":...}`, exit 0; silence allows |
157
+ | Codex | Stop | `.codex/hooks.json` | stdout `{"decision":"block","reason":...}`, exit 0 — plain text or empty stdout is documented as invalid for this event |
158
+ | Gemini CLI | AfterAgent | `.gemini/settings.json` | stdout `{"decision":"deny","reason":...}`, exit 0 — the documented preferred path over exit code 2 |
159
+
160
+ Codex's R9 exemption is best-effort, not silent failure: Codex's Stop payload
161
+ gives the assistant's final message directly but not the human's, so reading
162
+ the human side requires parsing a transcript file whose exact schema was not
163
+ confirmed against a real installation at the time this table was written. A
164
+ failed parse leaves the human side empty — detection still runs on the
165
+ assistant's message, only the R9 exemption for that one turn may be missed.
166
+
167
+ Cursor was evaluated and is not supported: whether its stop hook can actually
168
+ block a turn in the way this standard requires was unresolved as of this
169
+ writing, and shipping an adapter against an unverified contract would repeat
170
+ the exact failure R3 exists to prevent — an enforcement mechanism nobody has
171
+ confirmed enforces anything.
172
+
173
+ On any harness not in the table above, the check is inactive — the same
174
+ silent-by-default failure as an unsupported language (R8). `uds init
175
+ --with-hooks` reports which harnesses it wired a hook into; it does not
176
+ enumerate the rest here, because that list is a citation waiting to go stale
177
+ the next time a harness is added or dropped.
178
+
179
+ ---
180
+
149
181
  ## What the detector matches
150
182
 
151
183
  The shape is: **a first-person future marker, then an action verb, in the same
@@ -194,3 +226,5 @@ only because a corpus existed; the two that shipped were the ones no case covere
194
226
  - [ ] The check recognises its own block message and does not read it as the human's
195
227
  - [ ] The check recognises the itemized blocker ending R2 defines, and does not block it
196
228
  - [ ] The attribution search excludes the check's own headings and scaffolding
229
+ - [ ] Each supported harness's block contract (config file, event, output shape) is verified against that harness's own docs, not assumed from another harness
230
+ - [ ] The installer only writes a harness's hook config for a harness the adopter selected
@@ -0,0 +1,116 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * UDS Hook: Turn Completion Integrity — Codex adapter
4
+ *
5
+ * Runs at turn end. Blocks when the agent's final message states a first-person
6
+ * commitment to a next action that the turn then ended without taking.
7
+ *
8
+ * Contract (Codex Stop hook — https://developers.openai.com/codex/hooks,
9
+ * fetched 2026-09-25, cross-checked against the same content served at
10
+ * https://learn.chatgpt.com/docs/hooks):
11
+ * config — <repo>/.codex/hooks.json (or ~/.codex/hooks.json), a dedicated
12
+ * file, NOT config.toml's [hooks] table. Codex runs matching
13
+ * hooks from every location that defines them, so this adapter
14
+ * deliberately touches only hooks.json.
15
+ * stdin — JSON with session_id, transcript_path (string|null), cwd,
16
+ * hook_event_name, turn_id, stop_hook_active, permission_mode,
17
+ * and last_assistant_message (string|null) — the agent's final
18
+ * text for this turn, given directly, no transcript parse needed.
19
+ * output — "Stop expects JSON on stdout when it exits 0. Plain text output
20
+ * is invalid for this event." Block: {"decision":"block","reason":
21
+ * "..."}. Allow: any other JSON (this adapter always writes "{}"
22
+ * so stdout is valid JSON on every path, including R5 failures).
23
+ *
24
+ * R9 (exempt a human-directed stop) is best-effort here. Codex's Stop payload
25
+ * gives the assistant's last message directly but not the human's; this
26
+ * adapter tries transcript_path for it (rollout.jsonl), tolerantly, because
27
+ * its exact schema was not confirmed against a real Codex install at
28
+ * authoring time. If that read fails or the field is null, the user side is
29
+ * treated as empty — which means R9 cannot exempt that turn, not that the
30
+ * check goes silent (last_assistant_message still drives detection). This is
31
+ * the documented gap in core/turn-completion-integrity.md's Codex section.
32
+ *
33
+ * Judgement (packs, cooldown, rolling window, self-echo) lives in
34
+ * turn-completion/engine.mjs and is shared with every other adapter; this
35
+ * file only reads Codex's stdin shape and writes Codex's output shape.
36
+ *
37
+ * Usage: node check-turn-completion-codex.mjs (reads stdin)
38
+ * node check-turn-completion-codex.mjs --self-test
39
+ * node check-turn-completion-codex.mjs --languages
40
+ *
41
+ * @see core/turn-completion-integrity.md
42
+ */
43
+ import { readFileSync } from 'node:fs';
44
+ import { decide, isSelfEcho, runSelfTest, printLanguages } from './turn-completion/engine.mjs';
45
+
46
+ /**
47
+ * Best-effort extraction of the human's last message from a Codex transcript.
48
+ * Tolerant of several plausible JSONL shapes because the exact rollout.jsonl
49
+ * schema was not confirmed against a real install; any failure returns '',
50
+ * which is the same as "cannot tell" (R5) — it does not stop
51
+ * last_assistant_message from still being checked.
52
+ */
53
+ function bestEffortLastUserMessage(transcriptPath) {
54
+ if (typeof transcriptPath !== 'string' || !transcriptPath) return '';
55
+ try {
56
+ let user = '';
57
+ for (const line of readFileSync(transcriptPath, 'utf8').split('\n')) {
58
+ if (!line.trim()) continue;
59
+ let ev;
60
+ try { ev = JSON.parse(line); } catch { continue; }
61
+ // Try a few plausible shapes rather than committing to one: a nested
62
+ // { message: { role, content } } (Claude-Code-style rollout entry), or
63
+ // a flat { role, content }.
64
+ const msg = (ev && ev.message) || ev;
65
+ if (!msg || msg.role !== 'user') continue;
66
+ const c = msg.content;
67
+ const text = typeof c === 'string'
68
+ ? c
69
+ : Array.isArray(c)
70
+ ? c.filter((p) => p && (p.type === 'text' || typeof p.text === 'string')).map((p) => p.text || '').join('\n')
71
+ : '';
72
+ if (text && text.trim() && !isSelfEcho(text)) user = text;
73
+ }
74
+ return user;
75
+ } catch {
76
+ return '';
77
+ }
78
+ }
79
+
80
+ async function main() {
81
+ let out = {};
82
+ try {
83
+ const raw = readFileSync(0, 'utf8');
84
+ const data = JSON.parse(raw);
85
+ if (data && typeof data === 'object' && data.stop_hook_active !== true) {
86
+ const assistantText = typeof data.last_assistant_message === 'string' ? data.last_assistant_message : '';
87
+ const userText = bestEffortLastUserMessage(data.transcript_path);
88
+ const verdict = await decide({
89
+ sessionId: data.session_id,
90
+ assistantText,
91
+ userText,
92
+ });
93
+ if (verdict.fire) out = { decision: 'block', reason: verdict.reason };
94
+ }
95
+ } catch {
96
+ /* R5: fail open — stdout stays valid JSON, turn ends */
97
+ }
98
+ process.stdout.write(JSON.stringify(out));
99
+ }
100
+
101
+ const arg = process.argv[2];
102
+ if (arg === '--self-test') {
103
+ process.exit((await runSelfTest('turn-completion-codex')) ? 0 : 1);
104
+ } else if (arg === '--languages') {
105
+ await printLanguages();
106
+ } else {
107
+ try {
108
+ await main();
109
+ } catch {
110
+ // R5, belt and braces: main() already guards its own body, but a Codex
111
+ // Stop hook must exit 0 with valid JSON on stdout no matter what, and an
112
+ // uncaught throw here would instead crash with a non-zero exit and a
113
+ // stack trace on stderr.
114
+ process.stdout.write('{}');
115
+ }
116
+ }
@@ -0,0 +1,81 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * UDS Hook: Turn Completion Integrity — Gemini CLI adapter
4
+ *
5
+ * Runs once per turn, after the model's final response. Blocks when that
6
+ * response states a first-person commitment to a next action that the turn
7
+ * then ended without taking.
8
+ *
9
+ * Contract (Gemini CLI AfterAgent hook — https://geminicli.com/docs/hooks/reference/
10
+ * and https://geminicli.com/docs/hooks/, fetched 2026-09-25):
11
+ * config — .gemini/settings.json (project) / ~/.gemini/settings.json
12
+ * (user) / /etc/gemini-cli/settings.json (system), under a
13
+ * "hooks" key shared with the rest of Gemini CLI's settings.
14
+ * Project settings take precedence over user, which take
15
+ * precedence over system. AfterAgent does not use matchers
16
+ * (matchers apply only to Tool hooks).
17
+ * stdin — JSON with session_id, transcript_path, cwd, hook_event_name,
18
+ * timestamp, prompt, prompt_response, stop_hook_active.
19
+ * `prompt` is the human message that started THIS turn and
20
+ * `prompt_response` is the agent's final text for it — both given
21
+ * directly, so R9 (exempt a human-directed stop) needs no
22
+ * transcript parsing here, unlike the Codex adapter.
23
+ * output — exit codes are shared across every Gemini CLI hook type: "0:
24
+ * stdout is parsed as JSON. Preferred for all logic. 2: System
25
+ * Block, stderr is the rejection reason." AfterAgent's own JSON
26
+ * shape: {"decision":"deny","reason":"..."} rejects the response
27
+ * and forces a retry, with reason "sent to the agent as a new
28
+ * prompt". This adapter uses the preferred exit-0 + JSON path,
29
+ * not exit code 2, and always writes valid JSON on stdout
30
+ * (including on every R5 failure path) so exit 0 is never
31
+ * ambiguous with "nothing to say".
32
+ *
33
+ * Judgement (packs, cooldown, rolling window, self-echo) lives in
34
+ * turn-completion/engine.mjs and is shared with every other adapter; this
35
+ * file only reads Gemini CLI's stdin shape and writes Gemini CLI's output
36
+ * shape.
37
+ *
38
+ * Usage: node check-turn-completion-gemini.mjs (reads stdin)
39
+ * node check-turn-completion-gemini.mjs --self-test
40
+ * node check-turn-completion-gemini.mjs --languages
41
+ *
42
+ * @see core/turn-completion-integrity.md
43
+ */
44
+ import { readFileSync } from 'node:fs';
45
+ import { decide, runSelfTest, printLanguages } from './turn-completion/engine.mjs';
46
+
47
+ async function main() {
48
+ let out = {};
49
+ try {
50
+ const raw = readFileSync(0, 'utf8');
51
+ const data = JSON.parse(raw);
52
+ if (data && typeof data === 'object' && data.stop_hook_active !== true) {
53
+ const assistantText = typeof data.prompt_response === 'string' ? data.prompt_response : '';
54
+ const userText = typeof data.prompt === 'string' ? data.prompt : '';
55
+ const verdict = await decide({
56
+ sessionId: data.session_id,
57
+ assistantText,
58
+ userText,
59
+ });
60
+ if (verdict.fire) out = { decision: 'deny', reason: verdict.reason };
61
+ }
62
+ } catch {
63
+ /* R5: fail open — stdout stays valid JSON, turn ends */
64
+ }
65
+ process.stdout.write(JSON.stringify(out));
66
+ }
67
+
68
+ const arg = process.argv[2];
69
+ if (arg === '--self-test') {
70
+ process.exit((await runSelfTest('turn-completion-gemini')) ? 0 : 1);
71
+ } else if (arg === '--languages') {
72
+ await printLanguages();
73
+ } else {
74
+ try {
75
+ await main();
76
+ } catch {
77
+ // R5, belt and braces — see check-turn-completion-codex.mjs for why this
78
+ // outer guard exists alongside main()'s own.
79
+ process.stdout.write('{}');
80
+ }
81
+ }