gentle-pi 3.6.0 → 4.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (320) hide show
  1. package/README.md +37 -6
  2. package/assets/agents/gentle-ai-explore.md +4 -4
  3. package/assets/agents/gentle-ai-verify.md +6 -4
  4. package/assets/agents/gentle-ai-worker.md +7 -9
  5. package/assets/orchestrator-delegation.md +40 -41
  6. package/assets/orchestrator-memory.md +1 -22
  7. package/assets/orchestrator-skills.md +1 -1
  8. package/assets/orchestrator.md +12 -24
  9. package/assets/support/strict-tdd-verify.md +4 -266
  10. package/assets/support/strict-tdd.md +8 -360
  11. package/bin/gentle-shell.mjs +254 -29
  12. package/docs/delegated-verification.md +26 -1
  13. package/docs/gentle-agents-activity.md +24 -0
  14. package/docs/gentle-shell.md +95 -30
  15. package/docs/native-authority-architecture.md +2 -2
  16. package/docs/prompt-history.md +280 -0
  17. package/docs/readme-reference.md +117 -235
  18. package/docs/telemetry.md +1 -1
  19. package/docs/yolo-mode.md +86 -0
  20. package/extensions/child-context.ts +26 -0
  21. package/extensions/child-safety.ts +23 -0
  22. package/extensions/gentle-agents.ts +509 -275
  23. package/extensions/gentle-ai.ts +781 -683
  24. package/extensions/gentle-shell.ts +1375 -88
  25. package/extensions/gentle-stats.ts +101 -0
  26. package/extensions/gentle-todo.ts +18 -7
  27. package/extensions/history/atomic-write.ts +38 -0
  28. package/extensions/history/hide-prompts.ts +183 -0
  29. package/extensions/history/index.ts +1419 -0
  30. package/extensions/history/load-shared-history.ts +39 -0
  31. package/extensions/history/selector-helpers.ts +538 -0
  32. package/extensions/history/session-scan.ts +233 -0
  33. package/extensions/history/store.ts +1119 -0
  34. package/extensions/nan-provider.ts +6 -0
  35. package/extensions/quiet-tools.ts +179 -87
  36. package/extensions/resume-hint.ts +60 -0
  37. package/extensions/skill-registry.ts +16 -12
  38. package/extensions/startup-banner.ts +60 -31
  39. package/lib/agent-assets.ts +604 -0
  40. package/lib/agent-profile-pin.ts +12 -0
  41. package/lib/agents-message-delivery.ts +181 -0
  42. package/lib/agents-protocol.ts +60 -0
  43. package/lib/agents-runner.ts +105 -98
  44. package/lib/agents-view.ts +26 -3
  45. package/lib/agents-widget.ts +16 -9
  46. package/lib/append-system-prompt.ts +21 -0
  47. package/lib/bounded-writer-admission.ts +147 -0
  48. package/lib/card-style-policy.ts +60 -0
  49. package/lib/child-context-files.ts +166 -0
  50. package/lib/codemode-renderer.ts +185 -0
  51. package/lib/command-palette-catalog.ts +3 -9
  52. package/lib/command-palette.ts +25 -14
  53. package/lib/destructive-command-guard.ts +144 -0
  54. package/lib/foreign-target-grants.ts +32 -0
  55. package/lib/gentle-ai-elapsed-store.ts +87 -0
  56. package/lib/gentle-ai-renderer.ts +196 -39
  57. package/lib/gentle-shell-launcher.ts +24 -13
  58. package/lib/gentle-shell-resume-hint.ts +176 -0
  59. package/lib/history-capture-policy.ts +95 -0
  60. package/lib/model-routing-authority.ts +5 -1
  61. package/lib/nan-provider.ts +227 -0
  62. package/lib/native-review-cli.ts +53 -108
  63. package/lib/odd-phase-inference.ts +231 -0
  64. package/lib/odd-phase.ts +141 -0
  65. package/lib/overlay-repaint.ts +26 -0
  66. package/lib/pi-tui-keys.ts +53 -0
  67. package/lib/review-candidate-view-owner.ts +67 -17
  68. package/lib/review-candidate-view.ts +112 -25
  69. package/lib/review-reminder-receipt.ts +48 -8
  70. package/lib/review-risk-assessment.ts +156 -11
  71. package/lib/review-sidebar-state.ts +223 -0
  72. package/lib/selection-engine.ts +515 -0
  73. package/lib/session-change-capture.ts +10 -1
  74. package/lib/session-messaging-grants.ts +135 -0
  75. package/lib/session-worktree-registry.ts +14 -2
  76. package/lib/shell-bar.ts +179 -24
  77. package/lib/shell-card.ts +287 -18
  78. package/lib/shell-changes-view.ts +2 -1
  79. package/lib/shell-prompt.ts +102 -6
  80. package/lib/shell-sidebar-layout.ts +78 -14
  81. package/lib/shell-sidebar.ts +15 -1
  82. package/lib/shell-todo.ts +23 -13
  83. package/lib/shell-usage-view.ts +9 -4
  84. package/lib/shell-usage.ts +66 -10
  85. package/lib/stats-collector.ts +381 -0
  86. package/lib/stats-view.ts +431 -0
  87. package/lib/theme-customization.ts +52 -0
  88. package/lib/vim-editor-adapter.ts +379 -0
  89. package/lib/vim-normal-engine.ts +154 -0
  90. package/lib/vim-operator-engine.ts +416 -0
  91. package/lib/vim-policy.ts +49 -0
  92. package/lib/vim-visual-engine.ts +107 -0
  93. package/lib/visual-customization-policy.ts +108 -0
  94. package/lib/visual-customize-view.ts +330 -0
  95. package/lib/visual-profiles.ts +228 -0
  96. package/lib/yolo-session-policy.ts +240 -0
  97. package/package.json +20 -8
  98. package/runtime/gentle-shell-launcher.mjs +23 -12
  99. package/runtime/gentle-shell-resume-hint.mjs +177 -0
  100. package/runtime/native-review-cli.mjs +53 -108
  101. package/runtime/review-risk-assessment.mjs +154 -9
  102. package/scripts/build-runtime-modules.mjs +1 -0
  103. package/scripts/gentle-ai-installer.mjs +14 -13
  104. package/scripts/mirror-odd-routing.mjs +2 -2
  105. package/scripts/run-test-suite.mjs +76 -0
  106. package/scripts/test-packed-runner.mjs +31 -14
  107. package/scripts/verify-package-files.mjs +11 -22
  108. package/skills/branch-pr/SKILL.md +24 -52
  109. package/skills/chained-pr/SKILL.md +31 -15
  110. package/skills/chained-pr/references/chaining-details.md +31 -20
  111. package/skills/gentle-ai/SKILL.md +9 -15
  112. package/skills/issue-creation/SKILL.md +8 -2
  113. package/skills/issue-creation/references/delegated-workflow-actions.md +19 -0
  114. package/skills/work-unit-commits/SKILL.md +4 -3
  115. package/tests/agent-profiles.test.ts +18 -0
  116. package/tests/agents-fake-child.ts +2 -2
  117. package/tests/agents-message-delivery.test.ts +106 -0
  118. package/tests/agents-protocol.test.ts +40 -0
  119. package/tests/agents-runner.test.ts +400 -89
  120. package/tests/agents-view-thread-identity.test.ts +169 -0
  121. package/tests/agents-view.test.ts +8 -2
  122. package/tests/agents-widget.test.ts +154 -15
  123. package/tests/append-system-prompt-route.test.ts +160 -0
  124. package/tests/append-system-prompt.test.ts +46 -0
  125. package/tests/artifact-language.test.ts +19 -213
  126. package/tests/ask-user-question.test.ts +44 -1
  127. package/tests/asset-installation-runtime.test.ts +5 -16
  128. package/tests/autonomous-guard.test.ts +69 -1
  129. package/tests/bounded-writer-admission.test.ts +95 -0
  130. package/tests/branch-pr-skill.test.ts +43 -0
  131. package/tests/card-style-policy.test.ts +55 -0
  132. package/tests/chained-pr-skill.test.ts +124 -0
  133. package/tests/child-context-files.test.ts +255 -0
  134. package/tests/child-safety.test.ts +82 -0
  135. package/tests/codemode-rendering.test.ts +491 -0
  136. package/tests/command-palette.test.ts +39 -3
  137. package/tests/delegated-key-learnings-contract.test.ts +0 -76
  138. package/tests/destructive-command-guard.test.ts +84 -0
  139. package/tests/devbinary/native-review-parity.devtest.ts +170 -2
  140. package/tests/devbinary/non-git-subagent-bootstrap.devtest.ts +193 -0
  141. package/tests/fixtures/stats/sessions/--work-alpha--/2026-09-28T10-00-00-000Z_aaa.jsonl +7 -0
  142. package/tests/fixtures/stats/sessions/--work-alpha--/2026-09-29T23-00-00-000Z_bbb.jsonl +3 -0
  143. package/tests/fixtures/stats/sessions/--work-alpha--/2026-09-30T08-00-00-000Z_ddd.jsonl +2 -0
  144. package/tests/fixtures/stats/sessions/--work-alpha--/2026-09-30T09-00-00-000Z_eee.jsonl +3 -0
  145. package/tests/fixtures/stats/sessions/--work-alpha--/run-1/session.jsonl +2 -0
  146. package/tests/fixtures/stats/sessions/--work-beta--/2026-09-01T12-00-00-000Z_ccc.jsonl +2 -0
  147. package/tests/fixtures/stats/user-pi/sessions/--work-alpha--/2026-09-28T10-00-00-000Z_aaa.jsonl +2 -0
  148. package/tests/fixtures/stats/user-pi/sessions/--work-alpha--/2026-09-29T23-00-00-000Z_bbb.jsonl +4 -0
  149. package/tests/fixtures/stats/user-pi/sessions/--work-gamma--/2026-09-20T09-00-00-000Z_fff.jsonl +3 -0
  150. package/tests/foreign-target-grants.test.ts +58 -0
  151. package/tests/generic-agent-tools.test.ts +54 -0
  152. package/tests/gentle-agents.test.ts +1411 -256
  153. package/tests/gentle-ai-binary.test.ts +3 -3
  154. package/tests/gentle-ai-elapsed-store.test.ts +68 -0
  155. package/tests/gentle-ai-installer.test.ts +68 -54
  156. package/tests/gentle-ai-renderer.test.ts +487 -8
  157. package/tests/gentle-ai.test.ts +288 -65
  158. package/tests/gentle-card-text.ts +2 -1
  159. package/tests/gentle-shell-bin.test.ts +651 -114
  160. package/tests/gentle-shell-launcher.test.ts +99 -52
  161. package/tests/gentle-shell-resume-hint.test.ts +270 -0
  162. package/tests/gentle-shell.test.ts +3839 -187
  163. package/tests/gentle-stats.test.ts +152 -0
  164. package/tests/gentle-theme.test.ts +4 -1
  165. package/tests/gentle-todo.test.ts +80 -8
  166. package/tests/history-atomic-write.test.ts +57 -0
  167. package/tests/history-capture-policy.test.ts +102 -0
  168. package/tests/history-command-registration.test.ts +164 -0
  169. package/tests/history-dedupe-entries.test.ts +123 -0
  170. package/tests/history-delete-backfill.test.ts +190 -0
  171. package/tests/history-delete-confirm.test.ts +460 -0
  172. package/tests/history-dispatch.test.ts +180 -0
  173. package/tests/history-drain-hidden.test.ts +110 -0
  174. package/tests/history-drain-order.test.ts +98 -0
  175. package/tests/history-expanded-globals.test.ts +62 -0
  176. package/tests/history-gc.test.ts +832 -0
  177. package/tests/history-header-layout.test.ts +265 -0
  178. package/tests/history-hide-prompts.test.ts +275 -0
  179. package/tests/history-lazy-windowing.test.ts +508 -0
  180. package/tests/history-legacy-migrate-v2.test.ts +297 -0
  181. package/tests/history-load-shared-history.test.ts +53 -0
  182. package/tests/history-max-results-cap.test.ts +76 -0
  183. package/tests/history-multi-reader.test.ts +203 -0
  184. package/tests/history-off-path.test.ts +170 -0
  185. package/tests/history-openflow-integration.test.ts +173 -0
  186. package/tests/history-overlay-margin.test.ts +326 -0
  187. package/tests/history-preview-layout.test.ts +93 -0
  188. package/tests/history-registry.test.ts +143 -0
  189. package/tests/history-scope-delete.test.ts +411 -0
  190. package/tests/history-search-caret-keys.test.ts +142 -0
  191. package/tests/history-seed-bootstrap.test.ts +170 -0
  192. package/tests/history-seed-regen.test.ts +129 -0
  193. package/tests/history-selector-windowing.test.ts +94 -0
  194. package/tests/history-session-scan-directory.test.ts +87 -0
  195. package/tests/history-session-scan-extract.test.ts +583 -0
  196. package/tests/history-session-writer.test.ts +351 -0
  197. package/tests/history-store-paths.test.ts +79 -0
  198. package/tests/history-tombstone-exact.test.ts +139 -0
  199. package/tests/history-wheel-mouse.test.ts +242 -0
  200. package/tests/inprocess-reviewer.test.ts +29 -19
  201. package/tests/issue-creation-skill.test.ts +61 -0
  202. package/tests/model-routing-authority.test.ts +16 -0
  203. package/tests/nan-provider.test.ts +471 -0
  204. package/tests/native-review-capability-contract.test.ts +13 -1
  205. package/tests/native-review-cli.test.ts +6 -120
  206. package/tests/native-review-parity-runtime.test.ts +100 -3
  207. package/tests/odd-integration.test.ts +67 -0
  208. package/tests/odd-phase-inference.test.ts +213 -0
  209. package/tests/odd-phase-loader.test.ts +253 -0
  210. package/tests/odd-phase.test.ts +307 -0
  211. package/tests/odd-routing-canonical-ratchet.test.ts +11 -6
  212. package/tests/odd-routing-contract.test.ts +86 -35
  213. package/tests/orchestrator-budget.test.ts +14 -39
  214. package/tests/orchestrator-rdd-ownership.test.ts +3 -3
  215. package/tests/overlay-repaint.test.ts +74 -0
  216. package/tests/package-manifest.test.ts +252 -115
  217. package/tests/packed-runner-owned-path.test.ts +46 -0
  218. package/tests/persona-single-channel.test.ts +6 -6
  219. package/tests/provider-defect-handoff.test.ts +3 -11
  220. package/tests/quiet-bash-runtime.test.ts +76 -0
  221. package/tests/quiet-tool-rendering.test.ts +409 -184
  222. package/tests/rdd-aware-verification-contract.test.ts +76 -1
  223. package/tests/rdd-status-line.test.ts +9 -4
  224. package/tests/resume-hint-extension.test.ts +122 -0
  225. package/tests/review-agent-end-preflight.test.ts +176 -12
  226. package/tests/review-candidate-owner-retry.test.ts +22 -1
  227. package/tests/review-candidate-view.test.ts +298 -0
  228. package/tests/review-contract-prompt.test.ts +108 -43
  229. package/tests/review-controller-lock-status.test.ts +0 -1
  230. package/tests/review-controller-native-routing.test.ts +611 -5
  231. package/tests/review-controller-workspace-root.test.ts +163 -4
  232. package/tests/review-host-relay-routing.test.ts +338 -2
  233. package/tests/review-integration-v2-forward.test.ts +200 -0
  234. package/tests/review-ledger-contract.test.ts +10 -34
  235. package/tests/review-reminder-receipt.test.ts +47 -1
  236. package/tests/review-risk-assessment.test.ts +498 -6
  237. package/tests/review-sidebar-state.test.ts +402 -0
  238. package/tests/run-test-suite.test.ts +124 -0
  239. package/tests/runtime-harness.mjs +145 -786
  240. package/tests/runtime-metrics-children.test.ts +16 -23
  241. package/tests/selection-engine.test.ts +421 -0
  242. package/tests/session-change-capture.test.ts +12 -1
  243. package/tests/session-messaging-grants.test.ts +255 -0
  244. package/tests/session-worktree-registry.test.ts +77 -0
  245. package/tests/shell-bar.test.ts +382 -1
  246. package/tests/shell-card.test.ts +353 -1
  247. package/tests/shell-changes-view.test.ts +52 -0
  248. package/tests/shell-prompt.test.ts +94 -2
  249. package/tests/shell-sidebar-layout.test.ts +325 -21
  250. package/tests/shell-sidebar-scroll-benchmark.test.ts +255 -0
  251. package/tests/shell-todo.test.ts +87 -1
  252. package/tests/shell-usage-view.test.ts +27 -0
  253. package/tests/shell-usage.test.ts +73 -0
  254. package/tests/skill-registry.test.ts +50 -1
  255. package/tests/startup-banner.test.ts +130 -2
  256. package/tests/stats-collector.test.ts +195 -0
  257. package/tests/stats-view.test.ts +202 -0
  258. package/tests/telemetry-trigger.test.ts +81 -20
  259. package/tests/theme-customization.test.ts +72 -0
  260. package/tests/vim-editor-adapter-host-resolution.test.ts +37 -0
  261. package/tests/vim-editor-adapter.test.ts +804 -0
  262. package/tests/vim-normal-engine.test.ts +101 -0
  263. package/tests/vim-operator-engine.test.ts +215 -0
  264. package/tests/vim-policy.test.ts +19 -0
  265. package/tests/vim-visual-engine.test.ts +52 -0
  266. package/tests/visual-customization-policy.test.ts +110 -0
  267. package/tests/visual-customize-view.test.ts +418 -0
  268. package/tests/visual-profiles.test.ts +87 -0
  269. package/tests/yolo-customize.test.ts +256 -0
  270. package/tests/yolo-mode-runtime.test.ts +161 -0
  271. package/tests/yolo-mode.test.ts +261 -0
  272. package/tests/yolo-session-policy.test.ts +59 -0
  273. package/themes/Gentle.json +2 -1
  274. package/themes/Gentleman-Cute.json +2 -1
  275. package/themes/Gentleman-Sexy.json +2 -1
  276. package/assets/agents/sdd-apply.md +0 -159
  277. package/assets/agents/sdd-archive.md +0 -228
  278. package/assets/agents/sdd-design.md +0 -49
  279. package/assets/agents/sdd-explore.md +0 -48
  280. package/assets/agents/sdd-init.md +0 -56
  281. package/assets/agents/sdd-onboard.md +0 -52
  282. package/assets/agents/sdd-proposal.md +0 -64
  283. package/assets/agents/sdd-remediate.md +0 -37
  284. package/assets/agents/sdd-research.md +0 -49
  285. package/assets/agents/sdd-spec.md +0 -192
  286. package/assets/agents/sdd-status.md +0 -54
  287. package/assets/agents/sdd-tasks.md +0 -108
  288. package/assets/agents/sdd-verify.md +0 -124
  289. package/assets/chains/sdd-full.chain.md +0 -83
  290. package/assets/chains/sdd-plan.chain.md +0 -56
  291. package/assets/chains/sdd-verify.chain.md +0 -43
  292. package/assets/sdd-orchestrator-workflow.md +0 -319
  293. package/assets/support/sdd-status-contract.md +0 -77
  294. package/docs/assets/diagrams/sdd-cycle.svg +0 -14
  295. package/extensions/sdd-init.ts +0 -816
  296. package/lib/openspec-deltas.ts +0 -156
  297. package/lib/sdd-preflight.ts +0 -1066
  298. package/lib/sdd-research-capabilities.ts +0 -94
  299. package/lib/sdd-status.ts +0 -26
  300. package/tests/fixtures/legacy/sdd-research-v2.5.0.md +0 -54
  301. package/tests/fixtures/native-review-cli/v2.1.3/bind-sdd.json +0 -25
  302. package/tests/fixtures/v0.10.7/assets/agents/sdd-apply.md +0 -132
  303. package/tests/openspec-deltas.test.ts +0 -209
  304. package/tests/sdd-agent-tools.test.ts +0 -156
  305. package/tests/sdd-archive-replay.test.ts +0 -82
  306. package/tests/sdd-classical-continuation.test.ts +0 -74
  307. package/tests/sdd-execution-routing-contract.test.ts +0 -44
  308. package/tests/sdd-managed-runtime-settlement.test.ts +0 -155
  309. package/tests/sdd-native-managed-uptake.test.ts +0 -243
  310. package/tests/sdd-no-attempts-contract.test.ts +0 -15
  311. package/tests/sdd-odd-integration.test.ts +0 -33
  312. package/tests/sdd-optional-research.test.ts +0 -124
  313. package/tests/sdd-planning-routing-contract.test.ts +0 -45
  314. package/tests/sdd-preflight-rpc-input.test.ts +0 -125
  315. package/tests/sdd-preflight.test.ts +0 -541
  316. package/tests/sdd-research-capabilities.test.ts +0 -114
  317. package/tests/sdd-research-live.test.ts +0 -241
  318. package/tests/sdd-selection-transport.test.ts +0 -653
  319. package/tests/sdd-status.test.ts +0 -9
  320. package/tests/sdd-task-truth.test.ts +0 -43
@@ -1,6 +1,6 @@
1
1
  # el Gentleman Orchestrator
2
2
 
3
- Bind this to the parent Pi session only. Do not apply it to SDD executor phase agents.
3
+ Bind this to the parent Pi session only; subagents receive bounded task instructions.
4
4
 
5
5
  ## Identity Contract
6
6
 
@@ -18,11 +18,11 @@ Keep synthesis short by default: decision, outcome, next action. Expand only whe
18
18
 
19
19
  Reply-language style and the active persona's Spanish variant are defined once in the identity/harness section above (its `Current persona mode:` line). The rules below are delegation/artifact-scoped and not restated there:
20
20
 
21
- Generated technical artifacts — whether by the parent inline or by subagents — (code, code comments, UI copy, identifiers, commit messages, filenames, PR descriptions, tests, fixtures, SDD/OpenSpec files, delegated phase outputs, and repository-facing documentation) default to English, regardless of the user's conversation language or active persona. Override only when the user explicitly requests another language for that artifact, or when extending a project whose existing convention is non-English.
21
+ Generated technical artifacts — whether by the parent inline or by subagents — (code, code comments, UI copy, identifiers, commit messages, filenames, PR descriptions, tests, fixtures, delegated outputs, and repository-facing documentation) default to English, regardless of the user's conversation language or active persona. Override only when the user explicitly requests another language for that artifact, or when extending a project whose existing convention is non-English.
22
22
 
23
23
  Public/contextual comments and replies are different from technical artifacts. When using `comment-writer` or drafting a human-facing GitHub, PR review, Slack, Discord, or async comment, write in the target context language by default. Spanish issue/thread -> Spanish comment. English thread -> English comment. Mixed context -> target message language. Explicit user language or tone override wins. Spanish comments default to neutral/professional Spanish unless the user or target context clearly calls for regional tone.
24
24
 
25
- Subagent-facing English delegation and the quote/UI/SDD-artifact exceptions: `orchestrator-delegation.md`.
25
+ Subagent-facing English delegation and quote/UI exceptions: `orchestrator-delegation.md`.
26
26
 
27
27
  ## Mental Model
28
28
 
@@ -30,20 +30,18 @@ el Gentleman is an ecosystem configurator and harness layer. After installation,
30
30
 
31
31
  - Small request: do it directly.
32
32
  - Substantial authorized work: use ODD; track feature progress automatically.
33
- - User explicitly asks to use SDD: run the SDD flow.
34
33
  - Parent session orchestrates; phase agents execute.
35
34
 
36
35
  Delegation is not optional once complexity appears. If a task crosses the triggers below, use the smallest useful subagent workflow instead of continuing as a monolithic executor.
37
36
 
38
37
  ## Work Routing Ladder
39
38
 
40
- Route work through the smallest harness that is safe. Three tiers:
39
+ Route ODD work through the smallest safe harness:
41
40
 
42
- 1. **Inline Direct** — small, mechanical, parent has context (typo, one-file edit, read-only check of 1-3 known files, bash for state). No SDD ceremony; stop when it is no longer small.
43
- 2. **Simple Delegation** — generic non-SDD exploration → `gentle-ai-explore`; bounded implementation → `gentle-ai-worker`; command-running generic non-SDD verification → `gentle-ai-verify`. Try its package role; if missing/unusable, use native `Agent` under the same read-only mapping/verification constraints and report fallback. SDD roles stay inside SDD.
44
- 3. **SDD (optional)** — only by explicit request or accepted proposal, never size, file count, or risk. Resolve organic ambiguity with optional research, not SDD. Selected SDD commands and approval gates: `sdd-orchestrator-workflow.md`.
41
+ 1. **Inline Direct** — small, mechanical, parent has context (typo, one-file edit, read-only check within the evidence budget, bash for state); stop when it is no longer small.
42
+ 2. **Simple Delegation** — exploration → `gentle-ai-explore`; bounded implementation → `gentle-ai-worker`; command-running verification → `gentle-ai-verify`. Try its package role; if missing/unusable, use native `Agent` under the same read-only mapping/verification constraints and report fallback.
45
43
 
46
- ODD (Default Workflow, harness section above) is mandatory on every request; detail: `orchestrator-delegation.md`, `orchestrator-memory.md`.
44
+ ODD (Default Workflow, harness section above) is mandatory on every request; detail: `orchestrator-delegation.md`, `orchestrator-memory.md`. For behavior changes with applicable runnable deterministic tests and a clear expected outcome, use test-first by default: observed RED, GREEN, then refactor with checks. For passive documentation, non-testable changes, an unavailable runner or no meaningful RED, state why and run proportionate ordinary functional or structural verification instead. Test presence alone is not applicability; no chat or TUI toggle activates this policy.
47
45
 
48
46
  ## Delegation Rules
49
47
 
@@ -53,33 +51,23 @@ Before launching bounded writer (`gentle-ai-worker` or `worker`), task/context n
53
51
 
54
52
  Mandatory Delegation Triggers — once fired, delegate through the best available runtime (prefer `subagent_run`, else native `Agent`):
55
53
 
56
- 1. **4-file rule** — 4+ files to understand → delegate a scout/mapping task.
54
+ 1. **Evidence-budget rule** — read inline only if evidence fits one parallel batch (at most 3 calls, ~10k tokens; grep and line ranges, never whole large files). Larger reads, >~5 sequential lookups, or a long session ahead → one scout/explorer returning at most ~2k tokens with `path:line` evidence; re-read nothing it covered beyond one spot check. Never force delegation for a small targeted question.
57
55
  2. **Multi-file write rule** — 2+ non-trivial files touched → delegate one writer.
58
56
  3. **Incident rule** — diagnose wrong cwd/worktree/git/tooling incidents separately before resuming work.
59
- 4. **Long-session rule** — ~20 tool calls, 5 exploratory reads, or 2 non-mechanical edits without delegation → pause and delegate.
60
- 5. **Verification rule** — executing/delegating verification commands → `gentle-ai-verify`; only the 1-3-file read-only check stays inline.
57
+ 4. **Context backstop** — parent context past ~150k tokens → pause and delegate the next bounded unit of work. Keep command output bounded (counts, `--stat`, `tail`); full suites and builds go to a verifier.
58
+ 5. **Verification rule** — executing/delegating verification commands → `gentle-ai-verify`; only a read-only check within the evidence budget stays inline.
61
59
 
62
60
  {{GENTLE_PI_BACKGROUND_POLICY}}; rules: the background-subagents block in the delegation contract.
63
61
 
64
62
  Per-action table, Work Routing Ladder examples, Cost and Context Balance, Canonical Workflows, and the mirrored gentle-ai canon (blocking-prompt relays, language, delegation): `orchestrator-delegation.md`.
65
63
 
66
- ## SDD Workflow (lazy-loaded)
67
-
68
- The detailed SDD workflow is intentionally not embedded in this always-on parent prompt. Before handling any `/sdd-*` command, natural-language SDD request, SDD continuation/routing, apply/verify/archive work, or SDD/Judgment-Day phase delegation, read this package asset first:
69
-
70
- `sdd-orchestrator-workflow.md`
71
-
72
- That lazy surface contains the SDD phases, native dispatcher rules, status contract, preflight/init guards, artifact-store policy, execution mode, Strict TDD forwarding, phase result contract, and review workload guard.
73
-
74
- Hard preflight invariant: `openspec/config.yaml`, existing SDD changes, installed `.pi`/global SDD assets, or a todo named "preflight" are not session preflight. Do not mark SDD preflight complete, start `sdd-init`, launch SDD subagents/chains, or move to explore/proposal/spec/design/tasks until this session has an injected `## SDD Session Preflight` block or a canonical-authority resolution. Defaults and capability constraints may resolve fields without confirmation prompts; preserve unresolved-choice and safety gates.
75
-
76
64
  ## Memory Contract
77
65
 
78
- When memory is available, the parent selects context and subagents save discoveries before returning. Phase table and artifact keys: `orchestrator-memory.md`.
66
+ When memory is available, the parent selects context and subagents save discoveries before returning. ODD task continuity and memory lifecycle: `orchestrator-memory.md`.
79
67
 
80
68
  ## Skill Registry Protocol
81
69
 
82
- The parent resolves skill paths once per session under `## Skills to load before work`; subagents read those `SKILL.md` files first, or report unavailable paths. Fallback semantics (`paths-injected`/`fallback-registry`/`fallback-path`/`none`) and the SDD-executor distinction: `orchestrator-skills.md`.
70
+ The parent resolves skill paths once per session under `## Skills to load before work`; subagents read those `SKILL.md` files first, or report unavailable paths. Fallback semantics (`paths-injected`/`fallback-registry`/`fallback-path`/`none`): `orchestrator-skills.md`.
83
71
 
84
72
  ## Intent-Driven Skill Discovery
85
73
 
@@ -1,269 +1,7 @@
1
- # Strict TDD Module — Verify Phase
1
+ # Applicable Test-First Development — ODD Verification
2
2
 
3
- > **This module is loaded ONLY when Strict TDD Mode is enabled AND a test runner is available.**
4
- > If you are reading this, the orchestrator already verified both conditions. Follow every instruction.
3
+ This legacy filename is retained for package compatibility. It does not represent a Strict TDD switch.
5
4
 
6
- ## TDD Verification Philosophy
5
+ Assess applicability for each changed behavior: is there a runnable deterministic test with a clear expected outcome? Test existence alone does not establish applicability. For applicable behavior, inspect the reported observed RED (the intended failure before implementation), observed GREEN (focused passing execution after implementation), relevant alternate cases, and post-refactor checks. Do not infer RED from a test file existing or GREEN from a claim without execution. Report missing or contradictory evidence honestly; do not manufacture a lifecycle from the final diff.
7
6
 
8
- When Strict TDD Mode is active, verification goes beyond "does the code work?" to "was the code built correctly?" — meaning: was TDD actually followed? The apply phase reports TDD evidence; your job is to validate that evidence against reality.
9
-
10
- ## Step 5a: TDD Compliance Check (includes Assertion Quality Audit)
11
-
12
- Read the `apply-progress` artifact and verify that TDD was actually followed:
13
-
14
- ```
15
- Read apply-progress artifact:
16
- ├── Find the "TDD Cycle Evidence" table
17
- ├── FOR EACH task row:
18
- │ ├── RED column:
19
- │ │ ├── Must say "✅ Written"
20
- │ │ ├── Verify: test file EXISTS in the codebase
21
- │ │ └── Flag: CRITICAL if test file does not exist
22
- │ │
23
- │ ├── GREEN column:
24
- │ │ ├── Must say "✅ Passed"
25
- │ │ ├── Cross-reference with Step 5b test execution results:
26
- │ │ │ └── The test file listed must PASS when you run it
27
- │ │ └── Flag: CRITICAL if test fails now (was it really green?)
28
- │ │
29
- │ ├── TRIANGULATE column:
30
- │ │ ├── If "✅ N cases" → verify N test cases exist in the test file
31
- │ │ ├── If "➖ Single" → verify spec truly has only one scenario for this task
32
- │ │ └── Flag: WARNING if spec has multiple scenarios but only 1 test case
33
- │ │
34
- │ ├── SAFETY NET column:
35
- │ │ ├── If "✅ N/N" → existing tests were run before modification (good)
36
- │ │ ├── If "N/A (new)" → verify the file was actually NEW (not modified)
37
- │ │ └── Flag: WARNING if file was modified but safety net shows "N/A"
38
- │ │
39
- │ └── REFACTOR column:
40
- │ ├── Not strictly verifiable (subjective quality)
41
- │ └── Skip verification, trust the report
42
- │
43
- ├── If NO "TDD Cycle Evidence" table found:
44
- │ └── Flag: CRITICAL — apply phase did not report TDD evidence
45
- │ (Strict TDD was enabled but apply did not follow the protocol)
46
- │
47
- └── Summary: "{N}/{total} tasks have complete TDD evidence"
48
- ```
49
-
50
- ## Step 5 Expanded: Test Layer Validation
51
-
52
- Classify ALL test files related to this change by their testing layer:
53
-
54
- ```
55
- Scan test files created/modified by this change:
56
- ├── Classify each test file:
57
- │ ├── Unit test: tests a single function/class in isolation
58
- │ │ └── Indicators: no render(), no page., no HTTP calls, mocked dependencies
59
- │ ├── Integration test: tests component interaction or user behavior
60
- │ │ └── Indicators: render(), screen., userEvent., testing-library imports
61
- │ ├── E2E test: tests full system through real browser/HTTP
62
- │ │ └── Indicators: page.goto(), playwright/cypress imports, browser context
63
- │ └── Unknown: cannot classify → report as-is
64
- │
65
- ├── Report distribution:
66
- │ ├── Unit: {N} tests across {N} files
67
- │ ├── Integration: {N} tests across {N} files
68
- │ ├── E2E: {N} tests across {N} files
69
- │ └── Total: {N} tests
70
- │
71
- ├── Cross-reference with capabilities:
72
- │ ├── If integration tests exist but tools not in capabilities → how?
73
- │ ├── If E2E tests exist but tools not in capabilities → how?
74
- │ └── Flag: WARNING if tests use tools not detected in capabilities
75
- │
76
- └── For each spec scenario: note which layer covers it
77
- └── Flag: SUGGESTION if critical business logic only has unit tests
78
- (only if integration/E2E tools are available)
79
- ```
80
-
81
- ## Step 5d Expanded: Changed File Coverage
82
-
83
- When coverage tool is available, report coverage for CHANGED files specifically:
84
-
85
- ```
86
- IF coverage tool available (from cached capabilities):
87
- ├── Run: {test_command} --coverage (or equivalent)
88
- ├── Parse the coverage report
89
- ├── Filter to ONLY files created or modified in this change
90
- │ (get file list from apply-progress "Files Changed" table)
91
- ├── Report per-file:
92
- │ ├── File path
93
- │ ├── Line coverage %
94
- │ ├── Branch coverage % (if available)
95
- │ ├── Uncovered line ranges (specific lines, not just %)
96
- │ └── Flag per file:
97
- │ ├── ≥ 95% → ✅ Excellent
98
- │ ├── ≥ 80% → ⚠️ Acceptable
99
- │ └── < 80% → ⚠️ Low (list uncovered lines)
100
- ├── Report aggregate:
101
- │ ├── Average coverage of changed files
102
- │ ├── Total uncovered lines in changed files
103
- │ └── Compare to threshold if configured
104
- └── Flag: WARNING if any changed file < 80% coverage
105
-
106
- IF coverage tool NOT available:
107
- └── Report: "Coverage analysis skipped — no coverage tool detected"
108
- (NOT a failure — just not available)
109
- ```
110
-
111
- ## Step 5e: Quality Metrics (if tools available)
112
-
113
- Run quality checks ONLY on changed files, ONLY if tools are available:
114
-
115
- ```
116
- Read quality tools from cached capabilities:
117
-
118
- IF linter available:
119
- ├── Run linter on changed files only
120
- ├── Report: errors and warnings
121
- └── Flag: WARNING for errors, SUGGESTION for warnings
122
-
123
- IF type checker available:
124
- ├── Run type checker (usually whole-project, not per-file)
125
- ├── Filter output to changed files
126
- ├── Report: type errors in changed files
127
- └── Flag: WARNING for type errors
128
-
129
- IF neither available:
130
- └── Report: "Quality metrics skipped — no tools detected"
131
- ```
132
-
133
- ## Report Template Extension
134
-
135
- When Strict TDD Mode is active, your verification report MUST include these additional sections:
136
-
137
- ```markdown
138
- ### TDD Compliance
139
- | Check | Result | Details |
140
- |-------|--------|---------|
141
- | TDD Evidence reported | ✅ / ❌ | {Found in apply-progress / Missing} |
142
- | All tasks have tests | ✅ / ❌ | {N}/{total} tasks have test files |
143
- | RED confirmed (tests exist) | ✅ / ⚠️ | {N}/{total} test files verified |
144
- | GREEN confirmed (tests pass) | ✅ / ❌ | {N}/{total} tests pass on execution |
145
- | Triangulation adequate | ✅ / ⚠️ / ➖ | {N} tasks triangulated / {N} single-case |
146
- | Safety Net for modified files | ✅ / ⚠️ | {N}/{total} modified files had safety net |
147
-
148
- **TDD Compliance**: {N}/{total} checks passed
149
-
150
- ---
151
-
152
- ### Test Layer Distribution
153
- | Layer | Tests | Files | Tools |
154
- |-------|-------|-------|-------|
155
- | Unit | {N} | {N} | {tool} |
156
- | Integration | {N} | {N} | {tool or "not installed"} |
157
- | E2E | {N} | {N} | {tool or "not installed"} |
158
- | **Total** | **{N}** | **{N}** | |
159
-
160
- ---
161
-
162
- ### Changed File Coverage
163
- | File | Line % | Branch % | Uncovered Lines | Rating |
164
- |------|--------|----------|-----------------|--------|
165
- | `path/to/file.ext` | 95% | 90% | — | ✅ Excellent |
166
- | `path/to/other.ext` | 82% | 75% | L45-48, L62 | ⚠️ Acceptable |
167
- | `path/to/new.ext` | 100% | 100% | — | ✅ Excellent |
168
-
169
- **Average changed file coverage**: {N}%
170
- {or "Coverage analysis skipped — no coverage tool detected"}
171
-
172
- ---
173
-
174
- ### Assertion Quality
175
- | File | Line | Assertion | Issue | Severity |
176
- |------|------|-----------|-------|----------|
177
- | ... | ... | ... | ... | ... |
178
-
179
- **Assertion quality**: {N} CRITICAL, {N} WARNING
180
- {or "✅ All assertions verify real behavior"}
181
-
182
- ---
183
-
184
- ### Quality Metrics
185
- **Linter**: ✅ No errors / ⚠️ {N} warnings / ❌ {N} errors / ➖ Not available
186
- **Type Checker**: ✅ No errors / ❌ {N} errors / ➖ Not available
187
- ```
188
-
189
- ## Step 5f: Assertion Quality Audit (MANDATORY)
190
-
191
- Scan ALL test files created or modified by this change and check for trivial/meaningless assertions:
192
-
193
- ```
194
- FOR EACH test file related to the change:
195
- ├── Read the file content
196
- ├── Scan for BANNED assertion patterns:
197
- │ ├── Tautologies: expect(true).toBe(true), assert True, expect(1).toBe(1)
198
- │ ├── Orphan empty checks: expect(result).toEqual([]) or assert len(result) == 0
199
- │ │ └── UNLESS there is a companion test with same setup that asserts NON-EMPTY
200
- │ ├── Type-only assertions used alone: toBeDefined(), not.toBeNull(), typeof checks
201
- │ │ └── These are OK if COMBINED with value assertions in the same test
202
- │ ├── Assertions that never call production code (no function call, no render, no request)
203
- │ ├── Ghost loops: assertions inside for/forEach over queryAll/filter results
204
- │ │ └── Check if the collection could be empty — if so, the assertions NEVER RUN
205
- │ │ Flag: CRITICAL — a loop over an empty array is a test that ALWAYS passes
206
- │ ├── Incomplete TDD cycle: test passes because preconditions prevent code from running
207
- │ │ └── e.g., testing behavior of a component that is never rendered due to state
208
- │ │ Flag: CRITICAL — test must set up conditions where the code path IS exercised
209
- │ ├── Smoke-test-only: render() + toBeInTheDocument() without behavioral assertions
210
- │ │ └── "Renders without crash" is NOT a valid test — it must assert WHAT was rendered
211
- │ │ Flag: WARNING — smoke tests do not count toward TDD coverage
212
- │ ├── Implementation detail coupling: assertions on CSS classes, internal state, mock call counts
213
- │ │ └── expect(el.className).toContain("text-xs") or expect(mock.calls.length).toBe(3)
214
- │ │ Flag: WARNING — tests must assert behavior, not implementation
215
- │ └── Mock/assertion ratio: count vi.mock() calls vs expect() calls per test file
216
- │ └── If mocks > 2× assertions → Flag: WARNING — "Mock-heavy test ({N} mocks, {N} assertions)"
217
- │ Recommend: extract logic to pure function or move to higher test layer
218
- │
219
- ├── For each violation found:
220
- │ ├── Record: file, line number, the assertion, why it's trivial
221
- │ └── Classify:
222
- │ ├── CRITICAL: tautology (expect(true).toBe(true)) — test proves NOTHING
223
- │ ├── CRITICAL: assertion without production code call — test exercises nothing
224
- │ ├── CRITICAL: ghost loop — assertions inside loop over possibly-empty collection
225
- │ ├── WARNING: empty collection without companion non-empty test
226
- │ ├── WARNING: type-only assertion without value assertion
227
- │ ├── WARNING: smoke-test-only — render + toBeInTheDocument without behavioral check
228
- │ ├── WARNING: CSS class / implementation detail assertion
229
- │ └── WARNING: mock-heavy test (mocks > 2× assertions) — wrong test layer
230
- │
231
- ├── Check triangulation quality:
232
- │ ├── Count distinct test cases per behavior
233
- │ ├── If only 1 test case exists for a behavior with multiple spec scenarios:
234
- │ │ └── Flag: WARNING — "Insufficient triangulation for {behavior}"
235
- │ ├── If all test cases assert the SAME type of value (e.g., all check empty arrays):
236
- │ │ └── Flag: WARNING — "No variance in test expectations — all assert empty/trivial"
237
- │ └── A well-triangulated behavior has tests asserting DIFFERENT expected values
238
- │
239
- └── Summary: "{N} trivial assertions found across {N} files"
240
- ```
241
-
242
- ### Assertion Quality Report Table
243
-
244
- Include this table in the verification report when any issues are found:
245
-
246
- ```markdown
247
- ### Assertion Quality
248
- | File | Line | Assertion | Issue | Severity |
249
- |------|------|-----------|-------|----------|
250
- | `path/test.ts` | 15 | `expect(true).toBe(true)` | Tautology — proves nothing | CRITICAL |
251
- | `path/test.ts` | 23 | `expect(result).toEqual([])` | Empty without companion non-empty test | WARNING |
252
- | `path/test.ts` | 31 | `expect(result).toBeDefined()` | Type-only — no value asserted | WARNING |
253
-
254
- **Assertion quality**: {N} CRITICAL, {N} WARNING
255
- ```
256
-
257
- If zero issues found, report: "**Assertion quality**: ✅ All assertions verify real behavior"
258
-
259
- ## Rules (Strict TDD Verify specific)
260
-
261
- - ALWAYS check the TDD Cycle Evidence table from apply-progress — it's the primary artifact
262
- - ALWAYS cross-reference reported test files against actual execution — don't trust the report blindly
263
- - ALWAYS run the Assertion Quality Audit (Step 5f) — trivial tests are WORSE than missing tests
264
- - If apply-progress has no TDD evidence table, flag as CRITICAL — the protocol was not followed
265
- - If tautology assertions are found (expect(true).toBe(true)), flag as CRITICAL — these MUST be rewritten
266
- - Coverage and quality metrics are informational, NOT blocking — only flag as WARNING, never CRITICAL
267
- - Test layer distribution is informational — SUGGESTION level only
268
- - DO NOT fix issues — only report. The orchestrator decides.
269
- - If coverage/quality tools are not available, say so cleanly and move on — never flag missing tools as failures
7
+ For passive documentation, non-testable changes, an unavailable runner, or no meaningful RED, assess the stated exception and the proportionate ordinary functional or structural verification. Never demand a chat/TUI activation choice, and never skip all checks merely because test-first is inapplicable. Execute only exact commands authorized by the parent; report actual results, limitations and remaining uncertainty without editing code or overriding parent-owned RDD review.