@ngockhoale/ukit 3.3.2 → 3.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (89) hide show
  1. package/CHANGELOG.md +44 -0
  2. package/manifests/engineConformance.yaml +17 -1
  3. package/manifests/hostCapabilities.yaml +68 -1
  4. package/manifests/platform.full.yaml +138 -0
  5. package/manifests/platform.user.yaml +255 -3
  6. package/package.json +1 -1
  7. package/scripts/bench/subagent-orchestrator-corpus.mjs +275 -0
  8. package/scripts/bench/subagent-orchestrator-eval.mjs +565 -0
  9. package/scripts/probe/codex-capability-probe.mjs +169 -0
  10. package/src/cli/commands/doctor.js +168 -0
  11. package/src/cli/commands/indexTools.js +7 -0
  12. package/src/cli/commands/metrics.js +66 -2
  13. package/src/cli/commands/playbook.js +4 -4
  14. package/src/cli/commands/vm.js +49 -8
  15. package/src/core/agentRuntime/adapters.js +328 -27
  16. package/src/core/agentRuntime/artifacts.js +89 -0
  17. package/src/core/agentRuntime/context.js +345 -1
  18. package/src/core/agentRuntime/contract.js +296 -0
  19. package/src/core/agentRuntime/eventStore.js +176 -0
  20. package/src/core/agentRuntime/shadowRun.js +481 -5
  21. package/src/core/agentRuntime/telemetry.js +121 -0
  22. package/src/core/observability/emit/lifecycle.js +68 -1
  23. package/src/core/observability/emit/sessionBoot.js +393 -0
  24. package/src/core/observability/privacy/allowlist.js +10 -1
  25. package/src/core/observability/schema/registry.js +10 -0
  26. package/src/core/runtimeConfig.js +133 -0
  27. package/src/core/userPlaybooks.js +18 -3
  28. package/src/decision/registry.js +19 -0
  29. package/src/diagnostics/feedbackEvents.js +7 -4
  30. package/src/diagnostics/routeOutcomes.js +51 -6
  31. package/src/diagnostics/skillAccuracy.js +43 -3
  32. package/src/index/crossCheckMatrix.js +412 -0
  33. package/src/index/fixLoopEscalation.js +453 -0
  34. package/src/index/playbookRegistry.js +691 -0
  35. package/src/index/reviewPolicy.js +368 -0
  36. package/src/index/routeResolver.js +915 -0
  37. package/src/index/sessionHistoryExtractor.js +359 -0
  38. package/src/index/taskRouting.js +764 -581
  39. package/src/index/tierSelection.js +308 -0
  40. package/src/index/verificationMap.js +404 -0
  41. package/template_project/.claude/hooks/observability-emit.mjs +14 -0
  42. package/template_project/.claude/hooks/record-execution.mjs +19 -1
  43. package/template_project/.claude/hooks/skill-router.sh +691 -25
  44. package/template_project/.claude/hooks/verification-guard.sh +230 -1
  45. package/template_project/.claude/settings.json +2 -2
  46. package/template_project/.claude/ukit/index/cross-check-matrix.mjs +415 -0
  47. package/template_project/.claude/ukit/index/fix-loop-escalation.mjs +456 -0
  48. package/template_project/.claude/ukit/index/playbook-registry.mjs +690 -0
  49. package/template_project/.claude/ukit/index/review-panel-aggregate.mjs +20 -2
  50. package/template_project/.claude/ukit/index/review-policy.mjs +376 -0
  51. package/template_project/.claude/ukit/index/route-resolver.mjs +1059 -0
  52. package/template_project/.claude/ukit/index/route-task.mjs +1253 -846
  53. package/template_project/.claude/ukit/index/session-history-extractor.mjs +362 -0
  54. package/template_project/.claude/ukit/index/tier-selection.mjs +309 -0
  55. package/template_project/.claude/ukit/index/verification-map.mjs +403 -0
  56. package/template_project/.claude/ukit/index/worktree-sweep.mjs +195 -0
  57. package/template_project/.claude/ukit/runtime/execution-ledger.mjs +789 -11
  58. package/template_project/.claude/ukit/runtime/observability-emit.mjs +1102 -0
  59. package/template_project/.claude/ukit/runtime/reinject-context.mjs +9 -1
  60. package/template_project/.claude/ukit/runtime/resumable-run.mjs +149 -5
  61. package/template_project/.claude/ukit/runtime/stop-coordinator.mjs +323 -6
  62. package/template_project/.codex/README.md +8 -0
  63. package/template_project/.omp/hooks/pre/ukit-bridge.js +8 -1
  64. package/template_project/ukit/README.md +1 -1
  65. package/template_project/ukit/storage/config.json +20 -0
  66. package/template_user/playbooks/architecture-decision.md +28 -0
  67. package/template_user/playbooks/autonomous-run.md +43 -0
  68. package/template_user/playbooks/autopilot-full.md +59 -0
  69. package/template_user/playbooks/autopilot-stack.md +54 -0
  70. package/template_user/playbooks/babysit.md +39 -0
  71. package/template_user/playbooks/bug-fix.md +3 -1
  72. package/template_user/playbooks/{issue-implementation.md → feature-implementation.md} +4 -2
  73. package/template_user/playbooks/hillclimb.md +44 -0
  74. package/template_user/playbooks/investigation.md +21 -0
  75. package/template_user/playbooks/migration.md +21 -0
  76. package/template_user/playbooks/open-pr.md +48 -0
  77. package/template_user/playbooks/orchestrate.md +45 -0
  78. package/template_user/playbooks/performance.md +33 -0
  79. package/template_user/playbooks/prototype.md +28 -0
  80. package/template_user/playbooks/refactor.md +19 -0
  81. package/template_user/playbooks/release.md +28 -0
  82. package/template_user/playbooks/runtime-forensics.md +23 -0
  83. package/template_user/playbooks/session-pickup.md +31 -0
  84. package/template_user/playbooks/shipping.md +53 -0
  85. package/template_user/playbooks/skill-evaluation.md +48 -0
  86. package/template_user/playbooks/small-feature.md +20 -0
  87. package/template_user/playbooks/verification-map.json +153 -0
  88. package/template_user/playbooks/verification.md +22 -0
  89. package/template_user/playbooks/worktree-cleanup.md +37 -0
@@ -0,0 +1,53 @@
1
+ ---
2
+ id: shipping
3
+ lanes: [review-release]
4
+ ---
5
+ You own the landing of a stack of ready changes — the work after `babysit`
6
+ leaves each PR merge-ready. You verify each unit on its own evidence, confirm
7
+ the stack is contiguous, land in order behind per-action authorization, and
8
+ confirm the final head. Getting a unit to merge-ready is babysit's job; a unit
9
+ still failing CI or sitting under open review threads belongs back there.
10
+ If the user says "new task", re-route — do not treat the message as the next step.
11
+ 1. Confirm this is shipping territory: every unit in the stack is open and
12
+ merge-ready (the upstream PR-create and babysit steps, RCC-01 classes 3–5,
13
+ are done) and you can reach CI/review state plus git history. A single
14
+ unverified change still in development, or a repo where you lack land
15
+ permission, is a when-not — say so and stop. If `gh` is missing or
16
+ unauthenticated, report the CLI failure verbatim — never fabricate check,
17
+ review, or head state.
18
+ 2. Verify each unit independently (RCC-01 class 4): `gh pr checks`, `gh pr
19
+ view`, and review threads via
20
+ `gh api repos/{owner}/{repo}/pulls/{n}/comments` — plus the unit's own
21
+ test/build run where the repo defines one. Record a per-unit verification
22
+ receipt: commands run, output, result. Independent means a green unit does
23
+ not vouch for its neighbor.
24
+ 3. Confirm stack contiguity: each unit's base is the previous unit's head —
25
+ order them by `gh pr view --json baseRefName,headRefName` and `git`
26
+ branch inspection. A gap or a unit that fails verification breaks the
27
+ chain: skip it if the remainder still lands on its own base, otherwise
28
+ stop the stack there — you never land unverified or non-contiguous work,
29
+ and the skipped unit is named in the report.
30
+ 4. Land in order — each land is an irreversible write behind the RCC-01
31
+ §Land-authorization boundary. Merge/land to the default branch and
32
+ push-to-default are gated: stop and `Ask the human` before each one;
33
+ authorization is per-action and never carried forward — "land the stack"
34
+ authorizes the described sequence once, and a new merge is a new
35
+ authorization. The fallback land step is `gh pr merge` or `git merge`
36
+ into the default branch plus `git push` — no connector-only step. A
37
+ model never authorizes; the gate stays deterministic and human.
38
+ 5. Confirm the final head: `git log`/`git rev-parse` the default branch and
39
+ check every intended unit is in; attach the receipts and the head
40
+ confirmation as evidence. If authorization is missing, a unit fails, or
41
+ contiguity breaks mid-stack, report BLOCKED at that step naming the
42
+ boundary or the failed unit — never soft-fail around it and never land
43
+ the remainder silently.
44
+ Every step here is `gh`/git over Bash (RCC-01 §Capability matrix) — full on
45
+ Claude Code and omp; on Codex the gate posture is advisory /
46
+ instruction-mediated (RCC-01 §Hook-free enforcement): the model is asked, not
47
+ stopped.
48
+ Reply: the ordered unit list, each per-unit verification receipt, the land
49
+ authorizations asked and granted, the final-head confirmation, and the
50
+ terminal state — stack landed, or BLOCKED with the named cause.
51
+ Ask the human at every irreversible write — per-action, even mid-stack — and
52
+ for a genuine preference call no experiment settles or a real dead end.
53
+ Everything else: do it, report it.
@@ -0,0 +1,48 @@
1
+ ---
2
+ id: skill-evaluation
3
+ lanes: [local-build, review-release]
4
+ ---
5
+ You own the skill — whether authoring it or measuring whether it helped. Two
6
+ halves, one finish: a validated, discoverable skill, or a scored verdict backed
7
+ by recorded fixtures and arms.
8
+ If the user says "new task", re-route — do not treat the message as the next step.
9
+ Authoring ("make this a skill"):
10
+ 1. Shape the skill file — `.claude/skills/<name>/SKILL.md` with `name:` +
11
+ `description:` frontmatter that states *when* to load it. UKit's own skill
12
+ standard applies, not a foreign file format copied blindly.
13
+ 2. Scaffold its audit — `yarn skill:audit init <name>` creates
14
+ `docs/skill-audits/<name>/` with the pressure-scenario, rationalization-table,
15
+ and trigger-accuracy templates.
16
+ 3. Pressure-test the discipline — write one pressure scenario; RED: run it
17
+ without the skill, record the verbatim rationalization; GREEN: re-run with
18
+ the skill loaded and expect compliance; REFACTOR: a new rationalization gets
19
+ an explicit counter in the table, never a soft exception. Record each phase
20
+ via `yarn skill:audit record <name>`; `yarn skill:audit status <name>` is
21
+ clean only when open rationalizations = 0.
22
+ 4. Prove discoverability — run the trigger-accuracy check: prompts that should
23
+ load the skill do, prompts that should not do not, judged on the frontmatter
24
+ `description` alone. Fix the description until both hold.
25
+ Measurement ("did this change help"):
26
+ 5. Name the fixture set — a stable fixture file or list of prompts (fixture
27
+ IDs), the same set for both arms, decided before scoring.
28
+ 6. Split the arms — control = the baseline (`baselineRef`), changed = the new
29
+ skill/playbook/prompt/route. The arm assignment is recorded before any
30
+ scoring happens.
31
+ 7. Score the comparison — judge each fixture per arm blinded to which arm
32
+ produced it; deltas across quality, rework, latency, calls.
33
+ 8. Record the verdict — `emitExperimentScorecard` with `net-gain`, `neutral`,
34
+ or `negative`, carrying fixture IDs + per-arm scores + the arm assignment.
35
+ A bare vibe verdict with no scored comparison does not finish —
36
+ `emitExperimentScorecard` enforces the verdict enum and fixture/id shape;
37
+ an unscored claim fails the recorded-evidence bar. Promotion = `net-gain`
38
+ plus maintainer review; `neutral|negative` stays disabled.
39
+ Machinery: `yarn skill:audit init|status|record` (scaffolds + records only —
40
+ the model run is the pressure check itself); `src/core/experiments/
41
+ deliberation.js` `emitExperimentScorecard` for the verdict contract.
42
+ Reply: authoring — the skill path, audit status (entries, open
43
+ rationalizations), trigger-accuracy result; measurement — fixture IDs, per-arm
44
+ scores, arm assignment, verdict + promotion eligibility verbatim.
45
+ Ask the human only for: shipping a skill that failed its audit, promoting an
46
+ experiment (net-gain goes to maintainer review, never auto-promote), a genuine
47
+ preference call no fixture settles, or a real dead end. Everything else: do it,
48
+ report it.
@@ -0,0 +1,20 @@
1
+ ---
2
+ id: small-feature
3
+ lanes: [local-fix, local-build]
4
+ ---
5
+ You own this slice. Make it appear and work on the real surface — a created file
6
+ nothing reaches is not done.
7
+ If the user says "new task", re-route — do not treat the message as the next step.
8
+ 1. Infer scope and the done predicate from the request plus repo conventions —
9
+ routine choices (file name, style, component split, state pattern) are yours;
10
+ ask only when the choices diverge into materially different product behavior.
11
+ 2. Build the smallest usable slice — everything needed for the behavior to appear
12
+ and work. No spec phase, no architect lane on this path.
13
+ 3. Wire it into the usage flow — the unwired component is the most likely cheat;
14
+ an import plus a callsite is the minimum proof of wiring.
15
+ 4. Verify on the real surface: UI = run it and observe render/interaction;
16
+ function = concrete input→output cases. Build-pass proves nothing.
17
+ 5. Record the assumptions you made — one line each.
18
+ Reply: what was added, where it is wired, and the surface evidence it works.
19
+ Ask the human only for: irreversible writes, a genuine preference call no experiment
20
+ settles, or a real dead end. Everything else: do it, report it.
@@ -0,0 +1,153 @@
1
+ {
2
+ "version": 1,
3
+ "playbooks": {
4
+ "small-feature": {
5
+ "launch": "\"create X doing Y\" single-surface",
6
+ "doctor": "repo conventions identified; harness exists?",
7
+ "drive": "implement + wire + render/run",
8
+ "expected": "slice works on real surface",
9
+ "evidence": ["render-observation or IO-case receipt"],
10
+ "cleanup": "remove throwaway scaffolds"
11
+ },
12
+ "feature-implementation": {
13
+ "launch": "\"implement X\" w/ contract",
14
+ "doctor": "done predicate stated; analog found",
15
+ "drive": "predicate → smallest change → verify → widen once",
16
+ "expected": "predicate holds on real artifact",
17
+ "evidence": ["predicate evidence + impact-sweep output"],
18
+ "cleanup": "none"
19
+ },
20
+ "bug-fix": {
21
+ "launch": "\"fix X\" reproducible",
22
+ "doctor": "repro command/surface reachable",
23
+ "drive": "repro → hypothesis bisect → minimal fix → re-verify",
24
+ "expected": "original repro passes same-surface",
25
+ "evidence": ["repro-before + repro-after verbatim, root cause named"],
26
+ "cleanup": "rejected-hypothesis notes kept"
27
+ },
28
+ "investigation": {
29
+ "launch": "\"how/why X\"",
30
+ "doctor": "read-only confirmed",
31
+ "drive": "frame → gather → classify → answer",
32
+ "expected": "cited answer or falsified assumption",
33
+ "evidence": ["path:symbol/command-output anchors"],
34
+ "cleanup": "git status clean check"
35
+ },
36
+ "refactor": {
37
+ "launch": "\"restructure X\" behavior-preserved",
38
+ "doctor": "pin exists or characterization tests writable",
39
+ "drive": "pin → stepwise change → re-pin → migrate callers",
40
+ "expected": "pin green before+after; zero legacy callers",
41
+ "evidence": ["pin receipts + caller sweep"],
42
+ "cleanup": "old surface deleted"
43
+ },
44
+ "performance": {
45
+ "launch": "\"X is slow\" + symptom",
46
+ "doctor": "harness identified; baseline measurable",
47
+ "drive": "baseline → trace → improve → re-measure",
48
+ "expected": "measured delta ≥ stated margin",
49
+ "evidence": ["before/after numbers + commands"],
50
+ "cleanup": "harness left runnable"
51
+ },
52
+ "runtime-forensics": {
53
+ "launch": "live symptom / dropped artifact",
54
+ "doctor": "instrumentation tools present (sample, node --inspect, profilers)",
55
+ "drive": "attach/observe → capture → isolate → name mechanism",
56
+ "expected": "anomaly explained by captured evidence",
57
+ "evidence": ["heap/CPU/handle captures + derived readings"],
58
+ "cleanup": "captures stored under artifacts; processes released"
59
+ },
60
+ "architecture-decision": {
61
+ "launch": "\"A or B\" fork",
62
+ "doctor": "alternatives enumerable",
63
+ "drive": "gather constraints → compare → record",
64
+ "expected": "decision recorded w/ falsifiable rationale",
65
+ "evidence": ["decision record + alternatives"],
66
+ "cleanup": "none"
67
+ },
68
+ "prototype": {
69
+ "launch": "\"spike Y\"",
70
+ "doctor": "decision-to-settle stated",
71
+ "drive": "cheapest artifact → measure → verdict",
72
+ "expected": "fork settled; artifact dispositioned",
73
+ "evidence": ["experiment output + verdict + disposable/promoted flag"],
74
+ "cleanup": "throwaway flagged or removed"
75
+ },
76
+ "migration": {
77
+ "launch": "\"move to new schema/API\"",
78
+ "doctor": "target shape defined; rollback path exists",
79
+ "drive": "pin → migrate → verify on target → rollback demo",
80
+ "expected": "target verified + rollback demonstrated",
81
+ "evidence": ["migration verify + rollback receipts"],
82
+ "cleanup": "legacy path scheduled for removal"
83
+ },
84
+ "verification": {
85
+ "launch": "\"verify/check X\"",
86
+ "doctor": "the named check is runnable",
87
+ "drive": "run check → record verdict",
88
+ "expected": "verdict recorded (pass/fail/inconclusive)",
89
+ "evidence": ["check output verbatim"],
90
+ "cleanup": "none"
91
+ },
92
+ "skill-evaluation": {
93
+ "launch": "\"skill X\" / \"did change help\"",
94
+ "doctor": "fixture set named",
95
+ "drive": "author→validate→ship OR arms→score→verdict",
96
+ "expected": "skill discoverable OR scored comparison recorded",
97
+ "evidence": ["audit output / fixture IDs + scores"],
98
+ "cleanup": "fixtures versioned"
99
+ },
100
+ "session-pickup": {
101
+ "launch": "resume signal",
102
+ "doctor": "resumable record exists",
103
+ "drive": "load → reconstruct → verify vs worktree → continue",
104
+ "expected": "resumed at correct cursor, no re-done steps",
105
+ "evidence": ["state-verify receipt"],
106
+ "cleanup": "stale records pruned"
107
+ },
108
+ "autonomous-run": {
109
+ "launch": "bounded long goal",
110
+ "doctor": "finish condition fixable",
111
+ "drive": "loop: work → verify → re-check finish",
112
+ "expected": "finish verified or escape hatch + named blocker",
113
+ "evidence": ["progress ledger + final verification"],
114
+ "cleanup": "resumable record closed"
115
+ },
116
+ "release": {
117
+ "launch": "\"bump/ship/publish\"",
118
+ "doctor": "on UKit repo; scripts exist (verify-release.mjs)",
119
+ "drive": "version → test:release-core → push+tag → gh release → npm dry-run → publish → parity probes",
120
+ "expected": "3 version probes equal",
121
+ "evidence": ["probe outputs (npm view / gh release list / git tag)"],
122
+ "cleanup": "if parity fails: report \"not shipped\" (owner §4)"
123
+ }
124
+ },
125
+ "artifactClasses": {
126
+ "ui": {
127
+ "name": "UI surface verification",
128
+ "check": "render the changed surface or drive it through the project's UI harness (package.json script `test:ui`) and record the receipt",
129
+ "harness": { "packageScript": "test:ui" },
130
+ "fallback": "receipt class `render-observation` — attest the observed render/interaction via UKIT_EVIDENCE (class=render-observation;surface=<how observed>)",
131
+ "evidenceRequired": ["render-observation|io-case"]
132
+ },
133
+ "cli": {
134
+ "name": "CLI/IO surface verification",
135
+ "check": "run the CLI entry or its IO harness (package.json script `test:cli`) and record the receipt",
136
+ "harness": { "packageScript": "test:cli" },
137
+ "fallback": "receipt class `io-case` — attest one input/output case via UKIT_EVIDENCE (class=io-case;surface=<command>)",
138
+ "evidenceRequired": ["io-case"]
139
+ },
140
+ "docs": {
141
+ "name": "Docs/prose verification",
142
+ "check": "skim the rendered prose for link/format integrity; no test run required",
143
+ "fallback": "receipt class `doc-render` — optional; attest a docs render check via UKIT_EVIDENCE when available",
144
+ "evidenceRequired": []
145
+ },
146
+ "config": {
147
+ "name": "Config/settings verification",
148
+ "check": "load or lint the changed config once (project's own loader/lint command) and record the receipt",
149
+ "fallback": "receipt class `config-load` — attest a config load/lint observation via UKIT_EVIDENCE when no script exists",
150
+ "evidenceRequired": []
151
+ }
152
+ }
153
+ }
@@ -0,0 +1,22 @@
1
+ ---
2
+ id: verification
3
+ lanes: [review-release, local-fix]
4
+ ---
5
+ You own this check. Run the named check and record its verdict — a failed check is
6
+ a finding to report, not a task to fix.
7
+ If the user says "new task", re-route — do not treat the message as the next step.
8
+ 1. Name the check and its verdict predicate — the command, surface, or
9
+ observation that produces pass/fail.
10
+ 2. Run it and capture the output verbatim. If this engine cannot produce that
11
+ evidence class, fall back to the closest runnable check and mark the limit
12
+ `unverified` — never fabricate a pass.
13
+ 3. Visual channel (skippable — only when parity against a reference is the ask):
14
+ capture reference → render → vision-compare → iterate until equivalence or
15
+ named divergence; prefer DOM/snapshot assertions when they suffice — vision
16
+ passes are the expensive channel.
17
+ 4. Record the verdict: pass, fail, or inconclusive with what is missing. A fail
18
+ hands off to the owning group (bug-fix, performance) as a new routed task.
19
+ Reply: the check run, the verbatim output or comparison captures, and the recorded
20
+ verdict.
21
+ Ask the human only for: irreversible writes, a genuine preference call no experiment
22
+ settles, or a real dead end. Everything else: do it, report it.
@@ -0,0 +1,37 @@
1
+ ---
2
+ id: worktree-cleanup
3
+ lanes: [local-fix, local-build]
4
+ ---
5
+ You own this workspace-hygiene pass. Reclaim stale/merged worktrees — every
6
+ removal sits behind a per-entry safety check, never a blanket sweep. If the
7
+ user says "new task", re-route — do not treat the message as the next step.
8
+ 1. Confirm this is cleanup territory: accumulated `.worktrees/*` entries (or
9
+ linked worktrees anywhere) left by completed or abandoned runs. A repo
10
+ with no extra worktrees is already clean — report that and stop.
11
+ 2. Enumerate candidates and classify them. Run
12
+ `node .claude/ukit/index/worktree-sweep.mjs` (dry-run is the default — it
13
+ reads `git worktree list --porcelain`, classifies each entry
14
+ merged/stale/active, and prints the classification table with a per-entry
15
+ verdict). It removes nothing.
16
+ 3. Read the verdicts. Each non-main worktree is `safe` or a named skip —
17
+ `skip-dirty` (uncommitted work), `skip-locked` (git lock, deliberate),
18
+ `skip-active` (a resumable-run record or a `docs/AI_HANDOFF/RUN.md` cursor
19
+ whose phase is not `done`/`blocked`), `skip-detached-head`. The gate is
20
+ per entry: no uncommitted work AND no active run.
21
+ 4. Apply — removal is a destructive operation: a worktree's checkout is
22
+ deleted, so `Ask the human` before running `--apply`, even when every
23
+ candidate reads `safe`. With authorization, run
24
+ `node .claude/ukit/index/worktree-sweep.mjs --apply`; it runs
25
+ `git worktree remove` only on `safe` entries and prints each removal's
26
+ gate verdict. Skipped entries are named in the report and left untouched.
27
+ 5. Report and finish: all stale/merged worktrees reclaimed, nothing active or
28
+ dirty touched. Nothing safe to remove, or everything skipped, is a report
29
+ — print the verdicts, never exit silently.
30
+ Machinery: `worktree-sweep.mjs` over `git` porcelain — full on Claude Code
31
+ and omp via Bash; on Codex run the same commands manually — enforcement stays
32
+ advisory/instruction-mediated, never a live gate claim.
33
+ Reply: the classification list with every per-entry verdict, what was removed
34
+ (and under which authorization), and what was skipped.
35
+ Ask the human only for: irreversible writes (every `--apply` removal), a
36
+ genuine preference call no experiment settles, or a real dead end. Everything
37
+ else: do it, report it.