@ngockhoale/ukit 3.3.3 → 3.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +44 -0
- package/manifests/engineConformance.yaml +17 -1
- package/manifests/hostCapabilities.yaml +68 -1
- package/manifests/platform.full.yaml +138 -0
- package/manifests/platform.user.yaml +255 -3
- package/package.json +1 -1
- package/scripts/bench/subagent-orchestrator-corpus.mjs +275 -0
- package/scripts/bench/subagent-orchestrator-eval.mjs +565 -0
- package/scripts/probe/codex-capability-probe.mjs +169 -0
- package/src/cli/commands/doctor.js +168 -0
- package/src/cli/commands/indexTools.js +7 -0
- package/src/cli/commands/metrics.js +66 -2
- package/src/cli/commands/playbook.js +4 -4
- package/src/cli/commands/vm.js +49 -8
- package/src/core/agentRuntime/adapters.js +328 -27
- package/src/core/agentRuntime/artifacts.js +89 -0
- package/src/core/agentRuntime/context.js +345 -1
- package/src/core/agentRuntime/contract.js +296 -0
- package/src/core/agentRuntime/eventStore.js +176 -0
- package/src/core/agentRuntime/shadowRun.js +481 -5
- package/src/core/agentRuntime/telemetry.js +121 -0
- package/src/core/observability/emit/lifecycle.js +68 -1
- package/src/core/observability/emit/sessionBoot.js +393 -0
- package/src/core/observability/privacy/allowlist.js +10 -1
- package/src/core/observability/schema/registry.js +10 -0
- package/src/core/runtimeConfig.js +133 -0
- package/src/core/userPlaybooks.js +18 -3
- package/src/decision/registry.js +19 -0
- package/src/diagnostics/feedbackEvents.js +7 -4
- package/src/diagnostics/routeOutcomes.js +51 -6
- package/src/diagnostics/skillAccuracy.js +43 -3
- package/src/index/crossCheckMatrix.js +412 -0
- package/src/index/fixLoopEscalation.js +453 -0
- package/src/index/playbookRegistry.js +691 -0
- package/src/index/reviewPolicy.js +368 -0
- package/src/index/routeResolver.js +915 -0
- package/src/index/sessionHistoryExtractor.js +359 -0
- package/src/index/taskRouting.js +764 -581
- package/src/index/tierSelection.js +308 -0
- package/src/index/verificationMap.js +404 -0
- package/template_project/.claude/hooks/observability-emit.mjs +14 -0
- package/template_project/.claude/hooks/record-execution.mjs +19 -1
- package/template_project/.claude/hooks/skill-router.sh +691 -25
- package/template_project/.claude/hooks/verification-guard.sh +230 -1
- package/template_project/.claude/settings.json +2 -2
- package/template_project/.claude/ukit/index/cross-check-matrix.mjs +415 -0
- package/template_project/.claude/ukit/index/fix-loop-escalation.mjs +456 -0
- package/template_project/.claude/ukit/index/playbook-registry.mjs +690 -0
- package/template_project/.claude/ukit/index/review-panel-aggregate.mjs +20 -2
- package/template_project/.claude/ukit/index/review-policy.mjs +376 -0
- package/template_project/.claude/ukit/index/route-resolver.mjs +1059 -0
- package/template_project/.claude/ukit/index/route-task.mjs +1253 -846
- package/template_project/.claude/ukit/index/session-history-extractor.mjs +362 -0
- package/template_project/.claude/ukit/index/tier-selection.mjs +309 -0
- package/template_project/.claude/ukit/index/verification-map.mjs +403 -0
- package/template_project/.claude/ukit/index/worktree-sweep.mjs +195 -0
- package/template_project/.claude/ukit/runtime/execution-ledger.mjs +789 -11
- package/template_project/.claude/ukit/runtime/observability-emit.mjs +1102 -0
- package/template_project/.claude/ukit/runtime/reinject-context.mjs +9 -1
- package/template_project/.claude/ukit/runtime/resumable-run.mjs +149 -5
- package/template_project/.claude/ukit/runtime/stop-coordinator.mjs +323 -6
- package/template_project/.codex/README.md +8 -0
- package/template_project/.omp/hooks/pre/ukit-bridge.js +8 -1
- package/template_project/ukit/README.md +1 -1
- package/template_project/ukit/storage/config.json +20 -0
- package/template_user/playbooks/architecture-decision.md +28 -0
- package/template_user/playbooks/autonomous-run.md +43 -0
- package/template_user/playbooks/autopilot-full.md +59 -0
- package/template_user/playbooks/autopilot-stack.md +54 -0
- package/template_user/playbooks/babysit.md +39 -0
- package/template_user/playbooks/bug-fix.md +3 -1
- package/template_user/playbooks/{issue-implementation.md → feature-implementation.md} +4 -2
- package/template_user/playbooks/hillclimb.md +44 -0
- package/template_user/playbooks/investigation.md +21 -0
- package/template_user/playbooks/migration.md +21 -0
- package/template_user/playbooks/open-pr.md +48 -0
- package/template_user/playbooks/orchestrate.md +45 -0
- package/template_user/playbooks/performance.md +33 -0
- package/template_user/playbooks/prototype.md +28 -0
- package/template_user/playbooks/refactor.md +19 -0
- package/template_user/playbooks/release.md +28 -0
- package/template_user/playbooks/runtime-forensics.md +23 -0
- package/template_user/playbooks/session-pickup.md +31 -0
- package/template_user/playbooks/shipping.md +53 -0
- package/template_user/playbooks/skill-evaluation.md +48 -0
- package/template_user/playbooks/small-feature.md +20 -0
- package/template_user/playbooks/verification-map.json +153 -0
- package/template_user/playbooks/verification.md +22 -0
- package/template_user/playbooks/worktree-cleanup.md +37 -0
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
---
|
|
2
|
+
id: shipping
|
|
3
|
+
lanes: [review-release]
|
|
4
|
+
---
|
|
5
|
+
You own the landing of a stack of ready changes — the work after `babysit`
|
|
6
|
+
leaves each PR merge-ready. You verify each unit on its own evidence, confirm
|
|
7
|
+
the stack is contiguous, land in order behind per-action authorization, and
|
|
8
|
+
confirm the final head. Getting a unit to merge-ready is babysit's job; a unit
|
|
9
|
+
still failing CI or sitting under open review threads belongs back there.
|
|
10
|
+
If the user says "new task", re-route — do not treat the message as the next step.
|
|
11
|
+
1. Confirm this is shipping territory: every unit in the stack is open and
|
|
12
|
+
merge-ready (the upstream PR-create and babysit steps, RCC-01 classes 3–5,
|
|
13
|
+
are done) and you can reach CI/review state plus git history. A single
|
|
14
|
+
unverified change still in development, or a repo where you lack land
|
|
15
|
+
permission, is a when-not — say so and stop. If `gh` is missing or
|
|
16
|
+
unauthenticated, report the CLI failure verbatim — never fabricate check,
|
|
17
|
+
review, or head state.
|
|
18
|
+
2. Verify each unit independently (RCC-01 class 4): `gh pr checks`, `gh pr
|
|
19
|
+
view`, and review threads via
|
|
20
|
+
`gh api repos/{owner}/{repo}/pulls/{n}/comments` — plus the unit's own
|
|
21
|
+
test/build run where the repo defines one. Record a per-unit verification
|
|
22
|
+
receipt: commands run, output, result. Independent means a green unit does
|
|
23
|
+
not vouch for its neighbor.
|
|
24
|
+
3. Confirm stack contiguity: each unit's base is the previous unit's head —
|
|
25
|
+
order them by `gh pr view --json baseRefName,headRefName` and `git`
|
|
26
|
+
branch inspection. A gap or a unit that fails verification breaks the
|
|
27
|
+
chain: skip it if the remainder still lands on its own base, otherwise
|
|
28
|
+
stop the stack there — you never land unverified or non-contiguous work,
|
|
29
|
+
and the skipped unit is named in the report.
|
|
30
|
+
4. Land in order — each land is an irreversible write behind the RCC-01
|
|
31
|
+
§Land-authorization boundary. Merge/land to the default branch and
|
|
32
|
+
push-to-default are gated: stop and `Ask the human` before each one;
|
|
33
|
+
authorization is per-action and never carried forward — "land the stack"
|
|
34
|
+
authorizes the described sequence once, and a new merge is a new
|
|
35
|
+
authorization. The fallback land step is `gh pr merge` or `git merge`
|
|
36
|
+
into the default branch plus `git push` — no connector-only step. A
|
|
37
|
+
model never authorizes; the gate stays deterministic and human.
|
|
38
|
+
5. Confirm the final head: `git log`/`git rev-parse` the default branch and
|
|
39
|
+
check every intended unit is in; attach the receipts and the head
|
|
40
|
+
confirmation as evidence. If authorization is missing, a unit fails, or
|
|
41
|
+
contiguity breaks mid-stack, report BLOCKED at that step naming the
|
|
42
|
+
boundary or the failed unit — never soft-fail around it and never land
|
|
43
|
+
the remainder silently.
|
|
44
|
+
Every step here is `gh`/git over Bash (RCC-01 §Capability matrix) — full on
|
|
45
|
+
Claude Code and omp; on Codex the gate posture is advisory /
|
|
46
|
+
instruction-mediated (RCC-01 §Hook-free enforcement): the model is asked, not
|
|
47
|
+
stopped.
|
|
48
|
+
Reply: the ordered unit list, each per-unit verification receipt, the land
|
|
49
|
+
authorizations asked and granted, the final-head confirmation, and the
|
|
50
|
+
terminal state — stack landed, or BLOCKED with the named cause.
|
|
51
|
+
Ask the human at every irreversible write — per-action, even mid-stack — and
|
|
52
|
+
for a genuine preference call no experiment settles or a real dead end.
|
|
53
|
+
Everything else: do it, report it.
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
---
|
|
2
|
+
id: skill-evaluation
|
|
3
|
+
lanes: [local-build, review-release]
|
|
4
|
+
---
|
|
5
|
+
You own the skill — whether authoring it or measuring whether it helped. Two
|
|
6
|
+
halves, one finish: a validated, discoverable skill, or a scored verdict backed
|
|
7
|
+
by recorded fixtures and arms.
|
|
8
|
+
If the user says "new task", re-route — do not treat the message as the next step.
|
|
9
|
+
Authoring ("make this a skill"):
|
|
10
|
+
1. Shape the skill file — `.claude/skills/<name>/SKILL.md` with `name:` +
|
|
11
|
+
`description:` frontmatter that states *when* to load it. UKit's own skill
|
|
12
|
+
standard applies, not a foreign file format copied blindly.
|
|
13
|
+
2. Scaffold its audit — `yarn skill:audit init <name>` creates
|
|
14
|
+
`docs/skill-audits/<name>/` with the pressure-scenario, rationalization-table,
|
|
15
|
+
and trigger-accuracy templates.
|
|
16
|
+
3. Pressure-test the discipline — write one pressure scenario; RED: run it
|
|
17
|
+
without the skill, record the verbatim rationalization; GREEN: re-run with
|
|
18
|
+
the skill loaded and expect compliance; REFACTOR: a new rationalization gets
|
|
19
|
+
an explicit counter in the table, never a soft exception. Record each phase
|
|
20
|
+
via `yarn skill:audit record <name>`; `yarn skill:audit status <name>` is
|
|
21
|
+
clean only when open rationalizations = 0.
|
|
22
|
+
4. Prove discoverability — run the trigger-accuracy check: prompts that should
|
|
23
|
+
load the skill do, prompts that should not do not, judged on the frontmatter
|
|
24
|
+
`description` alone. Fix the description until both hold.
|
|
25
|
+
Measurement ("did this change help"):
|
|
26
|
+
5. Name the fixture set — a stable fixture file or list of prompts (fixture
|
|
27
|
+
IDs), the same set for both arms, decided before scoring.
|
|
28
|
+
6. Split the arms — control = the baseline (`baselineRef`), changed = the new
|
|
29
|
+
skill/playbook/prompt/route. The arm assignment is recorded before any
|
|
30
|
+
scoring happens.
|
|
31
|
+
7. Score the comparison — judge each fixture per arm blinded to which arm
|
|
32
|
+
produced it; deltas across quality, rework, latency, calls.
|
|
33
|
+
8. Record the verdict — `emitExperimentScorecard` with `net-gain`, `neutral`,
|
|
34
|
+
or `negative`, carrying fixture IDs + per-arm scores + the arm assignment.
|
|
35
|
+
A bare vibe verdict with no scored comparison does not finish —
|
|
36
|
+
`emitExperimentScorecard` enforces the verdict enum and fixture/id shape;
|
|
37
|
+
an unscored claim fails the recorded-evidence bar. Promotion = `net-gain`
|
|
38
|
+
plus maintainer review; `neutral|negative` stays disabled.
|
|
39
|
+
Machinery: `yarn skill:audit init|status|record` (scaffolds + records only —
|
|
40
|
+
the model run is the pressure check itself); `src/core/experiments/
|
|
41
|
+
deliberation.js` `emitExperimentScorecard` for the verdict contract.
|
|
42
|
+
Reply: authoring — the skill path, audit status (entries, open
|
|
43
|
+
rationalizations), trigger-accuracy result; measurement — fixture IDs, per-arm
|
|
44
|
+
scores, arm assignment, verdict + promotion eligibility verbatim.
|
|
45
|
+
Ask the human only for: shipping a skill that failed its audit, promoting an
|
|
46
|
+
experiment (net-gain goes to maintainer review, never auto-promote), a genuine
|
|
47
|
+
preference call no fixture settles, or a real dead end. Everything else: do it,
|
|
48
|
+
report it.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
---
|
|
2
|
+
id: small-feature
|
|
3
|
+
lanes: [local-fix, local-build]
|
|
4
|
+
---
|
|
5
|
+
You own this slice. Make it appear and work on the real surface — a created file
|
|
6
|
+
nothing reaches is not done.
|
|
7
|
+
If the user says "new task", re-route — do not treat the message as the next step.
|
|
8
|
+
1. Infer scope and the done predicate from the request plus repo conventions —
|
|
9
|
+
routine choices (file name, style, component split, state pattern) are yours;
|
|
10
|
+
ask only when the choices diverge into materially different product behavior.
|
|
11
|
+
2. Build the smallest usable slice — everything needed for the behavior to appear
|
|
12
|
+
and work. No spec phase, no architect lane on this path.
|
|
13
|
+
3. Wire it into the usage flow — the unwired component is the most likely cheat;
|
|
14
|
+
an import plus a callsite is the minimum proof of wiring.
|
|
15
|
+
4. Verify on the real surface: UI = run it and observe render/interaction;
|
|
16
|
+
function = concrete input→output cases. Build-pass proves nothing.
|
|
17
|
+
5. Record the assumptions you made — one line each.
|
|
18
|
+
Reply: what was added, where it is wired, and the surface evidence it works.
|
|
19
|
+
Ask the human only for: irreversible writes, a genuine preference call no experiment
|
|
20
|
+
settles, or a real dead end. Everything else: do it, report it.
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
{
|
|
2
|
+
"version": 1,
|
|
3
|
+
"playbooks": {
|
|
4
|
+
"small-feature": {
|
|
5
|
+
"launch": "\"create X doing Y\" single-surface",
|
|
6
|
+
"doctor": "repo conventions identified; harness exists?",
|
|
7
|
+
"drive": "implement + wire + render/run",
|
|
8
|
+
"expected": "slice works on real surface",
|
|
9
|
+
"evidence": ["render-observation or IO-case receipt"],
|
|
10
|
+
"cleanup": "remove throwaway scaffolds"
|
|
11
|
+
},
|
|
12
|
+
"feature-implementation": {
|
|
13
|
+
"launch": "\"implement X\" w/ contract",
|
|
14
|
+
"doctor": "done predicate stated; analog found",
|
|
15
|
+
"drive": "predicate → smallest change → verify → widen once",
|
|
16
|
+
"expected": "predicate holds on real artifact",
|
|
17
|
+
"evidence": ["predicate evidence + impact-sweep output"],
|
|
18
|
+
"cleanup": "none"
|
|
19
|
+
},
|
|
20
|
+
"bug-fix": {
|
|
21
|
+
"launch": "\"fix X\" reproducible",
|
|
22
|
+
"doctor": "repro command/surface reachable",
|
|
23
|
+
"drive": "repro → hypothesis bisect → minimal fix → re-verify",
|
|
24
|
+
"expected": "original repro passes same-surface",
|
|
25
|
+
"evidence": ["repro-before + repro-after verbatim, root cause named"],
|
|
26
|
+
"cleanup": "rejected-hypothesis notes kept"
|
|
27
|
+
},
|
|
28
|
+
"investigation": {
|
|
29
|
+
"launch": "\"how/why X\"",
|
|
30
|
+
"doctor": "read-only confirmed",
|
|
31
|
+
"drive": "frame → gather → classify → answer",
|
|
32
|
+
"expected": "cited answer or falsified assumption",
|
|
33
|
+
"evidence": ["path:symbol/command-output anchors"],
|
|
34
|
+
"cleanup": "git status clean check"
|
|
35
|
+
},
|
|
36
|
+
"refactor": {
|
|
37
|
+
"launch": "\"restructure X\" behavior-preserved",
|
|
38
|
+
"doctor": "pin exists or characterization tests writable",
|
|
39
|
+
"drive": "pin → stepwise change → re-pin → migrate callers",
|
|
40
|
+
"expected": "pin green before+after; zero legacy callers",
|
|
41
|
+
"evidence": ["pin receipts + caller sweep"],
|
|
42
|
+
"cleanup": "old surface deleted"
|
|
43
|
+
},
|
|
44
|
+
"performance": {
|
|
45
|
+
"launch": "\"X is slow\" + symptom",
|
|
46
|
+
"doctor": "harness identified; baseline measurable",
|
|
47
|
+
"drive": "baseline → trace → improve → re-measure",
|
|
48
|
+
"expected": "measured delta ≥ stated margin",
|
|
49
|
+
"evidence": ["before/after numbers + commands"],
|
|
50
|
+
"cleanup": "harness left runnable"
|
|
51
|
+
},
|
|
52
|
+
"runtime-forensics": {
|
|
53
|
+
"launch": "live symptom / dropped artifact",
|
|
54
|
+
"doctor": "instrumentation tools present (sample, node --inspect, profilers)",
|
|
55
|
+
"drive": "attach/observe → capture → isolate → name mechanism",
|
|
56
|
+
"expected": "anomaly explained by captured evidence",
|
|
57
|
+
"evidence": ["heap/CPU/handle captures + derived readings"],
|
|
58
|
+
"cleanup": "captures stored under artifacts; processes released"
|
|
59
|
+
},
|
|
60
|
+
"architecture-decision": {
|
|
61
|
+
"launch": "\"A or B\" fork",
|
|
62
|
+
"doctor": "alternatives enumerable",
|
|
63
|
+
"drive": "gather constraints → compare → record",
|
|
64
|
+
"expected": "decision recorded w/ falsifiable rationale",
|
|
65
|
+
"evidence": ["decision record + alternatives"],
|
|
66
|
+
"cleanup": "none"
|
|
67
|
+
},
|
|
68
|
+
"prototype": {
|
|
69
|
+
"launch": "\"spike Y\"",
|
|
70
|
+
"doctor": "decision-to-settle stated",
|
|
71
|
+
"drive": "cheapest artifact → measure → verdict",
|
|
72
|
+
"expected": "fork settled; artifact dispositioned",
|
|
73
|
+
"evidence": ["experiment output + verdict + disposable/promoted flag"],
|
|
74
|
+
"cleanup": "throwaway flagged or removed"
|
|
75
|
+
},
|
|
76
|
+
"migration": {
|
|
77
|
+
"launch": "\"move to new schema/API\"",
|
|
78
|
+
"doctor": "target shape defined; rollback path exists",
|
|
79
|
+
"drive": "pin → migrate → verify on target → rollback demo",
|
|
80
|
+
"expected": "target verified + rollback demonstrated",
|
|
81
|
+
"evidence": ["migration verify + rollback receipts"],
|
|
82
|
+
"cleanup": "legacy path scheduled for removal"
|
|
83
|
+
},
|
|
84
|
+
"verification": {
|
|
85
|
+
"launch": "\"verify/check X\"",
|
|
86
|
+
"doctor": "the named check is runnable",
|
|
87
|
+
"drive": "run check → record verdict",
|
|
88
|
+
"expected": "verdict recorded (pass/fail/inconclusive)",
|
|
89
|
+
"evidence": ["check output verbatim"],
|
|
90
|
+
"cleanup": "none"
|
|
91
|
+
},
|
|
92
|
+
"skill-evaluation": {
|
|
93
|
+
"launch": "\"skill X\" / \"did change help\"",
|
|
94
|
+
"doctor": "fixture set named",
|
|
95
|
+
"drive": "author→validate→ship OR arms→score→verdict",
|
|
96
|
+
"expected": "skill discoverable OR scored comparison recorded",
|
|
97
|
+
"evidence": ["audit output / fixture IDs + scores"],
|
|
98
|
+
"cleanup": "fixtures versioned"
|
|
99
|
+
},
|
|
100
|
+
"session-pickup": {
|
|
101
|
+
"launch": "resume signal",
|
|
102
|
+
"doctor": "resumable record exists",
|
|
103
|
+
"drive": "load → reconstruct → verify vs worktree → continue",
|
|
104
|
+
"expected": "resumed at correct cursor, no re-done steps",
|
|
105
|
+
"evidence": ["state-verify receipt"],
|
|
106
|
+
"cleanup": "stale records pruned"
|
|
107
|
+
},
|
|
108
|
+
"autonomous-run": {
|
|
109
|
+
"launch": "bounded long goal",
|
|
110
|
+
"doctor": "finish condition fixable",
|
|
111
|
+
"drive": "loop: work → verify → re-check finish",
|
|
112
|
+
"expected": "finish verified or escape hatch + named blocker",
|
|
113
|
+
"evidence": ["progress ledger + final verification"],
|
|
114
|
+
"cleanup": "resumable record closed"
|
|
115
|
+
},
|
|
116
|
+
"release": {
|
|
117
|
+
"launch": "\"bump/ship/publish\"",
|
|
118
|
+
"doctor": "on UKit repo; scripts exist (verify-release.mjs)",
|
|
119
|
+
"drive": "version → test:release-core → push+tag → gh release → npm dry-run → publish → parity probes",
|
|
120
|
+
"expected": "3 version probes equal",
|
|
121
|
+
"evidence": ["probe outputs (npm view / gh release list / git tag)"],
|
|
122
|
+
"cleanup": "if parity fails: report \"not shipped\" (owner §4)"
|
|
123
|
+
}
|
|
124
|
+
},
|
|
125
|
+
"artifactClasses": {
|
|
126
|
+
"ui": {
|
|
127
|
+
"name": "UI surface verification",
|
|
128
|
+
"check": "render the changed surface or drive it through the project's UI harness (package.json script `test:ui`) and record the receipt",
|
|
129
|
+
"harness": { "packageScript": "test:ui" },
|
|
130
|
+
"fallback": "receipt class `render-observation` — attest the observed render/interaction via UKIT_EVIDENCE (class=render-observation;surface=<how observed>)",
|
|
131
|
+
"evidenceRequired": ["render-observation|io-case"]
|
|
132
|
+
},
|
|
133
|
+
"cli": {
|
|
134
|
+
"name": "CLI/IO surface verification",
|
|
135
|
+
"check": "run the CLI entry or its IO harness (package.json script `test:cli`) and record the receipt",
|
|
136
|
+
"harness": { "packageScript": "test:cli" },
|
|
137
|
+
"fallback": "receipt class `io-case` — attest one input/output case via UKIT_EVIDENCE (class=io-case;surface=<command>)",
|
|
138
|
+
"evidenceRequired": ["io-case"]
|
|
139
|
+
},
|
|
140
|
+
"docs": {
|
|
141
|
+
"name": "Docs/prose verification",
|
|
142
|
+
"check": "skim the rendered prose for link/format integrity; no test run required",
|
|
143
|
+
"fallback": "receipt class `doc-render` — optional; attest a docs render check via UKIT_EVIDENCE when available",
|
|
144
|
+
"evidenceRequired": []
|
|
145
|
+
},
|
|
146
|
+
"config": {
|
|
147
|
+
"name": "Config/settings verification",
|
|
148
|
+
"check": "load or lint the changed config once (project's own loader/lint command) and record the receipt",
|
|
149
|
+
"fallback": "receipt class `config-load` — attest a config load/lint observation via UKIT_EVIDENCE when no script exists",
|
|
150
|
+
"evidenceRequired": []
|
|
151
|
+
}
|
|
152
|
+
}
|
|
153
|
+
}
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
---
|
|
2
|
+
id: verification
|
|
3
|
+
lanes: [review-release, local-fix]
|
|
4
|
+
---
|
|
5
|
+
You own this check. Run the named check and record its verdict — a failed check is
|
|
6
|
+
a finding to report, not a task to fix.
|
|
7
|
+
If the user says "new task", re-route — do not treat the message as the next step.
|
|
8
|
+
1. Name the check and its verdict predicate — the command, surface, or
|
|
9
|
+
observation that produces pass/fail.
|
|
10
|
+
2. Run it and capture the output verbatim. If this engine cannot produce that
|
|
11
|
+
evidence class, fall back to the closest runnable check and mark the limit
|
|
12
|
+
`unverified` — never fabricate a pass.
|
|
13
|
+
3. Visual channel (skippable — only when parity against a reference is the ask):
|
|
14
|
+
capture reference → render → vision-compare → iterate until equivalence or
|
|
15
|
+
named divergence; prefer DOM/snapshot assertions when they suffice — vision
|
|
16
|
+
passes are the expensive channel.
|
|
17
|
+
4. Record the verdict: pass, fail, or inconclusive with what is missing. A fail
|
|
18
|
+
hands off to the owning group (bug-fix, performance) as a new routed task.
|
|
19
|
+
Reply: the check run, the verbatim output or comparison captures, and the recorded
|
|
20
|
+
verdict.
|
|
21
|
+
Ask the human only for: irreversible writes, a genuine preference call no experiment
|
|
22
|
+
settles, or a real dead end. Everything else: do it, report it.
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
---
|
|
2
|
+
id: worktree-cleanup
|
|
3
|
+
lanes: [local-fix, local-build]
|
|
4
|
+
---
|
|
5
|
+
You own this workspace-hygiene pass. Reclaim stale/merged worktrees — every
|
|
6
|
+
removal sits behind a per-entry safety check, never a blanket sweep. If the
|
|
7
|
+
user says "new task", re-route — do not treat the message as the next step.
|
|
8
|
+
1. Confirm this is cleanup territory: accumulated `.worktrees/*` entries (or
|
|
9
|
+
linked worktrees anywhere) left by completed or abandoned runs. A repo
|
|
10
|
+
with no extra worktrees is already clean — report that and stop.
|
|
11
|
+
2. Enumerate candidates and classify them. Run
|
|
12
|
+
`node .claude/ukit/index/worktree-sweep.mjs` (dry-run is the default — it
|
|
13
|
+
reads `git worktree list --porcelain`, classifies each entry
|
|
14
|
+
merged/stale/active, and prints the classification table with a per-entry
|
|
15
|
+
verdict). It removes nothing.
|
|
16
|
+
3. Read the verdicts. Each non-main worktree is `safe` or a named skip —
|
|
17
|
+
`skip-dirty` (uncommitted work), `skip-locked` (git lock, deliberate),
|
|
18
|
+
`skip-active` (a resumable-run record or a `docs/AI_HANDOFF/RUN.md` cursor
|
|
19
|
+
whose phase is not `done`/`blocked`), `skip-detached-head`. The gate is
|
|
20
|
+
per entry: no uncommitted work AND no active run.
|
|
21
|
+
4. Apply — removal is a destructive operation: a worktree's checkout is
|
|
22
|
+
deleted, so `Ask the human` before running `--apply`, even when every
|
|
23
|
+
candidate reads `safe`. With authorization, run
|
|
24
|
+
`node .claude/ukit/index/worktree-sweep.mjs --apply`; it runs
|
|
25
|
+
`git worktree remove` only on `safe` entries and prints each removal's
|
|
26
|
+
gate verdict. Skipped entries are named in the report and left untouched.
|
|
27
|
+
5. Report and finish: all stale/merged worktrees reclaimed, nothing active or
|
|
28
|
+
dirty touched. Nothing safe to remove, or everything skipped, is a report
|
|
29
|
+
— print the verdicts, never exit silently.
|
|
30
|
+
Machinery: `worktree-sweep.mjs` over `git` porcelain — full on Claude Code
|
|
31
|
+
and omp via Bash; on Codex run the same commands manually — enforcement stays
|
|
32
|
+
advisory/instruction-mediated, never a live gate claim.
|
|
33
|
+
Reply: the classification list with every per-entry verdict, what was removed
|
|
34
|
+
(and under which authorization), and what was skipped.
|
|
35
|
+
Ask the human only for: irreversible writes (every `--apply` removal), a
|
|
36
|
+
genuine preference call no experiment settles, or a real dead end. Everything
|
|
37
|
+
else: do it, report it.
|