@christang/keel 5.66.0 → 5.68.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +63 -0
- package/assets/bootstrap/AGENTS.md +1 -1
- package/bin/keel.js +20 -0
- package/package.json +1 -1
- package/plugins/keel/.claude-plugin/plugin.json +1 -1
- package/plugins/keel/.codex-plugin/plugin.json +1 -1
- package/plugins/keel/skills/keel-run-single-task-goal/SKILL.md +6 -10
- package/plugins/keel/skills/keel-run-single-task-goal/guidance.md +30 -0
- package/plugins/keel/skills/keel-tdd-or-test-first/SKILL.md +4 -1
- package/scripts/validate_plugin.py +824 -8
- package/src/core/config.js +54 -0
- package/src/core/context.js +16 -0
- package/src/core/gates.js +185 -1
- package/src/core/task-contract.js +122 -1
package/README.md
CHANGED
|
@@ -295,6 +295,34 @@ decision either — routing decides whether a change exists, so there is nothing
|
|
|
295
295
|
to; what Keel does is make sure the rule and your exceptions are in front of the agent when it
|
|
296
296
|
decides.
|
|
297
297
|
|
|
298
|
+
### How much guidance the agent loads
|
|
299
|
+
|
|
300
|
+
Keel's skills carry two kinds of content, and their value moves in opposite directions. *How to do it*
|
|
301
|
+
— how to split a task, what order to run things in — matters less the stronger the executor is. *Make
|
|
302
|
+
yourself falsifiable* — red then green, the failure literal a check predicts, the fingerprint, whether
|
|
303
|
+
a `Durable owner:` reference actually exists — matters more, because a strong executor produces
|
|
304
|
+
confident work and those are the checks that can contradict it.
|
|
305
|
+
|
|
306
|
+
So the stepwise half of a skill lives in a `guidance.md` beside it, and a repository can say it does
|
|
307
|
+
not need that half:
|
|
308
|
+
|
|
309
|
+
```yaml
|
|
310
|
+
executor_tier: high
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
The default is `standard`, which reads the guidance; an absent or misspelled declaration reads it too,
|
|
314
|
+
so the worst an unconfigured repository does is pay for a read. `keel context` and `keel --doctor`
|
|
315
|
+
report the tier.
|
|
316
|
+
|
|
317
|
+
**The tier reaches guidance and nothing else** — no gate, criterion, evidence requirement, or Review
|
|
318
|
+
changes with it. That is not a promise in this README: a guidance file is checked to contain none of
|
|
319
|
+
the words Keel states criteria in, so a tier can only ever skip a file that decides nothing. Deciding
|
|
320
|
+
what to skip is a declaration rather than the agent's own call on purpose — "do I need this help?" is
|
|
321
|
+
the judgement a weak executor gets most wrong, and it would be answering it about itself.
|
|
322
|
+
|
|
323
|
+
One skill is split today, `keel-run-single-task-goal`, and its body is 11% smaller for it. The other
|
|
324
|
+
five are almost entirely criteria, so splitting them would move the half that has to stay.
|
|
325
|
+
|
|
298
326
|
## How the agent uses these
|
|
299
327
|
|
|
300
328
|
You rarely type the commands below. The point of Keel is that the discipline runs itself:
|
|
@@ -316,6 +344,41 @@ wiring, and `keel --init` whenever it tells you the repository is behind its ins
|
|
|
316
344
|
protocol version lives in your `AGENTS.md`, and updating the package does not move it.
|
|
317
345
|
Everything below is the vocabulary the agent uses on your behalf.
|
|
318
346
|
|
|
347
|
+
### When the right answer is "nothing changed"
|
|
348
|
+
|
|
349
|
+
A refactor, a move, a flow upgrade that claims the numbers hold — the correct evidence for these is
|
|
350
|
+
*zero difference*, and red-green has no shape for it. The repository that reported this had a task
|
|
351
|
+
whose `tasks.md` said "this one has no honest red" several times, and re-recorded its contract twice
|
|
352
|
+
trying to fit. The criterion was right; it had nowhere to live.
|
|
353
|
+
|
|
354
|
+
```
|
|
355
|
+
- Verify:
|
|
356
|
+
- Strategy: equivalence
|
|
357
|
+
- Base: origin/main
|
|
358
|
+
- Fields: wns, tns, cell_count
|
|
359
|
+
- M1: node compare.js --base --head reports every field equal
|
|
360
|
+
```
|
|
361
|
+
|
|
362
|
+
`equivalence` owes no red. Its criterion is that base and head agree on the fields you named, which is
|
|
363
|
+
*stronger* than red-green — it also catches the change that incidentally moved a result. What the gate
|
|
364
|
+
checks is every way that shape can look complete and compare nothing: a missing `Base:` or `Fields:`,
|
|
365
|
+
an empty field set, a ref that resolves to nothing, and a `Base:` that resolves to HEAD.
|
|
366
|
+
|
|
367
|
+
It is not a way out of red-green. A task declaring `equivalence` while covering a scenario its own
|
|
368
|
+
change *adds* is refused unless a sibling task covers that same entry under a red-green strategy —
|
|
369
|
+
behavior that is new is not behavior that is unchanged, and a task cannot prove both.
|
|
370
|
+
|
|
371
|
+
And Evidence no longer has to retell the output:
|
|
372
|
+
|
|
373
|
+
```
|
|
374
|
+
- M1: artifact openspec/changes/<change>/evidence/compare.json sha256:9f2c…
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
The gate checks the file is there and the digest matches, and refuses a path outside the change's own
|
|
378
|
+
directory, because archiving moves that directory and the pointer would break. Keel hashes the bytes
|
|
379
|
+
and reads nothing inside them — the claim stays yours, and what the digest buys is that the file your
|
|
380
|
+
reviewer opens is the file you meant.
|
|
381
|
+
|
|
319
382
|
## Verification layering
|
|
320
383
|
|
|
321
384
|
Keel splits verification into two layers so a slow suite never blocks your push:
|
package/bin/keel.js
CHANGED
|
@@ -49,6 +49,7 @@ const {
|
|
|
49
49
|
readPrecedentStore,
|
|
50
50
|
readStandingAuthorization,
|
|
51
51
|
readFullModePaths,
|
|
52
|
+
readExecutorTier,
|
|
52
53
|
fullModePathsUnreadableMessage,
|
|
53
54
|
readTriagePolicy,
|
|
54
55
|
triageIssue,
|
|
@@ -1731,6 +1732,7 @@ function runDoctor(options) {
|
|
|
1731
1732
|
printPrecedentSurface(repo);
|
|
1732
1733
|
printTriageSurface(repo);
|
|
1733
1734
|
printRoutingSurface(repo);
|
|
1735
|
+
printExecutorTierSurface(repo);
|
|
1734
1736
|
printFastPrePushSurface(repo);
|
|
1735
1737
|
printSourceRepoCliResolution(repo);
|
|
1736
1738
|
|
|
@@ -1855,6 +1857,24 @@ function printRoutingSurface(repo) {
|
|
|
1855
1857
|
for (const entry of paths) printDoctorLine(entry.path, "Full", entry.reason);
|
|
1856
1858
|
}
|
|
1857
1859
|
|
|
1860
|
+
// Reported whether or not it is declared, because the default is the state a
|
|
1861
|
+
// reader most needs to see: a repository that declared nothing is loading every
|
|
1862
|
+
// skill's guidance and has no other surface that says so.
|
|
1863
|
+
function printExecutorTierSurface(repo) {
|
|
1864
|
+
process.stdout.write("\nExecutor tier:\n");
|
|
1865
|
+
const { declared, tier, unknown, message } = readExecutorTier(repo);
|
|
1866
|
+
if (unknown.length > 0) {
|
|
1867
|
+
printDoctorLine("executor_tier", "unreadable", message);
|
|
1868
|
+
return;
|
|
1869
|
+
}
|
|
1870
|
+
printDoctorLine(
|
|
1871
|
+
"executor_tier",
|
|
1872
|
+
tier,
|
|
1873
|
+
(declared ? "declared in keel/config.yaml" : "undeclared; the default")
|
|
1874
|
+
+ " - affects which skill guidance is read and nothing else"
|
|
1875
|
+
);
|
|
1876
|
+
}
|
|
1877
|
+
|
|
1858
1878
|
function printTriageSurface(repo) {
|
|
1859
1879
|
process.stdout.write("\nUnattended triage:\n");
|
|
1860
1880
|
const { labels, issues, unreadable } = readTriagePolicy(repo);
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "keel",
|
|
3
|
-
"version": "5.
|
|
3
|
+
"version": "5.68.0",
|
|
4
4
|
"description": "Keel OpenSpec execution discipline: stateless continuity, task capsules, deterministic gates, and expectation alignment for Codex and Claude Code.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "TanglmChris",
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "keel",
|
|
3
|
-
"version": "5.
|
|
3
|
+
"version": "5.68.0",
|
|
4
4
|
"description": "Keel OpenSpec execution discipline: stateless continuity, task capsules, deterministic gates, and expectation alignment for Codex and Claude Code.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "TanglmChris",
|
|
@@ -13,16 +13,13 @@ metadata:
|
|
|
13
13
|
|
|
14
14
|
Activate a native goal or subagent runtime to execute exactly one authorized OpenSpec task end to end, while OpenSpec, Git, the task-capsule fingerprint, and deterministic Keel gates stay the only durable authority. The current agent remains the sole holder of write authority and owns Review, gate invocation, the task checkbox, and completion. Where delegation is declared, an authorized delegate may write inside the `Touch` boundary that authority already defined and acquires none of those decisions; the current agent re-runs each `M<n>` check itself before recording Evidence, because a delegate's reported result is a claim and the byte-identity check that validates a read-only helper cannot apply to a writer. A native evaluator declaring success never marks or reports the task complete.
|
|
15
15
|
|
|
16
|
-
##
|
|
16
|
+
## Guidance
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
Read `guidance.md` beside this file before proceeding, unless `keel/config.yaml` declares `executor_tier: high` — it holds the runtime references, the manual sequence, and the per-target detail. Every criterion is here, so an absent or unreadable declaration costs a read and nothing else.
|
|
19
19
|
|
|
20
|
-
|
|
21
|
-
- Codex subagents: https://developers.openai.com/codex/subagents
|
|
22
|
-
- Claude goal execution: https://code.claude.com/docs/en/goal
|
|
23
|
-
- Claude subagents: https://code.claude.com/docs/en/sub-agents
|
|
20
|
+
## License
|
|
24
21
|
|
|
25
|
-
License note:
|
|
22
|
+
License note: the Keel package license (UNLICENSED, all rights reserved by the author). Linking the official runtime docs relicenses nothing; their provenance is recorded beside the links and their content is never pasted into Keel artifacts.
|
|
26
23
|
|
|
27
24
|
## When to activate
|
|
28
25
|
|
|
@@ -60,9 +57,8 @@ Helpers are optional, read-only evidence producers and never a second writer. Co
|
|
|
60
57
|
|
|
61
58
|
## Manual fallback
|
|
62
59
|
|
|
63
|
-
When native activation is unavailable
|
|
60
|
+
When native activation is unavailable, do not fake activation: run the identical lifecycle by hand. The manual loop preserves the same single-task boundary and the same stop boundary; its steps are in `guidance.md`.
|
|
64
61
|
|
|
65
62
|
## Target activation
|
|
66
63
|
|
|
67
|
-
|
|
68
|
-
- Claude: activate one `/goal` whose condition stays within the 4,000-character budget; the evaluator is transcript-only, so surface command and gate evidence explicitly. If hooks are disabled, policy blocks activation, or trust is missing, report the manual fallback.
|
|
64
|
+
Claude and Codex only. Activate exactly one bounded goal for the selected task; a native evaluator declaring success never marks or reports the task complete. The per-target detail is in `guidance.md`.
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
# keel-run-single-task-goal — guidance
|
|
2
|
+
|
|
3
|
+
How to carry out the lifecycle. Every criterion is in `SKILL.md`; nothing here decides whether a task
|
|
4
|
+
may start, pass, or complete.
|
|
5
|
+
|
|
6
|
+
## Authoritative runtime references
|
|
7
|
+
|
|
8
|
+
Provenance: linked, not copied. Their text and trademarks belong to their owners, and Keel paraphrases
|
|
9
|
+
only the activation semantics it needs. License note: see `SKILL.md`.
|
|
10
|
+
|
|
11
|
+
- Codex goal-following: https://learn.chatgpt.com/use-cases/follow-goals
|
|
12
|
+
- Codex subagents: https://developers.openai.com/codex/subagents
|
|
13
|
+
- Claude goal execution: https://code.claude.com/docs/en/goal
|
|
14
|
+
- Claude subagents: https://code.claude.com/docs/en/sub-agents
|
|
15
|
+
|
|
16
|
+
## Running the lifecycle by hand
|
|
17
|
+
|
|
18
|
+
Native activation can be unavailable — no plugin, disabled hooks, managed policy, missing trust, or an
|
|
19
|
+
unsupported surface. Type the same numbered steps `SKILL.md` lists, in the same order: `keel gate task-start`, record the
|
|
20
|
+
fingerprint, `keel project goal … --json` for the view, implement inside `Touch`, surface every result
|
|
21
|
+
in the transcript, `keel gate task-complete`, check the box. What changes is who types them.
|
|
22
|
+
|
|
23
|
+
## Per-target activation notes
|
|
24
|
+
|
|
25
|
+
- **Codex**: where a callable goal or subagent surface exists, activate one bounded goal for the
|
|
26
|
+
selected task and use subagents only as bounded read-only helpers. Without a callable surface, paste
|
|
27
|
+
the exact `keel project goal` command and treat the capability as advisory.
|
|
28
|
+
- **Claude**: activate one `/goal` whose condition fits the 4,000-character budget. The evaluator sees
|
|
29
|
+
the transcript only, so command and gate evidence has to appear there explicitly. If hooks are
|
|
30
|
+
disabled, policy blocks activation, or trust is missing, the manual sequence above applies.
|
|
@@ -20,10 +20,13 @@ Read the selected task's compiled capsule: resolved Acceptance, Verify strategy
|
|
|
20
20
|
- `regression-first`: an observable defect; reproduce it through the public interface first, then prove the fix with the same check.
|
|
21
21
|
- `characterization` / `snapshot-characterization`: deterministic or generated outputs kept stable by byte or snapshot comparison; not downgraded to build success.
|
|
22
22
|
- `rendered-behavior`: interactive surfaces exercised through the real rendered interface; strict red-green optional by cost.
|
|
23
|
-
- `
|
|
23
|
+
- `equivalence`: work whose criterion is that a measured result does not change — a refactor, a move, a flow upgrade claiming the numbers hold. It declares `Base:` (a resolvable git ref) and `Fields:` (the compared field set) beside `Strategy:` and owes no red; the A/B command is an ordinary `M<n>` check. `keel gate task-start` refuses a missing `Base:` or `Fields:`, an empty field set, an unresolvable ref, and a `Base:` that is HEAD, because an A/B against itself always agrees. A task covering a scenario its own change *adds* is refused unless a sibling task covers that same entry under a red-green strategy: behavior that is new is not behavior that is unchanged.
|
|
24
|
+
- `evidence-first`: docs, configuration, or diagnosis work whose checks state the observable artifact or evidence instead of a red-green loop. It is scoped by an absence — nothing here can fail first — which is why it is not the home for `equivalence` work, whose criterion is stronger than red-green rather than missing.
|
|
24
25
|
|
|
25
26
|
Red-green strategies (`vertical-tdd`, `regression-first`) must record concrete per-label `.red` and `.green` Evidence entries for the same check; `keel gate task-complete` rejects absent or pending entries.
|
|
26
27
|
|
|
28
|
+
A check's Evidence may read `artifact <path> sha256:<digest>` instead of retelling the output. `keel gate task-complete` checks the file exists and the digest matches, and refuses a path outside the change's own directory — archiving moves that directory, so a pointer outside it breaks. Keel hashes the bytes and reads nothing inside them: the claim the artifact supports stays yours.
|
|
29
|
+
|
|
27
30
|
## Domain lenses
|
|
28
31
|
|
|
29
32
|
When the change's proposal/design/specs or the task's Touch extensions signal a domain, consult the matching lens's `Execution and review checks` section from `keel/lenses/` — the lens whose `Applies when:` header matches — before finalizing the strategy and the first check, and load only that one. When no lens matches, load nothing.
|