skill-harness 0.15.0 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +12 -0
- package/dist/index.js +14063 -6721
- package/docs/PRODUCT-GUIDE.md +190 -0
- package/docs/STATUS.md +56 -0
- package/package.json +4 -2
|
@@ -0,0 +1,190 @@
|
|
|
1
|
+
# Factory: work, review, learn
|
|
2
|
+
|
|
3
|
+
**Release candidate: skill-harness 0.17.0**, paired with **pi-daddy 0.27.3**.
|
|
4
|
+
Principal 3.2.0 is unchanged. The session retrospective feature is merged in PR82
|
|
5
|
+
with independent Sol review and green CI. This release adds explicit local session
|
|
6
|
+
import, timeline/usage coverage and evidence-linked notes. The producer adds a
|
|
7
|
+
visible count and toggle for completed ledger history.
|
|
8
|
+
|
|
9
|
+
Canonical packaging, installed verification and publication are recorded separately
|
|
10
|
+
in the release PR. Use a fresh Pi session after updating the harness: an existing
|
|
11
|
+
immutable dashboard bridge survives reload. Other running sessions need not stop.
|
|
12
|
+
|
|
13
|
+
## Who does what?
|
|
14
|
+
|
|
15
|
+
**pi-daddy** coordinates permitted work and owns execution, steering and the future-order
|
|
16
|
+
registry. **Herdr** displays its sessions and side panel. **skill-harness** retains
|
|
17
|
+
measurement and learning evidence. The panel is not an autonomous worker. Finished
|
|
18
|
+
execution is not human acceptance.
|
|
19
|
+
|
|
20
|
+
Use the producer's `/grants work` guided setup for multiple obligations, dependencies,
|
|
21
|
+
priority and explicit agent/model/thinking selections, then its supported run/host controls.
|
|
22
|
+
For existing single-task use, `pi-daddy work add --id report --outcome "Review the report"`
|
|
23
|
+
remains supported. Only listed controls are available. Pausing new dispatch does not
|
|
24
|
+
cancel running work; changing a next-order policy does not mutate active sessions.
|
|
25
|
+
Principal's ordinary current/stale/superseded reports remain unchanged.
|
|
26
|
+
|
|
27
|
+
## Open learning in Pi
|
|
28
|
+
|
|
29
|
+
With both matching extensions loaded, use **`/grants learning`** from the current host.
|
|
30
|
+
The producer binds the exact current snapshot, archive, population and local review author.
|
|
31
|
+
A changed scope must report a mismatch, not silently reuse an old learning workspace.
|
|
32
|
+
|
|
33
|
+
The harness also exposes **`/skill-harness learning --state "WORKSPACE"`** (use the
|
|
34
|
+
workspace path displayed by the host), or `skill-harness learning --state "WORKSPACE"`
|
|
35
|
+
in a terminal. These open the **same guided flow and durable records**, not another app.
|
|
36
|
+
|
|
37
|
+
1. **Readiness** explains what's retained, what's missing, and why automatic questions
|
|
38
|
+
are silent. Opening a review is deliberate: it neither reserves attention nor
|
|
39
|
+
represents an earned automatic question. It starts no model or delegation.
|
|
40
|
+
2. **Connect retained input** lists compatible case batches and qualified comparisons
|
|
41
|
+
already in the selected archive. Pick one and give it a short name and title.
|
|
42
|
+
No hand-written manifests or ledger JSON. No compatible evidence is an honest empty state.
|
|
43
|
+
3. **Review cases** shows nominations and lets you open their actual retained evidence
|
|
44
|
+
sources by selection (CLI: `learning evidence BATCH --item N --evidence N`). Redacted
|
|
45
|
+
or unavailable sources are marked, never called complete evidence. Record confirmed defect, expected
|
|
46
|
+
behavior, exemplar, uncertain or skip, with a reason. Corrections are append-only.
|
|
47
|
+
Confirming a case is not promoting a test or establishing a cause.
|
|
48
|
+
4. **Review comparisons** opens complete retained artifacts under A/B labels. Acknowledge
|
|
49
|
+
each complete artifact only after reviewing it. Then choose one, tie, none acceptable,
|
|
50
|
+
or insufficient evidence. **Reveal model/cost** is a separate action, enabled only after
|
|
51
|
+
the durable quality receipt. Artifact contents themselves may disclose identity.
|
|
52
|
+
5. **Record adopt / reject / defer** records a separate scope-bound decision. Reject and defer
|
|
53
|
+
are normal outcomes; neither forces a preference or reveal. Missing evidence can be
|
|
54
|
+
durably deferred. Adopt is intent, not activation.
|
|
55
|
+
6. **Propose hypothesis** records prediction, disproof, downside, rollback and a competing
|
|
56
|
+
explanation from an actual case. **Link original hypothesis** only accepts the comparison's
|
|
57
|
+
actual frozen origin, never a retrospective replacement.
|
|
58
|
+
|
|
59
|
+
Previously retained excerpt feedback concerned
|
|
60
|
+
**excerpts**. Keep that receipt unchanged and linked as prior feedback. It is not full-artifact
|
|
61
|
+
acceptance, detector calibration or a fresh blind experiment. An old unscoped quality choice
|
|
62
|
+
also does not automatically become a full-artifact acknowledgement.
|
|
63
|
+
|
|
64
|
+
## Configure trust without writing manifests
|
|
65
|
+
|
|
66
|
+
Open **Trust / independent labels** in the same guided flow:
|
|
67
|
+
|
|
68
|
+
- Before freezing, explicitly add any real unflagged incident evidence and select its split.
|
|
69
|
+
Only selected evidence files are retained; do not include private sessions or secrets.
|
|
70
|
+
- Select a retained case batch, detector and evaluation split. The product derives the
|
|
71
|
+
cohort and predictions from actual nominations and groups repeated target observations
|
|
72
|
+
into one incident. Detector version and population remain exact.
|
|
73
|
+
- Automatic questions default to **silent**. To configure them, explicitly select minimum
|
|
74
|
+
resolved independent incidents, the confidence lower bound, attention budget and expiry.
|
|
75
|
+
The form previews the exact frozen policy. It creates **no independent labels**.
|
|
76
|
+
- Review cases independently, then choose **Link current case label**. Unflagged sample
|
|
77
|
+
labels have their own explicit evidence/author step. Preference, model self-report,
|
|
78
|
+
panel agreement and exit zero are not independent correctness labels.
|
|
79
|
+
|
|
80
|
+
Missing calibration, missing unflagged samples, held-out absence, conflicting labels,
|
|
81
|
+
expired policy and exhausted attention stay visible. Reopening neither resets choices nor
|
|
82
|
+
refills attention. Correcting a linked case withdraws its old label from confidence.
|
|
83
|
+
The original producer closing path still requires presence, quiescence and its exposure gate.
|
|
84
|
+
|
|
85
|
+
## Scoped adoption and later work
|
|
86
|
+
|
|
87
|
+
Use **`/grants learning` in interactive Pi** for activation/rollback. The matching producer
|
|
88
|
+
implements two separate profiles: unchanged **`fixed-policy-v1`**, and
|
|
89
|
+
**`ordinary-work-policy-v1`** for model/effort-only changes to the same selected task IDs and
|
|
90
|
+
agent definitions. The latter does not change instructions, skills, grants, capabilities,
|
|
91
|
+
topology or assessment policy; do not disguise those changes as model/effort or fixed policy.
|
|
92
|
+
The producer prepares policy bytes and owns independent authority and the original registry.
|
|
93
|
+
Source/built integration has been checked with synthetic fixtures; final installed-release
|
|
94
|
+
verification remains pending. Bare legacy fixed-policy views lacking applied activation
|
|
95
|
+
proof defer; they must not be treated as activation receipts.
|
|
96
|
+
|
|
97
|
+
Activation needs the exact selected candidate, original confirmed case/hypothesis/comparison,
|
|
98
|
+
current independently verified eligible facts, explicit adoption authority, and registry
|
|
99
|
+
revision consent. The original registry revalidates these again. Only its actual activation
|
|
100
|
+
receipt advances the learning record to activated. New ordinary-work orders pin the policy
|
|
101
|
+
under fresh eligibility confirmation; active or history-bearing runs never migrate.
|
|
102
|
+
Rollback similarly needs its own authorized request and original applied registry receipt.
|
|
103
|
+
A stored lifecycle link, human preference or prepared receipt cannot substitute for either.
|
|
104
|
+
|
|
105
|
+
**Adoption / outcomes** can inspect linked later observations or record one using complete
|
|
106
|
+
artifact, original/current requirement and independent evidence files. Unknown is valid.
|
|
107
|
+
Acceptance requires both the exact accepted artifact and independent acceptance evidence.
|
|
108
|
+
Changed requirements, changed artifact, out-of-scope, success, caught defect and escape remain
|
|
109
|
+
distinct. No outcome triggers automatic rollback, and no single outcome proves improvement.
|
|
110
|
+
|
|
111
|
+
## CLI and local archive use
|
|
112
|
+
|
|
113
|
+
```sh
|
|
114
|
+
skill-harness learning status --state "WORKSPACE" # human readiness
|
|
115
|
+
skill-harness learning status --state "WORKSPACE" --json # integration/details
|
|
116
|
+
skill-harness learning review reports --state "WORKSPACE"
|
|
117
|
+
skill-harness learning help # all explicit operations
|
|
118
|
+
skill-harness learning guide # this guide, packaged offline
|
|
119
|
+
skill-harness learning current # packaged requirement register
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
Without a producer host, `skill-harness learning` guides connection to an explicitly selected
|
|
123
|
+
archive and selecting an actual retained case scope. Or use
|
|
124
|
+
`learning init --archive DIR --author NAME --confirm`; multiple retained scopes ask you to
|
|
125
|
+
select `--scope-item N`. Snapshot digest and population are derived from actual cases, not
|
|
126
|
+
typed hashes. For a comparison-only archive with no cases, `--scope TEXT --population TEXT`
|
|
127
|
+
can explicitly declare a local review scope; it cannot bind mismatching case snapshots.
|
|
128
|
+
Use the producer flow to connect a host. Noninteractive writes require explicit `--confirm`.
|
|
129
|
+
All learning commands are offline/no-model. Command completion is never a quality verdict.
|
|
130
|
+
Local author attribution and content hashes are not authentication or hostile-owner attestation.
|
|
131
|
+
|
|
132
|
+
Current requirements and remaining evidence: [STATUS.md](STATUS.md).
|
|
133
|
+
|
|
134
|
+
## Whole-session retrospective (0.17.0)
|
|
135
|
+
|
|
136
|
+
Use this after a real task to inspect the conversation, model/thinking settings,
|
|
137
|
+
tool use, individual delegation attempts and provider-reported usage. The importer
|
|
138
|
+
reads only explicitly selected files and never contacts or interrupts an agent.
|
|
139
|
+
|
|
140
|
+
Start with the terminal CLI (or prefix these arguments with
|
|
141
|
+
"/skill-harness learning" in Pi instead of "skill-harness learning"):
|
|
142
|
+
|
|
143
|
+
skill-harness learning session preview --session /absolute/parent.jsonl --ledger /absolute/grants.jsonl
|
|
144
|
+
skill-harness learning session import --session /absolute/parent.jsonl --ledger /absolute/grants.jsonl --name task --title "Task retrospective" --expected DIGEST_FROM_PREVIEW --confirm --state /private/learning/workspace --archive /private/learning/archive --author operator
|
|
145
|
+
skill-harness learning session show task --state /private/learning/workspace
|
|
146
|
+
skill-harness learning session source task --item 1 --state /private/learning/workspace
|
|
147
|
+
|
|
148
|
+
The first import can create a workspace. For an existing workspace omit --archive
|
|
149
|
+
and --author. Optional --feedback FILE and --artifact FILE retain selected external
|
|
150
|
+
feedback and artifact/source evidence. Use --sources FILE for a JSON array of
|
|
151
|
+
additional selections, for example:
|
|
152
|
+
|
|
153
|
+
[{"kind":"child","path":"/absolute/child.jsonl","executionId":"exec:10000000-0000-4000-8000-000000000001"}]
|
|
154
|
+
|
|
155
|
+
Every source must be a regular, unlinked file: at most 16 files, 8 MiB each,
|
|
156
|
+
32 MiB total; Pi/ledger JSONL is bounded to 4096 records with complete final lines.
|
|
157
|
+
Pi session format v3 and the existing frozen pi-daddy v3 ledger contract are
|
|
158
|
+
supported. Re-preview if any selected bytes change. The ledger association and
|
|
159
|
+
child execution mapping are operator declarations, not authenticated parentage.
|
|
160
|
+
|
|
161
|
+
**Privacy:** import retains exact original bytes locally, potentially including
|
|
162
|
+
credentials, reasoning, tool arguments or private text present in the source.
|
|
163
|
+
Preview prints paths, hashes and coverage, not a redacted copy. Inspect selected
|
|
164
|
+
inputs before confirming; no scanning, automatic redaction or publication occurs.
|
|
165
|
+
Full-source readback deliberately displays the selected source; terminal control
|
|
166
|
+
characters are escaped for display without altering archived bytes.
|
|
167
|
+
|
|
168
|
+
Open the existing learning wizard with the same --state path to choose **Import
|
|
169
|
+
session** or **Review sessions**. Editors display evidence; edits are discarded.
|
|
170
|
+
The ordinary producer "/grants learning" bridge also reaches those menu items
|
|
171
|
+
when its current host workspace is already connected. Reading does not reserve
|
|
172
|
+
attention, run a model or change a next-order policy.
|
|
173
|
+
|
|
174
|
+
Record a finding, proposal, acceptance observation or context note against a
|
|
175
|
+
specific retained source line:
|
|
176
|
+
|
|
177
|
+
skill-harness learning session note task --kind proposal --note "Ask for the capability map before implementation; compare this on a later task." --item 1 --line 2 --confirm --state /private/learning/workspace
|
|
178
|
+
|
|
179
|
+
Notes identify the configured local author. An acceptance note is an operator
|
|
180
|
+
annotation, not an authenticated user approval or qualified comparison.
|
|
181
|
+
The timeline is chronological and may include abandoned branches; excerpts are
|
|
182
|
+
labelled. It does not reconstruct the active branch or verify a test claimed in prose.
|
|
183
|
+
Child usage and missing evidence are reported separately; totals are provider
|
|
184
|
+
observations, not billing or a complete parent-plus-child cost.
|
|
185
|
+
|
|
186
|
+
This slice makes sessions reviewable. It does **not** automatically generate or
|
|
187
|
+
measure skill/tool/instruction improvements. A later experiment still needs a
|
|
188
|
+
frozen task and candidate, appropriate independent assessment and actual retained
|
|
189
|
+
measurements before claiming faster or equally good output. Existing adoption
|
|
190
|
+
supports its documented model/effort policy scope only; annotations do not expand it.
|
package/docs/STATUS.md
ADDED
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# Factory requirement register
|
|
2
|
+
|
|
3
|
+
Current entry point: [PRODUCT-GUIDE.md](PRODUCT-GUIDE.md). The dated
|
|
4
|
+
[CURRENT.md](CURRENT.md), build reports and transport decisions are retained history,
|
|
5
|
+
not current instructions to supply manifests or judge engineering repairs.
|
|
6
|
+
|
|
7
|
+
## 0.17.0 release candidate — session learning
|
|
8
|
+
|
|
9
|
+
Adds explicit Pi-session import, exact private retention, chronological evidence,
|
|
10
|
+
per-execution delegation history, parent/child usage coverage, and line-bound
|
|
11
|
+
operator annotations through the existing CLI/Pi learning workspace. No model
|
|
12
|
+
calls, automatic acceptance, calibrated improvement or expanded adoption authority.
|
|
13
|
+
See the whole-session section of PRODUCT-GUIDE.md. Feature PR82 is merged with independent review and green CI. Release metadata
|
|
14
|
+
pairs harness 0.17.0 with producer 0.27.3; canonical publication and installed user
|
|
15
|
+
testing remain separate. Entire-session optimization and measured skill/tool variants remain
|
|
16
|
+
follow-up work.
|
|
17
|
+
|
|
18
|
+
## Historical 0.16.0 release preparation
|
|
19
|
+
|
|
20
|
+
Release pair: **skill-harness 0.16.0 / pi-daddy 0.27.0**; Principal 3.2.0 is unchanged.
|
|
21
|
+
The prior audited baseline was harness 0.15.0 (`a725579`) / producer 0.26.1 (`93a9b01`).
|
|
22
|
+
Harness PR80 merged at `c6cd55d4db9ccf9a1ec790459c3cc5e6b1eac69f` with green CI and
|
|
23
|
+
Sol approval. Actual source/built cross-package learning/registry checks passed with
|
|
24
|
+
synthetic fixtures. This is **release preparation**, not publication, final installed
|
|
25
|
+
proof, human approval, calibration or measured improvement. Loaded resources, source
|
|
26
|
+
HEAD, canonical archives, published bytes and active sessions remain separate identities.
|
|
27
|
+
Canonical packing, publication and final installed verification are pending.
|
|
28
|
+
|
|
29
|
+
| Requirement | Current implementation / boundary | Remaining evidence |
|
|
30
|
+
|---|---|---|
|
|
31
|
+
| Understandable side panel | Producer owns selected-outcome-first status, observed agents, available controls and Details. | Matching producer's installed narrow-pane/host checks. No invented progress/ETA/acceptance. |
|
|
32
|
+
| Ordinary multi-obligation work | Producer owns work setup, dependencies, bounded parallel/sequential execution, model/effort choice and revision consent. | Installed checks of the matching producer's supported profiles. Existing Principal workflows remain unchanged. |
|
|
33
|
+
| Retained cases/comparisons | Harness learning workspace composes existing archive, case-review, blind-choice and CAS journal primitives. Named bindings and guided CLI/Pi entrypoints; missing/stale/mismatched inputs defer. | Final installed producer connection; source review/CI and synthetic source/built linkage passed. No new model evidence claimed. |
|
|
34
|
+
| Complete quality review | Complete-output acknowledgement, durable choice before reveal, immutable replay, separate adopt/reject/defer. Prior excerpt feedback stays separate. | A real full-artifact product choice remains absent where not actually recorded. Excerpt preference is not acceptance. |
|
|
35
|
+
| Scope-bound trust | Guided frozen cohort/policy, actual nomination predictions, incident grouping, explicit independent case/unflagged labels and pure exposure preview. Existing automatic gates unchanged. | Real independent labels, unflagged coverage and held-out evidence where absent. Implementation cannot manufacture them. |
|
|
36
|
+
| Hypothesis and comparison linkage | Human-authored falsifiable hypothesis form, confirmed-case linkage, monotonic optional context and retained lifecycle navigation. | Ongoing weekly model-backed retro is not demonstrated or automatically scheduled. Existing bounded weekly/model primitives remain separate authorized paths. |
|
|
37
|
+
| Adoption and rollback | Harness validates scope/choice/current independent authority and eligible facts. Producer source supports unchanged `fixed-policy-v1` plus separate `ordinary-work-policy-v1` (model/effort only, same task IDs/agent definitions and scope/assessment). Producer owns activation/rollback and fresh-eligibility next-order pins; profile-agnostic harness linkage retains original receipts. | Synthetic source/built original activation/rollback linkage passed. Configured production adoption/rollback and final installed integration remain unverified; bare legacy views without applied activation proof defer. No active/history-bearing run migration or instruction/skill/grant/capability/topology/assessment expansion. |
|
|
38
|
+
| Later outcomes | Explicit file-based CLI/Pi observation forms plus producer API linkage; independent reference attribution; changed requirements/artifacts/out-of-scope/unknown remain distinct. | Genuine later work, exact acceptance and held-out efficacy where absent. No automatic rollback or universal quality score. |
|
|
39
|
+
| Packaging / portability | Public adapter exports, additive frozen dashboard bridge, CLI export and bundled Pi entrypoint; guide/register shipped offline via `learning guide` / `learning current`. A pre-existing immutable bridge needs a fresh Pi session. | Release-metadata review/CI, exact merged canonical pack with retained archives, empty-prefix installed Pi/registry verification, then authorized publication. |
|
|
40
|
+
|
|
41
|
+
## Historical claims remain bounded
|
|
42
|
+
|
|
43
|
+
- The failure-derived abort-race case and the public-weather comparison are different chains.
|
|
44
|
+
- Previously retained feedback covered excerpts only. It does not advance full comparison
|
|
45
|
+
acceptance or detector trust.
|
|
46
|
+
- Earlier no-tool measured SDK observations do not establish universal ordinary-child Usage,
|
|
47
|
+
writable-effect containment or actual provider-internal reasoning effort.
|
|
48
|
+
- No OS sandbox or malicious-local-owner protection has been added.
|
|
49
|
+
- No historical session/evidence, auth or unrelated worktree was modified in this candidate.
|
|
50
|
+
|
|
51
|
+
Feature source review/CI and source/built cross-package checks are completed evidence,
|
|
52
|
+
not final released-install proof. Metadata preparation runs consistency/diff validation only:
|
|
53
|
+
no builds, installs, suites or model calls. Canonical pack and actual installed two-extension
|
|
54
|
+
verification remain release gates under `PUBLISHING.md`. Prior live route smokes remain
|
|
55
|
+
historical and are not replayed merely for a version bump. Human acceptance, calibration and
|
|
56
|
+
later efficacy still require their own real evidence.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "skill-harness",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.17.0",
|
|
4
4
|
"description": "Test/optimize loop for agent skills — run spec'd scenarios on pi, LLM-judge, score, review, re-run",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"keywords": [
|
|
@@ -28,11 +28,13 @@
|
|
|
28
28
|
"dist/prompt-capture-extension.js",
|
|
29
29
|
"assets/report.template.html",
|
|
30
30
|
"assets/report.grade.js",
|
|
31
|
+
"docs/PRODUCT-GUIDE.md",
|
|
32
|
+
"docs/STATUS.md",
|
|
31
33
|
"LICENSE",
|
|
32
34
|
"README.md"
|
|
33
35
|
],
|
|
34
36
|
"dependencies": {
|
|
35
|
-
"@skill-harness/cli": "0.
|
|
37
|
+
"@skill-harness/cli": "0.17.0",
|
|
36
38
|
"typebox": "^1.1.38"
|
|
37
39
|
},
|
|
38
40
|
"repository": {
|