skill-harness 0.15.0 → 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,190 @@
1
+ # Factory: work, review, learn
2
+
3
+ **Release candidate: skill-harness 0.17.0**, paired with **pi-daddy 0.27.3**.
4
+ Principal 3.2.0 is unchanged. The session retrospective feature is merged in PR82
5
+ with independent Sol review and green CI. This release adds explicit local session
6
+ import, timeline/usage coverage and evidence-linked notes. The producer adds a
7
+ visible count and toggle for completed ledger history.
8
+
9
+ Canonical packaging, installed verification and publication are recorded separately
10
+ in the release PR. Use a fresh Pi session after updating the harness: an existing
11
+ immutable dashboard bridge survives reload. Other running sessions need not stop.
12
+
13
+ ## Who does what?
14
+
15
+ **pi-daddy** coordinates permitted work and owns execution, steering and the future-order
16
+ registry. **Herdr** displays its sessions and side panel. **skill-harness** retains
17
+ measurement and learning evidence. The panel is not an autonomous worker. Finished
18
+ execution is not human acceptance.
19
+
20
+ Use the producer's `/grants work` guided setup for multiple obligations, dependencies,
21
+ priority and explicit agent/model/thinking selections, then its supported run/host controls.
22
+ For existing single-task use, `pi-daddy work add --id report --outcome "Review the report"`
23
+ remains supported. Only listed controls are available. Pausing new dispatch does not
24
+ cancel running work; changing a next-order policy does not mutate active sessions.
25
+ Principal's ordinary current/stale/superseded reports remain unchanged.
26
+
27
+ ## Open learning in Pi
28
+
29
+ With both matching extensions loaded, use **`/grants learning`** from the current host.
30
+ The producer binds the exact current snapshot, archive, population and local review author.
31
+ A changed scope must report a mismatch, not silently reuse an old learning workspace.
32
+
33
+ The harness also exposes **`/skill-harness learning --state "WORKSPACE"`** (use the
34
+ workspace path displayed by the host), or `skill-harness learning --state "WORKSPACE"`
35
+ in a terminal. These open the **same guided flow and durable records**, not another app.
36
+
37
+ 1. **Readiness** explains what's retained, what's missing, and why automatic questions
38
+ are silent. Opening a review is deliberate: it neither reserves attention nor
39
+ represents an earned automatic question. It starts no model or delegation.
40
+ 2. **Connect retained input** lists compatible case batches and qualified comparisons
41
+ already in the selected archive. Pick one and give it a short name and title.
42
+ No hand-written manifests or ledger JSON. No compatible evidence is an honest empty state.
43
+ 3. **Review cases** shows nominations and lets you open their actual retained evidence
44
+ sources by selection (CLI: `learning evidence BATCH --item N --evidence N`). Redacted
45
+ or unavailable sources are marked, never called complete evidence. Record confirmed defect, expected
46
+ behavior, exemplar, uncertain or skip, with a reason. Corrections are append-only.
47
+ Confirming a case is not promoting a test or establishing a cause.
48
+ 4. **Review comparisons** opens complete retained artifacts under A/B labels. Acknowledge
49
+ each complete artifact only after reviewing it. Then choose one, tie, none acceptable,
50
+ or insufficient evidence. **Reveal model/cost** is a separate action, enabled only after
51
+ the durable quality receipt. Artifact contents themselves may disclose identity.
52
+ 5. **Record adopt / reject / defer** records a separate scope-bound decision. Reject and defer
53
+ are normal outcomes; neither forces a preference or reveal. Missing evidence can be
54
+ durably deferred. Adopt is intent, not activation.
55
+ 6. **Propose hypothesis** records prediction, disproof, downside, rollback and a competing
56
+ explanation from an actual case. **Link original hypothesis** only accepts the comparison's
57
+ actual frozen origin, never a retrospective replacement.
58
+
59
+ Previously retained excerpt feedback concerned
60
+ **excerpts**. Keep that receipt unchanged and linked as prior feedback. It is not full-artifact
61
+ acceptance, detector calibration or a fresh blind experiment. An old unscoped quality choice
62
+ also does not automatically become a full-artifact acknowledgement.
63
+
64
+ ## Configure trust without writing manifests
65
+
66
+ Open **Trust / independent labels** in the same guided flow:
67
+
68
+ - Before freezing, explicitly add any real unflagged incident evidence and select its split.
69
+ Only selected evidence files are retained; do not include private sessions or secrets.
70
+ - Select a retained case batch, detector and evaluation split. The product derives the
71
+ cohort and predictions from actual nominations and groups repeated target observations
72
+ into one incident. Detector version and population remain exact.
73
+ - Automatic questions default to **silent**. To configure them, explicitly select minimum
74
+ resolved independent incidents, the confidence lower bound, attention budget and expiry.
75
+ The form previews the exact frozen policy. It creates **no independent labels**.
76
+ - Review cases independently, then choose **Link current case label**. Unflagged sample
77
+ labels have their own explicit evidence/author step. Preference, model self-report,
78
+ panel agreement and exit zero are not independent correctness labels.
79
+
80
+ Missing calibration, missing unflagged samples, held-out absence, conflicting labels,
81
+ expired policy and exhausted attention stay visible. Reopening neither resets choices nor
82
+ refills attention. Correcting a linked case withdraws its old label from confidence.
83
+ The original producer closing path still requires presence, quiescence and its exposure gate.
84
+
85
+ ## Scoped adoption and later work
86
+
87
+ Use **`/grants learning` in interactive Pi** for activation/rollback. The matching producer
88
+ implements two separate profiles: unchanged **`fixed-policy-v1`**, and
89
+ **`ordinary-work-policy-v1`** for model/effort-only changes to the same selected task IDs and
90
+ agent definitions. The latter does not change instructions, skills, grants, capabilities,
91
+ topology or assessment policy; do not disguise those changes as model/effort or fixed policy.
92
+ The producer prepares policy bytes and owns independent authority and the original registry.
93
+ Source/built integration has been checked with synthetic fixtures; final installed-release
94
+ verification remains pending. Bare legacy fixed-policy views lacking applied activation
95
+ proof defer; they must not be treated as activation receipts.
96
+
97
+ Activation needs the exact selected candidate, original confirmed case/hypothesis/comparison,
98
+ current independently verified eligible facts, explicit adoption authority, and registry
99
+ revision consent. The original registry revalidates these again. Only its actual activation
100
+ receipt advances the learning record to activated. New ordinary-work orders pin the policy
101
+ under fresh eligibility confirmation; active or history-bearing runs never migrate.
102
+ Rollback similarly needs its own authorized request and original applied registry receipt.
103
+ A stored lifecycle link, human preference or prepared receipt cannot substitute for either.
104
+
105
+ **Adoption / outcomes** can inspect linked later observations or record one using complete
106
+ artifact, original/current requirement and independent evidence files. Unknown is valid.
107
+ Acceptance requires both the exact accepted artifact and independent acceptance evidence.
108
+ Changed requirements, changed artifact, out-of-scope, success, caught defect and escape remain
109
+ distinct. No outcome triggers automatic rollback, and no single outcome proves improvement.
110
+
111
+ ## CLI and local archive use
112
+
113
+ ```sh
114
+ skill-harness learning status --state "WORKSPACE" # human readiness
115
+ skill-harness learning status --state "WORKSPACE" --json # integration/details
116
+ skill-harness learning review reports --state "WORKSPACE"
117
+ skill-harness learning help # all explicit operations
118
+ skill-harness learning guide # this guide, packaged offline
119
+ skill-harness learning current # packaged requirement register
120
+ ```
121
+
122
+ Without a producer host, `skill-harness learning` guides connection to an explicitly selected
123
+ archive and selecting an actual retained case scope. Or use
124
+ `learning init --archive DIR --author NAME --confirm`; multiple retained scopes ask you to
125
+ select `--scope-item N`. Snapshot digest and population are derived from actual cases, not
126
+ typed hashes. For a comparison-only archive with no cases, `--scope TEXT --population TEXT`
127
+ can explicitly declare a local review scope; it cannot bind mismatching case snapshots.
128
+ Use the producer flow to connect a host. Noninteractive writes require explicit `--confirm`.
129
+ All learning commands are offline/no-model. Command completion is never a quality verdict.
130
+ Local author attribution and content hashes are not authentication or hostile-owner attestation.
131
+
132
+ Current requirements and remaining evidence: [STATUS.md](STATUS.md).
133
+
134
+ ## Whole-session retrospective (0.17.0)
135
+
136
+ Use this after a real task to inspect the conversation, model/thinking settings,
137
+ tool use, individual delegation attempts and provider-reported usage. The importer
138
+ reads only explicitly selected files and never contacts or interrupts an agent.
139
+
140
+ Start with the terminal CLI (or prefix these arguments with
141
+ "/skill-harness learning" in Pi instead of "skill-harness learning"):
142
+
143
+ skill-harness learning session preview --session /absolute/parent.jsonl --ledger /absolute/grants.jsonl
144
+ skill-harness learning session import --session /absolute/parent.jsonl --ledger /absolute/grants.jsonl --name task --title "Task retrospective" --expected DIGEST_FROM_PREVIEW --confirm --state /private/learning/workspace --archive /private/learning/archive --author operator
145
+ skill-harness learning session show task --state /private/learning/workspace
146
+ skill-harness learning session source task --item 1 --state /private/learning/workspace
147
+
148
+ The first import can create a workspace. For an existing workspace omit --archive
149
+ and --author. Optional --feedback FILE and --artifact FILE retain selected external
150
+ feedback and artifact/source evidence. Use --sources FILE for a JSON array of
151
+ additional selections, for example:
152
+
153
+ [{"kind":"child","path":"/absolute/child.jsonl","executionId":"exec:10000000-0000-4000-8000-000000000001"}]
154
+
155
+ Every source must be a regular, unlinked file: at most 16 files, 8 MiB each,
156
+ 32 MiB total; Pi/ledger JSONL is bounded to 4096 records with complete final lines.
157
+ Pi session format v3 and the existing frozen pi-daddy v3 ledger contract are
158
+ supported. Re-preview if any selected bytes change. The ledger association and
159
+ child execution mapping are operator declarations, not authenticated parentage.
160
+
161
+ **Privacy:** import retains exact original bytes locally, potentially including
162
+ credentials, reasoning, tool arguments or private text present in the source.
163
+ Preview prints paths, hashes and coverage, not a redacted copy. Inspect selected
164
+ inputs before confirming; no scanning, automatic redaction or publication occurs.
165
+ Full-source readback deliberately displays the selected source; terminal control
166
+ characters are escaped for display without altering archived bytes.
167
+
168
+ Open the existing learning wizard with the same --state path to choose **Import
169
+ session** or **Review sessions**. Editors display evidence; edits are discarded.
170
+ The ordinary producer "/grants learning" bridge also reaches those menu items
171
+ when its current host workspace is already connected. Reading does not reserve
172
+ attention, run a model or change a next-order policy.
173
+
174
+ Record a finding, proposal, acceptance observation or context note against a
175
+ specific retained source line:
176
+
177
+ skill-harness learning session note task --kind proposal --note "Ask for the capability map before implementation; compare this on a later task." --item 1 --line 2 --confirm --state /private/learning/workspace
178
+
179
+ Notes identify the configured local author. An acceptance note is an operator
180
+ annotation, not an authenticated user approval or qualified comparison.
181
+ The timeline is chronological and may include abandoned branches; excerpts are
182
+ labelled. It does not reconstruct the active branch or verify a test claimed in prose.
183
+ Child usage and missing evidence are reported separately; totals are provider
184
+ observations, not billing or a complete parent-plus-child cost.
185
+
186
+ This slice makes sessions reviewable. It does **not** automatically generate or
187
+ measure skill/tool/instruction improvements. A later experiment still needs a
188
+ frozen task and candidate, appropriate independent assessment and actual retained
189
+ measurements before claiming faster or equally good output. Existing adoption
190
+ supports its documented model/effort policy scope only; annotations do not expand it.
package/docs/STATUS.md ADDED
@@ -0,0 +1,56 @@
1
+ # Factory requirement register
2
+
3
+ Current entry point: [PRODUCT-GUIDE.md](PRODUCT-GUIDE.md). The dated
4
+ [CURRENT.md](CURRENT.md), build reports and transport decisions are retained history,
5
+ not current instructions to supply manifests or judge engineering repairs.
6
+
7
+ ## 0.17.0 release candidate — session learning
8
+
9
+ Adds explicit Pi-session import, exact private retention, chronological evidence,
10
+ per-execution delegation history, parent/child usage coverage, and line-bound
11
+ operator annotations through the existing CLI/Pi learning workspace. No model
12
+ calls, automatic acceptance, calibrated improvement or expanded adoption authority.
13
+ See the whole-session section of PRODUCT-GUIDE.md. Feature PR82 is merged with independent review and green CI. Release metadata
14
+ pairs harness 0.17.0 with producer 0.27.3; canonical publication and installed user
15
+ testing remain separate. Entire-session optimization and measured skill/tool variants remain
16
+ follow-up work.
17
+
18
+ ## Historical 0.16.0 release preparation
19
+
20
+ Release pair: **skill-harness 0.16.0 / pi-daddy 0.27.0**; Principal 3.2.0 is unchanged.
21
+ The prior audited baseline was harness 0.15.0 (`a725579`) / producer 0.26.1 (`93a9b01`).
22
+ Harness PR80 merged at `c6cd55d4db9ccf9a1ec790459c3cc5e6b1eac69f` with green CI and
23
+ Sol approval. Actual source/built cross-package learning/registry checks passed with
24
+ synthetic fixtures. This is **release preparation**, not publication, final installed
25
+ proof, human approval, calibration or measured improvement. Loaded resources, source
26
+ HEAD, canonical archives, published bytes and active sessions remain separate identities.
27
+ Canonical packing, publication and final installed verification are pending.
28
+
29
+ | Requirement | Current implementation / boundary | Remaining evidence |
30
+ |---|---|---|
31
+ | Understandable side panel | Producer owns selected-outcome-first status, observed agents, available controls and Details. | Matching producer's installed narrow-pane/host checks. No invented progress/ETA/acceptance. |
32
+ | Ordinary multi-obligation work | Producer owns work setup, dependencies, bounded parallel/sequential execution, model/effort choice and revision consent. | Installed checks of the matching producer's supported profiles. Existing Principal workflows remain unchanged. |
33
+ | Retained cases/comparisons | Harness learning workspace composes existing archive, case-review, blind-choice and CAS journal primitives. Named bindings and guided CLI/Pi entrypoints; missing/stale/mismatched inputs defer. | Final installed producer connection; source review/CI and synthetic source/built linkage passed. No new model evidence claimed. |
34
+ | Complete quality review | Complete-output acknowledgement, durable choice before reveal, immutable replay, separate adopt/reject/defer. Prior excerpt feedback stays separate. | A real full-artifact product choice remains absent where not actually recorded. Excerpt preference is not acceptance. |
35
+ | Scope-bound trust | Guided frozen cohort/policy, actual nomination predictions, incident grouping, explicit independent case/unflagged labels and pure exposure preview. Existing automatic gates unchanged. | Real independent labels, unflagged coverage and held-out evidence where absent. Implementation cannot manufacture them. |
36
+ | Hypothesis and comparison linkage | Human-authored falsifiable hypothesis form, confirmed-case linkage, monotonic optional context and retained lifecycle navigation. | Ongoing weekly model-backed retro is not demonstrated or automatically scheduled. Existing bounded weekly/model primitives remain separate authorized paths. |
37
+ | Adoption and rollback | Harness validates scope/choice/current independent authority and eligible facts. Producer source supports unchanged `fixed-policy-v1` plus separate `ordinary-work-policy-v1` (model/effort only, same task IDs/agent definitions and scope/assessment). Producer owns activation/rollback and fresh-eligibility next-order pins; profile-agnostic harness linkage retains original receipts. | Synthetic source/built original activation/rollback linkage passed. Configured production adoption/rollback and final installed integration remain unverified; bare legacy views without applied activation proof defer. No active/history-bearing run migration or instruction/skill/grant/capability/topology/assessment expansion. |
38
+ | Later outcomes | Explicit file-based CLI/Pi observation forms plus producer API linkage; independent reference attribution; changed requirements/artifacts/out-of-scope/unknown remain distinct. | Genuine later work, exact acceptance and held-out efficacy where absent. No automatic rollback or universal quality score. |
39
+ | Packaging / portability | Public adapter exports, additive frozen dashboard bridge, CLI export and bundled Pi entrypoint; guide/register shipped offline via `learning guide` / `learning current`. A pre-existing immutable bridge needs a fresh Pi session. | Release-metadata review/CI, exact merged canonical pack with retained archives, empty-prefix installed Pi/registry verification, then authorized publication. |
40
+
41
+ ## Historical claims remain bounded
42
+
43
+ - The failure-derived abort-race case and the public-weather comparison are different chains.
44
+ - Previously retained feedback covered excerpts only. It does not advance full comparison
45
+ acceptance or detector trust.
46
+ - Earlier no-tool measured SDK observations do not establish universal ordinary-child Usage,
47
+ writable-effect containment or actual provider-internal reasoning effort.
48
+ - No OS sandbox or malicious-local-owner protection has been added.
49
+ - No historical session/evidence, auth or unrelated worktree was modified in this candidate.
50
+
51
+ Feature source review/CI and source/built cross-package checks are completed evidence,
52
+ not final released-install proof. Metadata preparation runs consistency/diff validation only:
53
+ no builds, installs, suites or model calls. Canonical pack and actual installed two-extension
54
+ verification remain release gates under `PUBLISHING.md`. Prior live route smokes remain
55
+ historical and are not replayed merely for a version bump. Human acceptance, calibration and
56
+ later efficacy still require their own real evidence.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "skill-harness",
3
- "version": "0.15.0",
3
+ "version": "0.17.0",
4
4
  "description": "Test/optimize loop for agent skills — run spec'd scenarios on pi, LLM-judge, score, review, re-run",
5
5
  "type": "module",
6
6
  "keywords": [
@@ -28,11 +28,13 @@
28
28
  "dist/prompt-capture-extension.js",
29
29
  "assets/report.template.html",
30
30
  "assets/report.grade.js",
31
+ "docs/PRODUCT-GUIDE.md",
32
+ "docs/STATUS.md",
31
33
  "LICENSE",
32
34
  "README.md"
33
35
  ],
34
36
  "dependencies": {
35
- "@skill-harness/cli": "0.15.0",
37
+ "@skill-harness/cli": "0.17.0",
36
38
  "typebox": "^1.1.38"
37
39
  },
38
40
  "repository": {