skill-harness 0.14.0 → 0.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +12 -0
- package/dist/index.js +6209 -2483
- package/docs/PRODUCT-GUIDE.md +133 -0
- package/docs/STATUS.md +45 -0
- package/package.json +7 -3
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
# Factory: work, review, learn
|
|
2
|
+
|
|
3
|
+
**Product guide for the 0.16.0 release candidate**, paired with pi-daddy **0.27.0**;
|
|
4
|
+
Principal **3.2.0** is unchanged. The prior audited baseline was harness 0.15.0 /
|
|
5
|
+
producer 0.26.1. Harness PR80 is merged and independently approved with green CI;
|
|
6
|
+
source/built cross-package checks passed with synthetic fixtures, not human evidence.
|
|
7
|
+
Canonical release packaging, publication and final installed-flow verification remain
|
|
8
|
+
separate pending gates. After installation, expect `skill-harness --version` **0.16.0**
|
|
9
|
+
and producer package **0.27.0**; check `skill-harness learning help` and loaded Pi resources.
|
|
10
|
+
Do not assume an existing session reloaded because a package was updated. Use a **fresh
|
|
11
|
+
Pi session** when upgrading an already-published immutable dashboard bridge: `/reload`
|
|
12
|
+
cannot replace the old global API. No other running session needs to be stopped.
|
|
13
|
+
|
|
14
|
+
## Who does what?
|
|
15
|
+
|
|
16
|
+
**pi-daddy** coordinates permitted work and owns execution, steering and the future-order
|
|
17
|
+
registry. **Herdr** displays its sessions and side panel. **skill-harness** retains
|
|
18
|
+
measurement and learning evidence. The panel is not an autonomous worker. Finished
|
|
19
|
+
execution is not human acceptance.
|
|
20
|
+
|
|
21
|
+
Use the producer's `/grants work` guided setup for multiple obligations, dependencies,
|
|
22
|
+
priority and explicit agent/model/thinking selections, then its supported run/host controls.
|
|
23
|
+
For existing single-task use, `pi-daddy work add --id report --outcome "Review the report"`
|
|
24
|
+
remains supported. Only listed controls are available. Pausing new dispatch does not
|
|
25
|
+
cancel running work; changing a next-order policy does not mutate active sessions.
|
|
26
|
+
Principal's ordinary current/stale/superseded reports remain unchanged.
|
|
27
|
+
|
|
28
|
+
## Open learning in Pi
|
|
29
|
+
|
|
30
|
+
With both matching extensions loaded, use **`/grants learning`** from the current host.
|
|
31
|
+
The producer binds the exact current snapshot, archive, population and local review author.
|
|
32
|
+
A changed scope must report a mismatch, not silently reuse an old learning workspace.
|
|
33
|
+
|
|
34
|
+
The harness also exposes **`/skill-harness learning --state "WORKSPACE"`** (use the
|
|
35
|
+
workspace path displayed by the host), or `skill-harness learning --state "WORKSPACE"`
|
|
36
|
+
in a terminal. These open the **same guided flow and durable records**, not another app.
|
|
37
|
+
|
|
38
|
+
1. **Readiness** explains what's retained, what's missing, and why automatic questions
|
|
39
|
+
are silent. Opening a review is deliberate: it neither reserves attention nor
|
|
40
|
+
represents an earned automatic question. It starts no model or delegation.
|
|
41
|
+
2. **Connect retained input** lists compatible case batches and qualified comparisons
|
|
42
|
+
already in the selected archive. Pick one and give it a short name and title.
|
|
43
|
+
No hand-written manifests or ledger JSON. No compatible evidence is an honest empty state.
|
|
44
|
+
3. **Review cases** shows nominations and lets you open their actual retained evidence
|
|
45
|
+
sources by selection (CLI: `learning evidence BATCH --item N --evidence N`). Redacted
|
|
46
|
+
or unavailable sources are marked, never called complete evidence. Record confirmed defect, expected
|
|
47
|
+
behavior, exemplar, uncertain or skip, with a reason. Corrections are append-only.
|
|
48
|
+
Confirming a case is not promoting a test or establishing a cause.
|
|
49
|
+
4. **Review comparisons** opens complete retained artifacts under A/B labels. Acknowledge
|
|
50
|
+
each complete artifact only after reviewing it. Then choose one, tie, none acceptable,
|
|
51
|
+
or insufficient evidence. **Reveal model/cost** is a separate action, enabled only after
|
|
52
|
+
the durable quality receipt. Artifact contents themselves may disclose identity.
|
|
53
|
+
5. **Record adopt / reject / defer** records a separate scope-bound decision. Reject and defer
|
|
54
|
+
are normal outcomes; neither forces a preference or reveal. Missing evidence can be
|
|
55
|
+
durably deferred. Adopt is intent, not activation.
|
|
56
|
+
6. **Propose hypothesis** records prediction, disproof, downside, rollback and a competing
|
|
57
|
+
explanation from an actual case. **Link original hypothesis** only accepts the comparison's
|
|
58
|
+
actual frozen origin, never a retrospective replacement.
|
|
59
|
+
|
|
60
|
+
Previously retained excerpt feedback concerned
|
|
61
|
+
**excerpts**. Keep that receipt unchanged and linked as prior feedback. It is not full-artifact
|
|
62
|
+
acceptance, detector calibration or a fresh blind experiment. An old unscoped quality choice
|
|
63
|
+
also does not automatically become a full-artifact acknowledgement.
|
|
64
|
+
|
|
65
|
+
## Configure trust without writing manifests
|
|
66
|
+
|
|
67
|
+
Open **Trust / independent labels** in the same guided flow:
|
|
68
|
+
|
|
69
|
+
- Before freezing, explicitly add any real unflagged incident evidence and select its split.
|
|
70
|
+
Only selected evidence files are retained; do not include private sessions or secrets.
|
|
71
|
+
- Select a retained case batch, detector and evaluation split. The product derives the
|
|
72
|
+
cohort and predictions from actual nominations and groups repeated target observations
|
|
73
|
+
into one incident. Detector version and population remain exact.
|
|
74
|
+
- Automatic questions default to **silent**. To configure them, explicitly select minimum
|
|
75
|
+
resolved independent incidents, the confidence lower bound, attention budget and expiry.
|
|
76
|
+
The form previews the exact frozen policy. It creates **no independent labels**.
|
|
77
|
+
- Review cases independently, then choose **Link current case label**. Unflagged sample
|
|
78
|
+
labels have their own explicit evidence/author step. Preference, model self-report,
|
|
79
|
+
panel agreement and exit zero are not independent correctness labels.
|
|
80
|
+
|
|
81
|
+
Missing calibration, missing unflagged samples, held-out absence, conflicting labels,
|
|
82
|
+
expired policy and exhausted attention stay visible. Reopening neither resets choices nor
|
|
83
|
+
refills attention. Correcting a linked case withdraws its old label from confidence.
|
|
84
|
+
The original producer closing path still requires presence, quiescence and its exposure gate.
|
|
85
|
+
|
|
86
|
+
## Scoped adoption and later work
|
|
87
|
+
|
|
88
|
+
Use **`/grants learning` in interactive Pi** for activation/rollback. The matching producer
|
|
89
|
+
implements two separate profiles: unchanged **`fixed-policy-v1`**, and
|
|
90
|
+
**`ordinary-work-policy-v1`** for model/effort-only changes to the same selected task IDs and
|
|
91
|
+
agent definitions. The latter does not change instructions, skills, grants, capabilities,
|
|
92
|
+
topology or assessment policy; do not disguise those changes as model/effort or fixed policy.
|
|
93
|
+
The producer prepares policy bytes and owns independent authority and the original registry.
|
|
94
|
+
Source/built integration has been checked with synthetic fixtures; final installed-release
|
|
95
|
+
verification remains pending. Bare legacy fixed-policy views lacking applied activation
|
|
96
|
+
proof defer; they must not be treated as activation receipts.
|
|
97
|
+
|
|
98
|
+
Activation needs the exact selected candidate, original confirmed case/hypothesis/comparison,
|
|
99
|
+
current independently verified eligible facts, explicit adoption authority, and registry
|
|
100
|
+
revision consent. The original registry revalidates these again. Only its actual activation
|
|
101
|
+
receipt advances the learning record to activated. New ordinary-work orders pin the policy
|
|
102
|
+
under fresh eligibility confirmation; active or history-bearing runs never migrate.
|
|
103
|
+
Rollback similarly needs its own authorized request and original applied registry receipt.
|
|
104
|
+
A stored lifecycle link, human preference or prepared receipt cannot substitute for either.
|
|
105
|
+
|
|
106
|
+
**Adoption / outcomes** can inspect linked later observations or record one using complete
|
|
107
|
+
artifact, original/current requirement and independent evidence files. Unknown is valid.
|
|
108
|
+
Acceptance requires both the exact accepted artifact and independent acceptance evidence.
|
|
109
|
+
Changed requirements, changed artifact, out-of-scope, success, caught defect and escape remain
|
|
110
|
+
distinct. No outcome triggers automatic rollback, and no single outcome proves improvement.
|
|
111
|
+
|
|
112
|
+
## CLI and local archive use
|
|
113
|
+
|
|
114
|
+
```sh
|
|
115
|
+
skill-harness learning status --state "WORKSPACE" # human readiness
|
|
116
|
+
skill-harness learning status --state "WORKSPACE" --json # integration/details
|
|
117
|
+
skill-harness learning review reports --state "WORKSPACE"
|
|
118
|
+
skill-harness learning help # all explicit operations
|
|
119
|
+
skill-harness learning guide # this guide, packaged offline
|
|
120
|
+
skill-harness learning current # packaged requirement register
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
Without a producer host, `skill-harness learning` guides connection to an explicitly selected
|
|
124
|
+
archive and selecting an actual retained case scope. Or use
|
|
125
|
+
`learning init --archive DIR --author NAME --confirm`; multiple retained scopes ask you to
|
|
126
|
+
select `--scope-item N`. Snapshot digest and population are derived from actual cases, not
|
|
127
|
+
typed hashes. For a comparison-only archive with no cases, `--scope TEXT --population TEXT`
|
|
128
|
+
can explicitly declare a local review scope; it cannot bind mismatching case snapshots.
|
|
129
|
+
Use the producer flow to connect a host. Noninteractive writes require explicit `--confirm`.
|
|
130
|
+
All learning commands are offline/no-model. Command completion is never a quality verdict.
|
|
131
|
+
Local author attribution and content hashes are not authentication or hostile-owner attestation.
|
|
132
|
+
|
|
133
|
+
Current requirements and remaining evidence: [STATUS.md](STATUS.md).
|
package/docs/STATUS.md
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# Factory requirement register — 0.16.0 release preparation
|
|
2
|
+
|
|
3
|
+
Current entry point: [PRODUCT-GUIDE.md](PRODUCT-GUIDE.md). The dated
|
|
4
|
+
[CURRENT.md](CURRENT.md), build reports and transport decisions are retained history,
|
|
5
|
+
not current instructions to supply manifests or judge engineering repairs.
|
|
6
|
+
|
|
7
|
+
## Baseline and candidate
|
|
8
|
+
|
|
9
|
+
Release pair: **skill-harness 0.16.0 / pi-daddy 0.27.0**; Principal 3.2.0 is unchanged.
|
|
10
|
+
The prior audited baseline was harness 0.15.0 (`a725579`) / producer 0.26.1 (`93a9b01`).
|
|
11
|
+
Harness PR80 merged at `c6cd55d4db9ccf9a1ec790459c3cc5e6b1eac69f` with green CI and
|
|
12
|
+
Sol approval. Actual source/built cross-package learning/registry checks passed with
|
|
13
|
+
synthetic fixtures. This is **release preparation**, not publication, final installed
|
|
14
|
+
proof, human approval, calibration or measured improvement. Loaded resources, source
|
|
15
|
+
HEAD, canonical archives, published bytes and active sessions remain separate identities.
|
|
16
|
+
Canonical packing, publication and final installed verification are pending.
|
|
17
|
+
|
|
18
|
+
| Requirement | Current implementation / boundary | Remaining evidence |
|
|
19
|
+
|---|---|---|
|
|
20
|
+
| Understandable side panel | Producer owns selected-outcome-first status, observed agents, available controls and Details. | Matching producer's installed narrow-pane/host checks. No invented progress/ETA/acceptance. |
|
|
21
|
+
| Ordinary multi-obligation work | Producer owns work setup, dependencies, bounded parallel/sequential execution, model/effort choice and revision consent. | Installed checks of the matching producer's supported profiles. Existing Principal workflows remain unchanged. |
|
|
22
|
+
| Retained cases/comparisons | Harness learning workspace composes existing archive, case-review, blind-choice and CAS journal primitives. Named bindings and guided CLI/Pi entrypoints; missing/stale/mismatched inputs defer. | Final installed producer connection; source review/CI and synthetic source/built linkage passed. No new model evidence claimed. |
|
|
23
|
+
| Complete quality review | Complete-output acknowledgement, durable choice before reveal, immutable replay, separate adopt/reject/defer. Prior excerpt feedback stays separate. | A real full-artifact product choice remains absent where not actually recorded. Excerpt preference is not acceptance. |
|
|
24
|
+
| Scope-bound trust | Guided frozen cohort/policy, actual nomination predictions, incident grouping, explicit independent case/unflagged labels and pure exposure preview. Existing automatic gates unchanged. | Real independent labels, unflagged coverage and held-out evidence where absent. Implementation cannot manufacture them. |
|
|
25
|
+
| Hypothesis and comparison linkage | Human-authored falsifiable hypothesis form, confirmed-case linkage, monotonic optional context and retained lifecycle navigation. | Ongoing weekly model-backed retro is not demonstrated or automatically scheduled. Existing bounded weekly/model primitives remain separate authorized paths. |
|
|
26
|
+
| Adoption and rollback | Harness validates scope/choice/current independent authority and eligible facts. Producer source supports unchanged `fixed-policy-v1` plus separate `ordinary-work-policy-v1` (model/effort only, same task IDs/agent definitions and scope/assessment). Producer owns activation/rollback and fresh-eligibility next-order pins; profile-agnostic harness linkage retains original receipts. | Synthetic source/built original activation/rollback linkage passed. Configured production adoption/rollback and final installed integration remain unverified; bare legacy views without applied activation proof defer. No active/history-bearing run migration or instruction/skill/grant/capability/topology/assessment expansion. |
|
|
27
|
+
| Later outcomes | Explicit file-based CLI/Pi observation forms plus producer API linkage; independent reference attribution; changed requirements/artifacts/out-of-scope/unknown remain distinct. | Genuine later work, exact acceptance and held-out efficacy where absent. No automatic rollback or universal quality score. |
|
|
28
|
+
| Packaging / portability | Public adapter exports, additive frozen dashboard bridge, CLI export and bundled Pi entrypoint; guide/register shipped offline via `learning guide` / `learning current`. A pre-existing immutable bridge needs a fresh Pi session. | Release-metadata review/CI, exact merged canonical pack with retained archives, empty-prefix installed Pi/registry verification, then authorized publication. |
|
|
29
|
+
|
|
30
|
+
## Historical claims remain bounded
|
|
31
|
+
|
|
32
|
+
- The failure-derived abort-race case and the public-weather comparison are different chains.
|
|
33
|
+
- Previously retained feedback covered excerpts only. It does not advance full comparison
|
|
34
|
+
acceptance or detector trust.
|
|
35
|
+
- Earlier no-tool measured SDK observations do not establish universal ordinary-child Usage,
|
|
36
|
+
writable-effect containment or actual provider-internal reasoning effort.
|
|
37
|
+
- No OS sandbox or malicious-local-owner protection has been added.
|
|
38
|
+
- No historical session/evidence, auth or unrelated worktree was modified in this candidate.
|
|
39
|
+
|
|
40
|
+
Feature source review/CI and source/built cross-package checks are completed evidence,
|
|
41
|
+
not final released-install proof. Metadata preparation runs consistency/diff validation only:
|
|
42
|
+
no builds, installs, suites or model calls. Canonical pack and actual installed two-extension
|
|
43
|
+
verification remain release gates under `PUBLISHING.md`. Prior live route smokes remain
|
|
44
|
+
historical and are not replayed merely for a version bump. Human acceptance, calibration and
|
|
45
|
+
later efficacy still require their own real evidence.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "skill-harness",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.16.0",
|
|
4
4
|
"description": "Test/optimize loop for agent skills — run spec'd scenarios on pi, LLM-judge, score, review, re-run",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"keywords": [
|
|
@@ -18,7 +18,9 @@
|
|
|
18
18
|
"skill-harness": "bin.js"
|
|
19
19
|
},
|
|
20
20
|
"pi": {
|
|
21
|
-
"extensions": [
|
|
21
|
+
"extensions": [
|
|
22
|
+
"./dist/index.js"
|
|
23
|
+
]
|
|
22
24
|
},
|
|
23
25
|
"files": [
|
|
24
26
|
"bin.js",
|
|
@@ -26,11 +28,13 @@
|
|
|
26
28
|
"dist/prompt-capture-extension.js",
|
|
27
29
|
"assets/report.template.html",
|
|
28
30
|
"assets/report.grade.js",
|
|
31
|
+
"docs/PRODUCT-GUIDE.md",
|
|
32
|
+
"docs/STATUS.md",
|
|
29
33
|
"LICENSE",
|
|
30
34
|
"README.md"
|
|
31
35
|
],
|
|
32
36
|
"dependencies": {
|
|
33
|
-
"@skill-harness/cli": "0.
|
|
37
|
+
"@skill-harness/cli": "0.16.0",
|
|
34
38
|
"typebox": "^1.1.38"
|
|
35
39
|
},
|
|
36
40
|
"repository": {
|