humanish 0.86.0 → 0.87.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/actor-contract.d.ts +4 -0
- package/dist/actor-contract.js.map +1 -1
- package/dist/actor-stop-cause.d.ts +7 -0
- package/dist/actor-stop-cause.js +36 -0
- package/dist/actor-stop-cause.js.map +1 -0
- package/dist/computer-use.d.ts +1 -1
- package/dist/computer-use.js +15 -8
- package/dist/computer-use.js.map +1 -1
- package/dist/concurrent-shared-world-lab.d.ts +1 -1
- package/dist/concurrent-shared-world-lab.js +4 -0
- package/dist/concurrent-shared-world-lab.js.map +1 -1
- package/dist/cua-actor-lab.d.ts +1 -1
- package/dist/cua-actor-lab.js +9 -0
- package/dist/cua-actor-lab.js.map +1 -1
- package/dist/e2b-terminal-lab.d.ts +1 -1
- package/dist/e2b-terminal-lab.js +4 -0
- package/dist/e2b-terminal-lab.js.map +1 -1
- package/dist/export.js +1 -1
- package/dist/export.js.map +1 -1
- package/dist/index.d.ts +1 -1
- package/dist/index.js.map +1 -1
- package/dist/lab-config.d.ts +4 -0
- package/dist/lab-config.js +18 -0
- package/dist/lab-config.js.map +1 -1
- package/dist/lab-engine.js +24 -1
- package/dist/lab-engine.js.map +1 -1
- package/dist/observer-app.html +9 -9
- package/dist/observer-data.d.ts +5 -0
- package/dist/observer-data.js +27 -4
- package/dist/observer-data.js.map +1 -1
- package/dist/observer.d.ts +3 -1
- package/dist/observer.js +9 -6
- package/dist/observer.js.map +1 -1
- package/dist/openai-responses-cu.js +1 -1
- package/dist/openai-responses-cu.js.map +1 -1
- package/dist/oss-lab.d.ts +1 -1
- package/dist/oss-meta-lab.d.ts +1 -1
- package/dist/oss-meta-lab.js.map +1 -1
- package/dist/run.d.ts +6 -3
- package/dist/run.js +22 -7
- package/dist/run.js.map +1 -1
- package/dist/scripted-browser-lab.d.ts +1 -1
- package/dist/scripted-browser-lab.js +9 -0
- package/dist/scripted-browser-lab.js.map +1 -1
- package/dist/shared-world-lab.d.ts +1 -1
- package/dist/shared-world-lab.js +4 -0
- package/dist/shared-world-lab.js.map +1 -1
- package/docs/architecture/task-protocol-support.md +35 -0
- package/docs/contracts/run-bundle.md +17 -0
- package/docs/contracts/schemas.md +1 -1
- package/docs/goals/current.md +151 -717
- package/docs/ramp/README.md +11 -3
- package/docs/release/0.86.1-task-preflight-saved-recordings.md +40 -0
- package/docs/release/0.87.0-participant-endings-and-phone-review.md +64 -0
- package/package.json +3 -2
package/docs/ramp/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Status: public-safe contributor and agent ramp.
|
|
4
4
|
|
|
5
|
-
Package/source version in this tree: `0.
|
|
5
|
+
Package/source version in this tree: `0.87.0` (2026-09-10). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
|
|
6
6
|
`0.15.1` remains in force: managed run and output paths bind to validated
|
|
7
7
|
physical filesystem identities, and stored provider IDs are evidence, not
|
|
8
8
|
cleanup authority. The bundled OSS meta-lab is dry-run only until
|
|
@@ -47,6 +47,14 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
|
|
|
47
47
|
|
|
48
48
|
## Current State
|
|
49
49
|
|
|
50
|
+
The [0.87.0 release note](../release/0.87.0-participant-endings-and-phone-review.md)
|
|
51
|
+
describes recorded interruption causes, phone playback controls beside fitted
|
|
52
|
+
captures, and credential and evidence-sharing guidance by execution route.
|
|
53
|
+
|
|
54
|
+
The [0.86.1 release note](../release/0.86.1-task-preflight-saved-recordings.md)
|
|
55
|
+
describes task preflight on unsupported execution paths and saved HTML recordings
|
|
56
|
+
that remain snapshots when served over HTTP.
|
|
57
|
+
|
|
50
58
|
The [0.86.0 release note](../release/0.86.0-participant-evidence.md) describes
|
|
51
59
|
participant assignments, action-specific replay links and bounded cleanup with
|
|
52
60
|
the current desktop SDK. Old runs without assignment fields remain readable;
|
|
@@ -75,7 +83,7 @@ Implemented:
|
|
|
75
83
|
- first-class lab manifest resolution through `humanish/labs/*.yaml` and
|
|
76
84
|
ignored `.humanish/labs/*.yaml` overlays — `humanish.lab.v2` compositions
|
|
77
85
|
(`src/lab-config.ts`), one engine, no hardcoded lab kinds;
|
|
78
|
-
- a first-party actor registry with
|
|
86
|
+
- a first-party actor registry with seven registered descriptors
|
|
79
87
|
(`src/actor-registry.ts`); `actors[0].type` is a real dispatch key on the
|
|
80
88
|
computer-use, scripted-browser, and terminal-product routes;
|
|
81
89
|
- a computer-use route and clone subject provider: `subject.source: app-url`
|
|
@@ -126,7 +134,7 @@ without changing its success standard.
|
|
|
126
134
|
- capability receipts are not adopter replacement: no first-party deletion
|
|
127
135
|
branch has yet removed a bespoke generic harness while preserving
|
|
128
136
|
decision-equivalent proof;
|
|
129
|
-
- the
|
|
137
|
+
- the seven actor descriptors are a closed first-party union, not a supported
|
|
130
138
|
out-of-tree actor-registration API;
|
|
131
139
|
- run storage and provider-resource lifecycle logic still spans several routes
|
|
132
140
|
instead of one `RunStore` and `ResourceLease` boundary;
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# Humanish 0.86.1: enforce declared tasks and preserve saved recordings
|
|
2
|
+
|
|
3
|
+
Humanish now refuses a task protocol when the chosen execution path cannot run
|
|
4
|
+
it. Previously, a shared-world, terminal, scripted or synthetic lab could accept
|
|
5
|
+
`actors[0].tasks` and then omit those tasks during execution. Preflight now names
|
|
6
|
+
the unsupported field before creating a run or starting hooks, processes,
|
|
7
|
+
desktops or model requests. Remove `tasks` only when you intend a mission-only
|
|
8
|
+
study, or choose a supported per-lane computer-use route.
|
|
9
|
+
|
|
10
|
+
The same check protects direct library callers. Per-lane computer-use tasks
|
|
11
|
+
remain supported, including `local-agent`, `local-app` and `desktop-cli` paths.
|
|
12
|
+
Tasks on a second or later actor are rejected because current runners consume
|
|
13
|
+
the first actor. The [support matrix](../architecture/task-protocol-support.md)
|
|
14
|
+
describes the execution boundaries and structured errors.
|
|
15
|
+
|
|
16
|
+
Portable HTML exports now identify themselves as **Saved recording**, including
|
|
17
|
+
when opened from a website. They do not request a live feed, show a retry warning
|
|
18
|
+
for a missing endpoint, or open a desktop preview from a recording captured
|
|
19
|
+
while the participant was running. Recorded outcomes, timestamps and images
|
|
20
|
+
remain intact. Ordinary `humanish observe` and `humanish serve` pages still
|
|
21
|
+
request updates and report real connection failures.
|
|
22
|
+
|
|
23
|
+
Install with `npm install -g humanish@0.86.1`. Regenerate an existing export with
|
|
24
|
+
`humanish export --run RUN --format html` to use the corrected viewer. Existing
|
|
25
|
+
HTML files keep the behavior of the version that produced them.
|
|
26
|
+
|
|
27
|
+
## Verification and limits
|
|
28
|
+
|
|
29
|
+
The task preflight proof exercises unsupported lab shapes through the compiled
|
|
30
|
+
CLI with provider, network and subprocess operations denied. Parser and direct
|
|
31
|
+
runner tests also cover supported CUA paths and mission-only execution. These
|
|
32
|
+
are refusal and compatibility checks; they do not establish task support in
|
|
33
|
+
the rejected engines.
|
|
34
|
+
|
|
35
|
+
Production browser checks use a retained real recording over HTTP and from a
|
|
36
|
+
local file at desktop and phone widths. Controls cover ordinary served pages,
|
|
37
|
+
legacy exports and a snapshot captured while running. Snapshot checks verify
|
|
38
|
+
that no feed, history or live desktop request occurs; served controls preserve
|
|
39
|
+
the visible connection failure. No run-data schema or original study result
|
|
40
|
+
changes, and this release adds no continuous video or physical-device proof.
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
# Humanish 0.87.0: see why a participant stopped
|
|
2
|
+
|
|
3
|
+
The Observer now distinguishes a provider response limit from elapsed time or
|
|
4
|
+
estimated spend. Participant cards, summaries and reports show the recorded
|
|
5
|
+
cause, including for participants with no screenshot frames. New computer-use
|
|
6
|
+
traces retain an optional `stopCause` alongside the existing status and
|
|
7
|
+
completion reason.
|
|
8
|
+
|
|
9
|
+
New computer-use sessions that hit their estimated spend cap before any action
|
|
10
|
+
are classified as interrupted, with no target-app feedback inferred from that
|
|
11
|
+
harness stop. The cap timing and recorded usage are unchanged.
|
|
12
|
+
|
|
13
|
+
Older recordings use exact machine-generated notices where available. A recorded
|
|
14
|
+
“provider token limit reached” notice remains **provider token limit**: it does
|
|
15
|
+
not establish whether the provider exhausted its output or context allowance.
|
|
16
|
+
When an older run records only a generic limit, the Observer says **limit
|
|
17
|
+
reached**. Updating a viewer preserves the original evidence and outcome counts;
|
|
18
|
+
it does not infer a cause from participant prose.
|
|
19
|
+
|
|
20
|
+
On phones, a saved landscape capture in Fit mode keeps playback controls beside
|
|
21
|
+
the fitted image. Portrait captures, zoomed views and live views retain their
|
|
22
|
+
bounded viewing area. The run banner states **Saved recording** once; the player
|
|
23
|
+
keeps the participant status and selected capture timestamp. Exact action links
|
|
24
|
+
and saved-moment recall retain their existing behavior.
|
|
25
|
+
|
|
26
|
+
Expanded assignments show the recorded lane focus before the mission, so a long
|
|
27
|
+
mission does not put those instructions below it in the phone panel. The full
|
|
28
|
+
mission and participant tasks remain available.
|
|
29
|
+
|
|
30
|
+
The website also explains credential placement by execution route. Computer-use
|
|
31
|
+
model keys remain on the host. Terminal actors receive a command-scoped runtime
|
|
32
|
+
key by default; optional `openai-egress` keeps the raw key outside the sandbox
|
|
33
|
+
while still allowing sandbox processes to make paid requests through its proxy.
|
|
34
|
+
Evidence stays local by default, and exposing Observer or sharing an export makes
|
|
35
|
+
it accessible to others. See the [credential and sharing guide](https://humanish.dev/docs/budgets-and-privacy#store-credentials).
|
|
36
|
+
|
|
37
|
+
Install with `npm install -g humanish@0.87.0`. Reopen a run with
|
|
38
|
+
`humanish observe --run RUN`, or regenerate its HTML export, to use the updated
|
|
39
|
+
viewer. Existing HTML files retain the renderer that produced them.
|
|
40
|
+
|
|
41
|
+
## Verification and limits
|
|
42
|
+
|
|
43
|
+
[Interruption-cause checks](https://github.com/danielgwilson/humanish/pull/747)
|
|
44
|
+
cover recorded provider output limits, legacy token-limit notices, elapsed time,
|
|
45
|
+
spend and incomplete provider responses. Tests preserve old outcome counts and
|
|
46
|
+
source evidence through served and exported projections. A retained real
|
|
47
|
+
recording was inspected at desktop and phone widths, including its participant
|
|
48
|
+
whose provider stopped before any app action.
|
|
49
|
+
|
|
50
|
+
[Zero-action spend-cap checks](https://github.com/danielgwilson/humanish/pull/752)
|
|
51
|
+
verify one provider turn, zero executed actions, retained usage, an interrupted
|
|
52
|
+
outcome and no inferred target-app feedback. Older bundle counts are unchanged.
|
|
53
|
+
|
|
54
|
+
[Phone review checks](https://github.com/danielgwilson/humanish/pull/746) cover
|
|
55
|
+
landscape and portrait capture/control spacing, keyboard seeking, focus, zoom,
|
|
56
|
+
action links, refresh and saved-moment recall against retained recordings.
|
|
57
|
+
These browser checks do not establish physical-phone behavior or new live-stream
|
|
58
|
+
coverage. Screenshot replay remains sparse, and no participant success rate or
|
|
59
|
+
incremental persona benefit is inferred from these changes.
|
|
60
|
+
|
|
61
|
+
The [assignment ordering change](https://github.com/danielgwilson/humanish/pull/754)
|
|
62
|
+
preserves recorded text while moving lane instructions to the top of expanded
|
|
63
|
+
assignments. This addresses visibility; it does not establish that every reader
|
|
64
|
+
will interpret the instructions correctly.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "humanish",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.87.0",
|
|
4
4
|
"description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
|
|
5
5
|
"author": "Daniel G Wilson <daniel@danielgwilson.com>",
|
|
6
6
|
"keywords": [
|
|
@@ -50,8 +50,9 @@
|
|
|
50
50
|
},
|
|
51
51
|
"scripts": {
|
|
52
52
|
"build": "tsc -p tsconfig.build.json && pnpm --filter humanish-observer build && pnpm --filter humanish-tui build && node scripts/finish-build.mjs",
|
|
53
|
-
"check": "pnpm typecheck && pnpm test && pnpm tui:test && pnpm build && pnpm cli:startup:test && pnpm tui:smoke",
|
|
53
|
+
"check": "pnpm typecheck && pnpm test && pnpm tui:test && pnpm build && pnpm cli:startup:test && pnpm cli:preflight:test && pnpm tui:smoke",
|
|
54
54
|
"cli:startup:test": "node scripts/terminal-startup-exit-proof.mjs",
|
|
55
|
+
"cli:preflight:test": "node scripts/task-route-preflight-proof.mjs",
|
|
55
56
|
"humanish": "tsx src/cli.ts",
|
|
56
57
|
"pack:dry-run": "npm pack --dry-run",
|
|
57
58
|
"prepack": "pnpm build",
|