cowork-harness 1.22.0 → 1.23.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +2 -1
- package/CHANGELOG.md +131 -0
- package/DESIGN.md +3 -3
- package/README.md +8 -8
- package/baselines/desktop-1.28929.0.json +13 -3
- package/baselines/desktop-1.30096.1.json +577 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +9 -0
- package/dist/run/cassette.js +75 -3
- package/dist/runtime/image-capabilities.js +33 -0
- package/dist/sync/cowork-sync.js +250 -5
- package/docs/fidelity-gaps.md +40 -1
- package/docs/maintenance.md +2 -2
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +1 -1
- package/examples/replays/hostloop-computer-links.cassette.json +1 -1
- package/package.json +1 -1
- package/schema/cassette.v10.json +18 -0
- package/schema/cassette.v11.json +18 -0
- package/scripts/check-versions.ts +198 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.23.0
|
|
7
|
+
tracks-harness: cowork-harness 1.23.0 (baseline desktop-1.30096.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
26
|
-
> `desktop-1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.23.0` (baseline
|
|
26
|
+
> `desktop-1.30096.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.23.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.23.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.23.0"`. **Pin `@>=1.23.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.23.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -32,7 +32,7 @@ jobs:
|
|
|
32
32
|
- uses: actions/checkout@v4
|
|
33
33
|
- name: Stage the agent binary (official channel, sha256-verified — see https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md)
|
|
34
34
|
run: |
|
|
35
|
-
V=2.1.
|
|
35
|
+
V=2.1.229 # match your scenario's pinned baseline's agentVersion
|
|
36
36
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
37
37
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
38
38
|
# verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
|
|
@@ -58,7 +58,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
58
58
|
GitHub-hosted runners, no token/Docker/agent:
|
|
59
59
|
|
|
60
60
|
```yaml
|
|
61
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
61
|
+
- run: npm i -g "cowork-harness@>=1.23.0"
|
|
62
62
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
63
63
|
# no silent false-greens. WITHOUT --strict this
|
|
64
64
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -279,7 +279,7 @@ jobs:
|
|
|
279
279
|
with: { node-version: '24' }
|
|
280
280
|
- uses: actions/setup-python@v5
|
|
281
281
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
282
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
282
|
+
- run: npm i -g "cowork-harness@>=1.23.0"
|
|
283
283
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
284
284
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
285
285
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -308,7 +308,7 @@ jobs:
|
|
|
308
308
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
309
309
|
fi
|
|
310
310
|
- if: steps.guard.outputs.live == 'true'
|
|
311
|
-
run: npm i -g "cowork-harness@>=1.
|
|
311
|
+
run: npm i -g "cowork-harness@>=1.23.0"
|
|
312
312
|
- if: steps.guard.outputs.live == 'true'
|
|
313
313
|
run: cowork-harness run scenarios/ --output-format json
|
|
314
314
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 1.
|
|
3
|
+
Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
4
|
-
(baseline `desktop-1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.23.0`
|
|
4
|
+
(baseline `desktop-1.30096.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 1.
|
|
5
|
+
Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -69,6 +69,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v11.json`](htt
|
|
|
69
69
|
| `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
|
|
70
70
|
| `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
|
|
71
71
|
| `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
|
|
72
|
+
| `environment.agentImage` | The `agentImage` sub-block records the rootfs image the recording ran against, stamped only for the tiers whose capabilities come from it (`container`, `hostloop`; `microvm` probes the Lima guest instead). `ref` is the resolved image ref, `configId` the local config id (built **and** pulled images, NOT comparable across machines), `registryDigest` the registry manifest digest (pulled only, and the only cross-machine-comparable identity). The image decides `missingCapabilityUse`, which `computeVerdict` fails on, so replaying against a different rootfs can move a verdict with nothing in the cassette having changed — `replay` compares `registryDigest` first and prints an advisory note, never a failure. Re-inspected at cassette-WRITE time, so a rebuild between run and record records the later image. **Absent** on cassettes recorded before this field existed, and never backfilled |
|
|
72
73
|
|
|
73
74
|
## Recipe 3 — Set up redaction BEFORE your first hostloop/protocol record
|
|
74
75
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,137 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.23.0] — 2026-08-14
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **Platform baseline for Claude Desktop 1.30096.1 (bundled agent ELF `2.1.229`).** The agent ELF's
|
|
14
|
+
`sha256` was verified against the official release manifest for 2.1.229. The modeled spawn contract is
|
|
15
|
+
unchanged across the bump — `spawn.tools` stays 20 entries, `allowedTools` 19, the egress allowlist 15
|
|
16
|
+
domains, the spawn-env key set 61 with 31 conditional spreads, and the Cowork system prompt is
|
|
17
|
+
byte-identical (its fingerprint is recorded for 1.30096.1 in
|
|
18
|
+
`baselines/prompts/cowork-system-prompt-fingerprints.json`). `sync` refused to write this baseline until
|
|
19
|
+
the two sentinel defects below were corrected.
|
|
20
|
+
|
|
21
|
+
- **Two drift sentinels pinned in the synced baseline's `provenance.gates`**: the artifact-mount gate
|
|
22
|
+
(`coworkArtifacts`) and the CIC `can_use_tool` handler (`cicCanUseToolEnabled`). Both are force-ON in
|
|
23
|
+
production, and the artifact-mount gap documented in `docs/fidelity-gaps.md` rests on that fact — which
|
|
24
|
+
was previously read from a live feature cache the baseline never recorded, so nothing would have
|
|
25
|
+
noticed it changing.
|
|
26
|
+
|
|
27
|
+
- **`check:versions` guards DESIGN.md's live-verification scope note (invariant 11).** That note is the
|
|
28
|
+
repo's disclosure of how much of the *current* baseline has actually been verified live, and every
|
|
29
|
+
figure in it is derivable from `baselines/desktop-*.json` — yet it sat in unguarded prose and had
|
|
30
|
+
drifted twice, the baseline list having been extended without recounting. Understating how much is
|
|
31
|
+
unverified is the doc error least worth shipping, so it is now checked. Two forms, selected by whether
|
|
32
|
+
the note's live-pass baseline is the newest one: with a gap, the listed baselines must run contiguously
|
|
33
|
+
from wherever the list starts through the newest baseline, and both counts must match the list and the
|
|
34
|
+
real `agentVersion` transitions; with no gap, the note must say so explicitly and carry no stale
|
|
35
|
+
enumeration. Either way the named agent must be the newest baseline's. Because shipping a baseline flips
|
|
36
|
+
the no-gap form into the gap form, a new release now forces the note to be rewritten rather than
|
|
37
|
+
silently overstating coverage. The list's *start* is deliberately not derived — the note omits baselines
|
|
38
|
+
covered by the live pass itself, and encoding that rule would only relocate the drift. A missing or
|
|
39
|
+
unrecognisable note is an error, never a skip.
|
|
40
|
+
|
|
41
|
+
- **Cassettes now record the rootfs image they were recorded against**, closing the last gap in the
|
|
42
|
+
agent-image provenance work. The image decides `missingCapabilityUse`, which `computeVerdict` fails
|
|
43
|
+
on, so the rootfs is verdict-affecting — yet no cassette field named it, and a recording silently
|
|
44
|
+
inherited whatever image happened to be on the machine. `environment.agentImage` now carries the
|
|
45
|
+
resolved `ref` plus whichever identities exist: `configId` (the local config id, present for built
|
|
46
|
+
**and** pulled images but not comparable across machines) and `registryDigest` (the registry manifest
|
|
47
|
+
digest, pulled images only, and the only identity comparable across machines).
|
|
48
|
+
|
|
49
|
+
The field is stamped only for the tiers whose capabilities actually come from that image —
|
|
50
|
+
`container` and `hostloop`. `microvm` probes the Lima guest instead, so it records nothing rather
|
|
51
|
+
than naming an image that had no bearing on the run. Additive: no `cassetteVersion` bump, absent on
|
|
52
|
+
cassettes recorded before the field existed, and never backfilled — the absence is meaningful.
|
|
53
|
+
|
|
54
|
+
`ref` is a verbatim `COWORK_AGENT_IMAGE` value, so a private registry ref (`registry.acme.corp/…`)
|
|
55
|
+
would otherwise be committed into a public fixture with `grep`-clean transcript text. It is scanned
|
|
56
|
+
by `verify-cassettes` and rewritten by redaction like every other user-controlled string; the digests
|
|
57
|
+
are content hashes and are deliberately left intact.
|
|
58
|
+
|
|
59
|
+
- **`replay` warns when the rootfs image differs from the recording.** It compares `registryDigest`
|
|
60
|
+
first — the only identity stable across machines, which is the case the field exists to serve — and
|
|
61
|
+
falls back to the local config id only when neither side has a registry digest. A recording made
|
|
62
|
+
against a pulled image and replayed against a local rebuild is reported as drift rather than passing
|
|
63
|
+
silently. Advisory only: a legitimately re-pulled image is the common case, so this names the
|
|
64
|
+
difference instead of failing the replay.
|
|
65
|
+
|
|
66
|
+
The current image is inspected at most once per `replay` invocation, and only when a cassette
|
|
67
|
+
actually recorded one — so replay stays usable with no container runtime present, and the
|
|
68
|
+
`verify-cassettes` privacy scan never shells out.
|
|
69
|
+
|
|
70
|
+
### Fixed
|
|
71
|
+
|
|
72
|
+
- **The host-loop `canUseTool` chain sentinel could be widened silently.** Desktop 1.30096.1 inserts a
|
|
73
|
+
fourth link into the chain — an allow carve-out that rewrites the tool's input and runs ahead of the
|
|
74
|
+
existing deny. The sentinel was a prefix match with no terminator, so it accepted any chain that merely
|
|
75
|
+
*started* with the expected calls: an inserted **synchronous** link would have passed unnoticed, and
|
|
76
|
+
the block that surfaced this release only happened because the new link is `await`ed.
|
|
77
|
+
|
|
78
|
+
The check now decomposes the chain with a real scanner (the assignment must be brace/paren-balanced,
|
|
79
|
+
and `??` legitimately occurs inside a template interpolation in the chain's own log line) and asserts
|
|
80
|
+
it end to end: the terminal operand must be a bare call to the saved original, every operand whose
|
|
81
|
+
callee is an `async function` must be awaited, no link that can return an allow may precede the
|
|
82
|
+
VM-path deny, and each link must resolve to a definition. The await rule is the load-bearing one — an
|
|
83
|
+
un-awaited async link returns a Promise, which is never nullish, so `??` short-circuits and every later
|
|
84
|
+
link *including the original callback* is skipped. Three and four link chains are both accepted so an
|
|
85
|
+
older Desktop still syncs; a fifth is reported for classification.
|
|
86
|
+
|
|
87
|
+
- **The early-allow ordering check had never fired.** It searched for a containment helper by two
|
|
88
|
+
readable names, neither of which occurs in any shipped asar — production mangles the call — so the
|
|
89
|
+
guard was permanently inert and its test passed only because the fixture hard-coded a token that does
|
|
90
|
+
not exist in the product. The helper is now identified by shape and resolved through the export map.
|
|
91
|
+
|
|
92
|
+
- **S6c no longer hard-blocks on a minifier rename.** It pinned the HIPAA-restriction call by its
|
|
93
|
+
minified *member name*, so the rename `A.r()` → `t.hu()` failed a predicate that is otherwise
|
|
94
|
+
byte-for-byte identical. The callee is now resolved through the chunk's export map and verified two
|
|
95
|
+
hops to the reader that consults the restriction, with resolution failure treated as a miss.
|
|
96
|
+
|
|
97
|
+
- **The `protocol`-tier live suite had been silently skipping itself for many releases.** `live-matrix`
|
|
98
|
+
required a staged agent binary for a baseline pinned years back, but protocol fidelity spawns the host
|
|
99
|
+
`claude` from `PATH` and never resolves a staged binary — the requirement was never real for that tier.
|
|
100
|
+
Because Claude Desktop prunes old staged agents on update, the gate went false as soon as a machine
|
|
101
|
+
moved past that agent version, and the suite dropped out on every developer machine and in CI without
|
|
102
|
+
naming itself. It now gates on what the tier actually uses and emits a skip notice identifying the
|
|
103
|
+
failing precondition, since this is the only protocol-tier live coverage there is.
|
|
104
|
+
|
|
105
|
+
- **The `hostloop` uploads-are-`Read`-able live case no longer fails on the model's choice of exploration
|
|
106
|
+
tool.** It asserted that neither native nor workspace bash ran at all, as a proxy for "the agent needed a
|
|
107
|
+
workaround" — so a run where the agent listed the uploads directory with `ls` went red, while the next
|
|
108
|
+
run, which used `Glob` instead, went green. In both the upload was `Read` directly at the advertised
|
|
109
|
+
path and no outputs-delete fired, so the regression the case guards was absent either way. It now reads
|
|
110
|
+
the recorded bash commands and fails only when one **names the uploaded file**, which is what reading or
|
|
111
|
+
copying it as a workaround requires and what listing its directory cannot do. Verified against both
|
|
112
|
+
recorded runs plus `cat`/`cp` mutations: the false red is gone and the workaround chain still trips it.
|
|
113
|
+
|
|
114
|
+
### Documentation
|
|
115
|
+
|
|
116
|
+
- **A full live end-to-end pass now covers baseline `desktop-1.30096.1` / agent 2.1.229**, across all
|
|
117
|
+
three tiers (`protocol`, `container`, `hostloop`), superseding the `desktop-1.20186.0` pin. DESIGN.md's
|
|
118
|
+
claim and its scope note are re-stamped accordingly, including the caveats: two `live-outputs-delete`
|
|
119
|
+
cases skipped as model-behaviour misses, the tiers were covered across two invocations because the
|
|
120
|
+
`protocol` suite was repaired mid-pass, and the suites are model-dependent enough that a single red is
|
|
121
|
+
evidence of variance until a re-run says otherwise.
|
|
122
|
+
|
|
123
|
+
- **`docs/fidelity-gaps.md`'s artifacts section records the agent-side consent floor.** The server-delivered
|
|
124
|
+
session flag no longer only decides Desktop's spawned tool list — the agent reads the corresponding
|
|
125
|
+
spawn-env key itself and uses it to select its own artifact publish surface and a read-only mode. The
|
|
126
|
+
agent also refuses artifact publishes, comment replies, comment-thread resolves and artifact database
|
|
127
|
+
writes outright in a session with no answerable approval surface, and fails closed if it cannot confirm
|
|
128
|
+
one. The section now also states why this is recorded rather than modeled: artifact operations are
|
|
129
|
+
server-backed, on hosts outside the sandbox egress allowlist, so supplying the flag would offer a tool
|
|
130
|
+
resolving against a service the sandbox cannot reach.
|
|
131
|
+
|
|
132
|
+
- **The same section's mount-kind list is corrected.** It named a synthetic root that is not a mount kind
|
|
133
|
+
and omitted one that is, which mattered because the artifact-mount gap is stated in terms of what the
|
|
134
|
+
harness does and does not mount.
|
|
135
|
+
|
|
136
|
+
- **`docs/maintenance.md` corrects what moves the floating agent-image `:2` tag.** It is a curated pointer
|
|
137
|
+
moved deliberately by a manual publish with `immutable_only` unchecked — explicitly *not* something a
|
|
138
|
+
release tag push moves.
|
|
139
|
+
|
|
9
140
|
## [1.22.0] — 2026-08-12
|
|
10
141
|
|
|
11
142
|
### Added
|
package/DESIGN.md
CHANGED
|
@@ -44,7 +44,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
44
44
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
45
45
|
|
|
46
46
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
47
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
47
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.229**, per `baselines/desktop-1.30096.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
|
|
48
48
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
49
49
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
50
50
|
|
|
@@ -171,9 +171,9 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
171
171
|
|
|
172
172
|
> **Machine-readable form:** the five shapes below are schema'd as `schema/protocol.v1.json`, with a golden vector pack at `fixtures/protocol/v1/` — see [docs/protocol.md](./docs/protocol.md) for scope, versioning, and how to conformance-test against them.
|
|
173
173
|
|
|
174
|
-
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS
|
|
174
|
+
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.229**, the native host app that `hostloop` runs is **2.1.229**, baseline **`desktop-1.30096.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-14, superseding the prior `1.20186.0` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
175
175
|
|
|
176
|
-
> **Scope of that claim, stated plainly.** `2026-
|
|
176
|
+
> **Scope of that claim, stated plainly.** `2026-08-14 / desktop-1.30096.1` is the last baseline carrying a **full live end-to-end pass**, and it is the newest committed baseline — **no baselines have shipped since**. The pass ran against agent **2.1.229** and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe and resume continuity on the native binary). Three caveats, so this is not read as more than it is: two `live-outputs-delete` cases **skipped** — the agent declined to issue the pinned command, a model-behaviour miss the suite itself classifies as not a guard defect — so this is *every suite exercised*, not *every live assertion green*; the tiers were covered across **two invocations**, not one, because the `protocol` suite was repaired mid-pass (it had been silently skip-gated on a staged binary its tier never uses); and a live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — the `live-outputs-delete` skip set differs run to run for exactly that reason, which is why those cases skip loudly rather than fail. One such flake was root-caused rather than retried: the `hostloop` uploads-are-`Read`-able case counted bash calls, so its verdict turned on which exploration tool the model happened to pick (an `ls` of the uploads directory failed it; a `Glob` in the next run passed), even though the upload was `Read` directly and no outputs-delete fired either time. It now inspects what bash actually did and fails only when a command names the uploaded file — the workaround it exists to catch — so exploration is free and `cat`/`cp` of the upload still trips it. Every assertion in the lane has passed at least once against this baseline. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
177
177
|
|
|
178
178
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
179
179
|
|
package/README.md
CHANGED
|
@@ -92,7 +92,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
92
92
|
|
|
93
93
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
94
94
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
95
|
-
> From a global install (`npm i -g "cowork-harness@>=1.
|
|
95
|
+
> From a global install (`npm i -g "cowork-harness@>=1.23.0"`), point at the package root instead:
|
|
96
96
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
97
97
|
> (or copy the cassette into your own project and pass that path).
|
|
98
98
|
|
|
@@ -102,7 +102,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
102
102
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
103
103
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
104
104
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
105
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.
|
|
105
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.23.0"`.
|
|
106
106
|
|
|
107
107
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
108
108
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -127,7 +127,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
127
127
|
claude plugin install cowork-harness@cowork-harness
|
|
128
128
|
```
|
|
129
129
|
|
|
130
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.
|
|
130
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.23.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
131
131
|
|
|
132
132
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
133
133
|
|
|
@@ -149,7 +149,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
|
|
|
149
149
|
To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
|
|
150
150
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
151
151
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
152
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.
|
|
152
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.23.0"` — see
|
|
153
153
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
154
154
|
|
|
155
155
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -461,7 +461,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
461
461
|
|
|
462
462
|
**Flags worth knowing** (the full list is always `<command> --help`):
|
|
463
463
|
- `run`: a decider (`--decider-cmd <helper>`/`--decider-dir <dir>`, or a scenario's `on_unanswered: llm` — a scenario-YAML key, not a CLI flag; `run` rejects `--on-unanswered llm`) can answer unscripted gates; `--repeat N` (2-100) runs each scenario N times and aggregates a variance rollup instead of a single pass/fail (`--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd` tune the batch verdict/loop — and `--max-budget-usd` applies without `--repeat` too, as a pre-flight refusal from the scenario's cost history); `--matrix <matrix.yaml>` runs ONE scenario across the cross-product of baseline/model/skill_dir axes (worked example: `examples/matrices/csv-metrics-matrix.yaml`; `--max-cells`/`--concurrency` tune the cap/pool — any cell failing, assertion or infra, fails the run); `--compact`/`--demo` trim `run`/`skill` output for shareable screenshots/GIFs; `--label <tag>` (both `run` and `skill`) stamps a generation name for the iterate-across-fixes loop — surfaced in `result.json` (`runLabel`), the run-index row, `inspect`, and `status.json`, alongside the auto-recorded `skillCommit` (git provenance) and the authoritative `fingerprint.skillHash` a harvest step should pair critiques by.
|
|
464
|
-
- `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `replay --best-effort-future-cassette` lets a cassette from a newer format version replay anyway (warn instead of the default hard error).
|
|
464
|
+
- `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `replay --best-effort-future-cassette` lets a cassette from a newer format version replay anyway (warn instead of the default hard error). `replay` also prints an **advisory** note — never a failure — when the rootfs agent image differs from the one the cassette recorded (`environment.agentImage`): the image decides which capabilities exist, so replaying against a different one can move a verdict with nothing in the cassette having changed. It compares the registry digest first (the only identity comparable across machines) and falls back to the local image config id only when neither side has one; cassettes recorded before that field existed simply carry no image to compare.
|
|
465
465
|
- `verify-cassettes`: privacy scan (email/currency/domain/path/machine-inventory) + staleness — exit 1 when verification RAN and found a real problem (a finding, a genuine drift, or scenario-prompt drift), exit 3 when it could NOT complete (an unverifiable-class staleness finding, a too-new cassette format, or a read error); whole-token allows via `--allow <regex>` (a pattern) / class-scoped `--allow-domain` / `--allow-email` / `--allow-path` / `--allow-machine-inventory` / `--allow-patterns-file <path>` (a **file** of patterns, one regex per line); `--skip-privacy` or `--skip-staleness` runs only part of the gate; a diverged scenario `prompt` vs. the cassette's frozen prompt is also a hard fail (its own `scenarioDrift` bucket), opt out with `--skip-scenario-drift`; `--margins` adds a per-cassette recorded-vs-budget report for count-bound assertions (a single-sample estimate — diagnostic only, never changes the gate verdict); `--allow-empty` makes an **existing but cassette-free directory** exit 0 instead of the default loud exit 2 — for a repo that deliberately commits no cassettes (a *missing* path still fails, so the flag can never green a typo'd path).
|
|
466
466
|
- `stats`: reads `<runsRoot>/index.jsonl`, written automatically at every result; `--since`/`--baseline`/`--branch` filter; `--skill-hash <prefix>`/`--label <tag>` narrow to one generation of the iterate-across-fixes loop and `--group-by scenario|skill-hash|label|fidelity` splits per generation — or per effective fidelity tier — instead of aggregating across them (a window spanning >1 generation or >1 tier warns; under `skill-hash`/`label` grouping, rows lacking the field are excluded from grouping and counted, never bucketed blank — the `fidelity` key is total, so nothing is ever excluded under that grouping); `--runs` lists the individual runs behind each summary with their `skillHash`/`runLabel` (each `runs[]` entry also carries `fidelity` — the tier that run actually executed at, `effectiveFidelity ?? fidelity` — in the `--output-format json` envelope only; the text listing is unchanged); `--last <n>` windows per-group; `--reindex` rebuilds the index from the physical run-dir tree (the migration path for pre-index runs), reconstructing each critique's cost roll-up from its run dir along the way.
|
|
467
467
|
- `diff`: `--changelog` renders known-field prose for a baseline diff; `--view tools|transcript|artifacts|meta` narrows a run/cassette diff to one section; normalization (default on) masks per-run noise (ids/timestamps/session markers/host paths) so two runs of the same scenario diff as identical — `--no-normalize` compares raw values.
|
|
@@ -714,7 +714,7 @@ jobs:
|
|
|
714
714
|
- uses: actions/checkout@v4
|
|
715
715
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
716
716
|
run: |
|
|
717
|
-
V=2.1.
|
|
717
|
+
V=2.1.229 # match your scenario's pinned baseline's agentVersion
|
|
718
718
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
719
719
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
720
720
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
@@ -725,7 +725,7 @@ jobs:
|
|
|
725
725
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
726
726
|
```
|
|
727
727
|
|
|
728
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.
|
|
728
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.23.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](./action.yml) for the full input/output reference.
|
|
729
729
|
|
|
730
730
|
The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
731
731
|
|
|
@@ -893,6 +893,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
893
893
|
## Status
|
|
894
894
|
|
|
895
895
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
896
|
-
**`desktop-1.
|
|
896
|
+
**`desktop-1.30096.1`**. Release-by-release verification notes (what was re-verified against
|
|
897
897
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
898
898
|
this section would otherwise duplicate lives in the sections above.
|
|
@@ -446,6 +446,11 @@
|
|
|
446
446
|
},
|
|
447
447
|
"note": "coworkWebFetchViaApi=true coworkWebFetchPrompt=true workspaceBashWaitLonger=true sessionsBridgePollBlockMs=30 — web_fetch is host/API-routed (POST /api/organizations/<org>/cowork/web_fetch), NOT container egress; gated by a separate web-fetch hostname allowlist + URL provenance."
|
|
448
448
|
},
|
|
449
|
+
"cicCanUseToolEnabled:2051942385": {
|
|
450
|
+
"on": true,
|
|
451
|
+
"source": "force",
|
|
452
|
+
"value": true
|
|
453
|
+
},
|
|
449
454
|
"cliPlugin:2307090146": {
|
|
450
455
|
"on": false,
|
|
451
456
|
"source": "defaultValue",
|
|
@@ -467,6 +472,11 @@
|
|
|
467
472
|
"source": "defaultValue",
|
|
468
473
|
"value": "## Sensitive personal information\n\nDo not save the following to memory unless the user explicitly asks you to remember it:\n\n- Protected attributes: race, ethnicity, national origin, religion, age, sex, sexual orientation, gender identity, immigration status, disability, serious illness, union membership\n- Government identifiers: Social Security numbers, driver's license numbers, passport numbers, government ID numbers\n- Financial account details: credit card numbers, bank account numbers\n- Health information: medical conditions, diagnoses, lab results, mental health details, therapy or counseling\n- Home or personal mailing addresses (work addresses are fine)\n- Account passwords, secret tokens, or secret keys\n\nIf any of the above appears in conversation context, complete the task but do not persist it to a memory file. If the user explicitly says \"remember my address is X\", saving it is acceptable — they've given consent."
|
|
469
474
|
},
|
|
475
|
+
"coworkArtifacts:2940196192": {
|
|
476
|
+
"on": true,
|
|
477
|
+
"source": "force",
|
|
478
|
+
"value": true
|
|
479
|
+
},
|
|
470
480
|
"canSaveSkill:3246569822": {
|
|
471
481
|
"on": true,
|
|
472
482
|
"source": "force",
|
|
@@ -558,9 +568,9 @@
|
|
|
558
568
|
],
|
|
559
569
|
"spawnEnvSpreadCount": 31,
|
|
560
570
|
"fcache": {
|
|
561
|
-
"content16": "
|
|
562
|
-
"embeddedTimestamp":
|
|
563
|
-
"featureCount":
|
|
571
|
+
"content16": "341351c3593dc6d8",
|
|
572
|
+
"embeddedTimestamp": 1786562513613,
|
|
573
|
+
"featureCount": 253
|
|
564
574
|
}
|
|
565
575
|
},
|
|
566
576
|
"requireFullVmSandbox": null
|