pi-crew 0.9.62 → 0.9.65
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +105 -0
- package/README.md +47 -1
- package/agents/critic.md +1 -1
- package/agents/explorer.md +1 -1
- package/agents/planner.md +1 -1
- package/agents/reviewer.md +1 -1
- package/agents/security-reviewer.md +1 -1
- package/agents/test-engineer.md +1 -1
- package/agents/writer.md +1 -1
- package/dist/index.mjs +478 -366
- package/package.json +7 -3
- package/scripts/analyze-run.mjs +1333 -0
- package/scripts/pty_probe.py +10 -8
- package/scripts/resource-sampler.mjs +482 -0
- package/skills/real-test-pi-crew/SKILL.md +6 -6
- package/src/agents/agent-config.ts +4 -0
- package/src/agents/agent-serializer.ts +1 -0
- package/src/agents/discover-agents.ts +8 -0
- package/src/config/role-tools.ts +49 -1
- package/src/observability/event-to-metric.ts +29 -0
- package/src/observability/metrics-primitives.ts +41 -3
- package/src/prompt/prompt-runtime.ts +6 -0
- package/src/prompt/scratchpad-lifecycle.ts +605 -0
- package/src/runtime/README.md +1 -1
- package/src/runtime/broker/crew-broker.ts +0 -16
- package/src/runtime/child-pi/child-pi-spawn.ts +42 -1
- package/src/runtime/child-pi/child-pi.ts +5 -0
- package/src/runtime/effectiveness.ts +23 -1
- package/src/runtime/merge-gate.ts +202 -0
- package/src/runtime/model/model-fallback.ts +11 -0
- package/src/runtime/model/pi-args.ts +1 -1
- package/src/runtime/model/provider-extensions.ts +31 -12
- package/src/runtime/output/progress-tracker.ts +3 -33
- package/src/runtime/recovery/crash-recovery.ts +1 -1
- package/src/runtime/scratchpad/README.md +184 -0
- package/src/runtime/scratchpad/engine.ts +648 -0
- package/src/runtime/scratchpad/guest.ts +360 -0
- package/src/runtime/scratchpad/index.ts +22 -0
- package/src/runtime/scratchpad/protocol.ts +88 -0
- package/src/runtime/scratchpad/snapshot-hmac.ts +161 -0
- package/src/runtime/scratchpad/snapshot-lookup.ts +74 -0
- package/src/runtime/scratchpad/transform.ts +363 -0
- package/src/runtime/task-runner/child-executor.ts +48 -31
- package/src/runtime/team-runner.ts +128 -203
- package/src/schema/team-tool-schema.ts +2 -0
- package/src/teams/discover-teams.ts +2 -0
- package/src/teams/team-config.ts +7 -0
- package/src/teams/team-serializer.ts +1 -0
- package/src/ui/mascot.ts +1 -14
- package/teams/default.team.md +1 -0
- package/teams/fast-fix.team.md +1 -0
- package/src/observability/event-bus.ts +0 -86
- package/src/plugins/plugin-define.ts +0 -6
- package/src/plugins/plugin-registry.ts +0 -32
- package/src/plugins/plugins/index.ts +0 -3
- package/src/plugins/plugins/nextjs.ts +0 -19
- package/src/plugins/plugins/vite.ts +0 -10
- package/src/plugins/plugins/vitest.ts +0 -9
- package/src/runtime/child-pi/child-pi-pool.ts +0 -68
- package/src/runtime/iteration-hooks.ts +0 -305
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,111 @@
|
|
|
2
2
|
|
|
3
3
|
> **Note:** `atomic-write-v2.ts` / `AtomicWriter` mentioned in historical entries below was consolidated into `atomic-write.ts` as of v0.9.42. This changelog is preserved as historical record — the migration was completed (the v2 class was never adopted; v1 won on simplicity + symlink-safety + link+unlink atomicity). See `docs/migration/atomic-write-v2-migration.md` for the decision rationale.
|
|
4
4
|
|
|
5
|
+
## [0.9.65] — team-tool schema empty-string guard (budgetTotal) + effectiveness empty-result guard + skill drift fix (2026-08-10)
|
|
6
|
+
|
|
7
|
+
### Fixes
|
|
8
|
+
- **`budgetTotal` empty-string unset marker accepted** (`src/schema/team-tool-schema.ts`): `budgetTotal` was the only numeric `TeamToolParams` field missing the `Literal("")` union branch that its siblings (`budgetWarning`, `budgetAbort`, `tokenBudget`, `interval`, `replyDeadline`) already had. Calling models that emit every schema key with defaults (documented behavior in `normalizeTeamParams`) were rejected by pi-ai's pre-handler `validateToolArguments` → `Validation failed for tool "team"` on every action. The `MISCONFIGURATION GUARD` (rejects 1-999) is preserved — only the unset marker is added. Caught by the Tier 9 feature battery in the `real-test-pi-crew` skill (Tiers 1-8 stayed green while the team tool was broken for emitting models).
|
|
9
|
+
- **Effectiveness empty-result guard** (`src/runtime/effectiveness.ts`): a completed task with an EMPTY result artifact (`resultArtifact.sizeBytes === 0`) is now treated as no-observed-work, closing the monitoring gap where a child worker absorbed by a 429 rate-limit or model-not-found failure still emitted transcript/usage events and the run completed with `consistency=1` (zero real work). Empty-result tasks flow through the existing `noObservedWork` escalation (warn → blocked for mutating roles). New regression tests: `test/unit/runtime/core/effectiveness-guard.test.ts` (8 tests).
|
|
10
|
+
- **`real-test-pi-crew` skill drift fix**: `PI_CREW_BROKER_DIAG_UI` / `src/ui/run-dashboard.ts:831` citation removed from `skills/real-test-pi-crew/SKILL.md` Tier 6 (the env var was removed in `e3ee6fe2`); Tier 6 now documents the screen-change-evidence replacement. `scripts/pty_probe.py` no longer sets the dead env var.
|
|
11
|
+
|
|
12
|
+
### Verified
|
|
13
|
+
- `npm run test:critical`: 101/101 pass (default, `PI_CREW_BROKER=0`, `PI_CREW_BROKER=1`).
|
|
14
|
+
- New `effectiveness-guard.test.ts`: 8/8 pass.
|
|
15
|
+
- `npm run typecheck` + `npm run build:bundle` exit 0; bundle md5 `e39373498d618de7e233c361ebb03b03`.
|
|
16
|
+
- Full 9-tier real-test re-run (2026-08-10): Tiers 1-8 + 9a (10/10) + 9b (5/5 spawn paths) pass; the previous run's failure modes (model-not-found spawn loop, 429-absorbed empty output) did not reproduce. Report: `docs/real-test/reports/real-test-2026-08-10-full-9-tier-f4-effectiveness-guard.md`.
|
|
17
|
+
|
|
18
|
+
## [0.9.64] — pi-rlm→pi-crew pattern transfer: worker scratchpad + crash-resume + cancellation + quick wins (2026-08-09)
|
|
19
|
+
|
|
20
|
+
### Quick Wins (patterns 17/19/20/11 + spike CI)
|
|
21
|
+
|
|
22
|
+
### Features
|
|
23
|
+
- **Schema-driven docs (QW17)**: `agents/*.md` frontmatter `tools:` now matches the
|
|
24
|
+
enforced `ROLE_TOOL_CONFIGS` (4 drifts fixed); a sync-test pins the derivation;
|
|
25
|
+
`scripts/gen-role-tools-docs.mjs` renders `docs/role-tools.md` from the source.
|
|
26
|
+
- **Retry-resume contract suite (QW19)**: pins `executeWithRetry` (`?`-glob, empty-
|
|
27
|
+
retryableErrors, maxAttempts:0, abort-during-sleep, default attemptId),
|
|
28
|
+
`FileCheckpointStore` (corrupt-file quarantine, list-skip, wrong-runId delete),
|
|
29
|
+
and exports `shouldRecoverTask` for direct testing.
|
|
30
|
+
- **Failure-mode inventory (QW20)**: `docs/failure-mode-inventory.md` maps the 7
|
|
31
|
+
pi-rlm failure modes to pi-crew handlers (wedge gap closed by Phase 1
|
|
32
|
+
ping-before-execute; EPIPE/timeout interplay declared as gaps).
|
|
33
|
+
- **Error-as-data contract (QW11)**: `evidenceStatusFor` + `attemptErrorFor`
|
|
34
|
+
extracted as pure functions from `runChildProcessTask`; contract test pins the
|
|
35
|
+
precedence (cancelled > failed > completed; E007 timedOut override; 429 gated).
|
|
36
|
+
- **`test:spike` script**: wires the scratchpad spike tests into CI.
|
|
37
|
+
|
|
38
|
+
|
|
39
|
+
### Phase 3 cancellation hardening (experimental)
|
|
40
|
+
|
|
41
|
+
### Features
|
|
42
|
+
|
|
43
|
+
- **Atomic snapshot write (D2')**: `engine.snapshotState` now writes via
|
|
44
|
+
`temp + rename` (same directory), eliminating the theoretical torn-write race
|
|
45
|
+
between the debounce timer and the F3 quit flush.
|
|
46
|
+
- **Kill-and-restore verified**: the existing SIGTERM → pi print-mode dispose →
|
|
47
|
+
`session_shutdown` reason:"quit" → F3 flush chain already captures a fresh
|
|
48
|
+
snapshot on worker kill (no new handler required). Documented + pinned by a
|
|
49
|
+
gated real-worker test.
|
|
50
|
+
- **EngineBusyError: deliberately skipped** — ping-before-execute already
|
|
51
|
+
detects a wedged guest; the FIFO execute queue makes a busy-reject redundant.
|
|
52
|
+
|
|
53
|
+
|
|
54
|
+
### Phase 2 crash-resume (experimental)
|
|
55
|
+
|
|
56
|
+
### Features
|
|
57
|
+
|
|
58
|
+
- **Cross-attempt restore (crash-resume)**: a scratchpad worker attempt N+1 (retry
|
|
59
|
+
/ crash-recovery re-queue / manual re-run) automatically revives the namespace
|
|
60
|
+
from the previous attempt's redacted snapshot artifact — via a single
|
|
61
|
+
spawn-time lookup (`findLatestScratchpadSnapshot`), no retry-loop / recovery
|
|
62
|
+
changes needed. Restores lazily on the first `execute` call (D7 lazy invariant
|
|
63
|
+
kept) with a one-line model-visible notice.
|
|
64
|
+
- Lookup: latest mtime (model-fallback index resets each retry round → number
|
|
65
|
+
is not write-order); tie-break lowest attempt = newest round.
|
|
66
|
+
- Security: read-time re-validation (containment + filename pattern + lstat +
|
|
67
|
+
size + mtime hint, D10); fail-open (D11) never breaks the worker; strict scan
|
|
68
|
+
(D12); base64 round-trip (D13); redacted secret → literal `"***"` (D4).
|
|
69
|
+
- **Guest zombie backstop (SEC-1)**: the engine now overrides
|
|
70
|
+
`PI_CREW_PARENT_PID=worker pid` for the guest (D5), so an orphaned guest
|
|
71
|
+
(worker SIGKILL'd) is flagged by the zombie scanner — pure inheritance left
|
|
72
|
+
guests LIVE forever holding provider keys + broker token.
|
|
73
|
+
- **Snapshot cap (D6)**: 4 MiB two-sided (write-side raw byteLength trim, read-
|
|
74
|
+
side file + per-var 256 KiB) bounds v8.deserialize amplification.
|
|
75
|
+
|
|
76
|
+
|
|
77
|
+
### Phase 1 worker scratchpad (experimental)
|
|
78
|
+
|
|
79
|
+
### Features
|
|
80
|
+
|
|
81
|
+
- **Worker stateful scratchpad (experimental)**: opt-in persistent Bun-free JS
|
|
82
|
+
evaluator (`execute` tool) for workers, ported from the pi-rlm pattern.
|
|
83
|
+
State (variables, parsed data) compounds across `execute` calls within a task
|
|
84
|
+
attempt — intermediate results live in a namespace instead of being re-derived
|
|
85
|
+
from transcript text. Snapshot per-attempt into the run artifact store
|
|
86
|
+
(redacted + atomic) prepares the ground for Phase 2 crash-resume.
|
|
87
|
+
- Opt-in per role: `executor`, `test-engineer`, `verifier` (default on); other
|
|
88
|
+
roles via agent frontmatter `scratchpad: true` (write roles only).
|
|
89
|
+
- Security: S-6 read-only roles are gated out regardless of frontmatter
|
|
90
|
+
(privilege-elevation guard); F6 `scratchpad: false` is an explicit kill-switch;
|
|
91
|
+
raw snapshots never land in the artifact root (temp dir → redacted
|
|
92
|
+
`writeArtifact`).
|
|
93
|
+
- Dormant by default: only active when `PI_CREW_SCRATCHPAD=1` is set by the
|
|
94
|
+
spawner; zero behavior change for non-opt-in workers.
|
|
95
|
+
|
|
96
|
+
|
|
97
|
+
## [0.9.63] — built-in performance observability + local-path provider-extension discovery (2026-08-08)
|
|
98
|
+
|
|
99
|
+
### Features
|
|
100
|
+
|
|
101
|
+
- **Built-in performance observability (byte-built-in, always-on).** Every team run now auto-attaches a detached resource sampler and auto-generates a detailed performance report on completion — no separate benchmark harness needed. Toggle per-team via frontmatter `observability: true|false` (default `true` for parsed team files; direct-object `TeamConfig` fixtures stay unset for test isolation).
|
|
102
|
+
- **Live resource sampler** (`scripts/resource-sampler.mjs`): samples CPU/RSS per-PID every 2s via ppid-tree attribution (root runner + all child workers, including respawns), with PID-reuse guard (`/proc` starttime), first-sample CPU exclusion, and **6 live warning categories** — `high_cpu` (≥300%), `rss_jump` (+200MB/interval), `rss_high` (≥1GB), `rss_leak` (window-30 monotonic +100MB), `proc_died`, `proc_zombie`. Rate-limited (10s/pid/category); `--no-live-warn` flag to silence.
|
|
103
|
+
- **Post-hoc analyzer** (`scripts/analyze-run.mjs`): combines `events.jsonl` + transcripts + `resources.jsonl` into a markdown report at `docs/perf-report-<runId>.md` with **22 anomaly categories** (task_failed, model_retry/model_cascade, large_gap, slow_phase, launch_delay, drain_stall, token_imbalance, no_cache, worker_respawn_churn, api_error_storm, zero_output_completion, sustained_cpu, transient_cpu/rss_spike, rss_growth, run_not_completed, high_failure_rate, run_idle, cost_unreported, missing_transcript, sampler_gap, tool_churn), a per-subagent timeline (launch / respawn / startup / active-work / drain / finalize), and token/cost/model attribution. Optional `--agents` flag emits per-agent breakdowns.
|
|
104
|
+
- **Runtime wiring** (`src/runtime/team-runner.ts`): `startPerfSampler` spawns the sampler detached + `unref`'d (death never affects the run); `schedulePerfAnalyze` runs the analyzer +3s after `after_run_complete` via an `unref`'d `setTimeout` (never blocks run completion). Strict `observability !== true` keeps test fixtures from spawning. The sampler auto-stops when the run manifest reaches a terminal status (`--watch-run` mode).
|
|
105
|
+
- **Overhead ≈ 0** (measured A/B): sampler ~0.05% of one core, 56MB RSS fixed; analyzer ~72ms one-shot after run; ~32KB artifacts/run. Verified end-to-end on real runs — see `docs/real-test/reports/real-test-2026-08-07-perf-obs-overhead.md`.
|
|
106
|
+
|
|
107
|
+
### Bug fixes
|
|
108
|
+
|
|
109
|
+
- **Local-path provider extensions were not discovered for child workers (oc-go and any `pi install <local-path>` provider went "Model not found").** `discoverProviderExtensions` only resolved `npm:` specs from `~/.pi/agent/settings.json` `packages`, skipping local-path specs (e.g. `../../source/my_pi/source/pi-other-provider`) on the assumption that local paths were the pi-crew extension itself. That assumption was wrong for local provider extensions: `--no-extensions` in `buildPiWorkerArgs` stripped the provider from every child spawn → every model from that provider hit `Error: Model "…" not found` → the fallback chain burned 5 spawn-fails (~10s) per task before landing on a builtin provider. Fix: resolve local-path specs (`./`, `../`, absolute) relative to the settings.json dir — same trust level as `npm:` (settings packages are a sanctioned channel written by `pi install`); skip the pi-crew package itself via `packageRoot()` (a worker must not re-load the orchestrator). **SEC-1 preserved** (project/project-pi AGENT extensions in `.crew/agents/*.md` frontmatter stay gated — that is a separate, untrusted channel). Tests in `test/unit/runtime/model/provider-extensions.test.ts`. Investigation + correction of the earlier "hidden models" mis-attribution in `docs/real-test/reports/real-test-2026-08-08-provider-ext-local-path.md`.
|
|
5
110
|
|
|
6
111
|
## [0.9.62] — provider-quota attribution per live-session agent + dead-worker alert re-fire fix (2026-08-06)
|
|
7
112
|
|
package/README.md
CHANGED
|
@@ -69,8 +69,10 @@ repo: https://github.com/baphuongna/pi-crew
|
|
|
69
69
|
- **Durable event replay** (L1, v0.9.8) — `RunEventBus.onWithReplay()` catches up a re-subscribing dashboard/overlay with events it missed during transient absence (toggle, reconnect), replaying from the durable JSONL log with seq-based dedup. No information loss even if the live subscriber was briefly gone.
|
|
70
70
|
- **Lossless-by-default output handling** (L4, v0.9.8) — worker output thresholds sized from measured data (100% of real outputs fit without compaction); when compaction is unavoidable it keeps head+tail (preserves closing code fences/headings) instead of head-only truncation. No more `[pi-crew compacted N chars]` markers eating the end of a worker's result.
|
|
71
71
|
- **Inter-pi broker** (v0.9.47, default-on) — a Unix-domain-socket message bus that lets concurrently-running Pi sessions pass messages, steering notes, and task-status events to each other. **On by default** on Linux + macOS; auto-disabled on native Windows (no unix socket). Three independent kill switches: `broker.enabled: false` (config), `PI_CREW_BROKER=0` (env, always wins), Windows auto-disable. See [docs/decisions/2026-07-22-broker-phase4-gated-on.md](docs/decisions/2026-07-22-broker-phase4-gated-on.md).
|
|
72
|
+
- **Worker stateful scratchpad** (experimental, opt-in per role) — `executor` / `test-engineer` / `verifier` workers get a `scratchpad` tool: a persistent Bun-free JS evaluator whose namespace **compounds across calls within a task attempt** (variables set in one cell are visible in the next), so intermediate results live in memory instead of being re-derived from the transcript. Snapshots are flushed (redacted, atomic) per-attempt into the artifact store; the next attempt (retry / crash-recovery re-queue / re-run) **automatically revives the namespace** from the latest snapshot. Dormant by default (armed only when the spawner sets `PI_CREW_SCRATCHPAD=1`); zero behavior change for non-opt-in workers. Ported from the `@shift-labs/pi-rlm` pattern. See [src/runtime/scratchpad/README.md](src/runtime/scratchpad/README.md) (Phase 1-3 design, env keys, guards, threat model).
|
|
72
73
|
- **`test:critical` + `real-test-pi-crew` skill** (v0.9.47) — a curated 14-file / 97-test subset (`npm run test:critical`, ~20s) for fast in-loop verification, plus a bundled skill distilling the full 8-tier end-to-end verification discipline (unit → 3-path kill-switch proof → typecheck/bundle → live TUI probing → smoke team run). Prevents the verifier-worker hang that full `npm test` (>4 min) caused against the 300s worker timeout.
|
|
73
|
-
- **Provider extensions in subagents** (v0.9.57) — pi-crew spawns child-pi workers with `--no-extensions` (security posture), which made extension-registered providers (e.g. `pi-commandcode-provider`) unresolvable inside subagents. pi-crew now **auto-discovers provider packages** from `~/.pi/agent/settings.json` `packages`
|
|
74
|
+
- **Provider extensions in subagents** (v0.9.57, local-path support v0.9.63) — pi-crew spawns child-pi workers with `--no-extensions` (security posture), which made extension-registered providers (e.g. `pi-commandcode-provider`, `pi-other-provider`) unresolvable inside subagents. pi-crew now **auto-discovers provider packages** from `~/.pi/agent/settings.json` `packages` — both `npm:` specs (v0.9.57) and **local-path specs** like `../../source/foo` (v0.9.63) — and loads them via `--extension` in every builtin/user subagent, so **all provider models work in subagents**. An explicit `runtime.agentExtensions: string[]` config is an optional extra allowlist on top of auto-discovery. **SEC-1 preserved:** project/project-pi agents never receive these (env-gate unchanged).
|
|
75
|
+
- **Built-in performance observability** (v0.9.63) — every team run auto-attaches a detached resource sampler (per-PID CPU/RSS via ppid-tree attribution, 6 live warning categories) and auto-generates a markdown performance report on completion (22 anomaly categories, per-subagent timeline, token/cost/model attribution). Toggle per-team via frontmatter `observability: true|false`. Overhead ≈ 0 (sampler ~0.05% CPU / 56MB RSS; analyzer ~72ms post-run). See [Built-in performance observability](#built-in-performance-observability) below.
|
|
74
76
|
|
|
75
77
|
---
|
|
76
78
|
|
|
@@ -287,6 +289,18 @@ The advisory is **informational only** — there is no `force:true` flag needed
|
|
|
287
289
|
|
|
288
290
|
## Recent changes
|
|
289
291
|
|
|
292
|
+
### v0.9.65: team-tool schema empty-string guard + effectiveness empty-result guard (2026-08-10)
|
|
293
|
+
|
|
294
|
+
- **`budgetTotal` empty-string unset marker accepted**: `budgetTotal` was the only numeric `TeamToolParams` field missing the `Literal("")` union branch its siblings had. Calling models that emit every schema key with defaults were rejected by pi-ai's pre-handler validation → `Validation failed for tool "team"` on every action. The `MISCONFIGURATION GUARD` (rejects 1-999) is preserved. Caught by the Tier 9 feature battery — Tiers 1-8 stayed green while the team tool was broken for emitting models.
|
|
295
|
+
- **Effectiveness empty-result guard**: a completed task with an empty result artifact (`sizeBytes === 0`) is now treated as no-observed-work — closing the monitoring gap where a 429-absorbed / model-not-found worker produced zero real content but the run still completed with `consistency=1`. Empty-result tasks flow through the existing `noObservedWork` escalation. Regression tests: `test/unit/runtime/core/effectiveness-guard.test.ts` (8 tests).
|
|
296
|
+
- Full 9-tier real-test re-run (2026-08-10): Tiers 1-8 + 9a (10/10) + 9b (5/5) pass; previous failure modes did not reproduce. See [CHANGELOG.md](CHANGELOG.md) §0.9.65 and `docs/real-test/reports/real-test-2026-08-10-full-9-tier-f4-effectiveness-guard.md`.
|
|
297
|
+
|
|
298
|
+
### v0.9.63: built-in performance observability + local-path provider-extension discovery
|
|
299
|
+
|
|
300
|
+
- **Built-in performance observability (always-on, toggle per team)**: every team run now auto-attaches a detached resource sampler (`scripts/resource-sampler.mjs` — per-PID CPU/RSS via ppid-tree attribution, 6 live warning categories: high_cpu / rss_jump / rss_high / rss_leak / proc_died / proc_zombie) and auto-generates a markdown performance report (`scripts/analyze-run.mjs` → `docs/perf-report-<runId>.md` — 22 anomaly categories, per-subagent launch/respawn/active/drain timeline, token/cost/model attribution). Runtime wiring in `src/runtime/team-runner.ts`: `startPerfSampler` (detached + `unref`'d; death never affects the run) + `schedulePerfAnalyze` (+3s after `after_run_complete`, `unref`'d `setTimeout`). Toggle: team frontmatter `observability: true|false` (default `true`). **Overhead ≈ 0** (A/B verified: sampler ~0.05% CPU / 56MB RSS, analyzer ~72ms post-run, ~32KB artifacts/run). New scripts: `scripts/resource-sampler.mjs`, `scripts/analyze-run.mjs`. Tests: `test/unit/scripts/{analyze-run,resource-sampler}-audit.test.ts`.
|
|
301
|
+
- **Local-path provider extensions now discovered for child workers**: `discoverProviderExtensions` previously resolved only `npm:` specs from `~/.pi/agent/settings.json` `packages`, skipping local-path specs on the assumption they were the pi-crew extension itself. That broke local provider extensions (e.g. `pi-other-provider` installed via `pi install <local-path>`) — every model from such a provider hit `Error: Model "…" not found` in child workers and burned ~10s/task of fallback churn. Fix: resolve `./`, `../`, and absolute specs relative to the settings.json dir (same sanctioned trust level as `npm:`), and skip pi-crew itself via `packageRoot()`. **SEC-1 preserved** (project/project-pi AGENT extensions stay gated). Tests in `test/unit/runtime/model/provider-extensions.test.ts`.
|
|
302
|
+
- See [CHANGELOG.md](CHANGELOG.md) §0.9.63 and `docs/real-test/reports/real-test-2026-08-08-provider-ext-local-path.md`.
|
|
303
|
+
|
|
290
304
|
### v0.9.57: team-tool schema repair + provider-extension auto-discovery + post-reorg repo-layout consolidation
|
|
291
305
|
|
|
292
306
|
- **Team tool repaired (was broken live while tests stayed green)**: calling models emit empty-string/boolean defaults for every schema key, which pi-ai's pre-handler `validateToolArguments` rejected (`Validation failed for tool team`) and `Type.Unsafe` schema fields without `[TypeBox.Kind]` made `Value.Check` throw (`Unknown type`). Schema now accepts unset markers natively; `SkillOverride`/`FreeformConfig` switched to TypeBox-native constructors; `normalizeTeamParams` drops empties in the handler. Chain-runner also fixed (quote-aware step splitting). Caught by a new **Tier 9 feature battery** in the [`real-test-pi-crew`](skills/real-test-pi-crew/SKILL.md) skill (live team-tool action coverage).
|
|
@@ -363,6 +377,38 @@ pi-crew supports multiple runtime modes for task execution:
|
|
|
363
377
|
{ "executeWorkers": false }
|
|
364
378
|
```
|
|
365
379
|
|
|
380
|
+
## Built-in performance observability
|
|
381
|
+
|
|
382
|
+
Every team run auto-attaches a **detached resource sampler** and auto-generates a **performance report** on completion — measuring real resource usage and surfacing anomalies from the run's actual events/transcripts, not a synthetic benchmark.
|
|
383
|
+
|
|
384
|
+
### Artifacts produced (per run, under `.crew/artifacts/<runId>/`)
|
|
385
|
+
|
|
386
|
+
| File | Contents |
|
|
387
|
+
|------|----------|
|
|
388
|
+
| `resources.jsonl` | Per-PID CPU/RSS samples every 2s (root runner + all child workers via ppid-tree attribution, including respawns). PID-reuse guarded by `/proc` starttime; first-sample CPU excluded from averages. |
|
|
389
|
+
| `perf-obs.log` | Sampler diagnostics: spawn marker, live warnings, terminal-stop confirmation. |
|
|
390
|
+
| `docs/perf-report-<runId>.md` | Markdown report: 22 anomaly categories, per-subagent timeline (launch/respawn/startup/active-work/drain/finalize), token/cost/model attribution. |
|
|
391
|
+
|
|
392
|
+
### Live warnings (written to `perf-obs.log` during the run)
|
|
393
|
+
|
|
394
|
+
`high_cpu` (≥300% one core) · `rss_jump` (+200MB/interval) · `rss_high` (≥1GB) · `rss_leak` (window-30 monotonic +100MB) · `proc_died` · `proc_zombie`. Rate-limited (10s/pid/category).
|
|
395
|
+
|
|
396
|
+
### Toggle
|
|
397
|
+
|
|
398
|
+
```yaml
|
|
399
|
+
# teams/my-team.team.md
|
|
400
|
+
---
|
|
401
|
+
name: my-team
|
|
402
|
+
observability: false # default: true — set false to skip sampler + report
|
|
403
|
+
---
|
|
404
|
+
```
|
|
405
|
+
|
|
406
|
+
`observability: true` is the default for parsed team files; direct-object `TeamConfig` fixtures (unit tests) stay unset so they never spawn the sampler. `schedulePerfAnalyze` runs the analyzer `+3s` after `after_run_complete` via an `unref`'d `setTimeout` — it never blocks run completion. The sampler auto-stops when the run manifest reaches a terminal status.
|
|
407
|
+
|
|
408
|
+
### Overhead
|
|
409
|
+
|
|
410
|
+
Measured A/B (same team/goal, observability on vs off): **no detectable wall-time difference** (delta inside the 429-storm noise). Sampler: ~0.05% of one core, 56MB RSS fixed, detached + `unref`'d. Analyzer: ~72ms one-shot after run. ~32KB artifacts/run. See `docs/real-test/reports/real-test-2026-08-07-perf-obs-overhead.md`.
|
|
411
|
+
|
|
366
412
|
## Async Runs
|
|
367
413
|
|
|
368
414
|
Async runs are **detached** from the session — they survive session switches and reloads. Pi-crew notifies when complete.
|
package/agents/critic.md
CHANGED
|
@@ -5,7 +5,7 @@ model: false
|
|
|
5
5
|
systemPromptMode: replace
|
|
6
6
|
inheritProjectContext: true
|
|
7
7
|
inheritSkills: false
|
|
8
|
-
tools: read, grep, find, ls
|
|
8
|
+
tools: read, grep, find, ls, glob
|
|
9
9
|
---
|
|
10
10
|
|
|
11
11
|
You are a critical reviewer. Find flaws, missing steps, unsafe assumptions, overengineering, underengineering, and verification gaps. Return concrete fixes to the plan.
|
package/agents/explorer.md
CHANGED
|
@@ -5,7 +5,7 @@ model: false
|
|
|
5
5
|
systemPromptMode: replace
|
|
6
6
|
inheritProjectContext: true
|
|
7
7
|
inheritSkills: false
|
|
8
|
-
tools: read, grep, find, ls
|
|
8
|
+
tools: read, grep, find, ls, glob, bash
|
|
9
9
|
---
|
|
10
10
|
|
|
11
11
|
You are a fast codebase explorer. Map relevant files, symbols, data flow, and constraints. Do not modify files. Return concise findings with paths and evidence.
|
package/agents/planner.md
CHANGED
|
@@ -5,7 +5,7 @@ model: false
|
|
|
5
5
|
systemPromptMode: replace
|
|
6
6
|
inheritProjectContext: true
|
|
7
7
|
inheritSkills: false
|
|
8
|
-
tools: read, grep, find, ls
|
|
8
|
+
tools: read, grep, find, ls, glob
|
|
9
9
|
---
|
|
10
10
|
|
|
11
11
|
You are a planning specialist. Convert the goal and discovery notes into a concrete, ordered plan. Identify dependencies, risks, validation steps, and handoff instructions for implementers.
|
package/agents/reviewer.md
CHANGED
|
@@ -5,7 +5,7 @@ model: false
|
|
|
5
5
|
systemPromptMode: replace
|
|
6
6
|
inheritProjectContext: true
|
|
7
7
|
inheritSkills: false
|
|
8
|
-
tools: read, grep, find, ls, bash
|
|
8
|
+
tools: read, grep, find, ls, glob, bash
|
|
9
9
|
---
|
|
10
10
|
|
|
11
11
|
You are a code reviewer. Review the implementation for bugs, regressions, maintainability issues, missing tests, and project-rule violations. Return prioritized findings with evidence.
|
|
@@ -5,7 +5,7 @@ model: false
|
|
|
5
5
|
systemPromptMode: replace
|
|
6
6
|
inheritProjectContext: true
|
|
7
7
|
inheritSkills: false
|
|
8
|
-
tools: read, grep, find
|
|
8
|
+
tools: read, grep, find
|
|
9
9
|
---
|
|
10
10
|
|
|
11
11
|
You are a security reviewer. Look for injection, authn/authz flaws, insecure defaults, secret exposure, unsafe filesystem/network behavior, and dependency risks. Return severity and remediation.
|
package/agents/test-engineer.md
CHANGED
|
@@ -5,7 +5,7 @@ model: false
|
|
|
5
5
|
systemPromptMode: replace
|
|
6
6
|
inheritProjectContext: true
|
|
7
7
|
inheritSkills: false
|
|
8
|
-
tools: read,
|
|
8
|
+
tools: read, edit, write, bash, ls
|
|
9
9
|
---
|
|
10
10
|
|
|
11
11
|
You are a test engineer. Identify the right test level, add or adjust tests when asked, detect flaky assumptions, and report exact validation commands and results.
|
package/agents/writer.md
CHANGED
|
@@ -5,7 +5,7 @@ model: false
|
|
|
5
5
|
systemPromptMode: replace
|
|
6
6
|
inheritProjectContext: true
|
|
7
7
|
inheritSkills: false
|
|
8
|
-
tools: read,
|
|
8
|
+
tools: read, edit, write, ls
|
|
9
9
|
---
|
|
10
10
|
|
|
11
11
|
You are a documentation specialist. Produce clear, concise, maintainable docs and summaries. Preserve technical accuracy and avoid marketing fluff.
|