pi-browser-use 0.10.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -179,7 +179,7 @@ class OwnedChromeProcess {
179
179
  waitForExit() {
180
180
  return this.exitPromise;
181
181
  }
182
- async shutdown(graceMs = 3_000) {
182
+ async shutdown(graceMs = 10_000) {
183
183
  if (this.exited)
184
184
  return;
185
185
  this.child.kill('SIGTERM');
@@ -57,3 +57,15 @@ npm run bench -- --startup-only --json
57
57
  ```
58
58
 
59
59
  `perf:audit` and `bench --json` are suitable for machine capture. Use at least seven cold imports and ten runtime iterations when updating the table. Record Node, OS, architecture, commit, and whether Chrome was headless; do not compare results collected with different modes or iteration counts.
60
+
61
+ ## Task-level evaluation
62
+
63
+ `npm run eval -- --json` runs deterministic browser tasks against fixed DOM fixtures in a fresh isolated session. Each task attempt gets its own runtime and records objective checks, startup/task/step timings, bounded failure text, evidence-capture error counts, and artifact paths in `eval-results/result.json`. Failed attempts save a screenshot by default; use `--evidence all` to capture evidence after every successful step or `--evidence none` to disable it. Failures are classified as `harness` (the browser environment never became usable: startup failed, no step succeeded, or cleanup broke) or `task` (the scenario regressed), and the summary reports `harnessFailures` separately so a broken browser environment is not mistaken for task regressions. Unknown CLI options are rejected. The report contract is documented in [Eval result schema](eval-results.md), including the `eval-budgets.json` gate that mirrors the Code Foundry eval contract.
64
+
65
+ ```sh
66
+ npm run eval -- --iterations 3 --json
67
+ npm run eval -- --task form-submit,extract-list --evidence all
68
+ npm run eval -- --list
69
+ ```
70
+
71
+ These tasks evaluate runtime and browser-tool behavior, not model quality. Keep them deterministic and network-free for local/CI regression checks. Model-agent evaluations should use the same task IDs and metadata but run in a separate harness, with the model, prompt, reasoning budget, dependency hash, and retry policy recorded. Do not silently truncate normal browser results; bounded evidence is only for judge-facing output.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-browser-use",
3
- "version": "0.10.0",
3
+ "version": "0.11.0",
4
4
  "description": "Opinionated browser automation via chrome-devtools-mcp: native Pi extension and portable Agent Plugins 1.0 skills + MCP server.",
5
5
  "keywords": [
6
6
  "automation",
@@ -54,6 +54,8 @@
54
54
  "typecheck": "tsc -p tsconfig.json --noEmit",
55
55
  "prepublishOnly": "npm run build",
56
56
  "bench": "node scripts/bench.mjs",
57
+ "eval": "node scripts/eval.mjs",
58
+ "preeval": "npm run build",
57
59
  "test": "node --test test/*.test.mjs",
58
60
  "preperf:audit": "npm run build",
59
61
  "perf:audit": "node scripts/performance.mjs",
@@ -73,7 +75,7 @@
73
75
  "@types/node": "^22.0.0",
74
76
  "oxfmt": "^0.66.0",
75
77
  "oxlint": "^1.81.0",
76
- "typescript": "^5.9.0"
78
+ "typescript": "^6.0.3"
77
79
  },
78
80
  "peerDependencies": {
79
81
  "@earendil-works/pi-ai": ">=0.80.10",
package/plugin.json CHANGED
@@ -1,12 +1,18 @@
1
1
  {
2
2
  "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
3
3
  "name": "pi-browser-use",
4
- "version": "0.10.0",
4
+ "version": "0.11.0",
5
5
  "description": "Managed persistent, fresh, and existing Chrome sessions with focus-safe browser tools and CLI-first skills.",
6
6
  "author": {
7
7
  "name": "0xPlayerOne"
8
8
  },
9
9
  "repository": "https://github.com/0xPlayerOne/pi-browser-use",
10
10
  "license": "MIT",
11
- "keywords": ["browser", "chrome", "mcp", "agent-skills", "pi"]
11
+ "keywords": [
12
+ "browser",
13
+ "chrome",
14
+ "mcp",
15
+ "agent-skills",
16
+ "pi"
17
+ ]
12
18
  }