@browserless/agent 0.0.0-stage → 14.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE.md ADDED
@@ -0,0 +1,21 @@
1
+ The MIT License (MIT)
2
+
3
+ Copyright © 2019 Microlink <hello@microlink.io> (microlink.io)
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in
13
+ all copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
21
+ THE SOFTWARE.
package/README.md CHANGED
@@ -1,3 +1,638 @@
1
- # Temporary Holding Version
1
+ # @browserless/agent
2
2
 
3
- This version is a temporary placeholder for this package. An operational version to replace this has been submitted for review and is awaiting a staged release.
3
+ Drive an existing Puppeteer page toward a natural-language goal. The caller owns
4
+ navigation, browser setup, screenshots and closing the browser. This package does
5
+ not change `@browserless/ai`, which uses Chrome's built-in AI APIs.
6
+
7
+ Experimental. It has completed a live Wallapop search (see below), but `done`
8
+ is the model's own judgment and runs are not fully reliable.
9
+
10
+ ## Usage
11
+
12
+ Requires Node.js >=24. After the package is published:
13
+
14
+ ```sh
15
+ npm install @browserless/agent puppeteer
16
+ export AI_GATEWAY_API_KEY=...
17
+ ```
18
+
19
+ ```js
20
+ import puppeteer from 'puppeteer'
21
+ import agent from '@browserless/agent'
22
+
23
+ const browser = await puppeteer.launch()
24
+ const page = agent(await browser.newPage())
25
+
26
+ await page.goto('https://wallapop.com', { waitUntil: 'networkidle2' })
27
+ const { status, profiling } = await page.goal('busca el bmw x3 más barato')
28
+ if (status === 'success') {
29
+ const { data } = await page.extract('get the search results', {
30
+ products: {
31
+ attr: { name: { type: 'string' }, price: { type: 'number' }, url: { type: 'url' } }
32
+ }
33
+ })
34
+ console.log(data.products)
35
+ }
36
+ const { accuracy, cost, timing } = await profiling() // optional: one extra request, for accuracy
37
+ await browser.close()
38
+ ```
39
+
40
+ `agent(page, defaults?)` adds two methods to the Puppeteer page it is given
41
+ and returns the same page:
42
+
43
+ - `page.goal(goal, options?)` drives the page, over as many steps as it takes,
44
+ until the model reports the goal is met.
45
+ - `page.extract(...)` returns data from the current page. It does not navigate.
46
+
47
+ `defaults` apply to both; options passed to a call override them, and an option
48
+ passed as `undefined` or `null` keeps the default. The same
49
+ functions are exported for use without touching the page: `agent.goal(page, …)`
50
+ and `agent.extract(page, …)`.
51
+
52
+ Import compatibility is CommonJS plus the Node ESM default import, with
53
+ TypeScript declarations.
54
+
55
+ ### From the browserless CLI
56
+
57
+ `browserless exec <file>` from `@browserless/cli` opens a page and passes it to
58
+ the function the file exports, so the script does not manage the browser:
59
+
60
+ ```js
61
+ const agent = require('@browserless/agent')
62
+
63
+ module.exports = async ({ page: browserPage, browserless }) => {
64
+ const page = agent(browserPage)
65
+ await browserless.goto(page, { url: 'https://wallapop.com' })
66
+ return page.goal('busca el bmw x3 más barato')
67
+ }
68
+ ```
69
+
70
+ ```sh
71
+ browserless exec agent-example.js
72
+ ```
73
+
74
+ A goal or an extraction that fails does not throw: the script gets
75
+ `status: 'error'` with `error` and decides what to do. The CLI exits with code 1
76
+ only when the script itself throws, so a script that should fail the command
77
+ throws that `error`.
78
+
79
+ ### Examples
80
+
81
+ Every file in `examples/` is an exec script: it navigates, gives
82
+ the page one or more goals, and returns the extracted data, which the CLI
83
+ prints. A goal or an extraction that fails makes the script throw its error. Jev is used through the gateway, so `AI_GATEWAY_API_KEY` is the only
84
+ key needed. With `DEBUG=browserless:agent`, each goal also logs one line with
85
+ its `status`, `accuracy`, `cost` and `timing`. The examples call `profiling()` for every goal
86
+ whether or not the line is printed, so each goal makes the one evaluation
87
+ request.
88
+
89
+ | Script | Steps |
90
+ | --- | --- |
91
+ | `examples/wallapop.js` | Search `bmw x3`; sort by lowest price; extract the products |
92
+ | `examples/wikipedia.js` | Search for an article and open it; extract title and first paragraph |
93
+ | `examples/google-flights.js` | Reject cookies if asked; find one-way Zurich to London flights 30 days from today; extract the flights |
94
+ | `examples/hacker-news.js` | Open the comments of the first story; extract title, points and comments |
95
+ | `examples/github.js` | Open the open issues of this repository; extract them |
96
+ | `examples/run.js` | Any page: `--url=<url> --goal=<goal> --extract=<what to get>` |
97
+
98
+ ```sh
99
+ browserless exec examples/wallapop.js --no-headless
100
+ DEBUG=browserless:agent browserless exec examples/hacker-news.js --no-headless
101
+ browserless exec examples/run.js --no-headless --url=https://en.wikipedia.org/wiki/Main_Page --goal='Find and open the Wikipedia article about Alan Turing.' --extract='get the article title'
102
+ ```
103
+
104
+ `--no-headless` opens a visible browser window, so every step the agent takes
105
+ can be watched. Without it the browser runs headless.
106
+
107
+ ```js
108
+ const agent = require('@browserless/agent')
109
+ const debug = require('debug-logfmt')('browserless:agent')
110
+ const { flat, succeeded } = require('./util')
111
+
112
+ module.exports = async ({ page, browserless }) => {
113
+ agent(page)
114
+
115
+ await browserless.goto(page, { url: 'https://wallapop.com' })
116
+ const searched = await page.goal('busca "bmw x3"')
117
+ debug('search', { status: searched.status, ...flat(await searched.profiling()) })
118
+ succeeded(searched)
119
+ const sorted = await page.goal('ordena los resultados de más barato a más caro')
120
+ debug('sort', { status: sorted.status, ...flat(await sorted.profiling()) })
121
+ succeeded(sorted)
122
+
123
+ const { data } = succeeded(
124
+ await page.extract('get the search results', {
125
+ products: {
126
+ attr: { name: { type: 'string' }, price: { type: 'number' }, url: { type: 'url' } }
127
+ }
128
+ })
129
+ )
130
+ return data
131
+ }
132
+ ```
133
+
134
+ `npm start` runs the Wallapop one. All six were run as shipped on a paid gateway
135
+ account and returned data matching their goal: 32 Wallapop products sorted by
136
+ lowest price, the Wikipedia title, 48 flights with prices, a Hacker News story
137
+ with its points and comments, 4 GitHub issues, and `Alan Turing`. That is a
138
+ check that they can work, not a reliability figure. What went wrong along the
139
+ way:
140
+
141
+ - Wikipedia's first paragraph was never extracted: the first `<p>` on the page
142
+ is empty and a rule reads the first match.
143
+ - Google Flights shows a cookie consent page in the EU, hence its first goal,
144
+ and navigates with the browserless adblocker off because the adblocker leaves
145
+ that page blank. One of its two final runs failed when the AI SDK refused a
146
+ decision whose chosen option was not the most probable.
147
+ - One Wallapop run failed when the gateway took over 25 seconds to answer.
148
+ - Model requests are not retried, so one slow or refused answer ends a run.
149
+
150
+ ## API
151
+
152
+ ### `await page.goal(goal, options?)`
153
+
154
+ Resolves to `{ status, steps, decisions, trace, profiling }`, plus `error` when
155
+ `status` is `'error'`:
156
+
157
+ ```js
158
+ const { status, error, profiling } = await page.goal('Open the comments page of the first story.')
159
+ ```
160
+
161
+ `status` is `'success'` when the model reports visible satisfaction of the goal,
162
+ and `'error'` when the run stopped for any other reason; `error` then says why.
163
+ `goal` does not throw for a run that fails. It throws only when it is called
164
+ wrong: no page, an empty goal, an invalid option, or a goal already running on
165
+ the page. Success is a model
166
+ judgment, not an independent proof of correctness. Each trace entry has the
167
+ operation, snapshot action ID, operation confidence/probabilities, selected
168
+ target confidence/probabilities, step/request counts, stale status and whether
169
+ the page changed. Text-entry values are recorded too; treat traces as private
170
+ data.
171
+
172
+ Call `profiling` to get what the run cost, how long it took, and a model's
173
+ verdict on whether the goal was met. It is there for failed runs too:
174
+
175
+ ```js
176
+ const { profiling } = await page.goal('Open the comments page of the first story on the front page.')
177
+ const { accuracy, cost, timing } = await profiling()
178
+ // example values:
179
+ // accuracy → { passed: true, probability: 0.92 }
180
+ // cost → { calls: 4, inputTokens: 8982, outputTokens: 500, cachedInputTokens: 1800,
181
+ // cacheWriteTokens: 600, reasoningTokens: 43, usd: 0.000306424,
182
+ // generationIds: ['gen_01M492ME3C…', …] }
183
+ // timing → { totalMs: 2410, modelMs: 1530, otherMs: 880, evaluationMs: 310 }
184
+ ```
185
+
186
+ `cost` and `timing` come from what each model call already returned, with no
187
+ extra request: tokens from the AI SDK `usage`, dollars from the AI Gateway
188
+ `providerMetadata.gateway.cost`, and `generationIds` to look calls up in the
189
+ gateway. `usd`, `inputTokens` and `outputTokens` are `undefined` when any call
190
+ did not report them, such as a model outside the gateway, so a total is never
191
+ understated. The cache and reasoning breakdowns count only the calls that report
192
+ them; decision models report none. `otherMs` is the run time not spent waiting
193
+ on a model: observing, acting, settling and waits.
194
+
195
+ `accuracy` is the one extra request: the `evaluator` model reads the final page,
196
+ its fields, and the actions taken, without the run's own DONE, and answers
197
+ whether the goal is met. `probability` is its probability that it is. It is a
198
+ second model judgment, not ground truth: the default evaluator,
199
+ `'typesafe-ai/jev'`, is the same model that takes the decisions by default. Its cost is included in `cost` and its time
200
+ in `evaluationMs`, not `totalMs`. Calling `profiling()` sends the page text and typed
201
+ values to the evaluator's provider.
202
+
203
+ The first call runs the evaluation and later calls return the same object. If
204
+ the evaluation fails, `profiling()` still resolves with `cost` and `timing`, `accuracy`
205
+ is `undefined`, `error` holds the reason, and the next call tries again. Pass
206
+ `profiling({ signal })` to cancel it. `JSON.stringify(result)` gives the result
207
+ without the function.
208
+
209
+ When the agent itself stops the run, `error` is an `agent.BlockedError` with
210
+ `code: 'BLOCKED'`, `reason` and `trace` (`accuracy` is `undefined` when no page
211
+ was observed). Any other failure, such as a model reply that was refused or a
212
+ provider that did not answer, is in `error` as it was thrown. The reasons:
213
+
214
+ | Reason | Meaning |
215
+ | --- | --- |
216
+ | `captcha` | Visible text suggests human verification; no bypass attempted |
217
+ | `login_wall` | A visible password field needs manual login |
218
+ | `unsupported_surface` | Unsupported DOM surface, no document body, or popup |
219
+ | `step_budget` | Action or decision-request budget exhausted |
220
+ | `no_change` | Three consecutive successful actions left the same observation. An action that changed nothing is not offered again until the page changes |
221
+ | `stale_target` | Three consecutive decisions chose the same target and it failed its freshness check each time |
222
+ | `model_blocked` | The provider selected BLOCKED |
223
+
224
+ Malformed responses, network/provider failures and aborts resolve with
225
+ `status: 'error'` and the original error in `error`; they never become invented
226
+ successful results. Something thrown that is not an `Error` is wrapped in one,
227
+ with the original as its `cause`. Invalid configuration rejects before the run
228
+ starts.
229
+ Stale decisions are discarded and re-observed without spending an action, but
230
+ each decision still spends a request.
231
+
232
+ | Option | Default | Meaning |
233
+ | --- | --- | --- |
234
+ | `decisions` | `'typesafe-ai/jev'` | Decision model: an AI Gateway model id or an AI SDK decision model. `false` makes the text model take the decisions |
235
+ | `text` | `'openai/gpt-6-luna'` | Language model: an AI Gateway model id or an AI SDK language model. Writes the value for TYPE_TEXT, and makes the decisions when `decisions` is `false` |
236
+ | `evaluator` | `'typesafe-ai/jev'` | Decision model that `profiling()` asks whether the goal was met: an AI Gateway model id or an AI SDK decision model |
237
+ | `reasoning` | `'none'` | Reasoning level for the text model: `provider-default`, `none`, `minimal`, `low`, `medium`, `high` or `xhigh` |
238
+ | `maxSteps` | `60` | Maximum successful input actions |
239
+ | `maxDecisions` | `120` | Maximum decision requests, including discarded stale decisions |
240
+ | `waitMs` | `100` | Duration of WAIT, accepts zero |
241
+ | `timeout` | `25000` | Milliseconds per model request |
242
+ | `signal` | none | AbortSignal; checked before decisions and input |
243
+
244
+ The default text model needs paid AI Gateway credits; the gateway's free tier
245
+ refuses `openai/gpt-6-luna`. On the free tier pass a model it allows, for
246
+ example `text: 'zai/glm-5.3-flash'`, at 5 requests per minute.
247
+
248
+ Model requests are never retried: this keeps request accounting exact. A reply
249
+ that arrives after `timeout` is discarded even if the model ignored the abort. A final
250
+ DONE observation can succeed after the last permitted action; no extra input is
251
+ allowed. Concurrent agent calls on the same Page are rejected.
252
+
253
+ ### `await page.extract(rules | instruction, options?)`
254
+
255
+ Data comes out of the page through rules, in the format of the Microlink
256
+ [`data`](https://microlink.io/docs/api/parameters/data) parameter. A rule has:
257
+
258
+ | Property | Meaning |
259
+ | --- | --- |
260
+ | `selector` | CSS selector for one element |
261
+ | `selectorAll` | CSS selector for every matching element; the field is a list |
262
+ | `attr` | What to read: `text`, `html` (default), `val`, `href`, `src`, any attribute name, a list of those to try in order, or an object of nested rules relative to the matched element |
263
+ | `type` | `number`, `url`, `image`, `date` or `string`. Without a type, text that is exactly a number or `true`/`false` is converted and everything else stays a string |
264
+
265
+ `extract` can be called at three levels:
266
+
267
+ ```js
268
+ // 1. Rules: no model involved.
269
+ await page.extract({
270
+ stories: {
271
+ selectorAll: '.athing',
272
+ attr: {
273
+ title: { selector: '.titleline > a', attr: 'text' },
274
+ href: { selector: '.titleline > a', attr: 'href', type: 'url' }
275
+ }
276
+ }
277
+ })
278
+
279
+ // 2. Instruction: the model writes the rules, then they run.
280
+ await page.extract('get the stories with their title and link')
281
+
282
+ // 3. Instruction plus rules without selectors: you name the fields and types, the model fills in the selectors.
283
+ await page.extract('get the stories', {
284
+ stories: { attr: { title: { type: 'string' }, href: { type: 'url' } } }
285
+ })
286
+ ```
287
+
288
+ Options go last: `page.extract(rules, options?)` and
289
+ `page.extract(instruction, rules?, options?)`. Rules passed with an instruction
290
+ that already have a selector for every field run as they are, with no model
291
+ request. Options passed where the rules belong throw.
292
+
293
+ `extract` resolves to `{ status, data, profiling }`, or to
294
+ `{ status: 'error', error, profiling }` when it failed:
295
+
296
+ ```js
297
+ const { status, data, error, profiling } = await page.extract('get the stories', {
298
+ stories: { attr: { title: { type: 'string' }, href: { type: 'url' } } }
299
+ })
300
+ const { rules, cost, timing } = await profiling()
301
+ ```
302
+
303
+ Like `goal`, it does not throw for an extraction that fails, only when it is
304
+ called wrong: no page, an empty instruction, an invalid option next to an
305
+ instruction, or rules the built-in engine cannot run. Rules passed on their own
306
+ to a custom engine are not checked here; rules passed with an instruction always
307
+ are, since this package fills them in. `data` is whatever the rules engine
308
+ returned.
309
+ `rules` are the rules that ran, so the ones a model wrote can be stored and
310
+ passed to `page.extract(rules)` next time without a request. `cost` and
311
+ `timing` have the same fields as for `goal`; with rules of your own they report
312
+ no request and no cost. `profiling` makes no extra request and has no
313
+ `accuracy`.
314
+
315
+ The values always come from the DOM. A language model only ever writes
316
+ selectors, so it cannot invent a value: a wrong selector gives a missing field,
317
+ `null` inside a list item, or the wrong element's text.
318
+
319
+ How rules are written:
320
+
321
+ - One request to the language model in `text`; `reasoning`, `timeout` and
322
+ `signal` apply. It receives an outline of the page: tags, ids, classes,
323
+ `data-*` and a few other attributes, and short text, indented by nesting. Of
324
+ several similar siblings it sees the first three and a count of the rest.
325
+ The outline is taken once it has stopped changing, waiting up to 3 seconds,
326
+ and is capped at 60,000 characters.
327
+ - A rule the model writes without `attr` reads the text, not the HTML.
328
+ - Rules from the model may only contain `selector`, `selectorAll`, `attr` and
329
+ `type`. Anything else, including `evaluate`, is rejected, so a model never
330
+ supplies code.
331
+ - With rules of your own, the result has exactly your field names and nesting. A rule you
332
+ wrote with its own selector is kept untouched; otherwise the model supplies
333
+ the selector, your `type` and `attr` win over the model's, and a reply that
334
+ omits a field or nests where you did not is an error.
335
+ - The model sees only what the outline shows: text is clipped at 80
336
+ characters, and of several siblings with the same classes and the same
337
+ children it sees three.
338
+ - The rules are tried on the page with the built-in engine before they are
339
+ returned. An invalid selector is the error
340
+ `The model did not write usable extraction rules.` and rules that give no
341
+ value at all are `The rules the model wrote matched nothing on the page.`
342
+ - A selector can still be wrong for one field, or brittle. On Wallapop the model
343
+ used generated class names such as `retrieval-item-card-module_…__ckj4h`,
344
+ which will break when the site is rebuilt. An instruction like "top 5" is not
345
+ honored: rules select every match.
346
+
347
+ The engine that runs rules is replaceable:
348
+
349
+ ```js
350
+ agent(page, { extractor: (page, rules) => myEngine(page, rules) })
351
+ ```
352
+
353
+ When the page already has an `extract` method, as it does inside a Microlink
354
+ function, `agent(page)` keeps it and uses it as the engine. Otherwise a small
355
+ built-in engine evaluates the rules in the page (`agent.applyRules`). It follows
356
+ the Microlink rule format for the properties above, with these differences:
357
+
358
+ - No `evaluate` rules, and `type` is one name, not a name with options.
359
+ - It reads the live DOM, not the fetched HTML, and does not look inside shadow
360
+ roots.
361
+ - Selectors are plain CSS: jQuery extensions such as `:contains()` and `:eq()`
362
+ are invalid.
363
+ - A rule without a selector reads the rendered text or the HTML of the whole
364
+ page; there is no readability pass, Markdown or JSON mode.
365
+ - `number` reads the first number in the text and accepts `.` or `,` as decimal
366
+ or thousands separator, decided by position: `16.690 €` is 16690, `1.234,50`
367
+ is 1234.5, `4.5` is 4.5, `0.125` is 0.125. A single separator followed by
368
+ exactly three digits is read as thousands, so `3.142` is 3142.
369
+ - `date` returns an ISO string and needs a four-digit year in the text.
370
+ - Only `number`, `url`, `image`, `date` and `string` are known types.
371
+
372
+ ## Models
373
+
374
+ Both model calls go through the [AI SDK](https://ai-sdk.dev) (`ai`). A model id
375
+ string is resolved by [Vercel AI Gateway](https://vercel.com/docs/ai-gateway),
376
+ which reads `AI_GATEWAY_API_KEY`:
377
+
378
+ ```js
379
+ const page = agent(await browser.newPage(), {
380
+ decisions: 'typesafe-ai/jev',
381
+ text: 'openai/gpt-6-luna'
382
+ })
383
+ ```
384
+
385
+ Those two are the defaults, so `agent(page)` does the same.
386
+
387
+ Any AI SDK model instance works too, so a provider can be called directly with
388
+ its own key. This runs Jev on a TypeSafe key and keeps the gateway for text:
389
+
390
+ ```js
391
+ import { createTypeSafeAi } from '@ai-sdk/typesafe-ai'
392
+
393
+ const typesafe = createTypeSafeAi({ apiKey: process.env.TYPESAFE_API_KEY })
394
+
395
+ const page = agent(await browser.newPage(), {
396
+ decisions: typesafe.decisionModel('jev-latest')
397
+ })
398
+ ```
399
+
400
+ The decision model answers typed questions with `experimental_decide`, an AI SDK
401
+ API that may still change. One request carries
402
+ `state: { page, elements, recent_actions }` and one `choice` question for the
403
+ operation plus one per operation that has targets. The AI SDK rejects answers
404
+ that miss a question or pick an option that was not offered. This package then
405
+ requires, for the operation and for the target head that operation selects, a
406
+ `probabilities` object with exactly the offered IDs, finite values in [0,1], a
407
+ winning choice and a sum close to 1. The AI SDK already checks the sum against
408
+ the rounding the provider declares, and rejects any deviation when none is
409
+ declared. This package allows 0.02, widened only for the rounding of a provider
410
+ that declares two or more decimals. Unused speculative heads cannot
411
+ drive input and are not consumed. `confidence` in the trace is the probability
412
+ of the chosen option.
413
+
414
+ ### Without a decision model
415
+
416
+ With `decisions: false`, the language model in `text` receives the
417
+ same state and the same questions and returns `{ "operation", "target" }`. The
418
+ operation must be one of the offered operations and the target one of the
419
+ targets offered for it, or the run stops before any input. A target sent for an
420
+ operation that has none, such as WAIT or DONE, is ignored. There are no
421
+ probabilities on this path, so `confidence` and `probabilities` are absent from
422
+ the trace.
423
+
424
+ ```js
425
+ await page.goal(goal) // Jev decides
426
+ await page.goal(goal, { decisions: false }) // language model decides
427
+ ```
428
+
429
+ Every trace entry has `decisionMs`, the time its decision took (the model
430
+ request plus building and checking it), and `textMs` when a value was generated,
431
+ so the two setups can be compared on the same goal.
432
+
433
+ ### Benchmark
434
+
435
+ `scripts/benchmark.js` runs the goal live with Jev, records every decision
436
+ request, and replays the same requests to Jev and to each language model, so
437
+ every decider answers identical inputs. It needs `TYPESAFE_API_KEY` and
438
+ `AI_GATEWAY_API_KEY`:
439
+
440
+ ```sh
441
+ npm run benchmark -- --samples=10 --models=amazon/nova-micro,mistral/mistral-nemo
442
+ ```
443
+
444
+ `--perMinute` (default 5) paces each language model for the gateway's free
445
+ tier. Result of one run on the Wallapop goal, 10 requests recorded from two live
446
+ Jev runs that both ended on the results sorted by lowest price:
447
+
448
+ | Decider | Answered | Median | p90 | Same operation as Jev | Same operation and target |
449
+ | --- | --- | --- | --- | --- | --- |
450
+ | `jev-latest` (direct) | 10/10 | 254 ms | 284 ms | reference | reference |
451
+ | `amazon/nova-micro` | 8/10 | 694 ms | 959 ms | 2 | 1 |
452
+ | `alibaba/qwen3.7-flash` | 7/10 | 1421 ms | 1666 ms | 6 | 6 |
453
+ | `mistral/mistral-nemo` | 10/10 | 1336 ms | 3291 ms | 5 | 2 |
454
+
455
+ The unanswered requests were free-tier rate-limit errors. Agreement with Jev is
456
+ not correctness: it shows how often a model picked the step of a run that is
457
+ known to have worked. These are the small models the free tier allows, one run
458
+ each; larger models were not measured.
459
+
460
+ ### Compare models
461
+
462
+ `scripts/compare.js` runs one task several times with each model setup, each
463
+ run in a fresh browser context, and ranks the setups:
464
+
465
+ ```sh
466
+ DEBUG=browserless:agent:compare npm run compare -- \
467
+ --url=https://news.ycombinator.com \
468
+ --goal='Open the comments page of the first story on the front page.' \
469
+ --extract='get the story title and its points' \
470
+ --decisions=typesafe-ai/jev,none \
471
+ --text=openai/gpt-6-luna,zai/glm-5.3-flash \
472
+ --runs=5
473
+ ```
474
+
475
+ `--decisions` and `--text` take comma-separated model ids and every decision
476
+ model is paired with every text model; `none` makes the text model take the
477
+ decisions. `--reasoning` takes comma-separated reasoning levels for the text
478
+ model (default `none`) and each level is one more setup per pair of models, so
479
+ `--reasoning=none,low` measures whether reasoning helps, and a model that
480
+ refuses `none` can be run with `minimal`. An unknown level is refused before
481
+ any run. `--goal` can be repeated for goals
482
+ that run one after another.
483
+ `--extract` is optional. `--out=<file.jsonl>` chooses where the records go;
484
+ the default is a new file in the temporary directory, printed as `file`.
485
+
486
+ One JSON line is appended per run, and with `DEBUG` set the same record is
487
+ logged as it finishes:
488
+
489
+ | Field | Meaning |
490
+ | --- | --- |
491
+ | `decisions`, `text`, `reasoning`, `run` | The setup and the run number |
492
+ | `status`, `error`, `goalsDone` | `done`, or the `BlockedError` reason, or `error` for anything else, including a run that could not start; how many goals finished |
493
+ | `navigationError` | Present when loading `--url` reported an error, such as a timeout. The run continues on whatever loaded |
494
+ | `path` | The operations and targets that ran, such as `CLICK e17 > DONE`; goals are separated by a vertical bar |
495
+ | `decisionRequests`, `staleDecisions`, `actionsWithoutPageChange` | How much work the run took and how much of it was wasted |
496
+ | `minConfidence` | The lowest confidence of any decision, in the operation or in its target; absent when the text model decides |
497
+ | `passed`, `probability` | The evaluator's judgment from `profiling()`: `passed` needs every goal judged met, `probability` is the lowest. Absent when the run did not finish or a goal was not judged |
498
+ | `calls`, `inputTokens`, `outputTokens`, `usd` | Cost of the goals from `profiling()`, which includes the evaluator's request, plus the extraction request with `--extract`. Absent when any request did not report it |
499
+ | `totalMs`, `modelMs` | Time of the goals from `profiling()`, without the evaluator's request |
500
+ | `finalUrl` | Where the run ended |
501
+ | `extracted`, `extractError`, `extractMs`, `extractUsd`, `dataValues`, `dataHash` | With `--extract`, after every goal finished: whether it returned at least one value, how many non-empty values, and a hash of the data that ignores key order |
502
+
503
+ The printed `ranking` has one entry per setup, best first: by `doneRate`, then
504
+ `passRate` (the share of all runs that finished and were judged met;
505
+ `judgedRuns` says how many were judged at all), then `pathAgreement` (the share
506
+ of finished runs that took the most common path), then `p90TotalMs`, then
507
+ `medianUsd`. `p90TotalMs` is the time that nine in ten finished runs did not
508
+ exceed; with fewer than ten finished runs it is the slowest one. `dataAgreement` is the
509
+ share of extractions that returned the most common data. Medians and
510
+ `p90TotalMs` cover finished runs only, and each is absent when no run finished
511
+ or one of them did not report the number. `totalUsd` is what every run of the setup cost, failed
512
+ ones included, and the `totalUsd` next to `ranking` is the cost of the whole
513
+ comparison; both are absent when any run did not report its cost. `reportedUsd` is the sum of the costs that
514
+ were reported and `runsWithoutCost` counts the runs that reported none, such
515
+ as a run that ended on an error. A run that fails is recorded and the remaining
516
+ runs go on.
517
+
518
+ Results of past comparisons are collected in [scripts/README.md](scripts/README.md).
519
+
520
+ What it does not measure: whether a finished run is correct. `passed` is a
521
+ second model's opinion, and on the same
522
+ path and final page it has answered both yes and no.
523
+
524
+ ### Text values
525
+
526
+ The text model is called with `generateText` and a JSON object output only for
527
+ TYPE_TEXT. It must return exactly `{ "text": "..." }` with a nonempty string of
528
+ at most 2000 characters. Null, extra keys, arrays, invalid JSON and a reply
529
+ with no output fail closed.
530
+ Reasoning models can spend the whole 1024-token reply budget thinking and
531
+ return no usable value, so reasoning is off unless `reasoning` says otherwise.
532
+ No personal information or missing value is guessed by the executor. Page text,
533
+ goals, fields and recent actions leave the browser for these providers. Use
534
+ only pages you may disclose to them.
535
+
536
+ ## How it works
537
+
538
+ 1. `page.evaluate` collects visible controls, viewport text, form state and live
539
+ DOM refs. Persistent DOM node IDs, snapshot action IDs (`e1`, `e2`) and model
540
+ indices are separate. At most 250 element actions and 6000 text characters
541
+ are offered per snapshot.
542
+ 2. One request proposes an operation plus speculative CLICK, TYPE_TEXT, SUBMIT and
543
+ SELECT target heads. The operation chooses which head to consume. The model
544
+ never supplies a selector, JavaScript or screen coordinates.
545
+ 3. Code resolves the selected node via a retained ElementHandle. Immediately
546
+ before input, it checks node identity, target attributes/value, nearby text,
547
+ same-form values, document identity, current visibility and hit-testing.
548
+ Unrelated changes outside the target context/form are tolerated. Moving
549
+ targets are re-resolved, not clicked at old coordinates. Text entry checks
550
+ again after focus and before inserting text, and requires the target to
551
+ still hold keyboard focus. A discarded stale decision is reported to the
552
+ next decision request as `stale`, so it is not mistaken for executed input.
553
+ 4. Puppeteer clicks/selects or inserts generated text, or scrolls/waits, then
554
+ collects a new observation. Three unchanged actions stop the loop.
555
+
556
+ There is no confidence threshold in v0. A finite action space prevents arbitrary
557
+ model-emitted code, but does NOT ensure that the chosen action is correct or
558
+ harmless. Do not use this autonomous prototype for purchases, sending messages,
559
+ deleting data or other irreversible actions without your own approval layer.
560
+ It is not a permission policy or a CAPTCHA-solving system.
561
+
562
+ ## Limits
563
+
564
+ Single top-level page only. Open shadow roots are traversed: their controls and
565
+ text are observed, slotted content names the control it is slotted into, and
566
+ hit-testing and focus checks follow the composed tree. Visible iframes, canvas,
567
+ file inputs and password inputs, in the light DOM or inside open shadow roots,
568
+ conservatively stop the entire run, even if they are unrelated to the goal.
569
+ The same-form and nearby-text freshness guards do not see changes inside shadow
570
+ roots, a `:host::after` overlay is not detected as a cover, and actions are
571
+ listed root by root, so the 250-action cap can drop shadow controls on very
572
+ large pages. Closed shadow roots cannot be detected from JS: their content is
573
+ invisible to the agent and never blocks or receives input. Popup tabs are reported, never
574
+ followed or automatically closed. A popup may already have opened before it can
575
+ be detected.
576
+
577
+ Login detection is limited to visible password fields, while CAPTCHA detection
578
+ is a text heuristic. Other login/challenge walls may end as model_blocked or
579
+ no_change. False positives are possible. There is no image/vision fallback,
580
+ multi-tab planner, hidden/offscreen target support, native browser dialog handler
581
+ or independent goal verifier. Control truncation, unsupported custom roles,
582
+ long context truncated at 6000 characters and asynchronous browser input races
583
+ remain v0 limits. Guards reduce stale-target risk, not all browser races.
584
+
585
+ ## Verification and publishing
586
+
587
+ `npm test` runs AVA with mocked model responses, a fake Puppeteer adapter and
588
+ jsdom DOM fixtures. No API key is needed for these tests. jsdom provides
589
+ synthetic visibility/layout, so `test/chrome.js` also runs text entry against
590
+ the Chrome that Puppeteer downloads: value replacement in inputs and editing
591
+ hosts, and the keyboard-focus guard. Clicks, selects, scrolling and the full
592
+ loop are not exercised in real Chrome. No paid-provider benchmark is claimed.
593
+
594
+ This directory is picked up by the existing `packages/*` workspace and release
595
+ globs. No CI workflow or root config is changed. `@browserless/cli` gains the
596
+ `exec` command and a real test script, so CI now tests that package too. The
597
+ owner must publish
598
+ `@browserless/agent` to npm with public access after review. No publishing or
599
+ merging is performed by this patch.
600
+
601
+ ## Attribution
602
+
603
+ The collector and finite-choice/speculative-head pattern are adapted from
604
+ [browser-use/jev-ultrafast](https://github.com/browser-use/jev-ultrafast), commit
605
+ `1231850a0bf1a0c0341fe408ef1668dbbfdfac46`, MIT, copyright 2026 Browser Use.
606
+ The upstream license is reproduced at the end of this section.
607
+ The Node executor/client are ports, not a Python wrapper: they use the caller's
608
+ Puppeteer Page without Browser Harness, a separate Python process or another
609
+ browser. No model weights or inference service are bundled.
610
+
611
+ <details>
612
+ <summary>browser-use/jev-ultrafast license</summary>
613
+
614
+ ```
615
+ MIT License
616
+
617
+ Copyright (c) 2026 Browser Use
618
+
619
+ Permission is hereby granted, free of charge, to any person obtaining a copy
620
+ of this software and associated documentation files (the "Software"), to deal
621
+ in the Software without restriction, including without limitation the rights
622
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
623
+ copies of the Software, and to permit persons to whom the Software is
624
+ furnished to do so, subject to the following conditions:
625
+
626
+ The above copyright notice and this permission notice shall be included in all
627
+ copies or substantial portions of the Software.
628
+
629
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
630
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
631
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
632
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
633
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
634
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
635
+ SOFTWARE.
636
+ ```
637
+
638
+ </details>