il-e2e-agent 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +333 -0
- package/bin/il-e2e-agent.mjs +35 -0
- package/bin/prune-bundled-claude.mjs +28 -0
- package/dist/agent.d.ts +19 -0
- package/dist/agent.js +332 -0
- package/dist/claude-binary.d.ts +4 -0
- package/dist/claude-binary.js +32 -0
- package/dist/config.d.ts +25 -0
- package/dist/config.js +16 -0
- package/dist/index.d.ts +6 -0
- package/dist/index.js +5 -0
- package/dist/network.d.ts +19 -0
- package/dist/network.js +52 -0
- package/dist/test.d.ts +7 -0
- package/dist/test.js +11 -0
- package/package.json +45 -0
package/README.md
ADDED
|
@@ -0,0 +1,333 @@
|
|
|
1
|
+
# il-e2e-agent
|
|
2
|
+
|
|
3
|
+
End-to-end browser tests written in plain English, executed by [tester-army/e2e](https://github.com/tester-army/e2e)
|
|
4
|
+
and driven by the **Claude Code installed on your machine**, signed in with your own Claude account.
|
|
5
|
+
No Anthropic API key is required.
|
|
6
|
+
|
|
7
|
+
```ts
|
|
8
|
+
import { test, expect } from 'il-e2e-agent';
|
|
9
|
+
|
|
10
|
+
test('a visitor reaches checkout', async ({ app, agent, browser }) => {
|
|
11
|
+
await app.open('/pricing');
|
|
12
|
+
await agent.act('Choose the monthly plan and continue to checkout.'); // Claude operates the page
|
|
13
|
+
await expect(browser).toHaveURL(/\/checkout/); // deterministic check, no model
|
|
14
|
+
});
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
## Contents
|
|
18
|
+
|
|
19
|
+
- [How it works](#how-it-works)
|
|
20
|
+
- [Requirements](#requirements)
|
|
21
|
+
- [Installation](#installation)
|
|
22
|
+
- [Framework setup](#framework-setup): [Next.js](#nextjs) · [Nuxt](#nuxt) · [React (Vite)](#react-vite) · [Create React App](#create-react-app)
|
|
23
|
+
- [Configuration](#configuration)
|
|
24
|
+
- [Writing tests](#writing-tests)
|
|
25
|
+
- [Running tests](#running-tests)
|
|
26
|
+
- [Replay cache](#replay-cache)
|
|
27
|
+
- [Keeping it out of production builds](#keeping-it-out-of-production-builds)
|
|
28
|
+
- [Troubleshooting](#troubleshooting)
|
|
29
|
+
- [Limitations](#limitations)
|
|
30
|
+
|
|
31
|
+
## How it works
|
|
32
|
+
|
|
33
|
+
`e2e` runs your tests in a real browser (Playwright). Steps written as `agent.act('…')` are handed to
|
|
34
|
+
an agent; this package supplies that agent.
|
|
35
|
+
|
|
36
|
+
For each agent step, il-e2e-agent starts your installed Claude Code and gives it a small set of tools:
|
|
37
|
+
read the screen, tap, type, select, check, scroll, and run several of those in one batch. Claude reads
|
|
38
|
+
the page as an accessibility listing, performs the step, and reports whether it succeeded. Every action
|
|
39
|
+
goes through e2e's own action layer, so it is recorded.
|
|
40
|
+
|
|
41
|
+
Once a step has passed and a later check has confirmed the result, e2e saves the recorded actions.
|
|
42
|
+
**Later runs replay them directly, without calling Claude**, until the page changes enough that the
|
|
43
|
+
replay no longer matches. Claude then takes over from the current screen and the new path is recorded.
|
|
44
|
+
|
|
45
|
+
Each Claude Code session is isolated from your personal setup: no `CLAUDE.md`, hooks, built-in tools,
|
|
46
|
+
personal MCP servers or connected apps are loaded, and no session history is written.
|
|
47
|
+
|
|
48
|
+
## Requirements
|
|
49
|
+
|
|
50
|
+
| Requirement | Notes |
|
|
51
|
+
| --- | --- |
|
|
52
|
+
| Claude Code, installed and signed in | `curl -fsSL https://claude.ai/install.sh \| bash`, then `claude auth login`. Found on `PATH` or in `~/.local/bin`; set `IL_E2E_CLAUDE_PATH` to use a different binary. |
|
|
53
|
+
| Node.js 20.19 or later | |
|
|
54
|
+
| Playwright's Chromium | Once per machine: `npx playwright install chromium` |
|
|
55
|
+
| `ANTHROPIC_API_KEY` **not** set | If it is set, Claude Code may bill that key instead of your plan. The CLI prints a warning. |
|
|
56
|
+
|
|
57
|
+
Agent steps count against **your own** Claude plan's usage limits.
|
|
58
|
+
|
|
59
|
+
## Installation
|
|
60
|
+
|
|
61
|
+
Install as a dev dependency of your app:
|
|
62
|
+
|
|
63
|
+
```bash
|
|
64
|
+
npm install --save-dev github:vishalmishraa22/il-e2e-agent
|
|
65
|
+
npx playwright install chromium
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Then create this layout in your project:
|
|
69
|
+
|
|
70
|
+
```
|
|
71
|
+
your-app/
|
|
72
|
+
├── package.json add a "test:e2e" script (below)
|
|
73
|
+
└── e2e/
|
|
74
|
+
├── il-e2e-agent.config.ts app URL and network policy
|
|
75
|
+
├── tests/
|
|
76
|
+
│ └── *.e2e.ts your tests
|
|
77
|
+
├── support/ optional shared helpers
|
|
78
|
+
└── .gitignore
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
`e2e/.gitignore`. Run output stays local; the replay cache is committed so teammates replay recorded steps:
|
|
82
|
+
|
|
83
|
+
```gitignore
|
|
84
|
+
.e2e/*
|
|
85
|
+
!.e2e/cache/
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
`package.json` scripts:
|
|
89
|
+
|
|
90
|
+
```json
|
|
91
|
+
{
|
|
92
|
+
"scripts": {
|
|
93
|
+
"test:e2e": "il-e2e-agent run --config e2e/il-e2e-agent.config.ts",
|
|
94
|
+
"test:e2e:fresh": "il-e2e-agent run --config e2e/il-e2e-agent.config.ts --no-cache"
|
|
95
|
+
}
|
|
96
|
+
}
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
## Framework setup
|
|
100
|
+
|
|
101
|
+
Tests run against your app's **dev server**. Pick a port for tests (3100 below) so a test run never
|
|
102
|
+
collides with the dev server you already use, and start it before running the tests. Alternatively,
|
|
103
|
+
let the runner start it with `app.command` (see [Configuration](#configuration)).
|
|
104
|
+
|
|
105
|
+
### Next.js
|
|
106
|
+
|
|
107
|
+
```json
|
|
108
|
+
{ "scripts": { "dev:e2e": "next dev -p 3100" } }
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
`next build` type-checks every `.ts` file that `tsconfig.json` includes. Exclude the test folder so
|
|
112
|
+
production builds never depend on test code:
|
|
113
|
+
|
|
114
|
+
```jsonc
|
|
115
|
+
// tsconfig.json
|
|
116
|
+
{ "exclude": ["node_modules", "e2e"] }
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
### Nuxt
|
|
120
|
+
|
|
121
|
+
```json
|
|
122
|
+
{ "scripts": { "dev:e2e": "nuxt dev --port 3100" } }
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Add the test folder to `.nuxtignore` so the dev server's file watcher skips it (test runs write files
|
|
126
|
+
under `e2e/.e2e/`):
|
|
127
|
+
|
|
128
|
+
```gitignore
|
|
129
|
+
e2e/**
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
If CI runs `nuxi typecheck`, also exclude `e2e/` there through `typescript.tsConfig.exclude` in
|
|
133
|
+
`nuxt.config.ts`.
|
|
134
|
+
|
|
135
|
+
### React (Vite)
|
|
136
|
+
|
|
137
|
+
```json
|
|
138
|
+
{ "scripts": { "dev:e2e": "vite --port 3100 --strictPort" } }
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
No further changes are needed with the default Vite templates: `tsconfig.app.json` only includes
|
|
142
|
+
`src/`, and Vite's watcher only reacts to files your app imports.
|
|
143
|
+
|
|
144
|
+
### Create React App
|
|
145
|
+
|
|
146
|
+
```json
|
|
147
|
+
{ "scripts": { "dev:e2e": "PORT=3100 BROWSER=none react-scripts start" } }
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
### Server-side tracking
|
|
151
|
+
|
|
152
|
+
The network policy below controls what the **test browser** can reach. Calls your **server** makes
|
|
153
|
+
(for example a Next.js route handler or Nuxt server route forwarding to an analytics or CRM API)
|
|
154
|
+
are outside its reach. If your app does this in development, turn it off in the `dev:e2e` script
|
|
155
|
+
with your app's own environment flags, and block the forwarding routes with `blockPaths`.
|
|
156
|
+
|
|
157
|
+
## Configuration
|
|
158
|
+
|
|
159
|
+
`e2e/il-e2e-agent.config.ts`:
|
|
160
|
+
|
|
161
|
+
```ts
|
|
162
|
+
import { defineConfig } from 'il-e2e-agent';
|
|
163
|
+
|
|
164
|
+
export default defineConfig({
|
|
165
|
+
app: { url: 'http://localhost:3100' },
|
|
166
|
+
timeout: 600_000,
|
|
167
|
+
actionTimeout: 8_000,
|
|
168
|
+
maxSteps: 40,
|
|
169
|
+
network: {
|
|
170
|
+
allow: ['*.stripe.com', '*.stripe.network', 'fonts.googleapis.com', 'fonts.gstatic.com'],
|
|
171
|
+
readOnly: ['api-staging.example.com'],
|
|
172
|
+
blockPaths: ['/api/analytics'],
|
|
173
|
+
log: Boolean(process.env.E2E_NETWORK_LOG),
|
|
174
|
+
},
|
|
175
|
+
});
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
`defineConfig` accepts every [e2e config option](https://e2e.tester.army/docs/reference/config)
|
|
179
|
+
except `targets` and `agents`, which it builds for you, plus:
|
|
180
|
+
|
|
181
|
+
| Option | Default | Description |
|
|
182
|
+
| --- | --- | --- |
|
|
183
|
+
| `app` | required | `url` of the app under test. Optionally `command` to let the runner start it, e.g. `{ executable: 'npm', args: ['run', 'dev:e2e'], cwd: '..', reuseExisting: true }`. |
|
|
184
|
+
| `model` | `'sonnet'` | Claude model for agent steps: `'sonnet'`, `'haiku'` or `'opus'`. |
|
|
185
|
+
| `effort` | `'low'` | Reasoning effort per agent step. Higher values think longer on every turn. |
|
|
186
|
+
| `maxSteps` | `25` | Actions a single `agent.act` may take. Raise it for long multi-screen steps. |
|
|
187
|
+
| `browser` | e2e defaults | Playwright web-engine options, e.g. `{ viewport: { width: 1280, height: 900 } }`. |
|
|
188
|
+
| `network` | localhost only | Which hosts the test browser may reach. `false` disables the guard. |
|
|
189
|
+
|
|
190
|
+
### Network policy
|
|
191
|
+
|
|
192
|
+
Every request the test browser makes is checked, and anything not explicitly allowed is aborted.
|
|
193
|
+
Development builds often send real analytics, advertising and CRM events. The policy keeps test runs
|
|
194
|
+
out of that data, and keeps tests from writing to shared backends.
|
|
195
|
+
|
|
196
|
+
| Field | Effect |
|
|
197
|
+
| --- | --- |
|
|
198
|
+
| `allow` | Hosts reachable with any method. `*.example.com` matches the domain and its subdomains. `localhost` and `127.0.0.1` are always allowed. |
|
|
199
|
+
| `readOnly` | Hosts reachable with `GET`, `HEAD` and `OPTIONS` only, e.g. a staging API your pages read from. |
|
|
200
|
+
| `blockPaths` | Path prefixes aborted on every host, including your own app's routes. |
|
|
201
|
+
| `log` | Prints each newly blocked host and path once. |
|
|
202
|
+
|
|
203
|
+
To build the list for a new project, start with only localhost allowed and run with
|
|
204
|
+
`E2E_NETWORK_LOG=1`. Each `[il-e2e-agent] blocked …` line names a host the page tried to reach:
|
|
205
|
+
- **Leave blocked:** trackers, pixels, tag managers and CRMs.
|
|
206
|
+
- **Add to `allow`:** what the page needs to work, such as payment widgets, fonts and image CDNs.
|
|
207
|
+
- **Add to `readOnly`:** APIs it only reads from.
|
|
208
|
+
|
|
209
|
+
## Writing tests
|
|
210
|
+
|
|
211
|
+
Tests live in `e2e/tests/*.e2e.ts`:
|
|
212
|
+
|
|
213
|
+
```ts
|
|
214
|
+
import { test, expect } from 'il-e2e-agent';
|
|
215
|
+
|
|
216
|
+
test('a returning visitor keeps their campaign parameters', async ({ app, agent, browser, screen }) => {
|
|
217
|
+
await app.open('/landing?utm_source=newsletter');
|
|
218
|
+
|
|
219
|
+
await agent.act('Start the sign-up flow using the main call-to-action.');
|
|
220
|
+
await expect(browser).toHaveURL(/utm_source=newsletter/);
|
|
221
|
+
|
|
222
|
+
await agent.act('Enter the email {email} and continue.', { params: { email: 'qa@example.com' } });
|
|
223
|
+
await expect(screen.getByText('Check your inbox', { exact: false })).toBeVisible();
|
|
224
|
+
});
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
Guidelines:
|
|
228
|
+
|
|
229
|
+
- **Follow every `agent.act` with an `expect`.** The check confirms the step did the right thing, and
|
|
230
|
+
only steps confirmed by a later check are recorded for replay.
|
|
231
|
+
- **One goal per `agent.act`.** Short, specific instructions are faster, cheaper and more reliable than
|
|
232
|
+
one instruction covering a whole flow.
|
|
233
|
+
- **Check with `expect` wherever possible.** `expect`, `screen.*` and `browser.*` never call a model.
|
|
234
|
+
`agent.assert`, `agent.waitFor` and `agent.extract` call Claude on every run, even when the steps
|
|
235
|
+
before them are replayed.
|
|
236
|
+
- **Pass data as `params`** and reference it as `{name}` in the instruction. Wrap values that change on
|
|
237
|
+
every run (timestamps, generated emails) in `unique()`, or the step is never replayed.
|
|
238
|
+
- **Prefer `exact: false` for copy checks** when text may come from a CMS or change case.
|
|
239
|
+
- **Never instruct the agent to submit real payments** or other irreversible actions against shared
|
|
240
|
+
environments.
|
|
241
|
+
|
|
242
|
+
`test`, `expect`, `unique`, `credentials` and `secrets` are re-exported from e2e. The full test API is
|
|
243
|
+
documented at [e2e.tester.army/docs](https://e2e.tester.army/docs).
|
|
244
|
+
|
|
245
|
+
## Running tests
|
|
246
|
+
|
|
247
|
+
```bash
|
|
248
|
+
npm run dev:e2e # terminal 1: the app on the test port
|
|
249
|
+
|
|
250
|
+
npm run test:e2e # terminal 2: all tests
|
|
251
|
+
npm run test:e2e -- e2e/tests/checkout.e2e.ts # one file
|
|
252
|
+
npm run test:e2e -- --headed # watch the browser
|
|
253
|
+
npm run test:e2e -- --reporter list,markdown # writes e2e/.e2e/summary.md and failure pages
|
|
254
|
+
npm run test:e2e:fresh # ignore recordings and run every step live
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
Failure details, screenshots and traces are written to `e2e/.e2e/`. Start with
|
|
258
|
+
`e2e/.e2e/failures/` when a test fails.
|
|
259
|
+
|
|
260
|
+
The CLI sets `E2E_TELEMETRY_DISABLED=1` unless you set it yourself.
|
|
261
|
+
|
|
262
|
+
## Replay cache
|
|
263
|
+
|
|
264
|
+
- **First run of a step:** Claude performs it live, and the actions are recorded in `e2e/.e2e/cache/`.
|
|
265
|
+
- **Later runs:** the recording is replayed with no model call.
|
|
266
|
+
- **The page changes:** the replay hands over to Claude where it stopped matching, and the step is
|
|
267
|
+
recorded again.
|
|
268
|
+
- **Changing a test's name, an instruction or its params** starts that step from scratch.
|
|
269
|
+
|
|
270
|
+
**Commit `e2e/.e2e/cache/`** so everyone on the team replays the same recordings instead of running
|
|
271
|
+
every step live on their own plan. Review cache changes in pull requests like any other test data.
|
|
272
|
+
|
|
273
|
+
## Keeping it out of production builds
|
|
274
|
+
|
|
275
|
+
This package is a development tool. Deployment and CI builds that install dev dependencies (for example
|
|
276
|
+
because the framework's build tooling lives there) should skip it. Remove it from `package.json` in the
|
|
277
|
+
build environment before installing:
|
|
278
|
+
|
|
279
|
+
```bash
|
|
280
|
+
npm pkg delete devDependencies.il-e2e-agent && npm install --include=dev
|
|
281
|
+
```
|
|
282
|
+
|
|
283
|
+
This changes the build's working copy only. Verified with both `npm install` and `npm ci` (npm 11).
|
|
284
|
+
|
|
285
|
+
AWS Amplify (`amplify.yml`):
|
|
286
|
+
|
|
287
|
+
```yaml
|
|
288
|
+
preBuild:
|
|
289
|
+
commands:
|
|
290
|
+
- npm pkg delete devDependencies.il-e2e-agent
|
|
291
|
+
- npm install --include=dev
|
|
292
|
+
```
|
|
293
|
+
|
|
294
|
+
Vercel (`vercel.json`):
|
|
295
|
+
|
|
296
|
+
```json
|
|
297
|
+
{ "installCommand": "npm pkg delete devDependencies.il-e2e-agent && npm install" }
|
|
298
|
+
```
|
|
299
|
+
|
|
300
|
+
GitHub Actions:
|
|
301
|
+
|
|
302
|
+
```yaml
|
|
303
|
+
- run: npm pkg delete devDependencies.il-e2e-agent
|
|
304
|
+
- run: npm ci
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
## Troubleshooting
|
|
308
|
+
|
|
309
|
+
| Symptom | Cause and fix |
|
|
310
|
+
| --- | --- |
|
|
311
|
+
| `Claude Code was not found` | Install Claude Code and run `claude auth login`, or set `IL_E2E_CLAUDE_PATH`. |
|
|
312
|
+
| `il-e2e-agent.config.ts not found` | Pass `--config e2e/il-e2e-agent.config.ts`, as the scripts above do, or run from the folder that holds it. |
|
|
313
|
+
| `APP_UNREACHABLE` | The app isn't running on the configured URL. Start `npm run dev:e2e` first, or configure `app.command`. |
|
|
314
|
+
| A page renders broken or incomplete | A resource it needs is blocked. Run with `E2E_NETWORK_LOG=1` and add the host to `network.allow`. |
|
|
315
|
+
| A step takes long or exhausts its steps | Split the instruction into smaller `agent.act` calls, or raise `maxSteps`. |
|
|
316
|
+
| A recorded step keeps re-running live | Its instruction, params or test name changed, or a value differs on every run. Use `unique()` for those. |
|
|
317
|
+
| The dev server misbehaves while tests run | Exclude `e2e/` from its file watcher (`.nuxtignore` for Nuxt). |
|
|
318
|
+
|
|
319
|
+
## Limitations
|
|
320
|
+
|
|
321
|
+
- **Local use.** Agent steps run on the signed-in developer's own Claude plan, for their own use. Don't
|
|
322
|
+
share sign-ins or run them on shared or automated infrastructure. Replays of committed recordings need
|
|
323
|
+
no model.
|
|
324
|
+
- **Text, not pixels.** The agent reads the page's accessibility tree. Layout and visual correctness
|
|
325
|
+
need separate visual testing.
|
|
326
|
+
- **`agent.assert`, `waitFor` and `extract` always run live**, so they use your plan on every run.
|
|
327
|
+
- **Pre-1.0 foundations.** e2e is still before 1.0, and this package pins exact versions of it.
|
|
328
|
+
|
|
329
|
+
## Disclaimer
|
|
330
|
+
|
|
331
|
+
il-e2e-agent is an independent project. It is not affiliated with, endorsed by or sponsored by
|
|
332
|
+
Anthropic. Claude and Claude Code are trademarks of Anthropic, PBC. Your use of Claude Code is governed
|
|
333
|
+
by Anthropic's terms for your plan.
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
import { spawn } from 'node:child_process';
|
|
3
|
+
import { existsSync } from 'node:fs';
|
|
4
|
+
import { dirname, join, resolve } from 'node:path';
|
|
5
|
+
import { fileURLToPath } from 'node:url';
|
|
6
|
+
|
|
7
|
+
const CONFIG_FILE = 'il-e2e-agent.config.ts';
|
|
8
|
+
const CONFIG_COMMANDS = new Set(['run', 'cache', 'mcp', 'explore']);
|
|
9
|
+
|
|
10
|
+
const argv = process.argv.slice(2);
|
|
11
|
+
const command = argv[0] && !argv[0].startsWith('-') && !argv[0].includes('.') ? argv.shift() : 'run';
|
|
12
|
+
|
|
13
|
+
if (process.env.ANTHROPIC_API_KEY) {
|
|
14
|
+
console.warn(
|
|
15
|
+
'[il-e2e-agent] ANTHROPIC_API_KEY is set, so Claude Code may bill that key instead of your Claude subscription. ' +
|
|
16
|
+
'Unset it to run on your own login.',
|
|
17
|
+
);
|
|
18
|
+
}
|
|
19
|
+
|
|
20
|
+
const args = [command, ...argv];
|
|
21
|
+
if (CONFIG_COMMANDS.has(command) && !argv.includes('--config')) {
|
|
22
|
+
const config = resolve(CONFIG_FILE);
|
|
23
|
+
if (!existsSync(config)) {
|
|
24
|
+
console.error(`[il-e2e-agent] ${CONFIG_FILE} not found in ${process.cwd()}. Run il-e2e-agent from the folder that holds it.`);
|
|
25
|
+
process.exit(2);
|
|
26
|
+
}
|
|
27
|
+
args.push('--config', config);
|
|
28
|
+
}
|
|
29
|
+
|
|
30
|
+
const e2eCli = join(dirname(fileURLToPath(import.meta.resolve('e2e'))), 'cli/bin.js');
|
|
31
|
+
const child = spawn(process.execPath, [e2eCli, ...args], {
|
|
32
|
+
stdio: 'inherit',
|
|
33
|
+
env: { E2E_TELEMETRY_DISABLED: '1', ...process.env },
|
|
34
|
+
});
|
|
35
|
+
child.on('exit', (code, signal) => (signal ? process.kill(process.pid, signal) : process.exit(code ?? 1)));
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// postinstall: the Claude Agent SDK installs its own ~215 MB Claude Code binary as an optional
|
|
3
|
+
// per-platform package. il-e2e-agent always runs the developer's installed Claude Code (it refuses to
|
|
4
|
+
// start without one), so the bundled copy is never used and is deleted. npm cannot skip just that
|
|
5
|
+
// one optional package (--omit=optional would also drop esbuild's platform binary, which e2e needs).
|
|
6
|
+
import { existsSync, statSync, unlinkSync } from 'node:fs';
|
|
7
|
+
import { dirname, join } from 'node:path';
|
|
8
|
+
import { fileURLToPath } from 'node:url';
|
|
9
|
+
|
|
10
|
+
let sdkDir;
|
|
11
|
+
try {
|
|
12
|
+
sdkDir = dirname(fileURLToPath(import.meta.resolve('@anthropic-ai/claude-agent-sdk')));
|
|
13
|
+
} catch {
|
|
14
|
+
process.exit(0);
|
|
15
|
+
}
|
|
16
|
+
|
|
17
|
+
const platformPackage = `claude-agent-sdk-${process.platform}-${process.arch}`;
|
|
18
|
+
const candidates = [
|
|
19
|
+
join(sdkDir, '..', platformPackage, 'claude'),
|
|
20
|
+
join(sdkDir, '..', `${platformPackage}-musl`, 'claude'),
|
|
21
|
+
join(sdkDir, 'node_modules', '@anthropic-ai', platformPackage, 'claude'),
|
|
22
|
+
];
|
|
23
|
+
|
|
24
|
+
for (const bundled of candidates.filter(existsSync)) {
|
|
25
|
+
const megabytes = Math.round(statSync(bundled).size / 1024 / 1024);
|
|
26
|
+
unlinkSync(bundled);
|
|
27
|
+
console.log(`[il-e2e-agent] Removed the Agent SDK's bundled Claude Code (${megabytes} MB); your installed one is used.`);
|
|
28
|
+
}
|
package/dist/agent.d.ts
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
import { type EffortLevel } from 'ai-sdk-provider-claude-code';
|
|
2
|
+
import type { StepExecutor } from 'e2e';
|
|
3
|
+
export interface ClaudeCodeAgentOptions {
|
|
4
|
+
/** Claude Code model alias or id: 'sonnet', 'haiku', 'opus'. */
|
|
5
|
+
readonly model?: string;
|
|
6
|
+
/** Claude Code turns allowed per agent step before it is reported as blocked. */
|
|
7
|
+
readonly maxTurns?: number;
|
|
8
|
+
/**
|
|
9
|
+
* Reasoning effort for agent steps. Defaults to 'low': driving a UI step by step needs little
|
|
10
|
+
* deliberation, and Claude Code's own default ('high') thinks on every turn, which is slow.
|
|
11
|
+
*/
|
|
12
|
+
readonly effort?: EffortLevel;
|
|
13
|
+
/** Claude Code binary to run. Defaults to the one installed on this machine (see claude-binary.ts). */
|
|
14
|
+
readonly claudeExecutable?: string;
|
|
15
|
+
}
|
|
16
|
+
export declare function claudeCodeAgent(options?: ClaudeCodeAgentOptions): {
|
|
17
|
+
executor: StepExecutor;
|
|
18
|
+
model: import("@ai-sdk/provider").LanguageModelV4;
|
|
19
|
+
};
|
package/dist/agent.js
ADDED
|
@@ -0,0 +1,332 @@
|
|
|
1
|
+
import { generateText } from 'ai';
|
|
2
|
+
import { z } from 'zod';
|
|
3
|
+
import { claudeCode, createAiSdkMcpServer, } from 'ai-sdk-provider-claude-code';
|
|
4
|
+
import { isAgentError } from 'e2e/agent';
|
|
5
|
+
import { requireInstalledClaude } from "./claude-binary.js";
|
|
6
|
+
// The provider warns on every step that maxTurns above 20 is high; long steps are expected here.
|
|
7
|
+
globalThis.AI_SDK_LOG_WARNINGS ??= false;
|
|
8
|
+
const SERVER = 'e2e';
|
|
9
|
+
const RUNTIME_STOPS = new Set(['STEP_BUDGET_EXHAUSTED', 'STEP_TIMEOUT', 'CANCELLED']);
|
|
10
|
+
const FAILED_CODES = ['ASSERTION_FAILED', 'ASSERTION_INCONCLUSIVE', 'ACTION_FAILED', 'LOCATOR_NOT_FOUND'];
|
|
11
|
+
const BLOCKED_CODES = [
|
|
12
|
+
'AUTOMATION_UNSUPPORTED',
|
|
13
|
+
'ENVIRONMENT_UNAVAILABLE',
|
|
14
|
+
'SEED_DATA_MISSING',
|
|
15
|
+
'TEST_SETUP_FAILED',
|
|
16
|
+
'AUTH_CREDENTIAL_INVALID',
|
|
17
|
+
'AUTH_CREDENTIAL_UNAVAILABLE',
|
|
18
|
+
];
|
|
19
|
+
// Above this share of changed lines a diff saves little, so the full listing is sent instead.
|
|
20
|
+
const FULL_LISTING_RATIO = 0.6;
|
|
21
|
+
// Runs Claude Code with nothing from ~/.claude (no CLAUDE.md, hooks, MCP servers or
|
|
22
|
+
// built-in tools) and without writing sessions to disk; it signs in with the local login.
|
|
23
|
+
// strictMcpConfig matters most: without it the user's own MCP servers and claude.ai
|
|
24
|
+
// connectors attach after the first tool call and add ~40k input tokens to every turn.
|
|
25
|
+
const ISOLATED = {
|
|
26
|
+
tools: [],
|
|
27
|
+
settingSources: [],
|
|
28
|
+
strictMcpConfig: true,
|
|
29
|
+
persistSession: false,
|
|
30
|
+
verbatimPrompts: true,
|
|
31
|
+
permissionPrompts: 'none',
|
|
32
|
+
};
|
|
33
|
+
const SYSTEM_PROMPT = `You are a QA agent testing a web app through tools. You see the app only as a listing with one element per line: "#id role "name" ...". Ids stay the same for as long as an element exists.
|
|
34
|
+
|
|
35
|
+
Screen updates:
|
|
36
|
+
- The step starts with the full listing.
|
|
37
|
+
- Every action returns only what changed: lines that appeared ("+") and ids that disappeared ("-"). Everything else is unchanged and its ids still work.
|
|
38
|
+
- When the path changes or most of the page changed, you get the full listing again. Call "observe" if you need it otherwise.
|
|
39
|
+
|
|
40
|
+
Working efficiently:
|
|
41
|
+
- Use "batch" to do several actions on one screen in a single call (fill every field, then click Continue). It stops at the first action that fails and tells you which.
|
|
42
|
+
|
|
43
|
+
Rules:
|
|
44
|
+
- Do exactly what the instruction asks, using only data the instruction or its params give. Never invent personal data.
|
|
45
|
+
- Prefer the visible, labelled control a real user would use.
|
|
46
|
+
- For an "assert" step, never change the app: only observe, then judge.
|
|
47
|
+
- When the step is done, or cannot be done, call "complete_step" exactly once:
|
|
48
|
+
- passed: the screen shows the goal was reached (act) or the statement is true (assert).
|
|
49
|
+
- failed: the app did not behave as required, or the statement is false. Use ASSERTION_INCONCLUSIVE when the screen does not prove it either way.
|
|
50
|
+
- blocked: something outside the product stopped you (an environment error, a control your tools cannot operate).
|
|
51
|
+
- Your summary is one sentence naming the evidence on screen.`;
|
|
52
|
+
export function claudeCodeAgent(options = {}) {
|
|
53
|
+
const modelId = options.model ?? 'sonnet';
|
|
54
|
+
const base = {
|
|
55
|
+
...ISOLATED,
|
|
56
|
+
pathToClaudeCodeExecutable: options.claudeExecutable ?? requireInstalledClaude(),
|
|
57
|
+
};
|
|
58
|
+
return {
|
|
59
|
+
executor: claudeCodeExecutor({
|
|
60
|
+
modelId,
|
|
61
|
+
maxTurns: options.maxTurns ?? 30,
|
|
62
|
+
base: { ...base, effort: options.effort ?? 'low' },
|
|
63
|
+
}),
|
|
64
|
+
// Judges agent.waitFor / agent.extract, which e2e runs through its own structured-output call.
|
|
65
|
+
model: claudeCode(modelId, base),
|
|
66
|
+
};
|
|
67
|
+
}
|
|
68
|
+
function claudeCodeExecutor(settings) {
|
|
69
|
+
return {
|
|
70
|
+
name: 'claude-code',
|
|
71
|
+
version: '2',
|
|
72
|
+
cache: 'inherit',
|
|
73
|
+
runStep: (ctx) => runStep(ctx, settings),
|
|
74
|
+
};
|
|
75
|
+
}
|
|
76
|
+
async function runStep(ctx, { modelId, maxTurns, base }) {
|
|
77
|
+
const controller = new AbortController();
|
|
78
|
+
const abort = () => controller.abort(ctx.signal.reason);
|
|
79
|
+
if (ctx.signal.aborted)
|
|
80
|
+
abort();
|
|
81
|
+
ctx.signal.addEventListener('abort', abort, { once: true });
|
|
82
|
+
const state = { transcript: [], screen: new ScreenTracker(), abort: () => controller.abort() };
|
|
83
|
+
const tools = buildTools(ctx, state);
|
|
84
|
+
const model = claudeCode(modelId, {
|
|
85
|
+
...base,
|
|
86
|
+
systemPrompt: SYSTEM_PROMPT,
|
|
87
|
+
maxTurns,
|
|
88
|
+
mcpServers: { [SERVER]: createAiSdkMcpServer(SERVER, tools) },
|
|
89
|
+
allowedTools: Object.keys(tools).map((name) => `mcp__${SERVER}__${name}`),
|
|
90
|
+
});
|
|
91
|
+
const startedAt = new Date();
|
|
92
|
+
try {
|
|
93
|
+
const result = await generateText({
|
|
94
|
+
model,
|
|
95
|
+
prompt: await buildPrompt(ctx, state.screen),
|
|
96
|
+
abortSignal: controller.signal,
|
|
97
|
+
});
|
|
98
|
+
ctx.budgets.recordModelCall({
|
|
99
|
+
...readUsage(result.usage),
|
|
100
|
+
startedAt: startedAt.toISOString(),
|
|
101
|
+
durationMs: Date.now() - startedAt.getTime(),
|
|
102
|
+
provider: 'claude-code',
|
|
103
|
+
modelId,
|
|
104
|
+
});
|
|
105
|
+
}
|
|
106
|
+
catch (error) {
|
|
107
|
+
if (state.hardStop)
|
|
108
|
+
throw state.hardStop;
|
|
109
|
+
if (state.verdict === undefined)
|
|
110
|
+
throw error;
|
|
111
|
+
}
|
|
112
|
+
finally {
|
|
113
|
+
ctx.signal.removeEventListener('abort', abort);
|
|
114
|
+
ctx.attachTranscript(state.transcript.join('\n'));
|
|
115
|
+
}
|
|
116
|
+
if (state.hardStop)
|
|
117
|
+
throw state.hardStop;
|
|
118
|
+
return (state.verdict ?? {
|
|
119
|
+
status: 'blocked',
|
|
120
|
+
errorCode: 'AUTOMATION_UNSUPPORTED',
|
|
121
|
+
summary: `Claude Code stopped without calling complete_step within ${maxTurns} turns.`,
|
|
122
|
+
});
|
|
123
|
+
}
|
|
124
|
+
/** Remembers the listing Claude last saw, so updates can carry only what changed. */
|
|
125
|
+
class ScreenTracker {
|
|
126
|
+
lines = new Set();
|
|
127
|
+
path;
|
|
128
|
+
full(screen) {
|
|
129
|
+
this.remember(screen);
|
|
130
|
+
return render(screen);
|
|
131
|
+
}
|
|
132
|
+
changes(screen) {
|
|
133
|
+
const current = listingLines(screen);
|
|
134
|
+
const added = current.filter((line) => !this.lines.has(line));
|
|
135
|
+
const currentIds = new Set(current.map(idOf));
|
|
136
|
+
const removed = [...this.lines].map(idOf).filter((id) => id !== undefined && !currentIds.has(id));
|
|
137
|
+
const pathChanged = screen.path !== this.path;
|
|
138
|
+
if (pathChanged || screen.treeUnavailable || added.length > current.length * FULL_LISTING_RATIO) {
|
|
139
|
+
return `${pathChanged ? 'Navigated. ' : ''}Full screen:\n${this.full(screen)}`;
|
|
140
|
+
}
|
|
141
|
+
this.remember(screen);
|
|
142
|
+
if (added.length === 0 && removed.length === 0)
|
|
143
|
+
return 'No visible change.';
|
|
144
|
+
return [
|
|
145
|
+
added.length > 0 ? `+ appeared or changed:\n${added.join('\n')}` : undefined,
|
|
146
|
+
removed.length > 0 ? `- gone: ${removed.map((id) => `#${id}`).join(' ')}` : undefined,
|
|
147
|
+
]
|
|
148
|
+
.filter((part) => part !== undefined)
|
|
149
|
+
.join('\n');
|
|
150
|
+
}
|
|
151
|
+
remember(screen) {
|
|
152
|
+
this.lines = new Set(listingLines(screen));
|
|
153
|
+
this.path = screen.path;
|
|
154
|
+
}
|
|
155
|
+
}
|
|
156
|
+
function listingLines(screen) {
|
|
157
|
+
return screen.text.split('\n').filter((line) => line.trim() !== '');
|
|
158
|
+
}
|
|
159
|
+
function idOf(line) {
|
|
160
|
+
return /#(\w+)/.exec(line)?.[1];
|
|
161
|
+
}
|
|
162
|
+
function render(screen) {
|
|
163
|
+
return [
|
|
164
|
+
screen.path === undefined ? undefined : `Path: ${screen.path}`,
|
|
165
|
+
screen.treeUnavailable ? 'The element list is unavailable right now; observe again.' : undefined,
|
|
166
|
+
screen.text,
|
|
167
|
+
screen.truncated ? '(listing truncated; scroll to see more)' : undefined,
|
|
168
|
+
]
|
|
169
|
+
.filter((line) => line !== undefined)
|
|
170
|
+
.join('\n');
|
|
171
|
+
}
|
|
172
|
+
const BATCH_ACTION = z.object({
|
|
173
|
+
action: z.enum(['tap', 'type', 'select', 'check', 'press']),
|
|
174
|
+
id: z.string(),
|
|
175
|
+
text: z.string().optional().describe('type: the text to enter'),
|
|
176
|
+
value: z.string().optional().describe('select: the option label or value'),
|
|
177
|
+
checked: z.boolean().optional().describe('check: true or false'),
|
|
178
|
+
key: z.string().optional().describe('press: a key such as "Enter"'),
|
|
179
|
+
});
|
|
180
|
+
function buildTools(ctx, state) {
|
|
181
|
+
const can = (verb) => ctx.target.verbs.has(verb);
|
|
182
|
+
const node = z.string().describe('Element id from the screen listing, e.g. "n12"');
|
|
183
|
+
const target = (id) => ({ id: id.replace(/^#/, '') });
|
|
184
|
+
const define = (name, description, inputSchema, body) => ({
|
|
185
|
+
description,
|
|
186
|
+
inputSchema,
|
|
187
|
+
execute: async (args) => {
|
|
188
|
+
state.transcript.push(`→ ${name} ${JSON.stringify(args)}`);
|
|
189
|
+
try {
|
|
190
|
+
const output = await body(args);
|
|
191
|
+
state.transcript.push(`← ${output.split('\n', 1)[0]}`);
|
|
192
|
+
return output;
|
|
193
|
+
}
|
|
194
|
+
catch (error) {
|
|
195
|
+
if (isAgentError(error) && RUNTIME_STOPS.has(error.code)) {
|
|
196
|
+
state.hardStop = error;
|
|
197
|
+
state.abort();
|
|
198
|
+
}
|
|
199
|
+
state.transcript.push(`✗ ${error instanceof Error ? error.message : String(error)}`);
|
|
200
|
+
throw error;
|
|
201
|
+
}
|
|
202
|
+
},
|
|
203
|
+
});
|
|
204
|
+
const changes = async () => state.screen.changes(await ctx.observe());
|
|
205
|
+
const act = (name, description, inputSchema, body) => define(name, description, inputSchema, async (args) => {
|
|
206
|
+
await body(args);
|
|
207
|
+
return `Done.\n${await changes()}`;
|
|
208
|
+
});
|
|
209
|
+
const tools = {
|
|
210
|
+
observe: define('observe', 'Read the full current screen listing.', z.object({}), async () => state.screen.full(await ctx.observe())),
|
|
211
|
+
complete_step: define('complete_step', 'Finish the step with a verdict. Call exactly once, last.', z.object({
|
|
212
|
+
status: z.enum(['passed', 'failed', 'blocked']),
|
|
213
|
+
summary: z.string().describe('One sentence naming the evidence on screen'),
|
|
214
|
+
errorCode: z
|
|
215
|
+
.enum([...FAILED_CODES, ...BLOCKED_CODES])
|
|
216
|
+
.optional()
|
|
217
|
+
.describe('Required for blocked; optional for failed; omit for passed'),
|
|
218
|
+
}), async ({ status, summary, errorCode }) => {
|
|
219
|
+
state.verdict = toVerdict(status, summary, errorCode);
|
|
220
|
+
return 'Verdict recorded. Stop now.';
|
|
221
|
+
}),
|
|
222
|
+
};
|
|
223
|
+
if (ctx.step.kind === 'assert')
|
|
224
|
+
return tools;
|
|
225
|
+
const runOne = async (step) => {
|
|
226
|
+
const t = target(step.id);
|
|
227
|
+
switch (step.action) {
|
|
228
|
+
case 'tap':
|
|
229
|
+
return ctx.actions.tap(t);
|
|
230
|
+
case 'type':
|
|
231
|
+
return ctx.actions.type(t, step.text ?? '');
|
|
232
|
+
case 'select':
|
|
233
|
+
return ctx.actions.select(t, step.value ?? '');
|
|
234
|
+
case 'check':
|
|
235
|
+
return ctx.actions.check(t, step.checked ?? true);
|
|
236
|
+
case 'press':
|
|
237
|
+
return ctx.actions.press(t, step.key ?? 'Enter');
|
|
238
|
+
}
|
|
239
|
+
};
|
|
240
|
+
tools.batch = define('batch', 'Run several actions in order on the current screen (e.g. fill fields, then tap Continue). Stops at the first failure.', z.object({ actions: z.array(BATCH_ACTION).min(1).max(15) }), async ({ actions }) => {
|
|
241
|
+
for (const [index, step] of actions.entries()) {
|
|
242
|
+
if (!can(step.action))
|
|
243
|
+
throw new Error(`Action ${index + 1} (${step.action}) is not supported on this target.`);
|
|
244
|
+
try {
|
|
245
|
+
await runOne(step);
|
|
246
|
+
}
|
|
247
|
+
catch (error) {
|
|
248
|
+
if (isAgentError(error) && RUNTIME_STOPS.has(error.code))
|
|
249
|
+
throw error;
|
|
250
|
+
const reason = error instanceof Error ? error.message : String(error);
|
|
251
|
+
return `Action ${index + 1} of ${actions.length} (${step.action} #${step.id}) failed: ${reason}\n${await changes()}`;
|
|
252
|
+
}
|
|
253
|
+
}
|
|
254
|
+
return `Done (${actions.length} actions).\n${await changes()}`;
|
|
255
|
+
});
|
|
256
|
+
if (can('tap')) {
|
|
257
|
+
tools.tap = act('tap', 'Click or tap one element.', z.object({ id: node }), ({ id }) => ctx.actions.tap(target(id)));
|
|
258
|
+
}
|
|
259
|
+
if (can('type')) {
|
|
260
|
+
tools.type = act('type', 'Fill a text field with a value.', z.object({ id: node, text: z.string() }), ({ id, text }) => ctx.actions.type(target(id), text));
|
|
261
|
+
}
|
|
262
|
+
if (can('typeSecret') && ctx.step.secrets.length > 0) {
|
|
263
|
+
tools.type_secret = act('type_secret', `Fill a field with a declared secret. Available: ${ctx.step.secrets.map((s) => `${s.name} (${s.purpose})`).join(', ')}.`, z.object({ id: node, name: z.string() }), ({ id, name }) => ctx.actions.typeSecret(target(id), name));
|
|
264
|
+
}
|
|
265
|
+
if (can('select')) {
|
|
266
|
+
tools.select = act('select', 'Choose an option in a select/dropdown by its visible label or value.', z.object({ id: node, value: z.string() }), ({ id, value }) => ctx.actions.select(target(id), value));
|
|
267
|
+
}
|
|
268
|
+
if (can('check')) {
|
|
269
|
+
tools.check = act('check', 'Set a checkbox, radio or switch to checked or unchecked.', z.object({ id: node, checked: z.boolean() }), ({ id, checked }) => ctx.actions.check(target(id), checked));
|
|
270
|
+
}
|
|
271
|
+
if (can('scroll')) {
|
|
272
|
+
tools.scroll = act('scroll', 'Scroll the page, or one scrollable element, by one screen.', z.object({ direction: z.enum(['up', 'down', 'left', 'right']), id: node.optional() }), ({ direction, id }) => ctx.actions.scroll(direction, id === undefined ? undefined : target(id)));
|
|
273
|
+
}
|
|
274
|
+
if (can('scrollUntil')) {
|
|
275
|
+
tools.scroll_until = act('scroll_until', 'Scroll until an element with this exact name or text is visible.', z.object({ text: z.string(), direction: z.enum(['up', 'down', 'left', 'right']) }), ({ text, direction }) => ctx.actions.scrollUntil(text, direction));
|
|
276
|
+
}
|
|
277
|
+
if (can('navigate')) {
|
|
278
|
+
tools.navigate = act('navigate', 'Go to an app-relative path or URL. Only when the instruction asks for it.', z.object({ url: z.string() }), ({ url }) => ctx.actions.navigate(url));
|
|
279
|
+
}
|
|
280
|
+
if (can('back')) {
|
|
281
|
+
tools.back = act('back', 'Go back one page.', z.object({}), () => ctx.actions.back());
|
|
282
|
+
}
|
|
283
|
+
return tools;
|
|
284
|
+
}
|
|
285
|
+
function toVerdict(status, summary, errorCode) {
|
|
286
|
+
if (status === 'passed')
|
|
287
|
+
return { status, summary };
|
|
288
|
+
if (status === 'blocked') {
|
|
289
|
+
const code = BLOCKED_CODES.includes(errorCode ?? '') ? errorCode : 'AUTOMATION_UNSUPPORTED';
|
|
290
|
+
return { status, summary, errorCode: code };
|
|
291
|
+
}
|
|
292
|
+
const code = FAILED_CODES.includes(errorCode ?? '') ? errorCode : undefined;
|
|
293
|
+
return code === undefined ? { status, summary } : { status, summary, errorCode: code };
|
|
294
|
+
}
|
|
295
|
+
async function buildPrompt(ctx, screen) {
|
|
296
|
+
const { step } = ctx;
|
|
297
|
+
return [
|
|
298
|
+
`Step kind: ${step.kind}`,
|
|
299
|
+
`Instruction: ${step.instruction}`,
|
|
300
|
+
step.params === undefined ? undefined : `Params: ${JSON.stringify(step.params)}`,
|
|
301
|
+
ctx.agentContext ? `Context: ${ctx.agentContext}` : undefined,
|
|
302
|
+
ctx.ledger ? `Earlier steps in this test:\n${ctx.ledger}` : undefined,
|
|
303
|
+
ctx.replayedPrefix === undefined
|
|
304
|
+
? undefined
|
|
305
|
+
: `A recorded replay already did part of this step and stopped (${ctx.replayedPrefix.stopReason}). Do not repeat these actions:\n- ${ctx.replayedPrefix.replayedActions.join('\n- ')}` +
|
|
306
|
+
(ctx.replayedPrefix.uncertainAction
|
|
307
|
+
? `\nThis action may or may not have taken effect; check the screen before retrying it: ${ctx.replayedPrefix.uncertainAction}`
|
|
308
|
+
: ''),
|
|
309
|
+
`Current screen:\n${screen.full(await ctx.observe())}`,
|
|
310
|
+
]
|
|
311
|
+
.filter((part) => part !== undefined)
|
|
312
|
+
.join('\n\n');
|
|
313
|
+
}
|
|
314
|
+
function readUsage(usage) {
|
|
315
|
+
const count = (value) => {
|
|
316
|
+
if (typeof value === 'number')
|
|
317
|
+
return value;
|
|
318
|
+
if (value !== null && typeof value === 'object' && typeof value.total === 'number') {
|
|
319
|
+
return value.total;
|
|
320
|
+
}
|
|
321
|
+
return undefined;
|
|
322
|
+
};
|
|
323
|
+
const details = usage
|
|
324
|
+
?.inputTokenDetails;
|
|
325
|
+
const fields = {
|
|
326
|
+
inputTokens: count(usage?.inputTokens),
|
|
327
|
+
outputTokens: count(usage?.outputTokens),
|
|
328
|
+
cacheReadTokens: details?.cacheReadTokens,
|
|
329
|
+
cacheWriteTokens: details?.cacheWriteTokens,
|
|
330
|
+
};
|
|
331
|
+
return Object.fromEntries(Object.entries(fields).filter(([, value]) => typeof value === 'number'));
|
|
332
|
+
}
|
|
@@ -0,0 +1,4 @@
|
|
|
1
|
+
export declare const CLAUDE_PATH_ENV = "IL_E2E_CLAUDE_PATH";
|
|
2
|
+
/** The developer's installed Claude Code: env override, then PATH, then the installer's default locations. */
|
|
3
|
+
export declare function findInstalledClaude(): string | undefined;
|
|
4
|
+
export declare function requireInstalledClaude(): string;
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
import { accessSync, constants } from 'node:fs';
|
|
2
|
+
import { homedir } from 'node:os';
|
|
3
|
+
import { delimiter, join } from 'node:path';
|
|
4
|
+
export const CLAUDE_PATH_ENV = 'IL_E2E_CLAUDE_PATH';
|
|
5
|
+
function executable(path) {
|
|
6
|
+
try {
|
|
7
|
+
accessSync(path, constants.X_OK);
|
|
8
|
+
return true;
|
|
9
|
+
}
|
|
10
|
+
catch {
|
|
11
|
+
return false;
|
|
12
|
+
}
|
|
13
|
+
}
|
|
14
|
+
/** The developer's installed Claude Code: env override, then PATH, then the installer's default locations. */
|
|
15
|
+
export function findInstalledClaude() {
|
|
16
|
+
const override = process.env[CLAUDE_PATH_ENV];
|
|
17
|
+
if (override)
|
|
18
|
+
return executable(override) ? override : undefined;
|
|
19
|
+
const fromPath = (process.env.PATH ?? '')
|
|
20
|
+
.split(delimiter)
|
|
21
|
+
.filter(Boolean)
|
|
22
|
+
.map((dir) => join(dir, 'claude'));
|
|
23
|
+
const defaults = [join(homedir(), '.local', 'bin', 'claude'), join(homedir(), '.claude', 'local', 'claude')];
|
|
24
|
+
return [...fromPath, ...defaults].find(executable);
|
|
25
|
+
}
|
|
26
|
+
export function requireInstalledClaude() {
|
|
27
|
+
const found = findInstalledClaude();
|
|
28
|
+
if (found)
|
|
29
|
+
return found;
|
|
30
|
+
throw new Error('il-e2e-agent: Claude Code was not found. Install it (curl -fsSL https://claude.ai/install.sh | bash), ' +
|
|
31
|
+
`sign in with \`claude auth login\`, or point ${CLAUDE_PATH_ENV} at the binary.`);
|
|
32
|
+
}
|
package/dist/config.d.ts
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
import type { E2EConfig } from 'e2e';
|
|
2
|
+
import { type WebOptions } from '@e2e-dev/web';
|
|
3
|
+
import { type ClaudeCodeAgentOptions } from './agent.ts';
|
|
4
|
+
import { type NetworkPolicy } from './network.ts';
|
|
5
|
+
type TargetApp = NonNullable<E2EConfig['targets'][number]['app']>;
|
|
6
|
+
export interface ClaudeE2EConfig extends Omit<E2EConfig, 'targets' | 'agents'> {
|
|
7
|
+
/** The app under test: its URL and, optionally, the command that starts it. */
|
|
8
|
+
readonly app: TargetApp;
|
|
9
|
+
/** Report label for the target; defaults to `web`. */
|
|
10
|
+
readonly name?: string;
|
|
11
|
+
/** Options for the Playwright web engine (viewport, browser, …). */
|
|
12
|
+
readonly browser?: WebOptions;
|
|
13
|
+
/** Claude Code model alias for agent steps: 'sonnet' (default), 'haiku' or 'opus'. */
|
|
14
|
+
readonly model?: string;
|
|
15
|
+
/** Reasoning effort for agent steps: 'low' (default), 'medium', 'high', … */
|
|
16
|
+
readonly effort?: ClaudeCodeAgentOptions['effort'];
|
|
17
|
+
/** Actions one `agent.act` may take; e2e's default is 25. */
|
|
18
|
+
readonly maxSteps?: number;
|
|
19
|
+
/** Hosts the test browser may reach; `false` turns the guard off. Default: localhost only. */
|
|
20
|
+
readonly network?: NetworkPolicy | false;
|
|
21
|
+
/** Extra named agents, passed through to e2e unchanged. */
|
|
22
|
+
readonly agents?: E2EConfig['agents'];
|
|
23
|
+
}
|
|
24
|
+
export declare function defineConfig(config: ClaudeE2EConfig): E2EConfig;
|
|
25
|
+
export {};
|
package/dist/config.js
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
import { web } from '@e2e-dev/web';
|
|
2
|
+
import { claudeCodeAgent } from "./agent.js";
|
|
3
|
+
import { registerNetworkPolicy } from "./network.js";
|
|
4
|
+
export function defineConfig(config) {
|
|
5
|
+
const { app, name, browser, model, effort, maxSteps, network, agents, ...rest } = config;
|
|
6
|
+
registerNetworkPolicy(network ?? {});
|
|
7
|
+
return {
|
|
8
|
+
...rest,
|
|
9
|
+
tests: rest.tests ?? ['e2e/**/*.e2e.ts', 'tests/**/*.e2e.ts'],
|
|
10
|
+
targets: [{ name: name ?? 'web', engine: web(browser), app }],
|
|
11
|
+
agents: {
|
|
12
|
+
default: { ...claudeCodeAgent({ model, effort }), ...(maxSteps === undefined ? {} : { maxSteps }) },
|
|
13
|
+
...agents,
|
|
14
|
+
},
|
|
15
|
+
};
|
|
16
|
+
}
|
package/dist/index.d.ts
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
export { defineConfig, type ClaudeE2EConfig } from './config.ts';
|
|
2
|
+
export { claudeCodeAgent, type ClaudeCodeAgentOptions } from './agent.ts';
|
|
3
|
+
export { test } from './test.ts';
|
|
4
|
+
export { decide, type NetworkPolicy } from './network.ts';
|
|
5
|
+
export { expect, unique, credentials, secrets } from 'e2e';
|
|
6
|
+
export type { Agent, Screen } from 'e2e';
|
package/dist/index.js
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
import type { WebRoute } from '@e2e-dev/web';
|
|
2
|
+
/**
|
|
3
|
+
* What the test browser may reach. Hosts accept a leading `*.` wildcard
|
|
4
|
+
* (`*.stripe.com` matches `js.stripe.com` and `stripe.com`). Everything not listed is aborted.
|
|
5
|
+
*/
|
|
6
|
+
export interface NetworkPolicy {
|
|
7
|
+
/** Hosts reachable with any method. `localhost` and `127.0.0.1` are always allowed. */
|
|
8
|
+
readonly allow?: readonly string[];
|
|
9
|
+
/** Hosts reachable with GET, HEAD and OPTIONS only (e.g. a staging API that must not be written to). */
|
|
10
|
+
readonly readOnly?: readonly string[];
|
|
11
|
+
/** Path prefixes aborted even on allowed hosts (e.g. the app's own `/api/facebook` proxy). */
|
|
12
|
+
readonly blockPaths?: readonly string[];
|
|
13
|
+
/** Print each newly blocked host/path once, to tune the policy. */
|
|
14
|
+
readonly log?: boolean;
|
|
15
|
+
}
|
|
16
|
+
/** Called by `defineConfig`; workers read it back through the environment. */
|
|
17
|
+
export declare function registerNetworkPolicy(policy: NetworkPolicy | false): void;
|
|
18
|
+
export declare function decide(policy: NetworkPolicy, url: string, method: string): 'continue' | 'abort';
|
|
19
|
+
export declare function createRouteGuard(): ((route: WebRoute) => Promise<void>) | undefined;
|
package/dist/network.js
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
const POLICY_ENV = 'IL_E2E_NETWORK_POLICY';
|
|
2
|
+
const ALWAYS_ALLOWED = ['localhost', '127.0.0.1'];
|
|
3
|
+
const READ_METHODS = new Set(['GET', 'HEAD', 'OPTIONS']);
|
|
4
|
+
/** Called by `defineConfig`; workers read it back through the environment. */
|
|
5
|
+
export function registerNetworkPolicy(policy) {
|
|
6
|
+
process.env[POLICY_ENV] = JSON.stringify(policy);
|
|
7
|
+
}
|
|
8
|
+
function currentPolicy() {
|
|
9
|
+
const raw = process.env[POLICY_ENV];
|
|
10
|
+
return raw === undefined ? {} : JSON.parse(raw);
|
|
11
|
+
}
|
|
12
|
+
function hostMatches(hostname, pattern) {
|
|
13
|
+
if (pattern.startsWith('*.')) {
|
|
14
|
+
const base = pattern.slice(2);
|
|
15
|
+
return hostname === base || hostname.endsWith(`.${base}`);
|
|
16
|
+
}
|
|
17
|
+
return hostname === pattern;
|
|
18
|
+
}
|
|
19
|
+
export function decide(policy, url, method) {
|
|
20
|
+
if (url.startsWith('data:') || url.startsWith('blob:'))
|
|
21
|
+
return 'continue';
|
|
22
|
+
const { hostname, pathname } = new URL(url);
|
|
23
|
+
if ((policy.blockPaths ?? []).some((prefix) => pathname.startsWith(prefix)))
|
|
24
|
+
return 'abort';
|
|
25
|
+
if ([...ALWAYS_ALLOWED, ...(policy.allow ?? [])].some((pattern) => hostMatches(hostname, pattern))) {
|
|
26
|
+
return 'continue';
|
|
27
|
+
}
|
|
28
|
+
if ((policy.readOnly ?? []).some((pattern) => hostMatches(hostname, pattern))) {
|
|
29
|
+
return READ_METHODS.has(method) ? 'continue' : 'abort';
|
|
30
|
+
}
|
|
31
|
+
return 'abort';
|
|
32
|
+
}
|
|
33
|
+
const reported = new Set();
|
|
34
|
+
export function createRouteGuard() {
|
|
35
|
+
const policy = currentPolicy();
|
|
36
|
+
if (policy === false)
|
|
37
|
+
return undefined;
|
|
38
|
+
return async (route) => {
|
|
39
|
+
const { url, method } = route.request;
|
|
40
|
+
if (decide(policy, url, method) === 'continue')
|
|
41
|
+
return route.continue();
|
|
42
|
+
if (policy.log) {
|
|
43
|
+
const { hostname, pathname } = new URL(url);
|
|
44
|
+
const key = `${method} ${hostname}${pathname.split('/').slice(0, 3).join('/')}`;
|
|
45
|
+
if (!reported.has(key)) {
|
|
46
|
+
reported.add(key);
|
|
47
|
+
console.warn(`[il-e2e-agent] blocked ${key}`);
|
|
48
|
+
}
|
|
49
|
+
}
|
|
50
|
+
return route.abort();
|
|
51
|
+
};
|
|
52
|
+
}
|
package/dist/test.d.ts
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
import type { Browser } from '@e2e-dev/web';
|
|
2
|
+
/** e2e's `test`, with the config's network policy applied before every test body. */
|
|
3
|
+
export declare const test: import("e2e").TestAPI<import("e2e").TestFixtures & {
|
|
4
|
+
browser: Browser;
|
|
5
|
+
} & {
|
|
6
|
+
networkPolicy: void;
|
|
7
|
+
}>;
|
package/dist/test.js
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
import { test as base } from 'e2e';
|
|
2
|
+
import { createRouteGuard } from "./network.js";
|
|
3
|
+
/** e2e's `test`, with the config's network policy applied before every test body. */
|
|
4
|
+
export const test = base.extend().extend({
|
|
5
|
+
networkPolicy: async ({ browser }, use) => {
|
|
6
|
+
const guard = createRouteGuard();
|
|
7
|
+
if (guard)
|
|
8
|
+
await browser.route('**/*', guard);
|
|
9
|
+
await use();
|
|
10
|
+
},
|
|
11
|
+
});
|
package/package.json
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "il-e2e-agent",
|
|
3
|
+
"version": "0.1.0",
|
|
4
|
+
"description": "Natural-language browser tests (tester-army/e2e) run by the Claude Code installed on each developer's machine, on their own login",
|
|
5
|
+
"type": "module",
|
|
6
|
+
"repository": {
|
|
7
|
+
"type": "git",
|
|
8
|
+
"url": "git+https://github.com/vishalmishraa22/il-e2e-agent.git"
|
|
9
|
+
},
|
|
10
|
+
"exports": {
|
|
11
|
+
".": {
|
|
12
|
+
"types": "./dist/index.d.ts",
|
|
13
|
+
"default": "./dist/index.js"
|
|
14
|
+
}
|
|
15
|
+
},
|
|
16
|
+
"bin": {
|
|
17
|
+
"il-e2e-agent": "./bin/il-e2e-agent.mjs"
|
|
18
|
+
},
|
|
19
|
+
"files": [
|
|
20
|
+
"dist",
|
|
21
|
+
"bin",
|
|
22
|
+
"README.md"
|
|
23
|
+
],
|
|
24
|
+
"scripts": {
|
|
25
|
+
"build": "tsc -p tsconfig.json",
|
|
26
|
+
"prepare": "tsc -p tsconfig.json",
|
|
27
|
+
"postinstall": "node bin/prune-bundled-claude.mjs",
|
|
28
|
+
"typecheck": "tsc -p tsconfig.json --noEmit"
|
|
29
|
+
},
|
|
30
|
+
"engines": {
|
|
31
|
+
"node": ">=20.19"
|
|
32
|
+
},
|
|
33
|
+
"dependencies": {
|
|
34
|
+
"@e2e-dev/web": "0.11.2",
|
|
35
|
+
"ai": "7.0.127",
|
|
36
|
+
"ai-sdk-provider-claude-code": "4.3.3",
|
|
37
|
+
"e2e": "0.16.0",
|
|
38
|
+
"playwright": "1.63.0",
|
|
39
|
+
"zod": "4.6.5"
|
|
40
|
+
},
|
|
41
|
+
"devDependencies": {
|
|
42
|
+
"@types/node": "^24.0.0",
|
|
43
|
+
"typescript": "^5.9.0"
|
|
44
|
+
}
|
|
45
|
+
}
|