@xoxoai/checkmate 0.4.1 → 0.4.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -8,38 +8,22 @@ AI test automation that actually works. Write tests in plain English, without lo
8
8
  ![openai](https://img.shields.io/badge/OpenAI-API-yellow.svg)
9
9
  [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
10
10
 
11
- ## Why?
12
-
13
- Spending countless hours building and maintaining E2E tests that look like this?
14
-
15
- ```
16
- await page.goto('https://www.google.com')
17
- const searchBox = page.getByRole('combobox', { name: 'Search', exact: true })
18
- await searchBox.fill('playwright test automation')
19
- await searchBox.press('Enter')
20
- await expect(page.getByRole('link', { name: 'playwright' })
21
- .filter({ hasText: 'playwright.dev' })
22
- .first(), 'playwright.dev link should be visible')
23
- .toBeVisible( { timeout: 30 * 1000 } )
24
- ```
25
-
26
- Try **_checkmate_**!
11
+ ##
27
12
 
28
13
  ```typescript
29
14
  await ai.run({
30
15
  action: `
31
- Navigate to google.com
32
- Type 'playwright test automation' in the search bar
33
- Press Enter key`,
16
+ Navigate to google.com
17
+ Type 'playwright test automation' in the search bar
18
+ Press Enter key`,
34
19
  expect: `
35
- Search results contain the playwright.dev link`,
20
+ Search results contain the playwright.dev link`,
36
21
  })
37
22
  ```
38
23
 
39
- ## What You Get
24
+ ##
40
25
 
41
26
  ✅ **Zero Locators** - Write tests in plain English
42
- ✅ **Self-Healing** - Tests adapt to UI changes automatically
43
27
  ✅ **Any Provider** - Gemini, Claude, Groq, GPT, xAI, or local models
44
28
  ✅ **Web & Salesforce** - Basic support out of the box
45
29
  ✅ **Cost Optimized** - Built-in token management and budgeting
@@ -97,41 +81,60 @@ npm run show:report
97
81
 
98
82
  ## Writing Tests
99
83
 
100
- Import `test` from `@xoxoai/checkmate/playwright` and use the `ai` fixture.
101
84
  **_checkmate_** tests are written using natural language by specifying `action` and `expect`:
102
85
 
103
86
  ```typescript
104
87
  import { test } from '@xoxoai/checkmate/playwright'
105
88
 
106
- test('google search', async ({ ai }) => {
107
- await ai.run({
108
- action: `
109
- Open the browser and navigate to google.com.
110
- Type 'playwright test automation' in the search bar.
111
- Press Enter key.`,
112
- expect: `
113
- Search results contain the 'playwright.dev' link`,
89
+ test.describe('multi-step : full AI mode', async () => {
90
+ test('purchase flow', async ({ ai }) => {
91
+ await test.step('Open Shop', async () => {
92
+ await ai.run({
93
+ action: `
94
+ Navigate to https://my-shop.com`,
95
+ expect: `
96
+ My Shop home page is loaded`,
97
+ })
98
+ })
99
+
100
+ await test.step('Select product', async () => {
101
+ await ai.run({
102
+ action: `
103
+ Click 'Shop Now' on 'Men's Outerwear' category
104
+ Click on the first Shell product in the list`,
105
+ expect: `
106
+ Product detail with title and price.`,
107
+ })
108
+ })
109
+
110
+ await test.step('Cart and checkout', async () => {
111
+ await ai.run({
112
+ action: `
113
+ Click 'Add to Cart'
114
+ Click 'Checkout' in the 'Added to cart' dialog`,
115
+ expect: `
116
+ Checkout with Order Summary and totals`,
117
+ })
118
+ })
114
119
  })
115
120
  })
116
121
  ```
117
122
 
118
123
  That's it. No page objects, no selectors. No locators. Peace on Earth.
119
124
 
120
- Browser settings (viewport, headless mode, video recording, timeouts, etc.) are configured in [playwright.config.ts](playwright.config.ts) using Playwright's [standard](https://playwright.dev/docs/test-configuration) configuration mechanism.
121
-
122
- See [guide](docs/GUIDE.md#best-practices) for detailed examples and best practices.
123
- See [guide](docs/GUIDE.md#core-concepts) for the main building blocks and [extensions](docs/EXTENSIONS.md) for customization.
125
+ Tests are orchestrated by [playwright](https://playwright.dev/docs/test-configuration) [config](playwright.config.ts).
124
126
 
125
- ### Programmatic API
127
+ ### API
126
128
 
127
- If you use **_checkmate_** without the fixture wrapper, compose a runner from `@xoxoai/checkmate/core` and extensions:
129
+ Compose your own **_checkmate_** using [extensions](docs/EXTENSIONS.md):
128
130
 
129
131
  ```typescript
130
132
  import { createRunner } from '@xoxoai/checkmate/core'
131
133
  import { web } from '@xoxoai/checkmate/playwright'
134
+ import { notion, database, api } from 'my-custom-extensions'
132
135
 
133
136
  const ai = createRunner({
134
- extensions: [web({ page })],
137
+ extensions: [web({ page }), notion(), database(), api()],
135
138
  })
136
139
 
137
140
  await ai.run({
@@ -140,55 +143,56 @@ await ai.run({
140
143
  })
141
144
  ```
142
145
 
143
- See [guide](docs/GUIDE.md#advanced-topics) for advanced topics and [extensions](docs/EXTENSIONS.md) for building custom tools, extensions, runners, and scaffolded starter projects.
146
+ ### Entry Points:
144
147
 
145
- Published entry points:
148
+ `@xoxoai/checkmate/core`: compose runner, tools, and extensions.
149
+ `@xoxoai/checkmate/playwright`: Web extension with Playwright `test` and `expect`.
150
+ `@xoxoai/checkmate/salesforce`: Salesforce extensions with the same `ai` fixture shape.
146
151
 
147
- `@xoxoai/checkmate/core`: Build your own runner with extensions.
148
- `@xoxoai/checkmate/playwright`: Use the built-in web extension with Playwright `test` and `expect`.
149
- `@xoxoai/checkmate/salesforce`: Use the built-in web + Salesforce extensions with the same `ai` fixture shape.
150
-
151
- The repository keeps runnable consumer-style examples under `test/examples/`.
152
+ See [guide](docs/GUIDE.md#best-practices) for tips on writing effective tests.
152
153
 
153
154
  ## Costs
154
155
 
155
- Costs vary based on model and provider, test complexity and number of steps.
156
- **_checkmate_** includes built-in token usage [monitoring](docs/GUIDE.md#cost-management).
156
+ They depend on the model, provider, test complexity, and number of steps.
157
157
 
158
- Cost estimates with [gpt-oss-20b hosted on groq.com](https://console.groq.com/docs/model/openai/gpt-oss-20b) for optimal balance:
158
+ Estimates for [gpt-oss-20b hosted on groq.com](https://console.groq.com/docs/model/openai/gpt-oss-20b):
159
159
 
160
160
  - Simple test (~5 steps): ~$0.001 - $0.01
161
161
  - Complex test (~20 steps): ~$0.01 - $0.05
162
162
  - Full E2E suite (~50 complex tests): ~$1.00 - $2.00
163
163
 
164
- See [guide](docs/GUIDE.md#cost-management) for detailed cost control and monitoring options.
164
+ **_checkmate_** includes built-in token usage [monitoring](docs/GUIDE.md#cost-management).
165
+
166
+ See [guide](docs/GUIDE.md#cost-management) for cost control and monitoring options.
165
167
 
166
168
  ## Common Issues
167
169
 
168
170
  **AI makes incorrect decisions**
169
171
 
170
- - Provide precise descriptions in `action` and more focused assertions in `expect`
171
- - Reference specific element identifiers and roles (for example: text, label, button, list)
172
- - Break complex workflows into single-action steps; use a step-by-step approach
172
+ - Provide precise descriptions in `action` and focused assertions in `expect`
173
+ - Reference specific element and roles, for example: text, label, button, list, etc.
174
+ - Break complex workflows into single-action steps and use a step-by-step approach
173
175
 
174
176
  **Tests loop during step execution**
175
177
 
176
178
  - Increase `OPENAI_TEMPERATURE` to encourage exploration
177
- - Use a reasoning/thinking model (if available) to improve planning and avoid repetitive loops
179
+ - Use a reasoning model if possible to improve accuracy
178
180
 
179
181
  **High token costs**
180
182
 
181
- - Enable [snapshot filtering](docs/GUIDE.md#using-snapshot-filtering-for-token-optimization) with `CHECKMATE_SNAPSHOT_FILTERING=true` to score and narrow the elements automatically from `action` and `expect`. Use `topPercent` to dial how much of the scored snapshot to keep for a step.
182
- - Set a lower reasoning effort: `OPENAI_REASONING_EFFORT`
183
- - Consider disabling `OPENAI_INCLUDE_SCREENSHOT_IN_SNAPSHOT`
184
- - Use a cheaper model, lower-end models often perform well (e.g., `gpt-5.4-nano` or `gpt-oss-20b`)
183
+ - Enable [snapshot filtering](docs/GUIDE.md#using-snapshot-filtering-for-token-optimization) with `CHECKMATE_SNAPSHOT_FILTERING=true` auto-filter elements
184
+ - Adjust reasoning effort: `OPENAI_REASONING_EFFORT`
185
+ - Consider disabling `OPENAI_INCLUDE_SCREENSHOT_IN_SNAPSHOT` if visuals are not needed
186
+ - Use a cheaper model, lower-end models often perform well: `gpt-5.4-nano` or `gpt-oss-20b`
185
187
 
186
- See [guide](docs/GUIDE.md#openai-api-settings) for detailed configuration options and troubleshooting tips.
188
+ See [guide](docs/GUIDE.md#openai-api-settings) for detailed configuration options and tips.
187
189
 
188
190
  ## FAQ
189
191
 
190
192
  **Which models work best?**
191
- You can use any model that was trained for tool use. Here are the best picks based on extensive testing:
193
+ You can use any model that was trained for tool use.
194
+
195
+ Here are the best picks based on extensive testing:
192
196
 
193
197
  - Highly recommended: [`gpt-oss-20b` hosted on groq.com](https://console.groq.com/docs/model/openai/gpt-oss-20b). Groq's infrastructure is optimized for minimal latency and fast inference, making it ideal for E2E test automation.
194
198
  - Google's `gemini-2.5-flash` offers an excellent balance of cost and performance if you prefer major cloud providers.
@@ -224,8 +228,9 @@ await ai.run({
224
228
 
225
229
  ## Documentation
226
230
 
227
- - [**_checkmate_**](docs/GUIDE.md)
228
- - [Playwright](https://playwright.dev/)
231
+ - [**_checkmate_** guide](docs/GUIDE.md)
232
+ - [**_checkmate_** extensions](docs/EXTENSIONS.md)
233
+ - [**playwright** official website](https://playwright.dev/)
229
234
 
230
235
  ## Contributing
231
236
 
@@ -238,9 +243,9 @@ I'd love your help! Key areas:
238
243
 
239
244
  See [roadmap](docs/ROADMAP.md) for future plans and development
240
245
 
241
- ## MIT License
246
+ ## License
242
247
 
243
- See [license](LICENSE) file for details
248
+ MIT [license](LICENSE)
244
249
 
245
250
  ## Why I build this?
246
251
 
@@ -15,29 +15,28 @@ Use it when you want to:
15
15
 
16
16
  The runner owns:
17
17
 
18
- - the model loop
19
- - retries
18
+ - the model loop with retries
20
19
  - pass/fail resolution
21
20
  - tool dispatch
22
21
 
23
- Extensions add domain-specific behavior such as:
22
+ Extensions add domain-specific skills such as:
24
23
 
25
24
  - tools
26
25
  - system instructions
27
26
  - initial step context
28
27
  - post-tool context
29
- - shared capabilities for other extensions
30
28
  - teardown logic
29
+ - shared capabilities across extensions
31
30
 
32
- That is how `web()` and `salesforce()` work, and it is the same model you use for your own extensions.
31
+ Pre-built `web()` and `salesforce()` are in fact extensions - made the same way as you would build your own.
33
32
 
34
33
  ## Tool vs Extension
35
34
 
36
- Use `defineTool()` when you need one new action the model can call.
35
+ Use `defineTool()` to create a new action the model can call.
37
36
 
38
- Use `defineExtension()` when you need to bundle tools with runtime behavior such as instructions, setup, or extra context.
37
+ Use `defineExtension()` to bundle different tools and instructions, setup, or extra context.
39
38
 
40
- Use `createRunner()` when you want to compose your own runtime from built-in and custom extensions.
39
+ Use `createRunner()` to compose your own runtime from extensions.
41
40
 
42
41
  ## Your First Tool
43
42
 
@@ -47,8 +46,8 @@ A tool is the smallest unit of behavior.
47
46
  import { defineTool } from '@xoxoai/checkmate/core'
48
47
  import { z } from 'zod/v4'
49
48
 
50
- export const apiHealthTool = defineTool({
51
- name: 'check_api_health',
49
+ export const health = defineTool({
50
+ name: 'api-health',
52
51
  description: 'Check whether the API is healthy',
53
52
  schema: z.object({ url: z.string().url() }).strict(),
54
53
  handler: async ({ url }) => {
@@ -69,21 +68,21 @@ Good tools are:
69
68
 
70
69
  ## Your First Extension
71
70
 
72
- An extension can be as small as a name, one tool, and one instruction.
71
+ An extension can bundle one or more tools and instructions.
73
72
 
74
73
  ```typescript
75
74
  import { createRunner, defineExtension } from '@xoxoai/checkmate/core'
76
75
  import { web } from '@xoxoai/checkmate/playwright'
77
- import { apiHealthTool } from './api-health-tool'
76
+ import { health, queryRecords } from './api-tools'
78
77
 
79
- export const apiHealth = defineExtension({
80
- name: 'api-health',
81
- tools: [apiHealthTool],
82
- instructions: ['Use check_api_health before relying on API-driven UI state.'],
78
+ export const apiExtension = defineExtension({
79
+ name: 'api',
80
+ tools: [health, queryRecords],
81
+ instructions: ['Use api tools to interact with xyz service.', 'Prefer api tools over web tools when possible.'],
83
82
  })
84
83
 
85
84
  const ai = createRunner({
86
- extensions: [web({ page }), apiHealth],
85
+ extensions: [web({ page }), apiExtension],
87
86
  })
88
87
  ```
89
88
 
package/docs/GUIDE.md CHANGED
@@ -21,10 +21,10 @@ Technical documentation for **_checkmate_** - AI test automation with Playwright
21
21
 
22
22
  Main building blocks:
23
23
 
24
- - **Runner**: The object that executes steps. The main programmatic entry point is `createRunner()` from `@xoxoai/checkmate/core`.
24
+ - **Runner**: The object that executes steps. The main API entry point is `createRunner()` from `@xoxoai/checkmate/core`.
25
25
  - **Step**: A plain object with `action` and `expect`. This is the main unit of execution.
26
26
  - **Extensions**: Composable modules that add tools and runtime behavior. Built-ins include `web()` and `salesforce()`.
27
- - **Fixtures**: Convenience Playwright entry points that provide an `ai` runner in tests.
27
+ - **Fixtures**: Convenience [Playwright](https://playwright.dev/docs/test-fixtures) entry points that provide an `ai` runner in tests.
28
28
 
29
29
  Published entry points:
30
30
 
@@ -39,22 +39,16 @@ import { test } from '@xoxoai/checkmate/playwright'
39
39
 
40
40
  test('search flow', async ({ ai }) => {
41
41
  await ai.run({
42
- action: 'Search for playwright documentation',
43
- expect: 'Search results are displayed',
42
+ action: `Type 'documentation' in the search bar and press Enter`,
43
+ expect: `At least 5 search results are displayed`,
44
44
  })
45
45
  })
46
46
  ```
47
47
 
48
- If you want a runnable sample project quickly, use the built-in scaffold command after installing the package:
49
-
50
- ```bash
51
- npx checkmate create-examples
52
- ```
53
-
54
- That command copies the repo example `playwright.config.ts`, example tests, and package scripts into your project.
55
-
56
48
  ## Configuration Reference
57
49
 
50
+ Tests are managed in [Playwright's](https://playwright.dev/docs/test-configuration) standard [config](playwright.config.ts).
51
+
58
52
  ### AI API Settings
59
53
 
60
54
  | Variable | Default | Description |
@@ -76,10 +70,6 @@ That command copies the repo example `playwright.config.ts`, example tests, and
76
70
  | `CHECKMATE_LOG_LEVEL` | `off` | Logging verbosity: debug, info, warn, error, off |
77
71
  | `CHECKMATE_SNAPSHOT_FILTERING` | `false` | Enable semantic page snapshot filtering before requests are sent to the model |
78
72
 
79
- ### Playwright Configuration
80
-
81
- Browser settings (viewport, headless mode, video recording, timeouts, etc.) are configured in [playwright.config.ts](../playwright.config.ts) using Playwright's [standard](https://playwright.dev/docs/test-configuration) configuration mechanism.
82
-
83
73
  ## Writing Effective Tests
84
74
 
85
75
  ### Best Practices
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@xoxoai/checkmate",
3
- "version": "0.4.1",
3
+ "version": "0.4.2",
4
4
  "description": "AI‑driven e2e test automation framework built on Playwright Test and OpenAI API",
5
5
  "homepage": "https://github.com/dawiddiwad/checkmate",
6
6
  "repository": {