ordel-cli 0.2.2__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- ordel_cli-0.3.0/CATALOG.md +98 -0
- ordel_cli-0.3.0/PKG-INFO +175 -0
- ordel_cli-0.3.0/README.md +151 -0
- ordel_cli-0.3.0/ordel_cli/SKILL.md +91 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/__init__.py +1 -1
- ordel_cli-0.3.0/ordel_cli/adopt.py +322 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/capture.cjs +3 -1
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/capture_extract.cjs +8 -3
- ordel_cli-0.3.0/ordel_cli/catalog.py +110 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/cli.py +292 -96
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/context.py +4 -1
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/doctor.py +35 -15
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/explore.py +43 -10
- ordel_cli-0.3.0/ordel_cli/human_record.cjs +196 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/ingest.py +20 -6
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/mcp_server.py +99 -56
- ordel_cli-0.3.0/ordel_cli/ordel_config.cjs +26 -0
- ordel_cli-0.3.0/ordel_cli/ordel_evidence.cjs +50 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/pom.py +1 -1
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/qa_plan.py +7 -2
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/record.cjs +12 -7
- ordel_cli-0.3.0/ordel_cli/record.py +309 -0
- ordel_cli-0.3.0/ordel_cli/runner.py +236 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/scenarios.py +7 -2
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/session.py +5 -10
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/session_driver.cjs +10 -2
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/store.py +114 -4
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/pyproject.toml +20 -5
- ordel_cli-0.2.2/.gitignore +0 -175
- ordel_cli-0.2.2/PKG-INFO +0 -97
- ordel_cli-0.2.2/README.md +0 -74
- ordel_cli-0.2.2/ordel_cli/record.py +0 -135
- ordel_cli-0.2.2/ordel_cli/runner.py +0 -137
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/LICENSE +0 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/PRIVACY.md +0 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/branding.py +0 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/coverage.py +0 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/heal_service.py +0 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/mcp_supervisor.cjs +0 -0
- {ordel_cli-0.2.2 → ordel_cli-0.3.0}/ordel_cli/py.typed +0 -0
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
# The Ordel element catalog
|
|
2
|
+
|
|
3
|
+
The catalog is what Ordel knows about your app's pages and elements.
|
|
4
|
+
It is the baseline a heal compares against, the source page objects are generated from, and the denominator coverage is measured against.
|
|
5
|
+
It lives in `.ordel/graph.json`, and it is plain JSON you can commit, diff and review.
|
|
6
|
+
|
|
7
|
+
## Where it comes from
|
|
8
|
+
|
|
9
|
+
- `ordel explore <url>` (or the `explore_app` MCP tool) captures one page.
|
|
10
|
+
- `capture_state` captures whatever page an agent-driven session is on.
|
|
11
|
+
- `ordel record` captures every page a person lands on while recording.
|
|
12
|
+
- `record_flow` adds a thin, element-less page for each URL a flow visits.
|
|
13
|
+
- `ordel adopt --apply` re-captures the page of a heal a human accepted, so the next break heals from the app as it is now.
|
|
14
|
+
|
|
15
|
+
## `graph.json`
|
|
16
|
+
|
|
17
|
+
```json
|
|
18
|
+
{
|
|
19
|
+
"schema_version": 1,
|
|
20
|
+
"pages": {
|
|
21
|
+
"HomePage": {
|
|
22
|
+
"url": "http://localhost:3000/",
|
|
23
|
+
"title": "Sign up",
|
|
24
|
+
"elements": {
|
|
25
|
+
"email": { "tag_canon": "input", "role_canon": "textbox", "testid_canon": "email", "...": "..." }
|
|
26
|
+
}
|
|
27
|
+
}
|
|
28
|
+
}
|
|
29
|
+
}
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
- `schema_version` is the format version, currently `1`.
|
|
33
|
+
A file without it is version `0`, which has the same shape and is read as-is.
|
|
34
|
+
A file with a version newer than the installed `ordel-cli` understands is refused with an error, never guessed at.
|
|
35
|
+
- `pages` maps a page id to the page.
|
|
36
|
+
A page id is 1 to 64 characters of `A-Z a-z 0-9 . _ -`, derived from the URL path (`/` is `HomePage`, `/cart.html` is `CartHtmlPage`).
|
|
37
|
+
It also names the page object: `pages/<PageId>.ts`.
|
|
38
|
+
- `url` is the page's URL without query string or fragment when it came from a recording.
|
|
39
|
+
`title` is the document title.
|
|
40
|
+
- `elements` maps an element key to its fingerprint.
|
|
41
|
+
|
|
42
|
+
### Element keys
|
|
43
|
+
|
|
44
|
+
A key is the element's test id when it has one (the project's `testIdAttribute`, `data-testid` by default).
|
|
45
|
+
Otherwise it is `<role>:<name-slug>` (for example `button:create-account`), and failing that `<tag>:<index>`.
|
|
46
|
+
A repeated key gets a `#2`, `#3` suffix.
|
|
47
|
+
The key is what a recorded step, a page object getter and a heal all refer to.
|
|
48
|
+
|
|
49
|
+
### Fingerprints
|
|
50
|
+
|
|
51
|
+
Every field is optional; the matcher weighs whatever is present.
|
|
52
|
+
The six `_canon` fields are the element's identity: they are Unicode-normalised, lowercased and whitespace-collapsed, and an exact heal is a match on them.
|
|
53
|
+
The other fields feed the fuzzy match that heals an element whose identity changed.
|
|
54
|
+
|
|
55
|
+
| Field | Meaning |
|
|
56
|
+
|---|---|
|
|
57
|
+
| `tag_canon` | tag name |
|
|
58
|
+
| `role_canon` | ARIA role, explicit or implicit |
|
|
59
|
+
| `accessible_name_canon` | accessible name (aria-label, label, text, placeholder, alt, title) |
|
|
60
|
+
| `testid_canon` | value of the project's test-id attribute |
|
|
61
|
+
| `name_attr_canon` | the `name` attribute |
|
|
62
|
+
| `id_canon` | the `id` attribute |
|
|
63
|
+
| `id_is_dynamic` | the id looks generated; such an id is left out of the identity |
|
|
64
|
+
| `visible_text` | rendered text |
|
|
65
|
+
| `neighbor_text` | text next to the element |
|
|
66
|
+
| `distinguishing_text` | text that tells repeated elements (table rows, cards) apart |
|
|
67
|
+
| `class_tokens` | class names, as a sorted list |
|
|
68
|
+
| `href`, `alt`, `title`, `type_attr`, `placeholder`, `aria_description` | the attributes of the same name |
|
|
69
|
+
| `xpath`, `id_xpath` | structural paths, when captured |
|
|
70
|
+
| `container_key`, `ordinal` | the repeated container it sits in, and its position there |
|
|
71
|
+
| `bbox` | `[x, y, width, height]` at capture time |
|
|
72
|
+
|
|
73
|
+
## Export and import
|
|
74
|
+
|
|
75
|
+
```bash
|
|
76
|
+
ordel catalog export catalog.json # write the portable file
|
|
77
|
+
ordel catalog import catalog.json # merge it into this project's catalog
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
An export is the catalog wrapped with what an importer needs to trust it:
|
|
81
|
+
|
|
82
|
+
```json
|
|
83
|
+
{
|
|
84
|
+
"format": "ordel.catalog",
|
|
85
|
+
"schema_version": 1,
|
|
86
|
+
"exported_at": "2026-09-28T02:40:13+00:00",
|
|
87
|
+
"test_id_attribute": "data-testid",
|
|
88
|
+
"pages": { "HomePage": { "url": "...", "title": "...", "elements": { } } }
|
|
89
|
+
}
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
Import merges and never overwrites.
|
|
93
|
+
A page this catalog lacks is added.
|
|
94
|
+
On a page it already has, new elements are added and an element it already has keeps its local fingerprint.
|
|
95
|
+
Each such conflict is listed, so you can re-capture the page if the import was the fresher one.
|
|
96
|
+
Every imported fingerprint is re-read through the same parser a local capture uses, so unknown fields are dropped.
|
|
97
|
+
Pages with an invalid id are skipped and listed.
|
|
98
|
+
If the export was captured with a different `test_id_attribute` than this project uses, the import says so, because test-id locators may then not match.
|
ordel_cli-0.3.0/PKG-INFO
ADDED
|
@@ -0,0 +1,175 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: ordel-cli
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: Ordel free CLI — local QA-automation for individual devs. Drives the deterministic ordel-engine and exposes it to a BYO coding agent over MCP. No account, no DB, no Ordel LLM.
|
|
5
|
+
Project-URL: Homepage, https://ordel.io
|
|
6
|
+
Author: Ordel
|
|
7
|
+
License-Expression: Apache-2.0
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Keywords: coding-agent,mcp,page-object,playwright,qa,self-healing,testing
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Environment :: Console
|
|
12
|
+
Classifier: Intended Audience :: Developers
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
14
|
+
Classifier: Topic :: Software Development :: Testing
|
|
15
|
+
Requires-Python: >=3.11
|
|
16
|
+
Requires-Dist: mcp<2.0,>=1.0
|
|
17
|
+
Requires-Dist: ordel-engine==0.3.0
|
|
18
|
+
Requires-Dist: ordel-tools==0.3.0
|
|
19
|
+
Requires-Dist: typer>=0.12
|
|
20
|
+
Provides-Extra: dev
|
|
21
|
+
Requires-Dist: hatchling; extra == 'dev'
|
|
22
|
+
Requires-Dist: pytest>=8; extra == 'dev'
|
|
23
|
+
Description-Content-Type: text/markdown
|
|
24
|
+
|
|
25
|
+
# ordel (free CLI)
|
|
26
|
+
|
|
27
|
+
Local QA automation for individual devs — a persistent, deterministic "QA brain" your
|
|
28
|
+
own coding agent (Claude Code / Cursor / Copilot) drives over MCP. **No account, no DB,
|
|
29
|
+
no Ordel LLM.** Your agent is the brain; Ordel is the memory + determinism.
|
|
30
|
+
|
|
31
|
+
> Status: **early build**. The engine loop (map → record → generate → run) works
|
|
32
|
+
> locally today; a single-command `npx ordel` distribution is planned but not started.
|
|
33
|
+
|
|
34
|
+
## What works today (local, anonymous)
|
|
35
|
+
|
|
36
|
+
```bash
|
|
37
|
+
ordel init # create .ordel/ + an ORDEL.md bridge file. No account.
|
|
38
|
+
ordel doctor # preflight: Node / npx / @playwright/test / Chromium (+ how to fix)
|
|
39
|
+
ordel status # local coverage: pages, elements, run history
|
|
40
|
+
ordel graph # print the local app graph
|
|
41
|
+
ordel adopt # existing Playwright/Cypress suite -> flake report + heal patches
|
|
42
|
+
ordel record <name> <url> # click through a flow yourself -> page objects + a spec using them
|
|
43
|
+
ordel run # run the Playwright suite; one evidence-backed verdict per test
|
|
44
|
+
ordel report # copy-pasteable Markdown report (Slack/Teams/Jira); --json too
|
|
45
|
+
ordel catalog export catalog.json # the element catalog as a portable file (CATALOG.md)
|
|
46
|
+
ordel catalog import catalog.json # merge one in; local elements win, conflicts listed
|
|
47
|
+
ordel mcp-install # print the MCP config to add to Claude Code / Cursor / Copilot
|
|
48
|
+
ordel mcp-install --skill # also install the agent skill (Claude Code, Codex)
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
### Recording a flow yourself
|
|
52
|
+
|
|
53
|
+
`ordel record signup http://localhost:3000/` opens a visible browser at that URL. Click
|
|
54
|
+
and type through the flow as a user would, then close the window to finish (Ctrl+C
|
|
55
|
+
works too, and keeps what was recorded).
|
|
56
|
+
|
|
57
|
+
To assert something, Alt+click it (Option+click on a Mac). That records
|
|
58
|
+
`toHaveText(<its text>)`, or `toBeVisible()` when it has no text. The final URL is
|
|
59
|
+
asserted automatically. A recording with no assertion runs `unjudged`, because it
|
|
60
|
+
proves nothing.
|
|
61
|
+
|
|
62
|
+
Ordel writes back:
|
|
63
|
+
- every page you landed on, into the catalog;
|
|
64
|
+
- a page object per page (`pages/<Class>.ts`, or its managed block updated);
|
|
65
|
+
- `tests/signup.spec.ts`, which drives those page objects (`homePage.submit.click()`).
|
|
66
|
+
|
|
67
|
+
A hand-written page object is never overwritten: steps on its page use a direct locator
|
|
68
|
+
instead, and the command says so. Passwords are never stored. The spec reads them from
|
|
69
|
+
`ORDEL_PASSWORD`, and query strings (where a GET form puts its fields) are dropped from
|
|
70
|
+
every recorded URL.
|
|
71
|
+
|
|
72
|
+
### Adopting an existing suite
|
|
73
|
+
|
|
74
|
+
`ordel adopt` points Ordel at the suite you already have (Playwright, or Cypress `*.cy.*`
|
|
75
|
+
specs) and reports on it without changing it:
|
|
76
|
+
|
|
77
|
+
- **Ingest:** every functional spec becomes a flow in `.ordel/flows/`. Cypress chains
|
|
78
|
+
are translated to the same flow format; a chain it cannot represent faithfully is
|
|
79
|
+
counted, never guessed.
|
|
80
|
+
- **Flake report:** it runs a Playwright suite until `.ordel/runs/` holds `--runs`
|
|
81
|
+
(default 3) judged runs, then classifies each test as `flaky` (passed only on retry,
|
|
82
|
+
or flipped twice), `regressed` (passed, then kept failing), `failing`, `recovered`,
|
|
83
|
+
`stable`, or too few runs to say.
|
|
84
|
+
- **Heal report:** each failing test whose test-id locator no longer resolves is healed
|
|
85
|
+
against a live capture of the page it was catalogued on. The page must have been
|
|
86
|
+
captured while the test was green: run `ordel explore <url>` first to build that
|
|
87
|
+
baseline. A confident heal becomes a unified diff under `.ordel/patches/<heal-id>.patch`.
|
|
88
|
+
An element that is simply gone is reported as a real failure, not healed to something else.
|
|
89
|
+
|
|
90
|
+
Nothing in your code changes until you review a patch:
|
|
91
|
+
|
|
92
|
+
```bash
|
|
93
|
+
ordel adopt --apply heal-0001 # write the fix, mark the heal accepted, refresh the catalog
|
|
94
|
+
ordel adopt --reject heal-0001 # write nothing, record the rejection
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
`--apply` refuses a patch whose file changed since it was made. Agents get the same
|
|
98
|
+
report through the `adopt_suite` MCP tool. They cannot apply or reject a heal: that is
|
|
99
|
+
the human review step.
|
|
100
|
+
|
|
101
|
+
### Honest verdicts
|
|
102
|
+
|
|
103
|
+
`ordel run` (and the `run_test` MCP tool) never reports a green it cannot back. Each
|
|
104
|
+
executed test gets one verdict: `pass`, `fail`, `flaky`, `skipped` or **`unjudged`**,
|
|
105
|
+
with its evidence: the assertions it actually ran and any heal it depends on.
|
|
106
|
+
|
|
107
|
+
A test that Playwright calls green is `unjudged`, not passed, when:
|
|
108
|
+
- it ran no assertion at all (a vacuous pass),
|
|
109
|
+
- no assertion evidence was captured for it,
|
|
110
|
+
- its spec depends on a heal no human has reviewed yet.
|
|
111
|
+
|
|
112
|
+
A run that executed no test is `no_tests`, never a pass. `ordel run --json` prints the
|
|
113
|
+
full `ordel.run/v1` document, which is also recorded under `.ordel/runs/`.
|
|
114
|
+
|
|
115
|
+
| Exit code | Meaning |
|
|
116
|
+
|---|---|
|
|
117
|
+
| 0 | at least one test ran, and every one passed with evidence |
|
|
118
|
+
| 1 | a test failed, no test ran, or the run itself errored |
|
|
119
|
+
| 3 | nothing failed, but at least one test is unjudged |
|
|
120
|
+
|
|
121
|
+
Then your coding agent, via MCP, drives the loop (needs Node + `@playwright/test` in the
|
|
122
|
+
project — run `ordel doctor` to check):
|
|
123
|
+
- `get_app_context` / `heal_selector` — what Ordel knows + **deterministic self-heal**
|
|
124
|
+
of a broken selector (fingerprint match, no LLM; ambiguous cases return `needs_agent`
|
|
125
|
+
with ranked candidates for your agent's *own* LLM to resolve — Ordel never spends inference).
|
|
126
|
+
Every resolved heal is recorded in `.ordel/heals.json` pending human review.
|
|
127
|
+
- `explore` / `record_flow` — drive the real browser to map the app + record a flow.
|
|
128
|
+
- `generate_scenarios` / `generate_invariant` / `generate_perf_check` / `generate_pom` —
|
|
129
|
+
turn artifacts into runnable specs (happy/negative/boundary, data-integrity, latency, POM).
|
|
130
|
+
- `run_test` — run a spec via `npx playwright test` and read the evidence-backed verdicts.
|
|
131
|
+
- QA-mind planning: `coverage_report` / `risk_rank` / `plan_tests` / `regression_set`.
|
|
132
|
+
|
|
133
|
+
## In progress (honest — no fake success)
|
|
134
|
+
|
|
135
|
+
- Single `npx ordel` distribution — the plan is a compiled Python engine binary wrapped
|
|
136
|
+
in one npm package; not started (today's install path is `pip`/`uv`, see Dev below).
|
|
137
|
+
- `eject` (one-command raw-Playwright export) — scaffold; but there's no lock-in today
|
|
138
|
+
either: generated tests are already plain `tests/*.spec.ts` + `pages/*.ts` on disk.
|
|
139
|
+
- Team sync + hosted dashboard = the paid upgrade (this CLI stays free & local).
|
|
140
|
+
|
|
141
|
+
## The pieces
|
|
142
|
+
|
|
143
|
+
- **`ordel-engine`** (sibling package) — the pure deterministic core: fingerprint
|
|
144
|
+
matching + self-heal, stdlib-only, zero backend.
|
|
145
|
+
- **`ordel_cli.store`** — the `.ordel/` file store (the local shell).
|
|
146
|
+
- **`ordel_cli.heal_service`** — heal + per-page circuit-breaker.
|
|
147
|
+
- **`ordel_cli.mcp_server`** — the local stdio MCP server your agent connects to.
|
|
148
|
+
|
|
149
|
+
## Dev
|
|
150
|
+
|
|
151
|
+
```bash
|
|
152
|
+
pip install -e packages/ordel-engine -e packages/ordel-cli
|
|
153
|
+
pip install pytest
|
|
154
|
+
pytest packages/ordel-cli/tests packages/ordel-engine/tests -q
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
## Known limitations (by design)
|
|
158
|
+
|
|
159
|
+
- **Single-writer.** The `.ordel/` store is for one dev on one machine. Writes are
|
|
160
|
+
*atomic* (temp-file + `os.replace`, so a crash can't corrupt a file), and a corrupted
|
|
161
|
+
`graph.json` is reported cleanly (never silently overwritten). But two processes
|
|
162
|
+
writing the *same* page concurrently is last-writer-wins — there is no file lock. That
|
|
163
|
+
is deliberate: a single-user local CLI doesn't warrant lock files / their failure
|
|
164
|
+
modes. Team-scale concurrency is the hosted product's job (`ordel push`).
|
|
165
|
+
- **Browser tools need a local Node + `@playwright/test`** (`explore`/`record_flow`/`run_test`).
|
|
166
|
+
They don't ship a browser; run `ordel doctor` — if Node/Playwright/Chromium are missing it
|
|
167
|
+
tells you the exact command to fix. Without them these tools report the missing dependency,
|
|
168
|
+
never fake success.
|
|
169
|
+
|
|
170
|
+
## Privacy
|
|
171
|
+
|
|
172
|
+
Local-first: **no Ordel LLM, no telemetry, no account, no data sent to Ordel.** The only
|
|
173
|
+
network requests are to your own target app and standard package/browser downloads you
|
|
174
|
+
initiate. Full policy: https://ordel.io/privacy (also ships as `PRIVACY.md` next to the
|
|
175
|
+
installed package, and at the repo root here).
|
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
# ordel (free CLI)
|
|
2
|
+
|
|
3
|
+
Local QA automation for individual devs — a persistent, deterministic "QA brain" your
|
|
4
|
+
own coding agent (Claude Code / Cursor / Copilot) drives over MCP. **No account, no DB,
|
|
5
|
+
no Ordel LLM.** Your agent is the brain; Ordel is the memory + determinism.
|
|
6
|
+
|
|
7
|
+
> Status: **early build**. The engine loop (map → record → generate → run) works
|
|
8
|
+
> locally today; a single-command `npx ordel` distribution is planned but not started.
|
|
9
|
+
|
|
10
|
+
## What works today (local, anonymous)
|
|
11
|
+
|
|
12
|
+
```bash
|
|
13
|
+
ordel init # create .ordel/ + an ORDEL.md bridge file. No account.
|
|
14
|
+
ordel doctor # preflight: Node / npx / @playwright/test / Chromium (+ how to fix)
|
|
15
|
+
ordel status # local coverage: pages, elements, run history
|
|
16
|
+
ordel graph # print the local app graph
|
|
17
|
+
ordel adopt # existing Playwright/Cypress suite -> flake report + heal patches
|
|
18
|
+
ordel record <name> <url> # click through a flow yourself -> page objects + a spec using them
|
|
19
|
+
ordel run # run the Playwright suite; one evidence-backed verdict per test
|
|
20
|
+
ordel report # copy-pasteable Markdown report (Slack/Teams/Jira); --json too
|
|
21
|
+
ordel catalog export catalog.json # the element catalog as a portable file (CATALOG.md)
|
|
22
|
+
ordel catalog import catalog.json # merge one in; local elements win, conflicts listed
|
|
23
|
+
ordel mcp-install # print the MCP config to add to Claude Code / Cursor / Copilot
|
|
24
|
+
ordel mcp-install --skill # also install the agent skill (Claude Code, Codex)
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
### Recording a flow yourself
|
|
28
|
+
|
|
29
|
+
`ordel record signup http://localhost:3000/` opens a visible browser at that URL. Click
|
|
30
|
+
and type through the flow as a user would, then close the window to finish (Ctrl+C
|
|
31
|
+
works too, and keeps what was recorded).
|
|
32
|
+
|
|
33
|
+
To assert something, Alt+click it (Option+click on a Mac). That records
|
|
34
|
+
`toHaveText(<its text>)`, or `toBeVisible()` when it has no text. The final URL is
|
|
35
|
+
asserted automatically. A recording with no assertion runs `unjudged`, because it
|
|
36
|
+
proves nothing.
|
|
37
|
+
|
|
38
|
+
Ordel writes back:
|
|
39
|
+
- every page you landed on, into the catalog;
|
|
40
|
+
- a page object per page (`pages/<Class>.ts`, or its managed block updated);
|
|
41
|
+
- `tests/signup.spec.ts`, which drives those page objects (`homePage.submit.click()`).
|
|
42
|
+
|
|
43
|
+
A hand-written page object is never overwritten: steps on its page use a direct locator
|
|
44
|
+
instead, and the command says so. Passwords are never stored. The spec reads them from
|
|
45
|
+
`ORDEL_PASSWORD`, and query strings (where a GET form puts its fields) are dropped from
|
|
46
|
+
every recorded URL.
|
|
47
|
+
|
|
48
|
+
### Adopting an existing suite
|
|
49
|
+
|
|
50
|
+
`ordel adopt` points Ordel at the suite you already have (Playwright, or Cypress `*.cy.*`
|
|
51
|
+
specs) and reports on it without changing it:
|
|
52
|
+
|
|
53
|
+
- **Ingest:** every functional spec becomes a flow in `.ordel/flows/`. Cypress chains
|
|
54
|
+
are translated to the same flow format; a chain it cannot represent faithfully is
|
|
55
|
+
counted, never guessed.
|
|
56
|
+
- **Flake report:** it runs a Playwright suite until `.ordel/runs/` holds `--runs`
|
|
57
|
+
(default 3) judged runs, then classifies each test as `flaky` (passed only on retry,
|
|
58
|
+
or flipped twice), `regressed` (passed, then kept failing), `failing`, `recovered`,
|
|
59
|
+
`stable`, or too few runs to say.
|
|
60
|
+
- **Heal report:** each failing test whose test-id locator no longer resolves is healed
|
|
61
|
+
against a live capture of the page it was catalogued on. The page must have been
|
|
62
|
+
captured while the test was green: run `ordel explore <url>` first to build that
|
|
63
|
+
baseline. A confident heal becomes a unified diff under `.ordel/patches/<heal-id>.patch`.
|
|
64
|
+
An element that is simply gone is reported as a real failure, not healed to something else.
|
|
65
|
+
|
|
66
|
+
Nothing in your code changes until you review a patch:
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
ordel adopt --apply heal-0001 # write the fix, mark the heal accepted, refresh the catalog
|
|
70
|
+
ordel adopt --reject heal-0001 # write nothing, record the rejection
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
`--apply` refuses a patch whose file changed since it was made. Agents get the same
|
|
74
|
+
report through the `adopt_suite` MCP tool. They cannot apply or reject a heal: that is
|
|
75
|
+
the human review step.
|
|
76
|
+
|
|
77
|
+
### Honest verdicts
|
|
78
|
+
|
|
79
|
+
`ordel run` (and the `run_test` MCP tool) never reports a green it cannot back. Each
|
|
80
|
+
executed test gets one verdict: `pass`, `fail`, `flaky`, `skipped` or **`unjudged`**,
|
|
81
|
+
with its evidence: the assertions it actually ran and any heal it depends on.
|
|
82
|
+
|
|
83
|
+
A test that Playwright calls green is `unjudged`, not passed, when:
|
|
84
|
+
- it ran no assertion at all (a vacuous pass),
|
|
85
|
+
- no assertion evidence was captured for it,
|
|
86
|
+
- its spec depends on a heal no human has reviewed yet.
|
|
87
|
+
|
|
88
|
+
A run that executed no test is `no_tests`, never a pass. `ordel run --json` prints the
|
|
89
|
+
full `ordel.run/v1` document, which is also recorded under `.ordel/runs/`.
|
|
90
|
+
|
|
91
|
+
| Exit code | Meaning |
|
|
92
|
+
|---|---|
|
|
93
|
+
| 0 | at least one test ran, and every one passed with evidence |
|
|
94
|
+
| 1 | a test failed, no test ran, or the run itself errored |
|
|
95
|
+
| 3 | nothing failed, but at least one test is unjudged |
|
|
96
|
+
|
|
97
|
+
Then your coding agent, via MCP, drives the loop (needs Node + `@playwright/test` in the
|
|
98
|
+
project — run `ordel doctor` to check):
|
|
99
|
+
- `get_app_context` / `heal_selector` — what Ordel knows + **deterministic self-heal**
|
|
100
|
+
of a broken selector (fingerprint match, no LLM; ambiguous cases return `needs_agent`
|
|
101
|
+
with ranked candidates for your agent's *own* LLM to resolve — Ordel never spends inference).
|
|
102
|
+
Every resolved heal is recorded in `.ordel/heals.json` pending human review.
|
|
103
|
+
- `explore` / `record_flow` — drive the real browser to map the app + record a flow.
|
|
104
|
+
- `generate_scenarios` / `generate_invariant` / `generate_perf_check` / `generate_pom` —
|
|
105
|
+
turn artifacts into runnable specs (happy/negative/boundary, data-integrity, latency, POM).
|
|
106
|
+
- `run_test` — run a spec via `npx playwright test` and read the evidence-backed verdicts.
|
|
107
|
+
- QA-mind planning: `coverage_report` / `risk_rank` / `plan_tests` / `regression_set`.
|
|
108
|
+
|
|
109
|
+
## In progress (honest — no fake success)
|
|
110
|
+
|
|
111
|
+
- Single `npx ordel` distribution — the plan is a compiled Python engine binary wrapped
|
|
112
|
+
in one npm package; not started (today's install path is `pip`/`uv`, see Dev below).
|
|
113
|
+
- `eject` (one-command raw-Playwright export) — scaffold; but there's no lock-in today
|
|
114
|
+
either: generated tests are already plain `tests/*.spec.ts` + `pages/*.ts` on disk.
|
|
115
|
+
- Team sync + hosted dashboard = the paid upgrade (this CLI stays free & local).
|
|
116
|
+
|
|
117
|
+
## The pieces
|
|
118
|
+
|
|
119
|
+
- **`ordel-engine`** (sibling package) — the pure deterministic core: fingerprint
|
|
120
|
+
matching + self-heal, stdlib-only, zero backend.
|
|
121
|
+
- **`ordel_cli.store`** — the `.ordel/` file store (the local shell).
|
|
122
|
+
- **`ordel_cli.heal_service`** — heal + per-page circuit-breaker.
|
|
123
|
+
- **`ordel_cli.mcp_server`** — the local stdio MCP server your agent connects to.
|
|
124
|
+
|
|
125
|
+
## Dev
|
|
126
|
+
|
|
127
|
+
```bash
|
|
128
|
+
pip install -e packages/ordel-engine -e packages/ordel-cli
|
|
129
|
+
pip install pytest
|
|
130
|
+
pytest packages/ordel-cli/tests packages/ordel-engine/tests -q
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
## Known limitations (by design)
|
|
134
|
+
|
|
135
|
+
- **Single-writer.** The `.ordel/` store is for one dev on one machine. Writes are
|
|
136
|
+
*atomic* (temp-file + `os.replace`, so a crash can't corrupt a file), and a corrupted
|
|
137
|
+
`graph.json` is reported cleanly (never silently overwritten). But two processes
|
|
138
|
+
writing the *same* page concurrently is last-writer-wins — there is no file lock. That
|
|
139
|
+
is deliberate: a single-user local CLI doesn't warrant lock files / their failure
|
|
140
|
+
modes. Team-scale concurrency is the hosted product's job (`ordel push`).
|
|
141
|
+
- **Browser tools need a local Node + `@playwright/test`** (`explore`/`record_flow`/`run_test`).
|
|
142
|
+
They don't ship a browser; run `ordel doctor` — if Node/Playwright/Chromium are missing it
|
|
143
|
+
tells you the exact command to fix. Without them these tools report the missing dependency,
|
|
144
|
+
never fake success.
|
|
145
|
+
|
|
146
|
+
## Privacy
|
|
147
|
+
|
|
148
|
+
Local-first: **no Ordel LLM, no telemetry, no account, no data sent to Ordel.** The only
|
|
149
|
+
network requests are to your own target app and standard package/browser downloads you
|
|
150
|
+
initiate. Full policy: https://ordel.io/privacy (also ships as `PRIVACY.md` next to the
|
|
151
|
+
installed package, and at the repo root here).
|
|
@@ -0,0 +1,91 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ordel-cli
|
|
3
|
+
description: Use when testing this repo's web app with Ordel's local MCP server (`ordel-cli`) - mapping pages, recording flows, generating Playwright specs, running them, adopting an existing Playwright or Cypress suite, or healing a broken locator. Ordel decides correctness; report its verdicts, never your own reading.
|
|
4
|
+
---
|
|
5
|
+
<!-- ordel:skill - written by `ordel mcp-install --skill`; re-run it to update. Delete this line to keep your own edits. -->
|
|
6
|
+
|
|
7
|
+
# Ordel: testing with the local engine
|
|
8
|
+
|
|
9
|
+
Tests here are backed by **Ordel**, a deterministic QA engine over MCP (zero Ordel LLM).
|
|
10
|
+
**Division of labor:** YOU navigate and judge ambiguity; **Ordel decides correctness.** Don't
|
|
11
|
+
eyeball whether the app "looks right" — route every correctness claim through an Ordel oracle
|
|
12
|
+
(a generated spec's pass/fail, an invariant, a fingerprint match). Agent-authored assertions
|
|
13
|
+
are notoriously weak; Ordel's are deterministic. Call `get_app_context` first.
|
|
14
|
+
|
|
15
|
+
> First run: enable the `ordel-cli` MCP server once — the client won't expose Ordel's tools
|
|
16
|
+
> until you approve it. Claude Code: `/mcp` → enable. Cursor / Copilot: enable it in the
|
|
17
|
+
> MCP/tools settings where you added `.mcp.json`. After a NEW Ordel tool ships, reconnect the
|
|
18
|
+
> server to pick it up (or use the equivalent `ordel <cmd>` CLI, which is always current).
|
|
19
|
+
|
|
20
|
+
## Verdicts: report only what is proven
|
|
21
|
+
`run_test` (and `ordel run`) returns one verdict per test with its evidence:
|
|
22
|
+
|
|
23
|
+
| verdict | meaning | how to report it |
|
|
24
|
+
|---|---|---|
|
|
25
|
+
| `pass` | ran, and its assertions (listed in `evidence.checks`) held | passed |
|
|
26
|
+
| `fail` | ran and failed; `message` says why | a defect, or a broken locator to heal |
|
|
27
|
+
| `flaky` | failed, then passed on retry | passed, and name it as flaky |
|
|
28
|
+
| `unjudged` | green in Playwright but NOT proven: no assertion ran, no evidence was captured, or it depends on a heal no human has reviewed (`reasons` says which) | **not proven**: never "passed" |
|
|
29
|
+
| `skipped` | did not run | skipped |
|
|
30
|
+
|
|
31
|
+
A run whose `status` is `no_tests` checked nothing: say so, never "all green". Before you call
|
|
32
|
+
a new test done, run it and confirm its verdict is `pass` with the checks you intended.
|
|
33
|
+
|
|
34
|
+
## The method — don't stop at "mapped", FIND BUGS
|
|
35
|
+
Mapping is the MEANS; the value is USING the map to run tests that expose defects.
|
|
36
|
+
|
|
37
|
+
orient ─▶ map ─▶ risk_rank ─▶ plan_tests ─▶ per-surface: record ─▶ generate_* ─▶ run_test
|
|
38
|
+
▲ (checklists) │
|
|
39
|
+
└───────────── regression_set ◀── coverage_report ◀───────────────────────────┘
|
|
40
|
+
|
|
41
|
+
1. **Orient** — `assess_context` (run-mode, lifecycle, existing assets + the REQUIRED next phase;
|
|
42
|
+
it detects any pre-existing FUNCTIONAL suite to reuse and flags unit tests as out of scope —
|
|
43
|
+
Ordel builds functional/e2e, not units) then `get_app_context` (what's already mapped).
|
|
44
|
+
An existing Playwright/Cypress suite? `adopt_suite` returns its flake report and heal
|
|
45
|
+
patches; show the human the patches, they apply or reject them (`ordel adopt --apply`).
|
|
46
|
+
2. **Map** — YOU crawl, Ordel captures (see below). Reach every surface, not just the landing.
|
|
47
|
+
3. **Prioritize** — `risk_rank` (where a bug would hurt most).
|
|
48
|
+
4. **Plan** — `plan_tests` classifies each page into a SURFACE (auth / form / table / search /
|
|
49
|
+
cart / detail) and returns that surface's **CHECKLIST** — the exact things to test, each
|
|
50
|
+
item naming the tool that does it. WORK THE CHECKLIST; don't improvise coverage from memory.
|
|
51
|
+
5. **Execute each checklist item** with the tool it names:
|
|
52
|
+
- `generate_scenarios <flow>` — happy + NEGATIVE (per constrained field) + boundary + replay.
|
|
53
|
+
A FAILING negative = a real bug (the app accepted bad input). RUN it; don't just read it.
|
|
54
|
+
- `generate_invariant <flow> parts total` — data-integrity (sum|min|max|count == total);
|
|
55
|
+
catches what URL tests can't (a discount applied twice: total != sum of line items).
|
|
56
|
+
- `generate_perf_check <flow> budget_ms` — median render <= budget (latency cliffs).
|
|
57
|
+
- `run_test` — execute and judge. This is where a regression is CAUGHT.
|
|
58
|
+
6. **Measure + maintain** — `coverage_report` (map ∩ specs, real gaps), `regression_set
|
|
59
|
+
<changed pages>` (after a change, the exact specs to re-run, not all).
|
|
60
|
+
|
|
61
|
+
The generated specs already classify gate-vs-filter, test each constrained field in isolation,
|
|
62
|
+
and default the run target from the recorded flow — trust their verdicts over your own reading.
|
|
63
|
+
|
|
64
|
+
## Map the app (YOU crawl, Ordel captures — a human can co-drive)
|
|
65
|
+
Reach auth-gated, wizard, and SPA states that have no unique URL:
|
|
66
|
+
- `browser_open` (once) -> `goto`/`click`/`fill`/`select`/`check` to navigate.
|
|
67
|
+
- **`browser_open(headed=true)`** opens a visible window that persists between calls, so a
|
|
68
|
+
HUMAN can watch, join, or take over the same session — they handle a CAPTCHA / a step that
|
|
69
|
+
needs judgment, you continue and `capture_state`. Offer headed when a human is available.
|
|
70
|
+
- `capture_state` on any page worth mapping (fingerprints the current DOM — whoever drove).
|
|
71
|
+
- `dialog(action=accept|dismiss, text)` BEFORE a click that raises an alert/confirm/prompt;
|
|
72
|
+
Ordel records what fired (click/capture responses carry a `dialog` field — never swallowed).
|
|
73
|
+
- Log in once, then `save_auth`; later `browser_open(use_auth=true)` skips login.
|
|
74
|
+
- `browser_close` when done.
|
|
75
|
+
- A human can record a flow themselves: `ordel record <name> <url>` (they click, Alt+click to
|
|
76
|
+
assert, close the window). It writes page objects and a spec that drives them.
|
|
77
|
+
|
|
78
|
+
## Heals: you propose, a human decides
|
|
79
|
+
`heal_selector`/`heal_element` recover a broken selector by fingerprint; pass `spec`, the file
|
|
80
|
+
you'll change. Every heal is recorded pending HUMAN review, and until then `run_test` reports
|
|
81
|
+
that spec's tests as `unjudged`. An `ambiguous`/`needs_agent` answer means pick by hand; a
|
|
82
|
+
`failed`/`gone` answer means the element is really gone: report a real failure, don't hunt
|
|
83
|
+
for a lookalike. You cannot apply or reject a heal yourself.
|
|
84
|
+
|
|
85
|
+
## Other tools
|
|
86
|
+
`generate_pom` (page objects from the map), `record_flow` (persist a step sequence + a
|
|
87
|
+
happy-path spec; add assertions, or it runs `unjudged`), `list_runs`. Prefer
|
|
88
|
+
getByRole/getByTestId; read the map before writing selectors, and heal (don't rewrite) when
|
|
89
|
+
one breaks.
|
|
90
|
+
|
|
91
|
+
Responses carry a `◉◯ ordel` signature — surface it so the work is attributed.
|