ordel-cli 0.2.2__tar.gz → 0.3.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. ordel_cli-0.3.1/.gitignore +5 -0
  2. ordel_cli-0.3.1/CATALOG.md +107 -0
  3. ordel_cli-0.3.1/PKG-INFO +182 -0
  4. ordel_cli-0.3.1/README.md +158 -0
  5. ordel_cli-0.3.1/ordel_cli/SKILL.md +104 -0
  6. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/__init__.py +1 -1
  7. ordel_cli-0.3.1/ordel_cli/adopt.py +326 -0
  8. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/capture.cjs +3 -1
  9. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/capture_extract.cjs +15 -3
  10. ordel_cli-0.3.1/ordel_cli/catalog.py +111 -0
  11. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/cli.py +356 -108
  12. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/context.py +5 -2
  13. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/coverage.py +1 -1
  14. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/doctor.py +47 -17
  15. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/explore.py +75 -12
  16. ordel_cli-0.3.1/ordel_cli/human_record.cjs +196 -0
  17. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/ingest.py +21 -7
  18. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/mcp_server.py +281 -75
  19. ordel_cli-0.3.1/ordel_cli/ordel_config.cjs +26 -0
  20. ordel_cli-0.3.1/ordel_cli/ordel_evidence.cjs +81 -0
  21. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/pom.py +6 -6
  22. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/qa_plan.py +13 -5
  23. ordel_cli-0.3.1/ordel_cli/record.cjs +256 -0
  24. ordel_cli-0.3.1/ordel_cli/record.py +428 -0
  25. ordel_cli-0.3.1/ordel_cli/runner.py +270 -0
  26. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/scenarios.py +19 -5
  27. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/session.py +37 -20
  28. ordel_cli-0.3.1/ordel_cli/session_driver.cjs +598 -0
  29. ordel_cli-0.3.1/ordel_cli/store.py +390 -0
  30. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/pyproject.toml +20 -5
  31. ordel_cli-0.2.2/.gitignore +0 -175
  32. ordel_cli-0.2.2/PKG-INFO +0 -97
  33. ordel_cli-0.2.2/README.md +0 -74
  34. ordel_cli-0.2.2/ordel_cli/record.cjs +0 -143
  35. ordel_cli-0.2.2/ordel_cli/record.py +0 -135
  36. ordel_cli-0.2.2/ordel_cli/runner.py +0 -137
  37. ordel_cli-0.2.2/ordel_cli/session_driver.cjs +0 -273
  38. ordel_cli-0.2.2/ordel_cli/store.py +0 -222
  39. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/LICENSE +0 -0
  40. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/PRIVACY.md +0 -0
  41. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/branding.py +0 -0
  42. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/heal_service.py +0 -0
  43. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/mcp_supervisor.cjs +0 -0
  44. {ordel_cli-0.2.2 → ordel_cli-0.3.1}/ordel_cli/py.typed +0 -0
@@ -0,0 +1,5 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ .venv/
4
+ dist/
5
+ *.egg-info/
@@ -0,0 +1,107 @@
1
+ # The Ordel element catalog
2
+
3
+ The catalog is what Ordel knows about your app's pages and elements.
4
+ It is the baseline a heal compares against, the source page objects are generated from, and the denominator coverage is measured against.
5
+ It lives in `.ordel/graph.json`, and it is plain JSON you can commit, diff and review.
6
+
7
+ ## Where it comes from
8
+
9
+ - `ordel explore <url>` (or the `explore_app` MCP tool) captures one page.
10
+ - `capture_state` captures whatever page an agent-driven session is on.
11
+ - `ordel record` captures every page a person lands on while recording.
12
+ - `record_flow` adds a thin, element-less page for each URL a flow visits.
13
+ - `ordel adopt --apply` re-captures the page of a heal a human accepted, so the next break heals from the app as it is now.
14
+
15
+ ## `graph.json`
16
+
17
+ ```json
18
+ {
19
+ "schema_version": 1,
20
+ "pages": {
21
+ "HomePage": {
22
+ "url": "http://localhost:3000/",
23
+ "title": "Sign up",
24
+ "elements": {
25
+ "email": { "tag_canon": "input", "role_canon": "textbox", "testid_canon": "email", "...": "..." }
26
+ }
27
+ }
28
+ }
29
+ }
30
+ ```
31
+
32
+ - `schema_version` is the format version, currently `1`.
33
+ A file without it is version `0`, which has the same shape and is read as-is.
34
+ A file with a version newer than the installed `ordel-cli` understands is refused with an error, never guessed at.
35
+ - `pages` maps a page id to the page.
36
+ A page id is 1 to 64 characters of `A-Z a-z 0-9 . _ -`, derived from the URL path (`/` is `HomePage`, `/cart.html` is `CartHtmlPage`).
37
+ It also names the page object: `pages/<PageId>.ts`.
38
+ - `url` is the page's URL without query string or fragment when it came from a recording.
39
+ `title` is the document title.
40
+ - `elements` maps an element key to its fingerprint.
41
+
42
+ ### Element keys
43
+
44
+ A key is the element's test id when it has one (the project's `testIdAttribute`, `data-testid` by default).
45
+ Otherwise it is `<role>:<name-slug>` (for example `button:create-account`), and failing that `<tag>:<index>`.
46
+ A repeated key gets a `#2`, `#3` suffix.
47
+ The key is what a recorded step, a page object getter and a heal all refer to.
48
+
49
+ ### Fingerprints
50
+
51
+ Every field is optional; the matcher weighs whatever is present.
52
+ The six `_canon` fields are the element's identity: they are Unicode-normalised, lowercased and whitespace-collapsed, and an exact heal is a match on them.
53
+ The other fields feed the fuzzy match that heals an element whose identity changed.
54
+
55
+ | Field | Meaning |
56
+ |---|---|
57
+ | `tag_canon` | tag name |
58
+ | `role_canon` | ARIA role, explicit or implicit |
59
+ | `accessible_name_canon` | accessible name (aria-label, label, text, placeholder, alt, title) |
60
+ | `testid_canon` | value of the project's test-id attribute |
61
+ | `name_attr_canon` | the `name` attribute |
62
+ | `id_canon` | the `id` attribute |
63
+ | `id_is_dynamic` | the id looks generated; such an id is left out of the identity |
64
+ | `visible_text` | rendered text |
65
+ | `neighbor_text` | text next to the element |
66
+ | `distinguishing_text` | text that tells repeated elements (table rows, cards) apart |
67
+ | `class_tokens` | class names, as a sorted list |
68
+ | `href`, `alt`, `title`, `type_attr`, `placeholder`, `aria_description` | the attributes of the same name |
69
+ | `xpath`, `id_xpath` | structural paths, when captured |
70
+ | `container_key`, `ordinal` | the repeated container it sits in, and its position there |
71
+ | `bbox` | `[x, y, width, height]` at capture time |
72
+
73
+ Beside the fingerprint, an element keeps three exact-case fields that the matcher ignores and `get_page` returns.
74
+ A catalog captured before they existed simply lacks them.
75
+
76
+ | Field | Meaning |
77
+ |---|---|
78
+ | `name_attr` | the `name` attribute as written |
79
+ | `testid` | the test id as written |
80
+ | `options` | a `<select>`'s options as `[{"value", "text"}]`, at most 100 |
81
+
82
+ ## Export and import
83
+
84
+ ```bash
85
+ ordel catalog export catalog.json # write the portable file
86
+ ordel catalog import catalog.json # merge it into this project's catalog
87
+ ```
88
+
89
+ An export is the catalog wrapped with what an importer needs to trust it:
90
+
91
+ ```json
92
+ {
93
+ "format": "ordel.catalog",
94
+ "schema_version": 1,
95
+ "exported_at": "2026-09-28T02:40:13+00:00",
96
+ "test_id_attribute": "data-testid",
97
+ "pages": { "HomePage": { "url": "...", "title": "...", "elements": { } } }
98
+ }
99
+ ```
100
+
101
+ Import merges and never overwrites.
102
+ A page this catalog lacks is added.
103
+ On a page it already has, new elements are added and an element it already has keeps its local fingerprint.
104
+ Each such conflict is listed, so you can re-capture the page if the import was the fresher one.
105
+ Every imported fingerprint is re-read through the same parser a local capture uses, so unknown fields are dropped.
106
+ Pages with an invalid id are skipped and listed.
107
+ If the export was captured with a different `test_id_attribute` than this project uses, the import says so, because test-id locators may then not match.
@@ -0,0 +1,182 @@
1
+ Metadata-Version: 2.5
2
+ Name: ordel-cli
3
+ Version: 0.3.1
4
+ Summary: Ordel free CLI — local QA-automation for individual devs. Drives the deterministic ordel-engine and exposes it to a BYO coding agent over MCP. No account, no DB, no Ordel LLM.
5
+ Project-URL: Homepage, https://ordel.io
6
+ Author: Ordel
7
+ License-Expression: Apache-2.0
8
+ License-File: LICENSE
9
+ Keywords: coding-agent,mcp,page-object,playwright,qa,self-healing,testing
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Environment :: Console
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Programming Language :: Python :: 3.11
14
+ Classifier: Topic :: Software Development :: Testing
15
+ Requires-Python: >=3.11
16
+ Requires-Dist: mcp<2.0,>=1.0
17
+ Requires-Dist: ordel-engine==0.3.1
18
+ Requires-Dist: ordel-tools==0.3.1
19
+ Requires-Dist: typer>=0.12
20
+ Provides-Extra: dev
21
+ Requires-Dist: hatchling; extra == 'dev'
22
+ Requires-Dist: pytest>=8; extra == 'dev'
23
+ Description-Content-Type: text/markdown
24
+
25
+ # ordel (free CLI)
26
+
27
+ Local QA automation for individual devs — a persistent, deterministic "QA brain" your
28
+ own coding agent (Claude Code / Cursor / Copilot) drives over MCP. **No account, no DB,
29
+ no Ordel LLM.** Your agent is the brain; Ordel is the memory + determinism.
30
+
31
+ > Status: **early build**. The engine loop (map → record → generate → run) works
32
+ > locally today; a single-command `npx ordel` distribution is planned but not started.
33
+
34
+ ## What works today (local, anonymous)
35
+
36
+ ```bash
37
+ ordel init # create .ordel/ + an ORDEL.md bridge file. No account.
38
+ ordel doctor # preflight: Node / npx / @playwright/test / Chromium (+ how to fix)
39
+ ordel status # local coverage: pages, elements, run history
40
+ ordel graph # print the local app graph
41
+ ordel adopt # existing Playwright/Cypress suite -> flake report + heal patches
42
+ ordel record <name> <url> # click through a flow yourself -> page objects + a spec using them
43
+ ordel run # run the Playwright suite; one evidence-backed verdict per test
44
+ ordel report # copy-pasteable Markdown report (Slack/Teams/Jira); --json too
45
+ ordel catalog export catalog.json # the element catalog as a portable file (CATALOG.md)
46
+ ordel catalog import catalog.json # merge one in; local elements win, conflicts listed
47
+ ordel mcp-install # print the MCP config to add to Claude Code / Cursor / Copilot
48
+ ordel mcp-install --skill # also install the agent skill (Claude Code, Codex)
49
+ ```
50
+
51
+ ### Recording a flow yourself
52
+
53
+ `ordel record signup http://localhost:3000/` opens a visible browser at that URL. Click
54
+ and type through the flow as a user would, then close the window to finish (Ctrl+C
55
+ works too, and keeps what was recorded).
56
+
57
+ To assert something, Alt+click it (Option+click on a Mac). That records
58
+ `toHaveText(<its text>)`, or `toBeVisible()` when it has no text. The final URL is
59
+ asserted automatically. A recording with no assertion runs `unjudged`, because it
60
+ proves nothing.
61
+
62
+ Ordel writes back:
63
+ - every page you landed on, into the catalog;
64
+ - a page object per page (`pages/<Class>.ts`, or its managed block updated);
65
+ - `tests/signup.spec.ts`, which drives those page objects (`homePage.submit.click()`).
66
+
67
+ A hand-written page object is never overwritten: steps on its page use a direct locator
68
+ instead, and the command says so. Passwords are never stored. The spec reads them from
69
+ `ORDEL_PASSWORD`, and query strings (where a GET form puts its fields) are dropped from
70
+ every recorded URL.
71
+
72
+ ### Adopting an existing suite
73
+
74
+ `ordel adopt` points Ordel at the suite you already have (Playwright, or Cypress `*.cy.*`
75
+ specs) and reports on it without changing it:
76
+
77
+ - **Ingest:** every functional spec becomes a flow in `.ordel/flows/`. Cypress chains
78
+ are translated to the same flow format; a chain it cannot represent faithfully is
79
+ counted, never guessed.
80
+ - **Flake report:** it runs a Playwright suite until `.ordel/runs/` holds `--runs`
81
+ (default 3) judged runs, then classifies each test as `flaky` (passed only on retry,
82
+ or flipped twice), `regressed` (passed, then kept failing), `failing`, `recovered`,
83
+ `stable`, or too few runs to say.
84
+ - **Heal report:** each failing test whose test-id locator no longer resolves is healed
85
+ against a live capture of the page it was catalogued on. The page must have been
86
+ captured while the test was green: run `ordel explore <url>` first to build that
87
+ baseline. A confident heal becomes a unified diff under `.ordel/patches/<heal-id>.patch`.
88
+ An element that is simply gone is reported as a real failure, not healed to something else.
89
+
90
+ Nothing in your code changes until you review a patch:
91
+
92
+ ```bash
93
+ ordel adopt --apply heal-0001 # write the fix, mark the heal accepted, refresh the catalog
94
+ ordel adopt --reject heal-0001 # write nothing, record the rejection
95
+ ```
96
+
97
+ `--apply` refuses a patch whose file changed since it was made. Agents get the same
98
+ report through the `adopt_suite` MCP tool. They cannot apply or reject a heal: that is
99
+ the human review step.
100
+
101
+ ### Honest verdicts
102
+
103
+ `ordel run` (and the `run_test` MCP tool) never reports a green it cannot back. Each
104
+ executed test gets one verdict: `pass`, `fail`, `flaky`, `skipped` or **`unjudged`**,
105
+ with its evidence: the assertions it actually ran and any heal it depends on.
106
+
107
+ A test that Playwright calls green is `unjudged`, not passed, when:
108
+ - it ran no assertion at all (a vacuous pass),
109
+ - no assertion evidence was captured for it,
110
+ - its spec depends on a heal no human has reviewed yet.
111
+
112
+ A run that executed no test is `no_tests`, never a pass. `ordel run --json` prints the
113
+ full `ordel.run/v1` document, which is also recorded under `.ordel/runs/`.
114
+
115
+ | Exit code | Meaning |
116
+ |---|---|
117
+ | 0 | at least one test ran, and every one passed with evidence |
118
+ | 1 | a test failed, no test ran, or the run itself errored |
119
+ | 3 | nothing failed, but at least one test is unjudged |
120
+
121
+ Then your coding agent, via MCP, drives the loop (needs Node + `@playwright/test` in the
122
+ project — run `ordel doctor` to check):
123
+ - `get_app_context` / `heal_selector` — what Ordel knows + **deterministic self-heal**
124
+ of a broken selector (fingerprint match, no LLM; ambiguous cases return `needs_agent`
125
+ with ranked candidates for your agent's *own* LLM to resolve — Ordel never spends inference).
126
+ Every resolved heal is recorded in `.ordel/heals.json` pending human review.
127
+ - `explore` / `record_flow` — drive the real browser to map the app + record a flow.
128
+ A value typed into a password-like field (type=password, or named/labelled password,
129
+ secret, token or api key; a step can also say `"secret": true`) is used for the recording
130
+ and never written to disk: the flow keeps a reference, and every generated spec reads it
131
+ from an environment variable such as `ORDEL_SECRET_PASSWORD`, listed in the result's
132
+ `needs_env`. Set it before running the spec; unset, the test fails naming it.
133
+ - `get_page` — read a mapped page back (inputs, buttons, links or all): each element's
134
+ test id, name, type, placeholder, href, `<select>` options and visible text, exact case.
135
+ - `generate_scenarios` / `generate_invariant` / `generate_perf_check` / `generate_pom` —
136
+ turn artifacts into runnable specs (happy/negative/boundary, data-integrity, latency, POM).
137
+ - `run_test` — run a spec via `npx playwright test` and read the evidence-backed verdicts.
138
+ - QA-mind planning: `coverage_report` / `risk_rank` / `plan_tests` / `regression_set`.
139
+
140
+ ## In progress (honest — no fake success)
141
+
142
+ - Single `npx ordel` distribution — the plan is a compiled Python engine binary wrapped
143
+ in one npm package; not started (today's install path is `pip`/`uv`, see Dev below).
144
+ - `eject` (one-command raw-Playwright export) — scaffold; but there's no lock-in today
145
+ either: generated tests are already plain `tests/*.spec.ts` + `pages/*.ts` on disk.
146
+ - Team sync + hosted dashboard = the paid upgrade (this CLI stays free & local).
147
+
148
+ ## The pieces
149
+
150
+ - **`ordel-engine`** (sibling package) — the pure deterministic core: fingerprint
151
+ matching + self-heal, stdlib-only, zero backend.
152
+ - **`ordel_cli.store`** — the `.ordel/` file store (the local shell).
153
+ - **`ordel_cli.heal_service`** — heal + per-page circuit-breaker.
154
+ - **`ordel_cli.mcp_server`** — the local stdio MCP server your agent connects to.
155
+
156
+ ## Dev
157
+
158
+ ```bash
159
+ pip install -e packages/ordel-engine -e packages/ordel-cli
160
+ pip install pytest
161
+ pytest packages/ordel-cli/tests packages/ordel-engine/tests -q
162
+ ```
163
+
164
+ ## Known limitations (by design)
165
+
166
+ - **Single-writer.** The `.ordel/` store is for one dev on one machine. Writes are
167
+ *atomic* (temp-file + `os.replace`, so a crash can't corrupt a file), and a corrupted
168
+ `graph.json` is reported cleanly (never silently overwritten). But two processes
169
+ writing the *same* page concurrently is last-writer-wins — there is no file lock. That
170
+ is deliberate: a single-user local CLI doesn't warrant lock files / their failure
171
+ modes. Team-scale concurrency is the hosted product's job (`ordel push`).
172
+ - **Browser tools need a local Node + `@playwright/test`** (`explore`/`record_flow`/`run_test`).
173
+ They don't ship a browser; run `ordel doctor` — if Node/Playwright/Chromium are missing it
174
+ tells you the exact command to fix. Without them these tools report the missing dependency,
175
+ never fake success.
176
+
177
+ ## Privacy
178
+
179
+ Local-first: **no Ordel LLM, no telemetry, no account, no data sent to Ordel.** The only
180
+ network requests are to your own target app and standard package/browser downloads you
181
+ initiate. Full policy: https://ordel.io/privacy (also ships as `PRIVACY.md` next to the
182
+ installed package, and at the repo root here).
@@ -0,0 +1,158 @@
1
+ # ordel (free CLI)
2
+
3
+ Local QA automation for individual devs — a persistent, deterministic "QA brain" your
4
+ own coding agent (Claude Code / Cursor / Copilot) drives over MCP. **No account, no DB,
5
+ no Ordel LLM.** Your agent is the brain; Ordel is the memory + determinism.
6
+
7
+ > Status: **early build**. The engine loop (map → record → generate → run) works
8
+ > locally today; a single-command `npx ordel` distribution is planned but not started.
9
+
10
+ ## What works today (local, anonymous)
11
+
12
+ ```bash
13
+ ordel init # create .ordel/ + an ORDEL.md bridge file. No account.
14
+ ordel doctor # preflight: Node / npx / @playwright/test / Chromium (+ how to fix)
15
+ ordel status # local coverage: pages, elements, run history
16
+ ordel graph # print the local app graph
17
+ ordel adopt # existing Playwright/Cypress suite -> flake report + heal patches
18
+ ordel record <name> <url> # click through a flow yourself -> page objects + a spec using them
19
+ ordel run # run the Playwright suite; one evidence-backed verdict per test
20
+ ordel report # copy-pasteable Markdown report (Slack/Teams/Jira); --json too
21
+ ordel catalog export catalog.json # the element catalog as a portable file (CATALOG.md)
22
+ ordel catalog import catalog.json # merge one in; local elements win, conflicts listed
23
+ ordel mcp-install # print the MCP config to add to Claude Code / Cursor / Copilot
24
+ ordel mcp-install --skill # also install the agent skill (Claude Code, Codex)
25
+ ```
26
+
27
+ ### Recording a flow yourself
28
+
29
+ `ordel record signup http://localhost:3000/` opens a visible browser at that URL. Click
30
+ and type through the flow as a user would, then close the window to finish (Ctrl+C
31
+ works too, and keeps what was recorded).
32
+
33
+ To assert something, Alt+click it (Option+click on a Mac). That records
34
+ `toHaveText(<its text>)`, or `toBeVisible()` when it has no text. The final URL is
35
+ asserted automatically. A recording with no assertion runs `unjudged`, because it
36
+ proves nothing.
37
+
38
+ Ordel writes back:
39
+ - every page you landed on, into the catalog;
40
+ - a page object per page (`pages/<Class>.ts`, or its managed block updated);
41
+ - `tests/signup.spec.ts`, which drives those page objects (`homePage.submit.click()`).
42
+
43
+ A hand-written page object is never overwritten: steps on its page use a direct locator
44
+ instead, and the command says so. Passwords are never stored. The spec reads them from
45
+ `ORDEL_PASSWORD`, and query strings (where a GET form puts its fields) are dropped from
46
+ every recorded URL.
47
+
48
+ ### Adopting an existing suite
49
+
50
+ `ordel adopt` points Ordel at the suite you already have (Playwright, or Cypress `*.cy.*`
51
+ specs) and reports on it without changing it:
52
+
53
+ - **Ingest:** every functional spec becomes a flow in `.ordel/flows/`. Cypress chains
54
+ are translated to the same flow format; a chain it cannot represent faithfully is
55
+ counted, never guessed.
56
+ - **Flake report:** it runs a Playwright suite until `.ordel/runs/` holds `--runs`
57
+ (default 3) judged runs, then classifies each test as `flaky` (passed only on retry,
58
+ or flipped twice), `regressed` (passed, then kept failing), `failing`, `recovered`,
59
+ `stable`, or too few runs to say.
60
+ - **Heal report:** each failing test whose test-id locator no longer resolves is healed
61
+ against a live capture of the page it was catalogued on. The page must have been
62
+ captured while the test was green: run `ordel explore <url>` first to build that
63
+ baseline. A confident heal becomes a unified diff under `.ordel/patches/<heal-id>.patch`.
64
+ An element that is simply gone is reported as a real failure, not healed to something else.
65
+
66
+ Nothing in your code changes until you review a patch:
67
+
68
+ ```bash
69
+ ordel adopt --apply heal-0001 # write the fix, mark the heal accepted, refresh the catalog
70
+ ordel adopt --reject heal-0001 # write nothing, record the rejection
71
+ ```
72
+
73
+ `--apply` refuses a patch whose file changed since it was made. Agents get the same
74
+ report through the `adopt_suite` MCP tool. They cannot apply or reject a heal: that is
75
+ the human review step.
76
+
77
+ ### Honest verdicts
78
+
79
+ `ordel run` (and the `run_test` MCP tool) never reports a green it cannot back. Each
80
+ executed test gets one verdict: `pass`, `fail`, `flaky`, `skipped` or **`unjudged`**,
81
+ with its evidence: the assertions it actually ran and any heal it depends on.
82
+
83
+ A test that Playwright calls green is `unjudged`, not passed, when:
84
+ - it ran no assertion at all (a vacuous pass),
85
+ - no assertion evidence was captured for it,
86
+ - its spec depends on a heal no human has reviewed yet.
87
+
88
+ A run that executed no test is `no_tests`, never a pass. `ordel run --json` prints the
89
+ full `ordel.run/v1` document, which is also recorded under `.ordel/runs/`.
90
+
91
+ | Exit code | Meaning |
92
+ |---|---|
93
+ | 0 | at least one test ran, and every one passed with evidence |
94
+ | 1 | a test failed, no test ran, or the run itself errored |
95
+ | 3 | nothing failed, but at least one test is unjudged |
96
+
97
+ Then your coding agent, via MCP, drives the loop (needs Node + `@playwright/test` in the
98
+ project — run `ordel doctor` to check):
99
+ - `get_app_context` / `heal_selector` — what Ordel knows + **deterministic self-heal**
100
+ of a broken selector (fingerprint match, no LLM; ambiguous cases return `needs_agent`
101
+ with ranked candidates for your agent's *own* LLM to resolve — Ordel never spends inference).
102
+ Every resolved heal is recorded in `.ordel/heals.json` pending human review.
103
+ - `explore` / `record_flow` — drive the real browser to map the app + record a flow.
104
+ A value typed into a password-like field (type=password, or named/labelled password,
105
+ secret, token or api key; a step can also say `"secret": true`) is used for the recording
106
+ and never written to disk: the flow keeps a reference, and every generated spec reads it
107
+ from an environment variable such as `ORDEL_SECRET_PASSWORD`, listed in the result's
108
+ `needs_env`. Set it before running the spec; unset, the test fails naming it.
109
+ - `get_page` — read a mapped page back (inputs, buttons, links or all): each element's
110
+ test id, name, type, placeholder, href, `<select>` options and visible text, exact case.
111
+ - `generate_scenarios` / `generate_invariant` / `generate_perf_check` / `generate_pom` —
112
+ turn artifacts into runnable specs (happy/negative/boundary, data-integrity, latency, POM).
113
+ - `run_test` — run a spec via `npx playwright test` and read the evidence-backed verdicts.
114
+ - QA-mind planning: `coverage_report` / `risk_rank` / `plan_tests` / `regression_set`.
115
+
116
+ ## In progress (honest — no fake success)
117
+
118
+ - Single `npx ordel` distribution — the plan is a compiled Python engine binary wrapped
119
+ in one npm package; not started (today's install path is `pip`/`uv`, see Dev below).
120
+ - `eject` (one-command raw-Playwright export) — scaffold; but there's no lock-in today
121
+ either: generated tests are already plain `tests/*.spec.ts` + `pages/*.ts` on disk.
122
+ - Team sync + hosted dashboard = the paid upgrade (this CLI stays free & local).
123
+
124
+ ## The pieces
125
+
126
+ - **`ordel-engine`** (sibling package) — the pure deterministic core: fingerprint
127
+ matching + self-heal, stdlib-only, zero backend.
128
+ - **`ordel_cli.store`** — the `.ordel/` file store (the local shell).
129
+ - **`ordel_cli.heal_service`** — heal + per-page circuit-breaker.
130
+ - **`ordel_cli.mcp_server`** — the local stdio MCP server your agent connects to.
131
+
132
+ ## Dev
133
+
134
+ ```bash
135
+ pip install -e packages/ordel-engine -e packages/ordel-cli
136
+ pip install pytest
137
+ pytest packages/ordel-cli/tests packages/ordel-engine/tests -q
138
+ ```
139
+
140
+ ## Known limitations (by design)
141
+
142
+ - **Single-writer.** The `.ordel/` store is for one dev on one machine. Writes are
143
+ *atomic* (temp-file + `os.replace`, so a crash can't corrupt a file), and a corrupted
144
+ `graph.json` is reported cleanly (never silently overwritten). But two processes
145
+ writing the *same* page concurrently is last-writer-wins — there is no file lock. That
146
+ is deliberate: a single-user local CLI doesn't warrant lock files / their failure
147
+ modes. Team-scale concurrency is the hosted product's job (`ordel push`).
148
+ - **Browser tools need a local Node + `@playwright/test`** (`explore`/`record_flow`/`run_test`).
149
+ They don't ship a browser; run `ordel doctor` — if Node/Playwright/Chromium are missing it
150
+ tells you the exact command to fix. Without them these tools report the missing dependency,
151
+ never fake success.
152
+
153
+ ## Privacy
154
+
155
+ Local-first: **no Ordel LLM, no telemetry, no account, no data sent to Ordel.** The only
156
+ network requests are to your own target app and standard package/browser downloads you
157
+ initiate. Full policy: https://ordel.io/privacy (also ships as `PRIVACY.md` next to the
158
+ installed package, and at the repo root here).
@@ -0,0 +1,104 @@
1
+ ---
2
+ name: ordel-cli
3
+ description: Use when testing this repo's web app with Ordel's local MCP server (`ordel-cli`) - mapping pages, recording flows, generating Playwright specs, running them, adopting an existing Playwright or Cypress suite, or healing a broken locator. Ordel decides correctness; report its verdicts, never your own reading.
4
+ ---
5
+ <!-- ordel:skill - written by `ordel mcp-install --skill`; re-run it to update. Delete this line to keep your own edits. -->
6
+
7
+ # Ordel: testing with the local engine
8
+
9
+ Tests here are backed by **Ordel**, a deterministic QA engine over MCP (zero Ordel LLM).
10
+ **Division of labor:** YOU navigate and judge ambiguity; **Ordel decides correctness.** Don't
11
+ eyeball whether the app "looks right" — route every correctness claim through an Ordel oracle
12
+ (a generated spec's pass/fail, an invariant, a fingerprint match). Agent-authored assertions
13
+ are notoriously weak; Ordel's are deterministic. Call `get_app_context` first.
14
+
15
+ > First run: enable the `ordel-cli` MCP server once — the client won't expose Ordel's tools
16
+ > until you approve it. Claude Code: `/mcp` → enable. Cursor / Copilot: enable it in the
17
+ > MCP/tools settings where you added `.mcp.json`. After a NEW Ordel tool ships, reconnect the
18
+ > server to pick it up (or use the equivalent `ordel <cmd>` CLI, which is always current).
19
+
20
+ ## Verdicts: report only what is proven
21
+ `run_test` (and `ordel run`) returns one verdict per test with its evidence:
22
+
23
+ | verdict | meaning | how to report it |
24
+ |---|---|---|
25
+ | `pass` | ran, and its assertions (listed in `evidence.checks`) held | passed |
26
+ | `fail` | ran and failed; `message` says why | a defect, or a broken locator to heal |
27
+ | `flaky` | failed, then passed on retry | passed, and name it as flaky |
28
+ | `unjudged` | green in Playwright but NOT proven: no assertion ran, no evidence was captured, or it depends on a heal no human has reviewed (`reasons` says which) | **not proven**: never "passed" |
29
+ | `skipped` | did not run | skipped |
30
+
31
+ A run whose `status` is `no_tests` checked nothing: say so, never "all green". Before you call
32
+ a new test done, run it and confirm its verdict is `pass` with the checks you intended.
33
+
34
+ ## The method — don't stop at "mapped", FIND BUGS
35
+ Mapping is the MEANS; the value is USING the map to run tests that expose defects.
36
+
37
+ orient ─▶ map ─▶ risk_rank ─▶ plan_tests ─▶ per-surface: record ─▶ generate_* ─▶ run_test
38
+ ▲ (checklists) │
39
+ └───────────── regression_set ◀── coverage_report ◀───────────────────────────┘
40
+
41
+ 1. **Orient** — `assess_context` (run-mode, lifecycle, existing assets + the REQUIRED next phase;
42
+ it detects any pre-existing FUNCTIONAL suite to reuse and flags unit tests as out of scope —
43
+ Ordel builds functional/e2e, not units) then `get_app_context` (what's already mapped).
44
+ An existing Playwright/Cypress suite? `adopt_suite` returns its flake report and heal
45
+ patches; show the human the patches, they apply or reject them (`ordel adopt --apply`).
46
+ 2. **Map** — YOU crawl, Ordel captures (see below). Reach every surface, not just the landing.
47
+ 3. **Prioritize** — `risk_rank` (where a bug would hurt most).
48
+ 4. **Plan** — `plan_tests` classifies each page into a SURFACE (auth / form / table / search /
49
+ cart / detail) and returns that surface's **CHECKLIST** — the exact things to test, each
50
+ item naming the tool that does it. WORK THE CHECKLIST; don't improvise coverage from memory.
51
+ 5. **Execute each checklist item** with the tool it names:
52
+ - `generate_scenarios <flow>` — happy + NEGATIVE (per constrained field) + boundary + replay.
53
+ A FAILING negative = a real bug (the app accepted bad input). RUN it; don't just read it.
54
+ - `generate_invariant <flow> parts total` — data-integrity (sum|min|max|count == total);
55
+ catches what URL tests can't (a discount applied twice: total != sum of line items).
56
+ - `generate_perf_check <flow> budget_ms` — median render <= budget (latency cliffs).
57
+ - `run_test` — execute and judge. This is where a regression is CAUGHT.
58
+ 6. **Measure + maintain** — `coverage_report` (map ∩ specs, real gaps), `regression_set
59
+ <changed pages>` (after a change, the exact specs to re-run, not all).
60
+
61
+ The generated specs already classify gate-vs-filter, test each constrained field in isolation,
62
+ and default the run target from the recorded flow — trust their verdicts over your own reading.
63
+
64
+ ## Map the app (YOU crawl, Ordel captures — a human can co-drive)
65
+ Reach auth-gated, wizard, and SPA states that have no unique URL:
66
+ - `browser_open` (once) -> `goto`/`click`/`fill`/`select`/`check` to navigate.
67
+ - Read each action's `outcome` before judging it: `summary` in one line, `submit_blocked` +
68
+ `invalid_fields` (the browser refused the form, with each field's message), `reloaded` (a
69
+ same-url server re-render, so it WAS submitted), `messages` (new alerts/toasts),
70
+ `network_errors` (4xx/5xx during the action), `appeared`/`disappeared` test ids.
71
+ "No visible change" is itself a finding when you expected one.
72
+ - **`browser_open(headed=true)`** opens a visible window that persists between calls, so a
73
+ HUMAN can watch, join, or take over the same session — they handle a CAPTCHA / a step that
74
+ needs judgment, you continue and `capture_state`. Offer headed when a human is available.
75
+ - `capture_state` on any page worth mapping (fingerprints the current DOM — whoever drove).
76
+ - `get_page(page_id, filter=inputs|buttons|links|all)` reads a mapped page back: each
77
+ element's testid, name, type, placeholder, href, `<select>` options (`value` is what
78
+ `select` takes) and visible text in its exact case. Use it; never parse `.ordel/graph.json`.
79
+ - `dialog(action=accept|dismiss, text)` BEFORE a click that raises an alert/confirm/prompt;
80
+ Ordel records what fired (click/capture responses carry a `dialog` field — never swallowed).
81
+ - Log in once, then `save_auth`; later `browser_open(use_auth=true)` skips login.
82
+ - `browser_close` when done.
83
+ - A human can record a flow themselves: `ordel record <name> <url>` (they click, Alt+click to
84
+ assert, close the window). It writes page objects and a spec that drives them.
85
+
86
+ ## Heals: you propose, a human decides
87
+ `heal_selector`/`heal_element` recover a broken selector by fingerprint; pass `spec`, the file
88
+ you'll change. Every heal is recorded pending HUMAN review, and until then `run_test` reports
89
+ that spec's tests as `unjudged`. An `ambiguous`/`needs_agent` answer means pick by hand; a
90
+ `failed`/`gone` answer means the element is really gone: report a real failure, don't hunt
91
+ for a lookalike. You cannot apply or reject a heal yourself.
92
+
93
+ ## Other tools
94
+ `generate_pom` (page objects from the map), `record_flow` (persist a step sequence + a
95
+ happy-path spec that asserts the exact final url and a key element of the landing page; add
96
+ your own assertions for what the flow must prove), `list_runs`. A value `record_flow`
97
+ fills into a password-like field (type=password, or named/labelled password, secret, token,
98
+ api key; mark any other with `"secret": true`) is never written: the flow and every spec read
99
+ it from the variable in `needs_env` (`ORDEL_SECRET_<FIELD>`), so tell the user to set it
100
+ before `run_test`, and never paste the value into a spec yourself. Prefer
101
+ getByRole/getByTestId; read the map (`get_page`) before writing selectors, and heal (don't rewrite) when
102
+ one breaks.
103
+
104
+ Responses carry a `◉◯ ordel` signature — surface it so the work is attributed.
@@ -5,4 +5,4 @@ from __future__ import annotations
5
5
  from ordel_cli.heal_service import CircuitOpenError, HealService
6
6
  from ordel_cli.store import OrdelStore
7
7
 
8
- __all__ = ["OrdelStore", "HealService", "CircuitOpenError"]
8
+ __all__ = ["CircuitOpenError", "HealService", "OrdelStore"]