spectoflow 0.17.1 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: write-e2e-tests
3
- description: Author durable, CI-runnable end-to-end tests with Playwright — locators over selectors, web-first assertions, no hard waits.
3
+ description: Author durable, CI-runnable Playwright end-to-end tests — run headed, directly in the browser, by default; falls back to MCP, native browser tooling or headless, always telling the user why.
4
4
  capability: testing
5
5
  inputs: The user-visible flow(s) under test (spec acceptance criteria or a plan task), the running app under test, and its existing `tests/e2e/` suite and Playwright config if any.
6
6
  outputs: Committed Playwright spec files under `tests/e2e/*.spec.ts`, runnable in CI.
@@ -44,22 +44,59 @@ Practices below are current Playwright guidance (see References for exact source
44
44
  8. **Keep specs scoped to one flow each**, named for the behavior under test, and placed under
45
45
  `tests/e2e/*.spec.ts` in the user's project.
46
46
 
47
- **Driving the browser (live repro + test generation) — use the best available, in this order:**
48
- 1. **Playwright MCP** (`@playwright/mcp`, wired into the project's `.mcp.json` by `spectoflow init`):
49
- the **agent-agnostic** way to drive a real browser and **generate** a spec from a recorded flow.
50
- Works in any MCP client (Claude Code, Codex, Cursor, …). `npx` fetches it on first use — nothing to
51
- install; if it isn't wired yet, add it or run `spectoflow init` again (idempotent).
52
- 2. **The client's native browser tooling** (for Claude Code, the Chrome extension / `claude-in-chrome`)
53
- for live/exploratory checks when MCP isn't wired.
54
- 3. **Local Playwright** — `npx playwright codegen` / headed mode for a quick, throwaway look;
55
- `npx playwright install` provides the browsers. Needs `@playwright/test` as the project's devDependency.
56
- 4. **If no browser can run at all** (restricted CI, no browsers installed): still **write the durable
57
- spec** (the artifact that lasts), then raise a `need` / Attention item with the exact commands to
58
- enable it — never report a pass you couldn't actually observe.
59
-
60
- **Live/exploratory verification is not this skill's output.** Whichever rung above you're on, the
61
- durable, CI-runnable artifact this skill produces is always the Playwright **spec file**, not the live
62
- session — the live drive is only the means to write and check it.
47
+ ## Running the tests — headed, in the real browser, by default
48
+
49
+ **Default: Playwright lib, headed.** When this skill's own agent runs the suite — while authoring it,
50
+ verifying a flow, or investigating a failure — run it **directly in a visible browser window**:
51
+ `npx playwright test --headed`. Watching the browser act out the flow, rather than reading a bare
52
+ pass/fail line, is the whole point: it is how you (and the user, if watching) catch a flow that
53
+ "passes" for the wrong reason. This is the default **local run mode** — not a suggestion to try once
54
+ in a while.
55
+
56
+ **`--ui` mode for authoring and debugging.** Reach for `npx playwright test --ui` when writing a new
57
+ flow or chasing a failure: it steps through each action with time-travel, showing the DOM/network at
58
+ every point, and is the fastest way to build a flow interactively before locking it into a spec.
59
+
60
+ **CI stays headless — that is not a fallback, it is a different job.** The committed suite still runs
61
+ `npx playwright test` (no `--headed`) in the project's CI pipeline, per the Quality bar below: most CI
62
+ runners have no real display, and headless is faster and the industry-standard way to gate a merge.
63
+ The headed-by-default rule governs *this skill's own local run loop*, not the CI config it authors.
64
+
65
+ **Switch away from headed only when:**
66
+ - **the user explicitly asked for something else** (headless, `--ui`, a specific project/browser) —
67
+ honor it, no need to justify; or
68
+ - **headed genuinely cannot run** (no display / sandboxed or remote environment / browsers not
69
+ installed) — then step down the ladder below.
70
+
71
+ Whenever you step down for the second reason, **say so** — don't silently swap to a quieter mode:
72
+
73
+ ```
74
+ ::spectoflow role=testing kind=progress msg=Running headless — no display available in this environment (would default to --headed)
75
+ ```
76
+
77
+ or, if it blocks the task entirely, raise it as a `need` (see Output contract).
78
+
79
+ ### The fallback ladder
80
+ 1. **Playwright lib, headed** (default) — `npx playwright test --headed`. Needs `@playwright/test` as
81
+ the project's devDependency and `npx playwright install` for the browsers.
82
+ 2. **Playwright lib, `--ui`** — for interactively authoring or debugging one flow before committing it.
83
+ 3. **Playwright lib, headless** — same lib, just invisible: use when headed can't launch, or the user
84
+ asked for headless. Still the lib — nothing else changes.
85
+ 4. **Playwright MCP** (`@playwright/mcp`, wired into the project's `.mcp.json` by `spectoflow init`):
86
+ the **agent-agnostic** way to drive a real browser and **generate** a spec from a recorded flow —
87
+ reach for this to explore an unfamiliar app or reproduce a bug, or when the project has no
88
+ Playwright lib set up yet. Works in any MCP client (Claude Code, Codex, Cursor, …); `npx` fetches
89
+ the server on first use — nothing to install; if it isn't wired, add it or re-run `spectoflow init`
90
+ (idempotent).
91
+ 5. **The client's native browser tooling** (for Claude Code, the Chrome extension / `claude-in-chrome`)
92
+ for live/exploratory checks when neither the local lib nor MCP is usable.
93
+ 6. **If no browser can run at all** (restricted CI, no browsers installed, nothing wired): still
94
+ **write the durable spec** (the artifact that lasts), then raise a `need` / Attention item with the
95
+ exact commands to enable it — never report a pass you couldn't actually observe.
96
+
97
+ **Live/exploratory verification is not this skill's output.** Whichever rung you're on, the durable,
98
+ CI-runnable artifact this skill produces is always the Playwright **spec file**, not the live session —
99
+ the live browser (headed, `--ui`, MCP, or the extension) is only the means to write and check it.
63
100
 
64
101
  **Playwright is a dependency of the user's project, never of spectoflow.** This skill authors tests
65
102
  against whatever Playwright version the target project has (or proposes adding `@playwright/test` as a
@@ -78,6 +115,14 @@ project devDependency when none exists) — spectoflow itself stays at zero runt
78
115
  ::spectoflow role=testing kind=result msg=<spec file> <pass|fail> (<n> tests)
79
116
  ::spectoflow role=testing kind=done msg=<flow name> e2e suite committed at tests/e2e/<file>
80
117
  ```
118
+ - Any step down the fallback ladder away from **Playwright lib, headed** (the default) is reported the
119
+ moment it happens, with the reason — never a silent switch:
120
+
121
+ ```
122
+ ::spectoflow role=testing kind=progress msg=Running headless — <reason> (would default to --headed)
123
+ ::spectoflow role=testing kind=progress msg=Falling back to Playwright MCP — <reason>
124
+ ::spectoflow role=testing kind=need msg=No browser can run here — spec written, needs <exact command> to verify
125
+ ```
81
126
 
82
127
  ## Quality bar
83
128
  - [ ] Each spec asserts user-visible behavior (text/role/state), not internal implementation.
@@ -90,6 +135,10 @@ project devDependency when none exists) — spectoflow itself stays at zero runt
90
135
  - [ ] `trace: 'on-first-retry'` (or equivalent) is set so a CI failure is debuggable without tracing
91
136
  every run.
92
137
  - [ ] No spectoflow file declares Playwright as a dependency — only the user project's `package.json`.
138
+ - [ ] The local run defaulted to **Playwright lib, headed** (`--headed`), not headless, unless the user
139
+ asked otherwise or headed genuinely could not launch.
140
+ - [ ] Any step away from the headed default (headless, MCP, native browser tooling) was **announced**
141
+ with its reason via the `::spectoflow` sentinel — never a silent switch.
93
142
 
94
143
  ## References
95
144
  - Playwright — Best Practices — https://playwright.dev/docs/best-practices (test user-visible behavior;
@@ -107,3 +156,9 @@ project devDependency when none exists) — spectoflow itself stays at zero runt
107
156
  matching configuration).
108
157
  - Playwright — Trace viewer — https://playwright.dev/docs/trace-viewer (`trace: 'on-first-retry'` for
109
158
  CI).
159
+ - Playwright — Running and debugging tests — https://playwright.dev/docs/running-tests (headed mode is
160
+ the default way to watch a run locally).
161
+ - Playwright — UI Mode — https://playwright.dev/docs/test-ui-mode (`--ui`, time-travel debugging for
162
+ authoring and chasing failures).
163
+ - Playwright MCP — https://github.com/microsoft/playwright-mcp (agent-driven browser automation and
164
+ spec generation, agent-agnostic via MCP).