@muggleai/works 4.12.2 → 4.12.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{chunk-I4VLYJ7M.js → chunk-CPF6AR2I.js} +499 -147
- package/dist/{chunk-2DVZ2LYO.js → chunk-JNI7INIO.js} +2 -2
- package/dist/cli.js +2 -2
- package/dist/index.js +2 -2
- package/dist/plugin/.claude-plugin/plugin.json +1 -1
- package/dist/plugin/.cursor-plugin/plugin.json +1 -1
- package/dist/plugin/skills/_shared/dev-server-readiness.md +39 -0
- package/dist/plugin/skills/_shared/failure-mode-handling.md +19 -4
- package/dist/plugin/skills/_shared/pr-branch-worktree.md +31 -0
- package/dist/plugin/skills/muggle-test/SKILL.md +7 -3
- package/dist/plugin/skills/muggle-test-feature-local/SKILL.md +6 -1
- package/dist/plugin/skills/muggle-test-prepare/SKILL.md +40 -256
- package/dist/plugin/skills/muggle-test-prepare/steps/check-running.md +31 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/env-file.md +18 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/fresh-install.md +20 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/identify-services.md +28 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/readiness-report.md +26 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/rebase-check.md +3 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/scope.md +11 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/smoke-test.md +26 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/start-commands.md +29 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/start-services.md +21 -0
- package/dist/plugin/skills/muggle-test-prepare/steps/viability-check.md +21 -0
- package/dist/release-manifest.json +4 -4
- package/dist/{src-ARTTHWNP.js → src-YR5UKLPC.js} +1 -1
- package/package.json +6 -6
- package/plugin/.claude-plugin/plugin.json +1 -1
- package/plugin/.cursor-plugin/plugin.json +1 -1
- package/plugin/skills/_shared/dev-server-readiness.md +39 -0
- package/plugin/skills/_shared/failure-mode-handling.md +19 -4
- package/plugin/skills/_shared/pr-branch-worktree.md +31 -0
- package/plugin/skills/muggle-test/SKILL.md +7 -3
- package/plugin/skills/muggle-test-feature-local/SKILL.md +6 -1
- package/plugin/skills/muggle-test-prepare/SKILL.md +40 -256
- package/plugin/skills/muggle-test-prepare/steps/check-running.md +31 -0
- package/plugin/skills/muggle-test-prepare/steps/env-file.md +18 -0
- package/plugin/skills/muggle-test-prepare/steps/fresh-install.md +20 -0
- package/plugin/skills/muggle-test-prepare/steps/identify-services.md +28 -0
- package/plugin/skills/muggle-test-prepare/steps/readiness-report.md +26 -0
- package/plugin/skills/muggle-test-prepare/steps/rebase-check.md +3 -0
- package/plugin/skills/muggle-test-prepare/steps/scope.md +11 -0
- package/plugin/skills/muggle-test-prepare/steps/smoke-test.md +26 -0
- package/plugin/skills/muggle-test-prepare/steps/start-commands.md +29 -0
- package/plugin/skills/muggle-test-prepare/steps/start-services.md +21 -0
- package/plugin/skills/muggle-test-prepare/steps/viability-check.md +21 -0
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
# Environment file sanity
|
|
2
|
+
|
|
3
|
+
The env file is **per-repo** — don't hardcode `.env.local`. Detect by scanning `package.json` `scripts/*` for `.env*` literals and known port vars (`PORT=`, `VITE_PORT=`); check framework config (`next.config.*`, `vite.config.*`).
|
|
4
|
+
|
|
5
|
+
When a dependency on an env file exists:
|
|
6
|
+
|
|
7
|
+
1. Check `<cwd>/<envfile>` exists — if yes, no-op.
|
|
8
|
+
2. If absent, `git worktree list --porcelain` and check each sibling for the same filename.
|
|
9
|
+
3. If found:
|
|
10
|
+
|
|
11
|
+
> "`<envfile>` is missing in this worktree but exists at `<sibling>/<envfile>`. Copy it before starting services?"
|
|
12
|
+
|
|
13
|
+
- Option 1: "Yes — copy from `<sibling>`"
|
|
14
|
+
- Option 2: "No — I'll provide it another way"
|
|
15
|
+
|
|
16
|
+
4. If not found anywhere, report and ask.
|
|
17
|
+
|
|
18
|
+
Skip silently when no env file is referenced.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# Fresh install (clean-start default)
|
|
2
|
+
|
|
3
|
+
Detect stack by indicator file, decide if install is missing/stale, run automatically (notify, don't ask). Only opt-out is aborting the skill.
|
|
4
|
+
|
|
5
|
+
| Indicator | Stack | Stale check | Install command |
|
|
6
|
+
|:----------|:------|:------------|:----------------|
|
|
7
|
+
| `package.json` | Node | `node_modules/` missing OR `package-lock.json` newer than `node_modules/.package-lock.json` | `npm install --prefer-offline --no-audit --no-fund` |
|
|
8
|
+
| `pyproject.toml` w/ `[tool.poetry]` | Poetry | `poetry.lock` newer than `.venv/pyvenv.cfg` (or `.venv/` missing) | `poetry install --no-interaction` |
|
|
9
|
+
| `pyproject.toml` (PEP 621) + `uv.lock` | uv | `uv.lock` newer than `.venv/pyvenv.cfg` | `uv sync` |
|
|
10
|
+
| `requirements.txt` | pip | `requirements.txt` newer than `.venv/pyvenv.cfg` (or `.venv/` missing) | `pip install -r requirements.txt` |
|
|
11
|
+
| `Gemfile` | Bundler | `Gemfile.lock` newer than `vendor/bundle/` mtime | `bundle install` |
|
|
12
|
+
| `composer.json` | Composer | `composer.lock` newer than `vendor/autoload.php` | `composer install --no-interaction` |
|
|
13
|
+
| `pom.xml` | Maven | always (heavy — opt-in via `AskUserQuestion`) | `mvn -DskipTests install` |
|
|
14
|
+
| `build.gradle*` | Gradle | always (heavy — opt-in via `AskUserQuestion`) | `gradle build -x test` |
|
|
15
|
+
| `go.mod` | Go | skip — `go run`/`go build` handle deps | |
|
|
16
|
+
| `Cargo.toml` | Rust | skip — `cargo run`/`cargo build` handle deps | |
|
|
17
|
+
|
|
18
|
+
Notify one-liner: `Installing <service-name> (<stack>: <missing|stale>)…`.
|
|
19
|
+
|
|
20
|
+
**Never symlink dep dirs** (`node_modules/`, `.venv/`, `vendor/bundle/`) from a sibling worktree — webpack rewrites paths via `resolve.symlinks: true` and Python/Ruby tooling has analogous issues. Run a real per-worktree install.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Identify required services & startup mode
|
|
2
|
+
|
|
3
|
+
List folder names in the **parent directory** of the current working directory:
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
ls -d "$(dirname "$PWD")"/*/ | xargs -I{} basename {}
|
|
7
|
+
```
|
|
8
|
+
|
|
9
|
+
Present folder names only (not contents). `AskUserQuestion` with `multiSelect: true`:
|
|
10
|
+
|
|
11
|
+
> "Which of these need to be running for your tests?"
|
|
12
|
+
|
|
13
|
+
Pre-check the ones matching the testing scope from [scope](./scope.md). Always include:
|
|
14
|
+
- "Just the current project (no other services needed)"
|
|
15
|
+
- "None of these — I'll tell you what I need"
|
|
16
|
+
|
|
17
|
+
Include the current working directory as a candidate. If the user provides manual paths, verify they exist.
|
|
18
|
+
|
|
19
|
+
**Immediately after selection**, ask startup mode:
|
|
20
|
+
|
|
21
|
+
> "How do you want to handle these?"
|
|
22
|
+
|
|
23
|
+
- Option 1: "Check what's running, start what's missing for me"
|
|
24
|
+
- Option 2: "I'll start them myself — just verify they're up when I'm done"
|
|
25
|
+
|
|
26
|
+
**Option 2**: skip [start-commands](./start-commands.md), [fresh-install](./fresh-install.md), [start-services](./start-services.md). Wait for ready signal, then [check-running](./check-running.md), run [smoke-test](./smoke-test.md) against everything (the user-started case is exactly where the smoke test matters most), then [readiness-report](./readiness-report.md).
|
|
27
|
+
|
|
28
|
+
**Option 1**: proceed through the normal flow.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Final readiness report
|
|
2
|
+
|
|
3
|
+
Only after every service passes [smoke-test](./smoke-test.md) or is skipped.
|
|
4
|
+
|
|
5
|
+
```
|
|
6
|
+
Service PID Port Status Smoke Test
|
|
7
|
+
─────────────────────────────────────────────────────────────────
|
|
8
|
+
backend-api 12345 3001 Running ✓
|
|
9
|
+
auth-service 12346 8080 Running ✓
|
|
10
|
+
frontend 12347 3000 Running ✓
|
|
11
|
+
─────────────────────────────────────────────────────────────────
|
|
12
|
+
All services verified. Ready for E2E.
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Surface skipped services so the caller knows the gap:
|
|
16
|
+
|
|
17
|
+
```
|
|
18
|
+
Skipped: payment-gateway — HTTP 500 on /
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
If you launched the services:
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
Logs: /tmp/muggle-prepare-*.log
|
|
25
|
+
Cleanup: say "stop services" or re-invoke this skill.
|
|
26
|
+
```
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
# Scope (what are you testing?)
|
|
2
|
+
|
|
3
|
+
`AskUserQuestion`:
|
|
4
|
+
|
|
5
|
+
> "What are you testing locally?"
|
|
6
|
+
|
|
7
|
+
- Option 1: "A frontend feature — I need the UI and its backend dependencies running"
|
|
8
|
+
- Option 2: "A backend API — I just need the API server running"
|
|
9
|
+
- Option 3: "The full stack — everything needs to be up"
|
|
10
|
+
|
|
11
|
+
Scopes the rest of the workflow. Record under `testing_scope` in the tracking file. Use it later to pre-check matching service candidates in [identify-services](./identify-services.md).
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Comprehensive smoke test
|
|
2
|
+
|
|
3
|
+
Runs for **every** service in the tracking file, including `external: true`. Port-listening is not proof a service works — a stale dev server binds and returns 200 with a webpack error overlay.
|
|
4
|
+
|
|
5
|
+
All three probes must pass:
|
|
6
|
+
|
|
7
|
+
1. **HTTP** — `GET <serviceUrl>`, 3 s timeout, accept `2xx`/`3xx` (one redirect).
|
|
8
|
+
2. **Body sniff** — match response body against broken-build markers in [`../../_shared/dev-server-readiness.md`](../../_shared/dev-server-readiness.md) → "Body sniff patterns".
|
|
9
|
+
3. **Log tail** — scan last 200 lines of `/tmp/muggle-prepare-<service-name>.log` for failure patterns after the latest ready signal. Skip for `external: true` (no log).
|
|
10
|
+
|
|
11
|
+
Use the primitives in `dev-server-readiness.md`. Don't re-implement.
|
|
12
|
+
|
|
13
|
+
## Diagnose-and-fix loop
|
|
14
|
+
|
|
15
|
+
On failure, show the concrete signal (HTTP code, sniff hit, or log line) and `AskUserQuestion`:
|
|
16
|
+
|
|
17
|
+
> "**<service-name>** isn't healthy: `<signal>`. How do you want to proceed?"
|
|
18
|
+
|
|
19
|
+
- Option 1: **Clean restart** (Recommended) — kill + [fresh-install](./fresh-install.md) + [start-services](./start-services.md) + re-run this step
|
|
20
|
+
- Option 2: **Restart only** — kill + start-services + re-run this step
|
|
21
|
+
- Option 3: **I'll fix it manually** — pause; re-run on user signal
|
|
22
|
+
- Option 4: **Skip** — append to `excluded_services` with reason, continue
|
|
23
|
+
|
|
24
|
+
Loop per service until pass or skip. Cap at **3 iterations** — then force a manual-intervention pause.
|
|
25
|
+
|
|
26
|
+
For `external: true`, only Options 3 and 4 apply.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Determine start commands
|
|
2
|
+
|
|
3
|
+
For each required service not already running, figure out how to start it. Read **only** the indicator file that exists.
|
|
4
|
+
|
|
5
|
+
| Indicator | Stack | Default command | What to check |
|
|
6
|
+
|:----------|:------|:----------------|:--------------|
|
|
7
|
+
| `package.json` | Node.js | `npm run dev` | Read `scripts`: prefer `dev` > `start` > `serve` |
|
|
8
|
+
| `Makefile` | Various | `make dev` | Existence; propose `make dev` or `make run` |
|
|
9
|
+
| `Cargo.toml` | Rust | `cargo run` | Existence |
|
|
10
|
+
| `go.mod` | Go | `go run .` | Existence |
|
|
11
|
+
| `pyproject.toml` | Python | Check for framework | Read `[project.scripts]` or `[tool.poetry.scripts]` if present |
|
|
12
|
+
| `requirements.txt` | Python | `python app.py` | Existence |
|
|
13
|
+
| `docker-compose.yml` | Docker | `docker compose up` | Existence |
|
|
14
|
+
|
|
15
|
+
If no indicator found, ask the user for the start command.
|
|
16
|
+
|
|
17
|
+
Present all commands in a single summary:
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
Service Directory Command
|
|
21
|
+
────────────────────────────────────────────────────────────────
|
|
22
|
+
backend-api ~/Github/backend-api npm run dev
|
|
23
|
+
auth-service ~/Github/auth-service go run .
|
|
24
|
+
frontend ~/Github/frontend npm run dev
|
|
25
|
+
────────────────────────────────────────────────────────────────
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
- Option 1: "Looks good, start them"
|
|
29
|
+
- Option 2: "I need to edit some commands"
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# Start services
|
|
2
|
+
|
|
3
|
+
For each service, launch in the background:
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
cd "<service-dir>" && nohup <command> > /tmp/muggle-prepare-<service-name>.log 2>&1 &
|
|
7
|
+
echo $!
|
|
8
|
+
```
|
|
9
|
+
|
|
10
|
+
Capture the PID. Write all service entries to `/tmp/muggle-test-prepare.json`.
|
|
11
|
+
|
|
12
|
+
**Startup verification** — confirm PID alive (`kill -0 <pid> 2>/dev/null`), then run the two-stage readiness probe per [`../../_shared/dev-server-readiness.md`](../../_shared/dev-server-readiness.md) against the log. Cap log-tail at 60 s. Halt on whatever surfaces.
|
|
13
|
+
|
|
14
|
+
If a PID dies immediately, show the last 20 log lines:
|
|
15
|
+
|
|
16
|
+
> "**backend-api** exited right after starting. Here's the tail of its log:"
|
|
17
|
+
|
|
18
|
+
- Option 1: "Skip it and continue with the others"
|
|
19
|
+
- Option 2: "Let me fix it — I'll re-invoke later"
|
|
20
|
+
|
|
21
|
+
**Port discovery** — if the port wasn't known upfront, re-scan listening ports after startup. Record in tracking file. If not found within ~10 s, mark port unknown.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
# Viability check
|
|
2
|
+
|
|
3
|
+
Some services can't run on a developer's machine by design — production secrets, HSMs, specific certificates, cloud-only infra. Don't waste time trying to start them.
|
|
4
|
+
|
|
5
|
+
If the user volunteered this in their initial message, acknowledge and skip the question. Otherwise:
|
|
6
|
+
|
|
7
|
+
> "Are there any services in your stack that **can't** run locally? (e.g., needs production secrets, specific certificates, or cloud-only infra)"
|
|
8
|
+
|
|
9
|
+
- Option 1: "All my services can run locally"
|
|
10
|
+
- Option 2: "Some can't — I'll tell you which"
|
|
11
|
+
|
|
12
|
+
If option 2, collect names and exclude from discovery.
|
|
13
|
+
|
|
14
|
+
If an excluded service is a hard dependency for the app under test, suggest remote testing:
|
|
15
|
+
|
|
16
|
+
> "Since **payment-gateway** can't run locally, you might get better coverage by merging and running `/muggle-test` against your preview environment. Want to continue with a partial local setup, or switch to remote testing?"
|
|
17
|
+
|
|
18
|
+
- Option 1: "Continue locally — I'll work around the missing service"
|
|
19
|
+
- Option 2: "Switch to remote — I'll merge and test on preview"
|
|
20
|
+
|
|
21
|
+
If remote, hand off to `/muggle-test` in remote mode and exit.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
|
-
"release": "4.12.
|
|
3
|
-
"buildId": "run-
|
|
4
|
-
"commitSha": "
|
|
5
|
-
"buildTime": "2026-05-
|
|
2
|
+
"release": "4.12.3",
|
|
3
|
+
"buildId": "run-41-1",
|
|
4
|
+
"commitSha": "9f6c0d27cf0e83976386db90321cf9cc35ac5328",
|
|
5
|
+
"buildTime": "2026-05-22T06:22:26Z",
|
|
6
6
|
"serviceName": "muggle-ai-works-mcp"
|
|
7
7
|
}
|
|
@@ -1 +1 @@
|
|
|
1
|
-
export { DEFAULT_PREFERENCES, PREFERENCES_FILE_NAME, PREFERENCES_PROJECT_DIR_NAME, PREFERENCES_SCHEMA, PREFERENCES_VERSION, PREFERENCE_ALLOWED_VALUES, PreferenceKey, PreferenceValue, buildElectronAppChecksumsUrl, buildElectronAppReleaseAssetUrl, buildElectronAppReleaseTag, calculateFileChecksum, createApiKeyWithToken, createChildLogger, deleteApiKeyData, deleteCredentials, e2e_exports as e2e, formatPreferencesOneLiner, getApiKey, getApiKeyFilePath, getAuthService, getBundledElectronAppVersion, getCallerCredentials, getCallerCredentialsAsync, getChecksumForPlatform, getConfig, getCredentialsFilePath, getDataDir, getDownloadBaseUrl, getElectronAppChecksums, getElectronAppDir, getElectronAppVersion, getElectronAppVersionSource, getLocalQaTools, getLogger, getPlatformKey, getQaTools, getValidApiKeyData, getValidCredentials, hasApiKey, isElectronAppInstalled, isFirstRun, loadApiKeyData, loadCredentials, local_exports as localQa, mcp_exports as mcp, openBrowserUrl, performLogin, performLogout, pollDeviceCode, e2e_exports as qa, readGlobalPreferences, readProjectPreferences, resetConfig, resetLogger, resetPreference, resolveElectronAppPathOrNull, resolvePreferences, saveApiKey, saveApiKeyData, saveCredentials, startDeviceCodeFlow, toolRequiresAuth, validatePreference, verifyFileChecksum, writePreferences } from './chunk-
|
|
1
|
+
export { DEFAULT_PREFERENCES, PREFERENCES_FILE_NAME, PREFERENCES_PROJECT_DIR_NAME, PREFERENCES_SCHEMA, PREFERENCES_VERSION, PREFERENCE_ALLOWED_VALUES, PreferenceKey, PreferenceValue, buildElectronAppChecksumsUrl, buildElectronAppReleaseAssetUrl, buildElectronAppReleaseTag, calculateFileChecksum, createApiKeyWithToken, createChildLogger, deleteApiKeyData, deleteCredentials, e2e_exports as e2e, formatPreferencesOneLiner, getApiKey, getApiKeyFilePath, getAuthService, getBundledElectronAppVersion, getCallerCredentials, getCallerCredentialsAsync, getChecksumForPlatform, getConfig, getCredentialsFilePath, getDataDir, getDownloadBaseUrl, getElectronAppChecksums, getElectronAppDir, getElectronAppVersion, getElectronAppVersionSource, getLocalQaTools, getLogger, getPlatformKey, getQaTools, getValidApiKeyData, getValidCredentials, hasApiKey, isElectronAppInstalled, isFirstRun, loadApiKeyData, loadCredentials, local_exports as localQa, mcp_exports as mcp, openBrowserUrl, performLogin, performLogout, pollDeviceCode, e2e_exports as qa, readGlobalPreferences, readProjectPreferences, resetConfig, resetLogger, resetPreference, resolveElectronAppPathOrNull, resolvePreferences, saveApiKey, saveApiKeyData, saveCredentials, startDeviceCodeFlow, toolRequiresAuth, validatePreference, verifyFileChecksum, writePreferences } from './chunk-CPF6AR2I.js';
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@muggleai/works",
|
|
3
3
|
"mcpName": "io.github.multiplex-ai/muggle",
|
|
4
|
-
"version": "4.12.
|
|
4
|
+
"version": "4.12.3",
|
|
5
5
|
"description": "Ship quality products with AI-powered E2E acceptance testing that validates your web app like a real user — from Claude Code and Cursor to PR.",
|
|
6
6
|
"type": "module",
|
|
7
7
|
"main": "dist/index.js",
|
|
@@ -41,14 +41,14 @@
|
|
|
41
41
|
"test:watch": "vitest"
|
|
42
42
|
},
|
|
43
43
|
"muggleConfig": {
|
|
44
|
-
"electronAppVersion": "1.0.
|
|
44
|
+
"electronAppVersion": "1.0.94",
|
|
45
45
|
"downloadBaseUrl": "https://github.com/multiplex-ai/muggle-ai-works/releases/download",
|
|
46
46
|
"runtimeTargetDefault": "production",
|
|
47
47
|
"checksums": {
|
|
48
|
-
"darwin-arm64": "
|
|
49
|
-
"darwin-x64": "
|
|
50
|
-
"linux-x64": "
|
|
51
|
-
"win32-x64": "
|
|
48
|
+
"darwin-arm64": "22731d8fac9b3ce7dc0490932a099fad0b7bab901e4efe3c2f8e44b4dd165bdb",
|
|
49
|
+
"darwin-x64": "0965f1a3f7b0db6a976650e26de024f0baea58102cf9acc93e8c69c6dfd8709f",
|
|
50
|
+
"linux-x64": "5280ae8ac73215f1c4f92d3b280eee0918341179485993493bd04f3f5a95b5b5",
|
|
51
|
+
"win32-x64": "a02bf08ba6f5463b1dad354824c6369738248e86ce5ef02644d2185fdc2d047f"
|
|
52
52
|
}
|
|
53
53
|
},
|
|
54
54
|
"dependencies": {
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "muggle",
|
|
3
3
|
"description": "Run real-browser end-to-end (E2E) acceptance tests on your web app from any AI coding agent. Generate test scripts from plain English, replay them on localhost, capture screenshots, and validate user flows like signup, checkout, and dashboards. Works across Claude Code, Cursor, Codex, and Windsurf.",
|
|
4
|
-
"version": "4.12.
|
|
4
|
+
"version": "4.12.3",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Muggle AI",
|
|
7
7
|
"email": "support@muggle-ai.com"
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"name": "muggle",
|
|
3
3
|
"displayName": "Muggle AI",
|
|
4
4
|
"description": "Ship quality products with AI-powered end-to-end (E2E) acceptance testing that validates your web app like a real user — from Claude Code and Cursor to PR.",
|
|
5
|
-
"version": "4.12.
|
|
5
|
+
"version": "4.12.3",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Muggle AI",
|
|
8
8
|
"email": "support@muggle-ai.com"
|
|
@@ -49,6 +49,45 @@ netstat -ano | findstr /R /C:":3000 " /C:":3001 " /C:":4200 " /C:":5173 " /C:":8
|
|
|
49
49
|
|
|
50
50
|
If the app declares a backend URL in its env file, probe the backend's health endpoint before treating the dev server as usable. 5xx or unreachable → halt; the frontend may render but its data layer is dead, so any query against it is meaningless.
|
|
51
51
|
|
|
52
|
+
## Body sniff patterns
|
|
53
|
+
|
|
54
|
+
A `200 OK` can still be a build-error overlay or stack trace. Search the response body (case-insensitive) for broken-build markers — a match means unhealthy regardless of status.
|
|
55
|
+
|
|
56
|
+
| Stack | Pattern (regex) |
|
|
57
|
+
|:------|:----------------|
|
|
58
|
+
| Next.js | `__next_error__\|Failed to compile\|webpack-internal://` |
|
|
59
|
+
| Vite | `vite-error-overlay\|Internal server error\|\[plugin:` |
|
|
60
|
+
| Node / Express | `MODULE_NOT_FOUND\|Cannot find module\|npm ERR!\|Cannot GET /\|Cannot POST /\|Error: ENOENT\|EACCES\|EADDRINUSE` |
|
|
61
|
+
| Django | `TemplateSyntaxError\|ProgrammingError at /\|<h1>Server Error \(500\)</h1>` |
|
|
62
|
+
| Flask | `Werkzeug Debugger\|werkzeug-debug` |
|
|
63
|
+
| FastAPI / Python | `Traceback \(most recent call last\)\|ModuleNotFoundError\|ImportError` |
|
|
64
|
+
| Rails | `Better Errors\|ActionController::RoutingError\|<title>Action Controller:` |
|
|
65
|
+
| Spring Boot | `Whitelabel Error Page` |
|
|
66
|
+
| Tomcat | `HTTP Status 500.*Apache Tomcat` |
|
|
67
|
+
| Laravel / PHP | `Whoops\\\\|<b>Fatal error</b>\|Parse error:\|Stack trace:` |
|
|
68
|
+
| JS stack frame | `at .*\(.*\.[jt]sx?:\d+:\d+\)` |
|
|
69
|
+
| Java stack frame | `at \w+(\.\w+)+\(\w+\.java:\d+\)` |
|
|
70
|
+
| Python stack frame | `File ".*", line \d+, in ` |
|
|
71
|
+
| Ruby stack frame | `\.rb:\d+:in ` |
|
|
72
|
+
|
|
73
|
+
The bash/PowerShell snippets below use the union of all patterns above. Trim per-stack when you know the target.
|
|
74
|
+
|
|
75
|
+
#### bash/zsh
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
BODY=$(curl -sS -L --max-redirs 1 --max-time 3 "$URL")
|
|
79
|
+
PATTERN='__next_error__|Failed to compile|webpack-internal://|vite-error-overlay|Internal server error|MODULE_NOT_FOUND|Cannot find module|Cannot GET /|npm ERR!|Error: ENOENT|EACCES|EADDRINUSE|TemplateSyntaxError|Werkzeug Debugger|Traceback \(most recent call last\)|ModuleNotFoundError|ImportError|Better Errors|ActionController::RoutingError|Whitelabel Error Page|Apache Tomcat|Whoops|<b>Fatal error</b>|Parse error:|Stack trace:|at .*\(.*\.[jt]sx?:[0-9]+:[0-9]+\)|at \w+(\.\w+)+\(\w+\.java:[0-9]+\)|File ".*", line [0-9]+, in |\.rb:[0-9]+:in '
|
|
80
|
+
echo "$BODY" | grep -qiE "$PATTERN" && { echo "BODY-SNIFF FAIL"; echo "$BODY" | grep -iE "$PATTERN" | head -3; exit 1; }
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
#### PowerShell
|
|
84
|
+
|
|
85
|
+
```powershell
|
|
86
|
+
$body = (Invoke-WebRequest -Uri $url -TimeoutSec 3 -MaximumRedirection 1 -ErrorAction Stop).Content
|
|
87
|
+
$pattern = '__next_error__|Failed to compile|webpack-internal://|vite-error-overlay|Internal server error|MODULE_NOT_FOUND|Cannot find module|Cannot GET /|npm ERR!|Error: ENOENT|EACCES|EADDRINUSE|TemplateSyntaxError|Werkzeug Debugger|Traceback \(most recent call last\)|ModuleNotFoundError|ImportError|Better Errors|ActionController::RoutingError|Whitelabel Error Page|Apache Tomcat|Whoops|<b>Fatal error</b>|Parse error:|Stack trace:|at .*\(.*\.[jt]sx?:\d+:\d+\)|at \w+(\.\w+)+\(\w+\.java:\d+\)|File ".*", line \d+, in |\.rb:\d+:in '
|
|
88
|
+
if ($body -imatch $pattern) { Write-Host "BODY-SNIFF FAIL"; [regex]::Matches($body, $pattern, 'IgnoreCase') | Select-Object -First 3 | ForEach-Object { $_.Value }; exit 1 }
|
|
89
|
+
```
|
|
90
|
+
|
|
52
91
|
## Two-stage readiness — after starting a dev server
|
|
53
92
|
|
|
54
93
|
Network reachability is necessary but not sufficient. Many dev servers bind to a port before build/startup work is complete. Wait for **both** network readiness and application readiness before issuing requests.
|
|
@@ -109,9 +109,15 @@ Triggered when `muggle-local-execute-replay` returns `status: "failed"` (or non-
|
|
|
109
109
|
| **stale-script** | The test script no longer matches the live UI (selectors moved, label paths changed, page renamed). The product still works; the script is out of date. |
|
|
110
110
|
| **product-defect** | The script and infra are fine; the user's app actually misbehaved (assertion failure on previously-passing step, unexpected error, wrong page after action). This is the failure mode acceptance testing exists to catch. |
|
|
111
111
|
|
|
112
|
-
###
|
|
112
|
+
### Where to read signals
|
|
113
|
+
|
|
114
|
+
Call `muggle-local-run-result-get` (local) or the remote equivalent and read **structured fields**, not `execute`'s response stdout tail (it's a truncated display excerpt and routinely cuts off mid-sentence). Order:
|
|
113
115
|
|
|
114
|
-
|
|
116
|
+
1. `Status` + `Error` — the verdict and the one-line cause.
|
|
117
|
+
2. `Artifacts` section, when present — opens `artifactsDir`. Read `results.md` (step-by-step + screenshot links) for the per-step verdict, then `action-script.json` for what the agent attempted.
|
|
118
|
+
3. `stdout.log` / `stderr.log` only when the Artifacts section is absent or `results.md` doesn't exist (e.g. early Electron failure).
|
|
119
|
+
|
|
120
|
+
### Initial signal heuristics
|
|
115
121
|
|
|
116
122
|
- **infra** signals: `electron-crash`, `chromium-error`, `click-no-effect-on-clickable-element`, `timeout-on-trivial-wait`, `internal-error-in-mcp-output`.
|
|
117
123
|
- **stale-script** signals: `element-not-found`, `selector-timeout`, `label-path-mismatch`, `nav-target-404`, `aria-label-changed`.
|
|
@@ -189,9 +195,18 @@ Triggered when `muggle-local-execute-test-generation` (or the remote equivalent)
|
|
|
189
195
|
| **agent-course** | The generation agent went down a wrong path (chose the wrong button, misread the goal, looped on a blocking modal). The product is fine and the test case is fine — the agent's *course* needs steering. |
|
|
190
196
|
| **product-uxux** | The product itself blocks the test (broken page, missing element, server error). Agent can't proceed because the feature doesn't actually work. |
|
|
191
197
|
|
|
192
|
-
###
|
|
198
|
+
### Where to read signals
|
|
193
199
|
|
|
194
|
-
|
|
200
|
+
Same rule as section B: read **structured fields** from `muggle-local-run-result-get`, not `execute`'s response stdout tail. The `Artifacts` section is present on failed regen too — `action-script.json` is included when generation reached the step-emission stage (typical for `goal_not_achievable`: the agent's attempted steps + halt summary). `results.md` and per-step screenshots are absent on failure (electron-app emits those only on the successful completion path).
|
|
201
|
+
|
|
202
|
+
Order:
|
|
203
|
+
|
|
204
|
+
1. `Status` + `Error` — the verdict and one-line cause. `Error: Electron exited with code 26` typically means `goal_not_achievable`.
|
|
205
|
+
2. `action-script.json` in `artifactsDir` when present — read the steps the agent attempted and the `summaryStep` (halt reason, goal-not-achievable verdict).
|
|
206
|
+
3. `stdout.log` / `stderr.log` at `artifactsDir/` — last 100 lines is usually enough; look for the final structured summary the generation agent emitted (it appears near the end as a JSON-ish block, not in the truncated execute tail).
|
|
207
|
+
4. Remote regen — fetch the workflow run with `muggle-remote-wf-get-ts-gen-latest-run`; signals live in `summaryStep` and the per-step list there.
|
|
208
|
+
|
|
209
|
+
### Initial signal heuristics
|
|
195
210
|
|
|
196
211
|
- **transient**: `network-error`, `llm-rate-limit`, `single-tool-call-error`, run had partial progress then died.
|
|
197
212
|
- **infra**: `electron-mcp-handler-crash`, `internal-validation-error`, `pipeline-stuck`, identical failure repeated more than twice.
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# PR-Branch Worktree — Shared Reference
|
|
2
|
+
|
|
3
|
+
> Source of truth for materializing a PR's branch in an isolated worktree so the user's main checkout is never disturbed. Used by `muggle-test` (and any future skill that takes a GitHub PR URL). Skills MUST link here rather than restate the steps.
|
|
4
|
+
|
|
5
|
+
## When this applies
|
|
6
|
+
|
|
7
|
+
A skill receives a GitHub PR URL of the form `github.com/<org>/<repo>/pull/<n>` and needs the PR's branch checked out to test against it locally.
|
|
8
|
+
|
|
9
|
+
## Steps
|
|
10
|
+
|
|
11
|
+
1. **Resolve the PR's head branch:**
|
|
12
|
+
`gh pr view <n> --repo <org>/<repo> --json headRefName -q .headRefName`
|
|
13
|
+
2. **Sanitize the branch name** for filesystem use — replace `/` and other path separators with `-`. Example: `claude/regen-test-replay-flow-ZSScQ` → `claude-regen-test-replay-flow-ZSScQ`.
|
|
14
|
+
3. **Build the target worktree path:** `<repo>/.claude/worktrees/<sanitized-branch>`.
|
|
15
|
+
4. **Materialize the worktree:**
|
|
16
|
+
- If the target path does NOT exist:
|
|
17
|
+
- `git -C <repo> fetch origin <branch>`
|
|
18
|
+
- `git -C <repo> worktree add <target-path> <branch>`
|
|
19
|
+
- If the target path EXISTS (reused from a prior run):
|
|
20
|
+
- `git -C <target-path> fetch`
|
|
21
|
+
- `git -C <target-path> reset --hard origin/<branch>` — picks up new pushes, drops any local cruft.
|
|
22
|
+
5. **Use the worktree path as the working directory** for the rest of the run, including:
|
|
23
|
+
- Passing it as the **`cwd` parameter** to `muggle-local-execute-test-generation` and `muggle-local-execute-replay`. This is required, not optional — see `_shared/failure-mode-handling.md` and the lock identity discussion in those tools' MCP source.
|
|
24
|
+
- Resolving any `npm install` / dev-server start commands inside the worktree (it has its own `node_modules/` and `.env*` files).
|
|
25
|
+
6. **Tell the user** where the worktree lives so they can clean it up later with `git -C <repo> worktree remove <target-path>`.
|
|
26
|
+
|
|
27
|
+
## Invariants
|
|
28
|
+
|
|
29
|
+
- **Never switch the user's main checkout.** The whole point of this flow is isolation; `git checkout <branch>` on the main checkout is forbidden.
|
|
30
|
+
- **Never share `node_modules/` via symlink** across worktrees. Each worktree runs its own `npm install` (or `pnpm install`) — webpack's `resolve.symlinks: true` rewrites paths and breaks asset-identity tracking.
|
|
31
|
+
- **`.env*` files do not propagate.** A freshly created worktree has no env files unless the repo commits them. If the parent skill's dev server fails to boot, check whether `.env.local` (or framework equivalent) needs to be copied from the main checkout before launching.
|
|
@@ -96,7 +96,7 @@ Analyze the changes to understand what's impacted. Two sources, picked by what t
|
|
|
96
96
|
**PR URL** (user passed `github.com/<org>/<repo>/pull/<n>`):
|
|
97
97
|
1. `gh pr diff <n> --repo <org>/<repo> --name-only` for the changed file list
|
|
98
98
|
2. `gh pr diff <n> --repo <org>/<repo>` for the actual diff
|
|
99
|
-
3.
|
|
99
|
+
3. Materialize the PR branch in a dedicated worktree per [`_shared/pr-branch-worktree.md`](../_shared/pr-branch-worktree.md). Use that worktree path as the `cwd` for the rest of the run (including the `cwd` parameter on local execute tools).
|
|
100
100
|
|
|
101
101
|
Either way:
|
|
102
102
|
1. Identify impacted feature areas:
|
|
@@ -273,6 +273,7 @@ Execution itself **must** be sequential because there is only one local Electron
|
|
|
273
273
|
1. Call `muggle-local-execute-test-generation`:
|
|
274
274
|
- `testCase`: Full test case object from the parallel fetch above
|
|
275
275
|
- `localUrl`: User's local URL from the pre-flight question
|
|
276
|
+
- `cwd`: Absolute path of the active working directory — the PR-branch worktree if one was created in Step 2, otherwise the user's repo root. Drives the cross-worktree single-flight lock so concurrent muggle-test runs from different branches serialize.
|
|
276
277
|
- `showUi`: from the `showElectronBrowser` resolution — omit (default visible) for `always`, pass `false` for `never`
|
|
277
278
|
- `freshSession`: `true` if the test case requires a clean browser state (see above), omit otherwise
|
|
278
279
|
2. Store the returned `runId` and tag the result `mode: "regen"`.
|
|
@@ -282,14 +283,16 @@ Execution itself **must** be sequential because there is only one local Electron
|
|
|
282
283
|
2. Call `muggle-local-execute-replay`:
|
|
283
284
|
- `testScript`: from `muggle-remote-test-script-get`
|
|
284
285
|
- `actionScript`: from `muggle-remote-action-script-get`
|
|
285
|
-
- `localUrl`, `showUi`, `freshSession`: same resolution as regen
|
|
286
|
+
- `localUrl`, `cwd`, `showUi`, `freshSession`: same resolution as regen
|
|
286
287
|
3. Store the returned `runId` and tag the result `mode: "replay"`.
|
|
287
288
|
|
|
288
289
|
If a run fails, log it and continue to the next — do not abort the batch. Failures are routed through Step 7C's post-failure handler after the batch completes.
|
|
289
290
|
|
|
290
291
|
### Collect results (in parallel)
|
|
291
292
|
|
|
292
|
-
For every `runId`, issue all `muggle-local-run-result-get` calls in parallel. Extract
|
|
293
|
+
For every `runId`, issue all `muggle-local-run-result-get` calls in parallel. Extract from the **structured response only** (not from `execute`'s stdout tail, which is a truncated display excerpt): `Status`, `Error`, `Duration`, and the `Artifacts` section (always present after a run completes — names `artifactsDir` and lists the files actually on disk).
|
|
294
|
+
|
|
295
|
+
For passed runs, `results.md` inside `artifactsDir` is the step-by-step verdict — read it before summarizing. For failed runs, `stdout.log` + `stderr.log` are always present and `action-script.json` is present when generation reached the step-emission stage (typical for `goal_not_achievable`); use `Error` as the headline verdict and route through Step 7C.
|
|
293
296
|
|
|
294
297
|
### Publish each run to cloud (gated by `autoPublishLocalResults`)
|
|
295
298
|
|
|
@@ -459,6 +462,7 @@ This is a suggestion, not automatic invocation. Skip silently if every test pass
|
|
|
459
462
|
## Guardrails
|
|
460
463
|
|
|
461
464
|
- **Always confirm intent first** — never assume local vs remote without asking
|
|
465
|
+
- **PR URLs always run in a dedicated worktree** — never switch the user's main checkout. Create or reuse `<repo>/.claude/worktrees/<sanitized-branch>` and pass that path as the `cwd` parameter to local execute tools. The cross-worktree single-flight lock relies on this to serialize concurrent runs from different branches.
|
|
462
466
|
- **User MUST select project** — present clickable options via `AskUserQuestion`, wait for explicit choice, never auto-select
|
|
463
467
|
- **Best-effort shortlist use cases** — use the change summary to narrow the list to the most relevant 1–5 use cases and pre-check them; never dump every use case in the project on the user. Always leave an escape hatch to reveal the full list.
|
|
464
468
|
- **Best-effort shortlist test cases** — same idea: pre-check the test cases most relevant to the change summary; never enumerate every test case attached to a use case. Always leave an escape hatch to reveal the full list.
|
|
@@ -195,8 +195,12 @@ If publish rejects with `has no generated actionScript steps to publish` (true z
|
|
|
195
195
|
|
|
196
196
|
### 9. Report
|
|
197
197
|
|
|
198
|
+
**Do not diagnose from `execute`'s response stdout tail.** That tail is a truncated excerpt for human display and routinely cuts off mid-sentence. The only ground truth is the run record.
|
|
199
|
+
|
|
198
200
|
- `muggle-local-run-result-get` with the run id from execute.
|
|
199
|
-
-
|
|
201
|
+
- **Read in this order:** `Status` → `Error` → **`Artifacts` section** (always present after a run completes; names `artifactsDir` and lists the files actually on disk: `action-script.json`, `results.md`, `screenshots/`, `stdout.log`, `stderr.log`). On a `passed` run, `results.md` is the step-by-step verdict with screenshot links — read it before summarizing.
|
|
202
|
+
- **On failure**, the `Artifacts` section is still present. `stdout.log` + `stderr.log` are always there. `action-script.json` is there when generation got far enough to emit it (typical for `goal_not_achievable` / mid-progress crashes — the file holds the agent's attempted steps + halt summary). `results.md` and per-step screenshots are absent on failure (electron-app only emits those on the successful completion path) — don't hunt elsewhere on disk for them.
|
|
203
|
+
- Include in the report: status, duration, pass/fail summary, per-step summary (passed runs), artifact paths, errors if failed, and script view URL when publishing ran.
|
|
200
204
|
|
|
201
205
|
### 9a. Route failures through the failure-mode handler
|
|
202
206
|
|
|
@@ -242,6 +246,7 @@ After reporting results:
|
|
|
242
246
|
|
|
243
247
|
- No silent auth skip.
|
|
244
248
|
- **Never prompt for Electron launch approval** before execution — invoking this skill is the approval. Just run.
|
|
249
|
+
- **Never diagnose a failed run from `execute`'s response stdout tail.** Always call `muggle-local-run-result-get` first; classify only from its structured fields and (when present) the artifacts it names. The execute tail is an excerpt and routinely truncates the failure cause.
|
|
245
250
|
- If replayable scripts exist, do not default to generation without user choice.
|
|
246
251
|
- No hiding failures: surface errors and artifact paths.
|
|
247
252
|
- **Always offer the agent-guidance reminder after every Electron run** (Step 9b) — pass or fail — unless 9a already routed the user into `muggle-feedback`. Never silently end a run without giving the user a one-click path to flag what was wrong.
|