forgeqe-cli 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,351 @@
1
+ Metadata-Version: 2.4
2
+ Name: forgeqe-cli
3
+ Version: 1.0.0
4
+ Summary: pip launcher for the ForgeQE browser agent (npm package: @capagents/forgeqe)
5
+ Author: CLI-Agents
6
+ License: MIT
7
+ Keywords: forgeqe,browser,agent,cli,opentui
8
+ Requires-Python: >=3.9
9
+ Description-Content-Type: text/markdown
10
+ Provides-Extra: dev
11
+ Requires-Dist: pytest>=8.0; extra == "dev"
12
+
13
+ # ForgeQE
14
+
15
+ Browser agent CLI. Say what you want in plain language: a goal to carry out in a browser, test cases to run from Jira, Azure DevOps, Xray or a file, a test to write, or a CI pipeline to set up. ForgeQE's agent works out what that needs and calls ForgeQE's tools. The ForgeQE engine drives Chrome, and a run that passes becomes a test in your repo's own language and framework, written by **Scribe**, ForgeQE's coding agent. Skills (`SKILL.md` folders, built in or in `.forgeqe/skills`) teach the agents your repo's conventions and the apps you test.
16
+
17
+ The engine is the browser service from test-agent-nexus, copied into this package (`py/forgeqe_engine`). It keeps the nexus behaviour: error-recovery system prompt, fast-mode agent tuning, Chrome launch flags, the search/select loop breaker, the 600 s run timeout, compact workflow YAML with semantic locators. There is no dependency on test-agent-nexus, FastAPI, or a database.
18
+
19
+ The shell is [OpenTUI](https://opentui.com). `forgeqe run` does the same without the shell, and `forgeqe suite` is the fixed command pipelines call. See [HOW_TO_USE.md](HOW_TO_USE.md), the interactive [architecture page](forgeqe-architecture.html) (`forgeqe docs --open`), and [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md).
20
+
21
+ ## Install
22
+
23
+ The CLI is published on **npm** as `@capagents/forgeqe`. **PyPI** package `forgeqe-cli` is a pip launcher for the same `forgeqe` command. Either one is the only thing to install, on Windows, macOS or Linux:
24
+
25
+ ```bash
26
+ npm install -g @capagents/forgeqe
27
+ # or
28
+ pip install forgeqe-cli
29
+
30
+ forgeqe setup # optional: the first live run does this too
31
+ ```
32
+
33
+ Everything else is downloaded on first use into `~/.forgeqe/tools` (`%USERPROFILE%\.forgeqe\tools` on Windows), checksum verified, without admin rights and without touching PATH or shell profiles:
34
+
35
+ - **Bun** (the CLI and the OpenTUI shell run on it). A Bun 1.3+ already on PATH or in `~/.bun/bin` is used instead. `FORGEQE_BUN` points at a specific one.
36
+ - **The agent runtime** that ForgeQE's agent and Scribe, the coding agent, run on, at the version ForgeQE is tested with. `code.runtime` points at a specific binary.
37
+ - **uv**, which then fetches **Python** for the engine. A uv already installed is used instead; when uv cannot be downloaded, a Python 3.11+ on PATH (or the Windows `py` launcher) is the fallback.
38
+ - **The ForgeQE engine** and its Python libraries in `~/.forgeqe/engine/.venv`. Installed Google Chrome is used, or Chromium is installed when Chrome is missing. `FORGEQE_PYTHON` points at an existing Python that already has the engine's libraries; when that Python cannot run the engine (missing, too old, or without them), ForgeQE says so and uses its own environment. A broken environment is rebuilt on the next run.
39
+
40
+ `forgeqe setup` does all of this up front; otherwise the first command that needs a piece fetches it. Upgrading is `forgeqe update`: it updates every `forgeqe` on PATH (bun, npm and pip installs) and waits while the registries catch up with a new release. A release that pins a newer Bun, agent runtime, uv or engine library downloads it on the next run and removes the old one. Deleting `~/.forgeqe` removes everything ForgeQE downloaded. Behind a proxy, set `HTTPS_PROXY`. The npm launcher's Bun download also needs `NODE_USE_ENV_PROXY=1` (Node 24+); on older Node, use the pip launcher or install Bun yourself.
41
+
42
+ From this repo:
43
+
44
+ ```bash
45
+ cd ForgeQE
46
+ bun install
47
+ bun src/index.ts init
48
+ bun src/index.ts start
49
+ ```
50
+
51
+ One-shot release (same version on both registries; assumes `npm` and `twine` are already logged in):
52
+
53
+ ```bash
54
+ ./scripts/publish.sh # current version
55
+ ./scripts/publish.sh 0.2.0 # bump + publish
56
+ ./scripts/publish.sh --dry-run # pack only
57
+ ```
58
+
59
+ ## Privacy
60
+
61
+ ForgeQE sends no telemetry and turns it off in everything it starts:
62
+ - The engine: PostHog telemetry, cloud sync, and the price-list download are off, and the blank-tab logo its library loads from a CDN is replaced by ForgeQE's start page, which loads nothing.
63
+ - The agent runtime: auto-update, session sharing, OpenTelemetry spans, and the model catalogue download are off; it uses its bundled model list.
64
+ - The repo's test command: `DO_NOT_TRACK=1`, plus the opt-outs for Next.js, Nuxt, Astro, Gatsby, Storybook, Turborepo, Angular, .NET, and Cypress.
65
+
66
+ `OTEL_EXPORTER_OTLP_*` variables are removed from those processes, and your environment cannot turn any of this back on. The only traffic is to your model endpoint and the sites the agent visits. The engine also downloads its ad-block and cookie-banner extensions from the Chrome Web Store once.
67
+
68
+ ## Configure
69
+
70
+ `init` writes:
71
+
72
+ | File | Purpose |
73
+ |------|---------|
74
+ | `forgeqe.yaml` | Browser, prompts, export, active profile |
75
+ | `llms.json` | Named LLM profiles (azure, openai, ollama, openai_compatible) |
76
+ | `.forgeqe/prompts/` | Every agent prompt as a markdown file, to change how an agent works (`forgeqe prompts` lists them) |
77
+ | `.forgeqe/skills/` | The built-in skills, to adapt; add your own beside them |
78
+ | `.forgeqe/knowledge/` | Product knowledge templates: product, roles, glossary, business rules, test data, environments |
79
+
80
+ An unchanged copy keeps following ForgeQE's updates; once you edit it, yours wins.
81
+
82
+ ```bash
83
+ forgeqe models
84
+ forgeqe run "Read the homepage of example.com" --plain
85
+ ```
86
+
87
+ The active profile is `llm.active` in `forgeqe.yaml`, else `default` in `llms.json`, else the first profile; Tab switches in the shell. With none set up, ForgeQE says so and how to add one.
88
+
89
+ ## CLI
90
+
91
+ ```bash
92
+ forgeqe start
93
+ forgeqe start -g "Go to booking.com and search hotels in Mumbai" -p azure-gpt4o
94
+
95
+ forgeqe run "Get a quote on example.com for a family of four" --profile azure-gpt4o --headed
96
+ forgeqe run "Run SHOP-12 and SHOP-14, 2 browsers at once"
97
+ forgeqe run "Write a Playwright test for the login flow on staging.app.test" --repo ../app
98
+ forgeqe run "Test the refund rules on https://acme.atlassian.net/wiki/spaces/QA/pages/123"
99
+ forgeqe login shop # sign in by hand once; runs reuse the session
100
+ forgeqe run --direct "Read the homepage" --url https://example.com --out ./.forgeqe/reports/demo # one browser goal, no agent
101
+ ```
102
+
103
+ There are no commands to pick a source or a mode. ForgeQE's agent reads the request and decides: a browser goal, test cases (Jira keys, Azure DevOps ids or test plans, Xray, TestRail, Zephyr Scale or qTest cases, files in the working directory, or cases pasted into the message), requirements to turn into cases (a Confluence page, a story), a test, or a pipeline. When a reference could mean more than one thing, it asks. The agent picks the site to open; `--url` only pins a start page.
104
+
105
+ Inside the shell, type the request and press Enter (Ctrl+J starts a new line). The only commands are for the session: `/stop`, `/headed`, `/details`, `/profile`, `/logs`, `/skills`, `/clear`, `/help`, `/exit`. Everything the other commands do can be asked for in the shell too, since each one is a tool the agent calls: "heal the checkout tests", "check the smoke cases once", "monitor them every 10 minutes", "sign in to shop", "forget the shop session", "which profiles do I have?". While the agents work, the shell shows each step's result, the files the coding agent writes, test results, and problems, with what is happening now on the activity line; `/details on` (or `--verbose`) adds the browser agent's reasoning, page memory, engine logs, and every file the coding agent reads. Results appear as cards in the transcript, and when a step needs you (signing in by hand), a **Your turn** prompt waits until you pick **Done, continue** or **Cancel**.
106
+
107
+ ForgeQE's agent needs the agent runtime (downloaded automatically, see Install) and a live profile. It uses the active `llms.json` profile, or `code.model: provider/model`. Without the agent runtime, the request runs as one browser goal.
108
+
109
+ `forgeqe run` exits 0 when the job is done, 1 when it failed, 2 when it was stopped, and 3 when the agent needs more information.
110
+
111
+ ## Page memory
112
+
113
+ Every live run records each page it visits in `.forgeqe/memory` next to `forgeqe.yaml` (`~/.forgeqe/memory` when there is no config file): every interactive element, with all of its locators (test id, id, role and name, label, placeholder, alt, title, href, text, CSS, XPath), a confidence % for each, and when it was first seen, last seen, and last used. Confidence rises when a locator keeps matching exactly one element and works when used, and falls when it goes missing or fails.
114
+
115
+ A goal that passed before is replayed from memory with no model calls. Each step finds its element by the highest-confidence locator that still matches, so a renamed button or a changed id heals itself. The agent then checks the result with one model call (`replay: verify`), or the run finishes without the model (`replay: trust`). If a step cannot be found, the agent takes over from that point. Generated specs use the best locator from memory.
116
+
117
+ ```bash
118
+ forgeqe memory # hosts, pages, elements, flows
119
+ forgeqe memory show booking.com # a host's pages
120
+ forgeqe memory show https://app.test/login --locators # every locator with confidence
121
+ forgeqe memory flows --steps # recorded goals and their steps
122
+ forgeqe memory clear booking.com --yes
123
+ forgeqe run "..." --replay trust # or --replay off, --no-memory
124
+ ```
125
+
126
+ ## Skills
127
+
128
+ Skills teach the agents what a prompt can't know: your repo's test conventions, a framework, or how to drive an app. A skill is a folder with a `SKILL.md` (the Agent Skills format: `name`, `description`, then instructions), plus ForgeQE's `agents: [forgeqe, scribe, browser]` to say who uses it. ForgeQE ships `playwright-typescript` and `dotnet-nunit-playwright` for Scribe, `dynamics-365` for the browser agent, and `writing-skills` for ForgeQE's agent. Your own go in `.forgeqe/skills` next to `forgeqe.yaml` (commit them) or `~/.forgeqe/skills`; a project skill replaces a built-in one with the same name. Skills in `~/.claude/skills` and `~/.agents/skills` are used only with `skills.external: true`.
129
+
130
+ Each agent sees the skills meant for it and loads the ones whose description fits the job. For browser runs, ForgeQE's agent passes the skills a goal needs; in `forgeqe suite`, `monitor` and `run --direct` the engine picks them from the goal once, and `--skill NAME` (or `--skill none`) fixes the list. Reports and `run.json` show which skills a run used.
131
+
132
+ ```bash
133
+ forgeqe skills # who can use what, and where each skill comes from
134
+ forgeqe skills new contoso-grid --agent browser --description "Use when the site is the Contoso ERP (erp.contoso.com)"
135
+ forgeqe skills show contoso-grid
136
+ forgeqe skills check # validate every skill
137
+ forgeqe skills disable dynamics-365 # edits skills.disable, keeping comments
138
+ ```
139
+
140
+ Or ask in the shell: "make a skill for our Dynamics date picker".
141
+
142
+ ## Tests
143
+
144
+ Point ForgeQE at a repo and a run that passes becomes a test in that repo. The agent gets the run's steps and the locators and confidence from page memory. It studies the repo's framework, page objects, fixtures, and helpers, reuses them, and writes the test in the repo's language and test framework (C# NUnit, TypeScript, Python, Java, …). The browser agent records what it checked with the exact text on the page, and the test asserts those checks with locators the Playwright replay found. ForgeQE runs the test command and hands any failure back (the error first, the page dump saved to a file), until the test passes or the attempts run out. What Scribe learned about the repo is kept, so the next job there starts from its notes instead of reading the repo again. The test counts only when ForgeQE's own run of it passes.
145
+
146
+ ```bash
147
+ forgeqe run "Log in and open the invoices page on staging.app.test" --repo ../app
148
+ forgeqe code --repo ../app # write a test for the latest run
149
+ forgeqe code .forgeqe/reports/20260928-172700 --repo ../app --attempts 3
150
+ ```
151
+
152
+ In the shell and `forgeqe run`, every passed run becomes a test unless you say otherwise: in `--repo`, in `code.repo` from `forgeqe.yaml`, or else in the current folder. `--no-code` turns tests off. The transcript shows the handover to Scribe, the coding agent, the files it reads and writes, and ForgeQE's run of the test, and each request ends with a card listing the run's report and the test with its result (or why no test was written).
153
+
154
+ ### Fixing broken tests
155
+
156
+ When the app changes, tests break. `forgeqe heal` runs the repo's tests, and for each failing test the agent re-runs its flow in the real browser. If the flow still works, the test is out of date: the agent updates its steps and locators (in the page object, when that's where they live) from what the browser just did and from page memory. If the browser fails at the same step, the app is broken there: the test is left alone and the failure is reported as an app bug. Tests are never deleted, skipped, loosened, or given longer timeouts to make them pass.
157
+
158
+ ForgeQE re-runs every healed test and then the whole command itself; a test counts as healed only when ForgeQE's own run of it passes.
159
+
160
+ ```bash
161
+ forgeqe heal --repo ../app # code.test_command, or the agent finds the command
162
+ forgeqe heal "npx playwright test tests/checkout" --repo ../app
163
+ forgeqe heal -- npx playwright test --project=chromium
164
+ forgeqe run "The login tests broke after the redesign, fix them" --repo ../app # the agent heals them the same way
165
+ ```
166
+
167
+ The report is `heal.md` (and `heal.json`) in `.forgeqe/reports/heal/<time>/`, with each test's verdict, the files changed, the browser runs, and the test output before and after. `forgeqe heal` exits 0 when the tests pass (or nothing needed healing), 1 when some still fail, and 2 when stopped.
168
+
169
+ ### Running existing tests
170
+
171
+ "Run all tests related to login" or "run the tests added this sprint" finds the repo's existing tests and runs them once you agree. Scribe reads the tests (and git history for a date range or sprint) and decides each one from what it does, not from its name. ForgeQE runs the framework's list command to confirm the selection and shows the list. Nothing runs until you pick **Run**, and then ForgeQE runs exactly those tests and shows a pass or fail for each. For a sprint, ForgeQE fetches the sprints from Azure DevOps and Jira and asks which one you mean.
172
+
173
+ ```bash
174
+ forgeqe run "Run all test cases related to login" --repo ../app
175
+ forgeqe run "Which tests were added this sprint?" --repo ../app
176
+ ```
177
+
178
+ ### API tests
179
+
180
+ REST, GraphQL, SOAP, and gRPC. Every browser run records the app's own API calls (secrets redacted), and "turn the last run's API calls into tests" makes contracts from them: status, schema, and key fields per operation. Or start from an OpenAPI or Swagger link, a WSDL, a GraphQL endpoint or schema, or .proto files. ForgeQE runs the contracts straight against the API, with no browser, and Scribe writes them as API tests in your repo's framework.
181
+
182
+ ```bash
183
+ forgeqe api contracts --spec https://petstore3.swagger.io/api/v3/openapi.json
184
+ forgeqe api contracts --run && forgeqe api run
185
+ ```
186
+
187
+ End-to-end scenarios chain calls, passing values from one answer to the next request: from a browser run (the values that flowed are found for you), from a description (each resource's lifecycle, links, and negative cases), or planned from test cases or plain English. Scribe writes them on your repo's existing API helpers.
188
+
189
+ ```bash
190
+ forgeqe api scenarios --text "create a user, add two items, check out" --spec openapi.yaml
191
+ forgeqe api run-scenarios --allow-writes
192
+ ```
193
+
194
+ ## Test cases and suites
195
+
196
+ Run existing test cases instead of typing goals: ask for them by name ("run SHOP-12", "run test plan 12 suite 34", "run the cases in smoke.yaml") or paste them into the message. Each case becomes one goal: its steps are what the browser agent does and its expected results are what it checks, and the agent's own verdict decides pass or fail. With a repo set, every passed case also gets a test.
197
+
198
+ `forgeqe suite` runs cases the same way every time, with no agent choosing anything, which is what pipelines need. It takes explicit references:
199
+
200
+ ```bash
201
+ forgeqe cases cases.yaml # what a source gives, without running anything
202
+ forgeqe suite cases.yaml --parallel 3 # run them, 3 browsers at once
203
+ forgeqe suite ado:plan=12/suite=34 --repo ../app
204
+ forgeqe suite jira:jql="project = SHOP AND labels = smoke"
205
+ forgeqe suite xray:plan=SHOP-100
206
+ forgeqe suite "Open example.com and check the title says Example Domain"
207
+ ```
208
+
209
+ | Source | Reference |
210
+ |--------|-----------|
211
+ | Plain text | the text itself, quoted |
212
+ | Text or Markdown file | `.txt` / `.md`, one case per block between `---` lines |
213
+ | CSV | `.csv` with `id, title, url, description, preconditions, tags, action, data, expected` columns |
214
+ | YAML / JSON | `.yaml` / `.json`, a list of `{id, title, url, preconditions, steps: [{action, data, expected}]}` |
215
+ | Gherkin | `.feature`, one case per scenario (and per Examples row) |
216
+ | Azure DevOps | `ado:123,456` (test case ids), `ado:plan=12`, `ado:plan=12/suite=34` |
217
+ | Jira | `jira:SHOP-1,SHOP-2`, `jira:jql=<query>` (Cloud and Server/Data Center) |
218
+ | Xray | `xray:SHOP-1`, `xray:jql=<query>`, `xray:plan=SHOP-100` (Cloud and Server/Data Center) |
219
+ | TestRail | `testrail:C12,C13`, `testrail:run=45`, `testrail:plan=7`, `testrail:suite=3` |
220
+ | Zephyr Scale (Cloud) | `zephyr:SHOP-T1,SHOP-T2`, `zephyr:cycle=SHOP-R5`, `zephyr:folder=12` |
221
+ | qTest | `qtest:1234` (test case ids), `qtest:cycle=88`, `qtest:suite=9` |
222
+
223
+ For a case with numbered steps the browser agent reports a verdict for every step (passed, failed, or not run, with what it actually saw), not only for the case.
224
+
225
+ Results go back where the cases came from, with the step verdicts: an Azure DevOps test run for a plan (or a comment on the work item), a Jira comment, an Xray Test Execution with the failure screenshot, a TestRail run (the run the cases came from, or a new one) with step results and the screenshot, a Zephyr Scale test cycle, or qTest test logs. `--no-write-back` or `sources.write_back: false` turns it off.
226
+
227
+ Test cases the agent writes from an Azure DevOps or Jira story are saved there, linked to the story: Azure DevOps Test Case work items with their steps (Tests / Tested By, in the story's area and iteration), or Jira issues of type Test. The reply lists each case as a link, and saved Azure DevOps cases are run by id so results post back to them. A case already linked to the story under the same title is reused rather than duplicated. `sources.save_cases: false` turns this off; the agent then shows the cases formatted in its reply.
228
+
229
+ ### Accessibility and performance
230
+
231
+ Every run checks each page against WCAG (axe-core; WCAG 2.1 AA by default, or 2.2, AAA, Section 508, EN 301 549) and records its Core Web Vitals (LCP, CLS, INP, FCP, TTFB). Every run also compares each page with its approved screenshot (masked regions hidden) and shows baseline, this run, and the difference in the report; `forgeqe visual approve` accepts an intended change. Set budgets under `quality.budgets` and a run fails only when one is exceeded. `forgeqe quality trends` shows what got worse or better per page over time.
232
+
233
+ ### Exploratory testing
234
+
235
+ `forgeqe explore ado:1234` (or "explore story 1234" in the shell) reads the story and its scripted test cases, plans charters beyond them (edge cases, invalid input, boundaries, role differences, refresh and back), and runs them in the browser. A defect counts only when a second run reproduces it; everything else is reported as an observation. You get a report, proposed test cases to save to the story, and bugs for confirmed defects only, after you agree.
236
+
237
+ ### Bugs
238
+
239
+ With `sources.bugs`, each failed case gets a bug in Jira or Azure DevOps with the steps and verdicts, where it ended, the build link, and the last screenshot, linked to the test case. ForgeQE tags the bug with a fingerprint of the case, so the next failure adds a "still failing" comment to the open bug instead of filing another, and a pass adds a note that it passes again.
240
+
241
+ ```yaml
242
+ sources:
243
+ bugs: { system: jira, project: SHOP } # or system: azure_devops (uses azure_devops.project)
244
+ ```
245
+
246
+ ### Reports
247
+
248
+ Every suite writes `report.html`, `junit.xml`, and `suite.json` to `.forgeqe/reports/suites/<time>/`, with each case's run folder beside them; every single run writes its own `report.html` too. The report is one HTML file with the screenshots embedded, so you can open it straight from disk, attach it to a build, or mail it; no server is needed. It has an overview (verdict, totals, step checks, tokens, cost, page memory savings, browser errors, accessibility issues, bugs filed, what changed since the last run, results over earlier runs, the cases that need attention, results by source and tag, the slowest cases, and what was posted where), a test results page (each case's step table with expected and actual results, why it failed, its last passing screenshot beside this run's, browser errors per step, page memory with healed locators, the browser agent's steps with a screenshot each and the clicked element outlined, a timing waterfall, accessibility issues, and the trace and workflow to read in place), a failures page, since last run, trends and flaky tests from the earlier suite runs in the same `.forgeqe/reports/suites/` folder, time and cost (model vs browser time, a parallel worker timeline, cost per case from the model's `price` in `llms.json`), coverage by requirement and group from the test system, accessibility checks for every page visited, and the environment (model, browser, ForgeQE version, CI build, branch, commit). It exports Markdown and JSON and prints cleanly. See `samples/reports/nightly/20260929-020005/report.html` for an example. `report.screenshots: failures` keeps screenshots for failed runs only, `off` drops them; `report.accessibility: false` skips the page checks. On GitHub Actions the suite also writes a job summary.
249
+
250
+ ## Signing in
251
+
252
+ Sites that need a login go under `logins:`. The browser agent types credentials from the environment without ever seeing them (it only sees placeholders such as `shop_password`), fills authenticator codes from a TOTP secret, and keeps the signed-in session in `.forgeqe/sessions/` so the next run starts signed in. Replays read the same environment variables, and so do the tests the coding agent writes.
253
+
254
+ ```yaml
255
+ logins:
256
+ shop:
257
+ url: https://staging.shop.test/login
258
+ username_env: SHOP_USER
259
+ password_env: SHOP_PASSWORD
260
+ totp_env: SHOP_TOTP_SECRET # optional
261
+ ```
262
+
263
+ Credentials in a message or a test case work too: ForgeQE's agent remembers them for the session as a login for the site and its sign-in hosts, so goals and reports never contain the password, the exact value is typed, and follow-up runs start signed in.
264
+
265
+ For single sign-on, captchas, or security keys, sign in by hand once: `forgeqe login shop` opens a browser, you sign in, press Enter, and the session is saved. `forgeqe logins` lists logins and sessions; `forgeqe logins clear shop` forgets one. Session files are private to your user and have their own `.gitignore`.
266
+
267
+ ## Azure DevOps, Jira, and other systems
268
+
269
+ ForgeQE's agent connects to each system's MCP server, configured from the same `sources:` settings in `forgeqe.yaml`: organization, site URL, and the names of the environment variables that hold the tokens. You don't write any MCP config yourself. Azure DevOps's server is the organization, so a project is optional and only limits ForgeQE's own test, sprint, and bug calls.
270
+
271
+ | System | MCP server | Needs |
272
+ |--------|------------|-------|
273
+ | Azure DevOps Services | Microsoft's [Azure DevOps MCP server](https://github.com/microsoft/azure-devops-mcp) (`npx @azure-devops/mcp`), with all its tools: work items, boards, test plans, pipelines, repos, wiki pages, search, and security alerts | Node.js, and `ADO_PAT` or `az login` |
274
+ | Jira Cloud, Server/Data Center | [mcp-atlassian](https://github.com/sooperset/mcp-atlassian) (`uvx`), with issues, JQL search, comments, and transitions. `mcp: rovo` uses [Atlassian's hosted server](https://github.com/atlassian/atlassian-mcp-server) instead (Cloud, and an admin must allow API tokens) | uv; `JIRA_EMAIL` and `JIRA_API_TOKEN` (Cloud) or a personal access token (Server/DC) |
275
+ | Confluence | mcp-atlassian, shared with Jira (or its own when Jira isn't set up); Atlassian's hosted server covers it on the same site | `confluence: true` on a Jira Cloud site, or `confluence.url` + tokens |
276
+ | Xray, TestRail, Zephyr Scale, qTest | none; ForgeQE's API | `XRAY_CLIENT_ID` / `XRAY_CLIENT_SECRET`, `TESTRAIL_EMAIL` / `TESTRAIL_API_KEY`, `ZEPHYR_API_TOKEN`, `QTEST_TOKEN` |
277
+ | Anything else | servers you add under `mcp:` | |
278
+
279
+ ```yaml
280
+ sources:
281
+ azure_devops: { org_url: https://dev.azure.com/acme } # every project; token in ADO_PAT. project: Shop limits ForgeQE's own calls
282
+ jira: { url: https://acme.atlassian.net } # JIRA_EMAIL + JIRA_API_TOKEN
283
+ mcp:
284
+ github:
285
+ url: https://api.githubcopilot.com/mcp/
286
+ headers: { Authorization: "Bearer {env:GITHUB_TOKEN}" }
287
+ ```
288
+
289
+ Tokens never go into the config: servers get them through `{env:NAME}` references that the agent runtime fills in, and the agent never sees them. `forgeqe mcp list` shows what your settings give. When the agent starts, the shell and `forgeqe run` print which servers connected.
290
+
291
+ Whatever has no MCP server, or whose server didn't connect, goes through ForgeQE's own tools, which call the REST APIs directly: `find_test_cases` and `run_test_cases` for test cases and posting results back, `read_confluence` for Confluence pages, and `query_api` to read anything else from Azure DevOps, Jira, Xray, or Confluence. `query_api` only reads (GET requests and Xray GraphQL queries) and only reaches the configured hosts. `sources.azure_devops.mcp: false` or `sources.jira.mcp: false` uses the API only.
292
+
293
+ Microsoft's Azure DevOps server can't create dashboards, so "create a dashboard in QI-Payments with all defects and this sprint's stories" goes through ForgeQE's `create_ado_dashboard`: one shared query per widget under `Shared Queries/ForgeQE`, then a dashboard with a work item list or count tile for each. Your `ADO_PAT` needs Work Items (read and write) and permission to edit dashboards. Chart widgets aren't made through the API; for those the agent uses the browser, signed in to `dev.azure.com` with a saved login.
294
+
295
+ Its wiki tools can only write into a wiki that already exists, so wiki pages go through ForgeQE's `write_ado_wiki_page`. If the project has no wiki yet, that tool creates the project wiki first, along with any missing parent pages. The PAT needs Wiki (read and write).
296
+
297
+ Requirements work too: "write and run test cases for the checkout rules page in Confluence" makes the agent read the page, write one case per rule or acceptance criterion (with the main negative paths), show them, and run them. It saves them into a test system only when you ask.
298
+
299
+ The agent creates or changes items (a bug, a comment, a status change) only when you ask it to; bugs for failed cases come from `sources.bugs`, not from the agent.
300
+
301
+ ## Pipelines
302
+
303
+ Ask for the pipeline you want and the agent builds it:
304
+
305
+ ```bash
306
+ forgeqe run "Add a GitHub Actions workflow that runs the smoke cases in cases.yaml on every pull request"
307
+ forgeqe run "Set up an Azure DevOps pipeline that runs test plan 12 every night and posts results back"
308
+ forgeqe run "The ForgeQE pipeline on main is failing, fix it"
309
+ ```
310
+
311
+ It works in the repo (the working directory, or `--repo`): it reads the repo's existing CI, writes or updates the workflow so it installs Bun, uv and ForgeQE, calls `forgeqe suite <source...> --no-code --plain`, and publishes `junit.xml` and the report, and validates it (actionlint when it is installed). Then it commits to a new branch named `forgeqe/ci-<topic>`, stores the keys the pipeline needs as secrets, triggers the run with `gh` or `az`, watches it, and fixes failures until the run passes (up to `code.max_attempts` failed runs).
312
+
313
+ Guard rails:
314
+ - Pushing goes through ForgeQE's `push_branch` tool, which refuses the default branch, `main`, and `master`, and never forces. `git push` itself is blocked for the agent.
315
+ - Secrets go through `set_ci_secret`: the agent names an environment variable and ForgeQE passes its value straight to `gh secret set` or `az pipelines variable`. The agent never sees the value.
316
+ - It needs `gh` (GitHub, logged in) or `az` with the `azure-devops` extension (Azure DevOps, logged in). Without them it still writes, validates, and commits the pipeline, then tells you what to run.
317
+
318
+ ## Monitoring
319
+
320
+ `forgeqe monitor` runs checks on a schedule and tells your team when one breaks. A check is any test case `forgeqe suite` takes: a YAML file, a Jira key, an Azure DevOps plan. Each round replays the flows from page memory, so a flow that passed before needs one model call to confirm the result (or none with `replay: trust`).
321
+
322
+ ```yaml
323
+ monitor:
324
+ checks: [checks/smoke.yaml, "jira:jql=labels = monitor"]
325
+ every: 15m
326
+ fail_after: 2 # two failing rounds in a row before the first alert, to ride out a flaky run
327
+ remind_every: 2h # repeat while it keeps failing
328
+ alerts:
329
+ teams: { webhook_env: TEAMS_WEBHOOK_URL }
330
+ email: { host: smtp.office365.com, port: 587, to: [qa-team@example.com] } # SMTP_USERNAME, SMTP_PASSWORD
331
+ ```
332
+
333
+ ```bash
334
+ forgeqe monitor --test-alerts # a sample alert, to check the webhook and SMTP settings
335
+ forgeqe monitor # every 15 minutes until Ctrl+C
336
+ forgeqe monitor --once # one round, for cron or a scheduled pipeline
337
+ ```
338
+
339
+ An alert goes out when a check starts failing, again every `remind_every` while it fails, and once when it recovers. One round sends one message covering every check that changed. Teams gets an Adaptive Card (both incoming webhooks and Workflows webhooks accept it) and email gets the failing checks' screenshots attached. Each alert links to the round's report: the local file, or `<report_url>/<round>/report.html` when you publish `.forgeqe/reports/monitor/` somewhere. A round that cannot run at all (a test system is down, the engine will not start) alerts as the check "ForgeQE monitor".
340
+
341
+ Rounds go to `.forgeqe/reports/monitor/<time>/`, the same as suites, so the report shows trends and flaky checks over the rounds; the last `keep` rounds (100) are kept. `state.json` remembers which checks are failing and since when, so with `--once` in a pipeline, cache that folder between runs. Only one monitor can use a folder at a time. `--once` exits 0 when every check passed and 1 otherwise. Results are posted back to the test system only with `write_back: true`. In the shell, "monitor the smoke cases every 10 minutes" runs the rounds in the background while you keep working, and each round shows as a card; they stop when the shell closes.
342
+
343
+ ## MCP
344
+
345
+ `forgeqe mcp` serves ForgeQE's tools over stdio for other agents such as Claude Code, Cursor, or Codex:
346
+
347
+ ```json
348
+ { "mcpServers": { "forgeqe": { "command": "forgeqe", "args": ["mcp", "--repo", "/path/to/app"] } } }
349
+ ```
350
+
351
+ Tools: `browser_run`, `find_test_cases`, `save_test_cases`, `run_test_cases`, `run_details`, `run_test`, `page_memory`, `read_confluence`, `query_api`, `push_branch`, `set_ci_secret`, `find_tests`, `run_tests` (runs only with `confirmed: true`, once the user agreed), `review_pull_request`, `post_pr_review` (posts only with `confirmed: true`), `list_sprints`, and one tool per command: `write_test`, `heal_tests`, `monitor`, `logins`, `forgeqe_setup`. Browser runs and suites log their progress to stderr, and long calls send progress notifications.