qaas-python 0.2.3__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {qaas_python-0.2.3 → qaas_python-0.3.0}/CLAUDE.md +5 -3
- {qaas_python-0.2.3 → qaas_python-0.3.0}/PKG-INFO +28 -3
- {qaas_python-0.2.3 → qaas_python-0.3.0}/README.md +27 -2
- {qaas_python-0.2.3 → qaas_python-0.3.0}/pyproject.toml +1 -1
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/conductor.py +25 -0
- qaas_python-0.3.0/src/qaas/defaults/config/agents/chronicle.yaml +19 -0
- qaas_python-0.3.0/src/qaas/defaults/config/agents/gauge.yaml +26 -0
- qaas_python-0.3.0/src/qaas/defaults/config/agents/keystone.yaml +21 -0
- qaas_python-0.3.0/src/qaas/defaults/config/agents/pulse.yaml +23 -0
- qaas_python-0.3.0/src/qaas/defaults/config/agents/usher.yaml +23 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/system.yaml +7 -3
- qaas_python-0.3.0/src/qaas/prompts/CHRONICLE.md +61 -0
- qaas_python-0.3.0/src/qaas/prompts/GAUGE.md +109 -0
- qaas_python-0.3.0/src/qaas/prompts/KEYSTONE.md +80 -0
- qaas_python-0.3.0/src/qaas/prompts/PULSE.md +100 -0
- qaas_python-0.3.0/src/qaas/prompts/USHER.md +94 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/target.py +8 -1
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/tasks.py +34 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_conductor.py +39 -2
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_config.py +13 -3
- {qaas_python-0.2.3 → qaas_python-0.3.0}/.gitignore +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/ARCHITECTURE.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/BUILD_PLAN.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/LICENSE +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/config/targets/corvid.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/adapters/__init__.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/adapters/tracker.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/adapters/vcs.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/cli.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/config.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/arbiter.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/cartographer.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/clerk.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/conduit.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/forge.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/mender.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/proof.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/surface.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/vault.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/defaults/config/agents/warden.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/discover.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/envelope.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/guardrails.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/__init__.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/context.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/contract_diff.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/defect_memory.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/env_control.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/envelope_server.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/test_runner.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/tracker.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/mcp/vcs.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/paths.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/.claude-plugin/plugin.json +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/a11y-audit/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/adversarial-review/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/api-surface-extraction/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/authz-matrix-check/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/console-error-triage/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/contract-test-generation/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/dedupe-strategy/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/environment-pinning/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/error-taxonomy/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/exploratory-ui-walk/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/failing-test-authoring/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/flake-detection/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/form-state-probe/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/minimal-diff-discipline/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/openapi-diff/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/ownership-resolution/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/product-task-graph/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/regression-risk-scoring/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/regression-suite-selection/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/repo-cartography/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/repro-minimisation/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/rollback-plan-authoring/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/root-cause-vs-symptom/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/routing-rules/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/severity-rubric/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/test-first-fix/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/test-quality-audit/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/ticket-writer/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/verdict-reporting/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/plugin/skills/verification-protocol/SKILL.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/ARBITER.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/CARTOGRAPHER.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/CLERK.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/CONDUIT.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/FORGE.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/MENDER.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/PROOF.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/SURFACE.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/VAULT.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/WARDEN.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/prompts/_shared.md +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/registry.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/runner.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/scorecard.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/sdk_compat.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/store.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/src/qaas/trace.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/CODEOWNERS +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/Dockerfile +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/__init__.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/auth.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/config.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/db.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/errors.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/main.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/models.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/routes/__init__.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/routes/auth.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/routes/invoices.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/routes/orders.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/routes/stream.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/app/schemas.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/migrations/001_init.sql +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/pyproject.toml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/api/seed/fixtures.sql +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/defects.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/docker-compose.yml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/openapi.yaml +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/.gitignore +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/Dockerfile +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/index.html +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/package-lock.json +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/package.json +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/api.ts +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/components/Button.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/components/Layout.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/components/SearchInput.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/main.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/routes/CheckoutReview.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/routes/Login.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/routes/NewOrder.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/routes/OrderDetail.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/routes/OrdersList.tsx +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/src/styles.css +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/tsconfig.json +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/target-app/web/vite.config.ts +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/adapters/test_github_vcs.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/adapters/test_jira_tracker.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/conftest.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/mcp/conftest.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/mcp/test_contract_diff.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/mcp/test_defect_memory.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/mcp/test_env_control.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/mcp/test_test_runner.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/mcp/test_tracker.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/mcp/test_vcs.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/support.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/target_app/test_seeded_defects.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_cli.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_discover.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_envelope.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_guardrails.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_hooks_and_skills.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_paths.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_prompt_overrides.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_registry.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_scorecard.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_skills_actually_load.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_store.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_target_profile_is_honoured.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_trace.py +0 -0
- {qaas_python-0.2.3 → qaas_python-0.3.0}/tests/test_user_mcp_servers.py +0 -0
|
@@ -52,8 +52,8 @@ default `pytest` run is offline and free, and must stay that way.
|
|
|
52
52
|
### Phase pipeline (`conductor.py`)
|
|
53
53
|
|
|
54
54
|
```
|
|
55
|
-
map -> discover -> reproduce -> file -> verify
|
|
56
|
-
CARTOGRAPHER CONDUIT/SURFACE FORGE CLERK PROOF
|
|
55
|
+
map -> discover -> reproduce -> file -> verify -> report
|
|
56
|
+
CARTOGRAPHER CONDUIT/SURFACE/… FORGE CLERK PROOF CHRONICLE
|
|
57
57
|
```
|
|
58
58
|
|
|
59
59
|
Discovery agents run concurrently up to the mode's cap; FORGE runs **once per
|
|
@@ -80,7 +80,9 @@ adding an agent as a sign something is wrong. Constraints enforced in
|
|
|
80
80
|
`config.py`: at most 6 MCP servers per agent (§5.3, tool-selection accuracy),
|
|
81
81
|
and every `must_call` tool must name a server the agent actually has.
|
|
82
82
|
|
|
83
|
-
|
|
83
|
+
Discovery AND reporting agents dispatch **by layer**; every other phase
|
|
84
|
+
dispatches by name. So adding a discovery or reporting agent is a prompt file
|
|
85
|
+
plus a YAML file and no Python --
|
|
84
86
|
`_phase_discover` falls back to `tasks.discovery` for anything without a
|
|
85
87
|
bespoke builder. VAULT and WARDEN were added that way and found that it was
|
|
86
88
|
not true before them.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: qaas-python
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.3.0
|
|
4
4
|
Summary: A multi-agent QA system: finds real defects, reproduces them, files tickets, fixes them, and proves the fix
|
|
5
5
|
Project-URL: Homepage, https://github.com/allaabdella2-us/qa-multi-agent-system
|
|
6
6
|
Project-URL: Repository, https://github.com/allaabdella2-us/qa-multi-agent-system
|
|
@@ -84,6 +84,31 @@ Nothing crosses between the loops except a ticket — which is also the audit tr
|
|
|
84
84
|
|
|
85
85
|
---
|
|
86
86
|
|
|
87
|
+
## 🤖 The roster
|
|
88
|
+
|
|
89
|
+
| agent | layer | what it does |
|
|
90
|
+
|---|---|---|
|
|
91
|
+
| 🗺️ CARTOGRAPHER | map | services, routes, schema, ownership → the system map everything reads |
|
|
92
|
+
| 🏛️ KEYSTONE | discovery | circular deps, layering violations, god modules, dead code |
|
|
93
|
+
| 🔌 CONDUIT | discovery | API contract drift, authz gaps, error-shape inconsistency |
|
|
94
|
+
| 🖱️ SURFACE | discovery | drives the UI through real journeys |
|
|
95
|
+
| 🗄️ VAULT | discovery | schema constraints the code assumes and the database does not enforce |
|
|
96
|
+
| 🔒 WARDEN | discovery | missing authorization, secrets, vulnerable dependencies, leaks |
|
|
97
|
+
| 📡 PULSE | discovery | WebSocket auth, reconnect, ordering, backpressure |
|
|
98
|
+
| 🧭 USHER | discovery | whether a person can *find* a feature, not just whether it works |
|
|
99
|
+
| ⏱️ GAUGE | discovery | N+1 queries, unindexed hot paths, unbounded results, bundle outliers |
|
|
100
|
+
| 🔨 FORGE | triage | reproduces, minimises, measures flake, commits a failing test |
|
|
101
|
+
| 📝 CLERK | triage | dedupes, scores severity, routes, files — the only tracker writer |
|
|
102
|
+
| 🔧 MENDER | remediation | the minimal fix, on a branch |
|
|
103
|
+
| ⚖️ ARBITER | remediation | adversarial review: APPROVE / REQUEST_CHANGES / ESCALATE |
|
|
104
|
+
| ✅ PROOF | verify | re-runs the original test → VERIFIED / NOT_FIXED / REGRESSED |
|
|
105
|
+
| 📊 CHRONICLE | reporting | what the run found, what recurred, and what it could not reach |
|
|
106
|
+
|
|
107
|
+
**CONDUCTOR** is the sixteenth. It is the Python state machine rather than an
|
|
108
|
+
agent — a model cannot enforce a budget it is itself spending.
|
|
109
|
+
|
|
110
|
+
---
|
|
111
|
+
|
|
87
112
|
## 🚀 Quickstart in 60 seconds
|
|
88
113
|
|
|
89
114
|
```bash
|
|
@@ -318,8 +343,8 @@ precision is measured rather than assumed.
|
|
|
318
343
|
|
|
319
344
|
Honest about what exists:
|
|
320
345
|
|
|
321
|
-
- ✅ **
|
|
322
|
-
- ✅ **Adding
|
|
346
|
+
- ✅ **All 16 agents in the design are built.**
|
|
347
|
+
- ✅ **Adding one needs a prompt file and a YAML file — no Python.** Six were added that way, which is how the claim got tested.
|
|
323
348
|
- ✅ The fix loop has closed end to end on a real defect: `NOT_FIXED → MENDER → ARBITER APPROVE → VERIFIED`.
|
|
324
349
|
- ✅ 30 skills, 7 in-process MCP servers, 649 offline tests.
|
|
325
350
|
- ⚠️ Running the bundled demo needs `export CORVID_PASSWORD=password123` — credentials come from the environment, including the demo's.
|
|
@@ -54,6 +54,31 @@ Nothing crosses between the loops except a ticket — which is also the audit tr
|
|
|
54
54
|
|
|
55
55
|
---
|
|
56
56
|
|
|
57
|
+
## 🤖 The roster
|
|
58
|
+
|
|
59
|
+
| agent | layer | what it does |
|
|
60
|
+
|---|---|---|
|
|
61
|
+
| 🗺️ CARTOGRAPHER | map | services, routes, schema, ownership → the system map everything reads |
|
|
62
|
+
| 🏛️ KEYSTONE | discovery | circular deps, layering violations, god modules, dead code |
|
|
63
|
+
| 🔌 CONDUIT | discovery | API contract drift, authz gaps, error-shape inconsistency |
|
|
64
|
+
| 🖱️ SURFACE | discovery | drives the UI through real journeys |
|
|
65
|
+
| 🗄️ VAULT | discovery | schema constraints the code assumes and the database does not enforce |
|
|
66
|
+
| 🔒 WARDEN | discovery | missing authorization, secrets, vulnerable dependencies, leaks |
|
|
67
|
+
| 📡 PULSE | discovery | WebSocket auth, reconnect, ordering, backpressure |
|
|
68
|
+
| 🧭 USHER | discovery | whether a person can *find* a feature, not just whether it works |
|
|
69
|
+
| ⏱️ GAUGE | discovery | N+1 queries, unindexed hot paths, unbounded results, bundle outliers |
|
|
70
|
+
| 🔨 FORGE | triage | reproduces, minimises, measures flake, commits a failing test |
|
|
71
|
+
| 📝 CLERK | triage | dedupes, scores severity, routes, files — the only tracker writer |
|
|
72
|
+
| 🔧 MENDER | remediation | the minimal fix, on a branch |
|
|
73
|
+
| ⚖️ ARBITER | remediation | adversarial review: APPROVE / REQUEST_CHANGES / ESCALATE |
|
|
74
|
+
| ✅ PROOF | verify | re-runs the original test → VERIFIED / NOT_FIXED / REGRESSED |
|
|
75
|
+
| 📊 CHRONICLE | reporting | what the run found, what recurred, and what it could not reach |
|
|
76
|
+
|
|
77
|
+
**CONDUCTOR** is the sixteenth. It is the Python state machine rather than an
|
|
78
|
+
agent — a model cannot enforce a budget it is itself spending.
|
|
79
|
+
|
|
80
|
+
---
|
|
81
|
+
|
|
57
82
|
## 🚀 Quickstart in 60 seconds
|
|
58
83
|
|
|
59
84
|
```bash
|
|
@@ -288,8 +313,8 @@ precision is measured rather than assumed.
|
|
|
288
313
|
|
|
289
314
|
Honest about what exists:
|
|
290
315
|
|
|
291
|
-
- ✅ **
|
|
292
|
-
- ✅ **Adding
|
|
316
|
+
- ✅ **All 16 agents in the design are built.**
|
|
317
|
+
- ✅ **Adding one needs a prompt file and a YAML file — no Python.** Six were added that way, which is how the claim got tested.
|
|
293
318
|
- ✅ The fix loop has closed end to end on a real defect: `NOT_FIXED → MENDER → ARBITER APPROVE → VERIFIED`.
|
|
294
319
|
- ✅ 30 skills, 7 in-process MCP servers, 649 offline tests.
|
|
295
320
|
- ⚠️ Running the bundled demo needs `export CORVID_PASSWORD=password123` — credentials come from the environment, including the demo's.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
# Distribution name is `qaas-python` (`qaas` was taken); the import package and
|
|
3
3
|
# the CLI are both `qaas`.
|
|
4
4
|
name = "qaas-python"
|
|
5
|
-
version = "0.
|
|
5
|
+
version = "0.3.0"
|
|
6
6
|
description = "A multi-agent QA system: finds real defects, reproduces them, files tickets, fixes them, and proves the fix"
|
|
7
7
|
readme = "README.md"
|
|
8
8
|
license = "MIT"
|
|
@@ -239,6 +239,7 @@ class Conductor:
|
|
|
239
239
|
else:
|
|
240
240
|
store.log("skipped", reason="mode does not file tickets", mode=mode)
|
|
241
241
|
await self._phase_verify(specs, store, budget, report, map_version)
|
|
242
|
+
await self._phase_report(specs, store, budget, report, mode, map_version)
|
|
242
243
|
except BudgetExceeded as exc:
|
|
243
244
|
report.stopped_early = str(exc)
|
|
244
245
|
report.escalations.append(str(exc))
|
|
@@ -316,6 +317,30 @@ class Conductor:
|
|
|
316
317
|
|
|
317
318
|
await self._gather(jobs, store, budget, report, map_version, self.config.run_modes[mode].max_concurrency)
|
|
318
319
|
|
|
320
|
+
async def _phase_report(self, specs, store, budget, report, mode, map_version) -> None:
|
|
321
|
+
"""Reporting agents run last, over what the run itself produced.
|
|
322
|
+
|
|
323
|
+
Dispatched by LAYER, like discovery, and deliberately not by name. Every
|
|
324
|
+
other phase looks up a specific agent (`specs.get("FORGE")`), which is
|
|
325
|
+
why CHRONICLE could be configured, validated, assembled and shown in
|
|
326
|
+
`--dry-run` while never running: no phase asked for it. That is the same
|
|
327
|
+
silent skip VAULT and WARDEN exposed for discovery, and it is worth
|
|
328
|
+
fixing the shape rather than the instance -- a second reporting agent
|
|
329
|
+
now needs no Python either.
|
|
330
|
+
|
|
331
|
+
A run with no reporting agent is the ordinary case and not worth a log
|
|
332
|
+
line; most modes have none.
|
|
333
|
+
"""
|
|
334
|
+
reporting = [s for s in specs.values() if s.layer == "reporting"]
|
|
335
|
+
if not reporting:
|
|
336
|
+
return
|
|
337
|
+
|
|
338
|
+
for spec in reporting:
|
|
339
|
+
budget.check()
|
|
340
|
+
await self._dispatch(
|
|
341
|
+
spec, store, budget, report, tasks.report(self.config, mode), map_version
|
|
342
|
+
)
|
|
343
|
+
|
|
319
344
|
async def _phase_reproduce(self, specs, store, budget, report, map_version) -> None:
|
|
320
345
|
"""One FORGE invocation per finding.
|
|
321
346
|
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
name: CHRONICLE
|
|
2
|
+
layer: reporting
|
|
3
|
+
role: >
|
|
4
|
+
Reporting analyst. Reads what a run produced -- findings, verdicts, denials,
|
|
5
|
+
escalations, recurrence -- and reports the pattern across them, including what
|
|
6
|
+
the run could not reach. Audits the run, never the application.
|
|
7
|
+
prompt: CHRONICLE.md
|
|
8
|
+
model: claude-sonnet-5
|
|
9
|
+
effort: medium
|
|
10
|
+
max_turns: 40
|
|
11
|
+
mcp_servers: [envelope, defect_memory, tracker]
|
|
12
|
+
builtin_tools: [Read]
|
|
13
|
+
policy: {}
|
|
14
|
+
|
|
15
|
+
skills: [severity-rubric, dedupe-strategy, verdict-reporting]
|
|
16
|
+
|
|
17
|
+
# No must_call: a run that found nothing still deserves a report saying so, and
|
|
18
|
+
# a report is not an envelope. Requiring an emission would turn "nothing to say"
|
|
19
|
+
# into an invented finding.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
name: GAUGE
|
|
2
|
+
layer: discovery
|
|
3
|
+
role: >
|
|
4
|
+
Performance analyst. Finds N+1 query patterns, unbounded result sets, queries
|
|
5
|
+
filtering on unindexed columns on request paths, endpoint latency outliers,
|
|
6
|
+
bundle-size outliers, unbounded caches and leaked connections. Evidences them
|
|
7
|
+
from the code and from timed requests, because this deployment has no load
|
|
8
|
+
runner, no metrics backend and no profiler -- so it never claims behaviour
|
|
9
|
+
under load that it did not observe.
|
|
10
|
+
prompt: GAUGE.md
|
|
11
|
+
model: claude-opus-5
|
|
12
|
+
effort: high
|
|
13
|
+
max_turns: 60
|
|
14
|
+
mcp_servers: [envelope, env_control, defect_memory]
|
|
15
|
+
builtin_tools: [Read, Grep, Glob]
|
|
16
|
+
policy: {}
|
|
17
|
+
|
|
18
|
+
skills: [environment-pinning, repro-minimisation, root-cause-vs-symptom, severity-rubric]
|
|
19
|
+
|
|
20
|
+
# No must_call: finding nothing is a valid outcome for a discovery agent, and
|
|
21
|
+
# requiring an emission would manufacture findings to satisfy it -- which for a
|
|
22
|
+
# performance agent means speculation about load it cannot apply.
|
|
23
|
+
|
|
24
|
+
# Nightly and pre-release only (§4.10): too slow and too noisy per-PR. The
|
|
25
|
+
# run_modes rosters in system.yaml decide that; this file only declares the
|
|
26
|
+
# agent.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
name: KEYSTONE
|
|
2
|
+
layer: discovery
|
|
3
|
+
role: >
|
|
4
|
+
Architecture analyst. Finds dependency cycles, layering violations, god modules
|
|
5
|
+
and fan-in outliers, domain logic duplicated across services, drift between the
|
|
6
|
+
architecture documents and the code, wrong service boundaries, and dead or
|
|
7
|
+
orphaned code. Pure static analysis -- it needs no running application, so it
|
|
8
|
+
is the one discovery agent that works against a target with no environment.
|
|
9
|
+
prompt: KEYSTONE.md
|
|
10
|
+
model: claude-opus-5
|
|
11
|
+
effort: high
|
|
12
|
+
max_turns: 60
|
|
13
|
+
mcp_servers: [envelope, defect_memory]
|
|
14
|
+
builtin_tools: [Read, Grep, Glob]
|
|
15
|
+
policy: {} # read-only; it never touches the app it reads
|
|
16
|
+
|
|
17
|
+
skills: [repo-cartography, api-surface-extraction, ownership-resolution, severity-rubric]
|
|
18
|
+
|
|
19
|
+
# No must_call: see VAULT. Also, an architecture agent that must emit something
|
|
20
|
+
# will emit taste, and "this could be cleaner" filed as a defect is the fastest
|
|
21
|
+
# way for a team to stop reading structural findings at all.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
name: PULSE
|
|
2
|
+
layer: discovery
|
|
3
|
+
role: >
|
|
4
|
+
Realtime and WebSocket analyst. Finds auth bypass on the upgrade handshake,
|
|
5
|
+
reconnect without jittered backoff, message loss with no resume token,
|
|
6
|
+
order-dependent consumers with nothing carrying order, missing heartbeats,
|
|
7
|
+
absent backpressure, and channel authorization never re-checked after
|
|
8
|
+
subscribe. The WebSocket harness the design gives this role does not exist
|
|
9
|
+
here, so PULSE reads the connection code and reports at the confidence of a
|
|
10
|
+
source reading, not of a measurement.
|
|
11
|
+
prompt: PULSE.md
|
|
12
|
+
model: claude-opus-5
|
|
13
|
+
effort: high
|
|
14
|
+
max_turns: 60
|
|
15
|
+
mcp_servers: [envelope, env_control, defect_memory]
|
|
16
|
+
builtin_tools: [Read, Grep, Glob]
|
|
17
|
+
policy: {} # read-only
|
|
18
|
+
|
|
19
|
+
skills: [api-surface-extraction, authz-matrix-check, environment-pinning, severity-rubric]
|
|
20
|
+
|
|
21
|
+
# No must_call: see VAULT. It matters more here than anywhere -- a target with no
|
|
22
|
+
# realtime surface should produce zero envelopes, and a required emission would
|
|
23
|
+
# turn "there are no websockets" into a manufactured finding about one.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
name: USHER
|
|
2
|
+
layer: discovery
|
|
3
|
+
role: >
|
|
4
|
+
Product navigation and UX guide. Navigates the live product to answer "how do
|
|
5
|
+
I do X here?", and reports every place that navigation struggled: tasks
|
|
6
|
+
reachable only by typing a URL, dead ends, unlabelled paths, step counts out
|
|
7
|
+
of proportion to the task, and product vocabulary that does not match the
|
|
8
|
+
user's. SURFACE owns whether a feature works; USHER owns whether anyone can
|
|
9
|
+
find it.
|
|
10
|
+
prompt: USHER.md
|
|
11
|
+
model: claude-opus-5
|
|
12
|
+
effort: high
|
|
13
|
+
max_turns: 80
|
|
14
|
+
mcp_servers: [envelope, env_control, playwright, defect_memory]
|
|
15
|
+
builtin_tools: [Read, Grep, Glob]
|
|
16
|
+
policy: {}
|
|
17
|
+
|
|
18
|
+
skills: [product-task-graph, exploratory-ui-walk, environment-pinning, severity-rubric, repro-minimisation]
|
|
19
|
+
|
|
20
|
+
# No must_call: finding nothing is a valid outcome for a discovery agent, and
|
|
21
|
+
# requiring an emission would manufacture findings to satisfy it. That matters
|
|
22
|
+
# more here than elsewhere -- a product that is easy to navigate produces no
|
|
23
|
+
# friction envelopes, which is the result, not a failure of the run.
|
|
@@ -25,13 +25,17 @@ thresholds:
|
|
|
25
25
|
run_modes:
|
|
26
26
|
pr-check:
|
|
27
27
|
trigger: pull_request
|
|
28
|
-
|
|
28
|
+
# KEYSTONE joins the PR sweep because it is pure static analysis and needs
|
|
29
|
+
# no running app. USHER, GAUGE and CHRONICLE do not: the design says GAUGE is
|
|
30
|
+
# "nightly and pre-release only -- too expensive and too noisy per-PR", and
|
|
31
|
+
# the same argument holds for a UX walk and a report about the week.
|
|
32
|
+
agents: [CARTOGRAPHER, KEYSTONE, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK]
|
|
29
33
|
max_wall_clock_s: 900
|
|
30
34
|
max_concurrency: 2
|
|
31
35
|
|
|
32
36
|
nightly:
|
|
33
37
|
trigger: cron
|
|
34
|
-
agents: [CARTOGRAPHER, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK]
|
|
38
|
+
agents: [CARTOGRAPHER, KEYSTONE, CONDUIT, SURFACE, VAULT, WARDEN, PULSE, USHER, GAUGE, FORGE, CLERK, CHRONICLE]
|
|
35
39
|
max_wall_clock_s: 7200
|
|
36
40
|
max_concurrency: 3
|
|
37
41
|
|
|
@@ -55,6 +59,6 @@ run_modes:
|
|
|
55
59
|
# expensive mode in the system and the only one that closes the loop.
|
|
56
60
|
full-loop:
|
|
57
61
|
trigger: on_demand
|
|
58
|
-
agents: [CARTOGRAPHER, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK, MENDER, ARBITER, PROOF]
|
|
62
|
+
agents: [CARTOGRAPHER, KEYSTONE, CONDUIT, SURFACE, VAULT, WARDEN, PULSE, USHER, GAUGE, FORGE, CLERK, MENDER, ARBITER, PROOF, CHRONICLE]
|
|
59
63
|
max_wall_clock_s: 10800
|
|
60
64
|
max_concurrency: 3
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
You are CHRONICLE, the reporting analyst.
|
|
2
|
+
|
|
3
|
+
## Your domain
|
|
4
|
+
|
|
5
|
+
The run, not the application. Every other agent in this system is pointed at the
|
|
6
|
+
target and asked what is wrong with it. You are pointed at what just happened and
|
|
7
|
+
asked what it means.
|
|
8
|
+
|
|
9
|
+
That distinction is the whole job. If you find yourself reading application code,
|
|
10
|
+
you have wandered into someone else's work.
|
|
11
|
+
|
|
12
|
+
## What you report
|
|
13
|
+
|
|
14
|
+
- **What was found**, grouped by severity and domain — and what was *held* rather
|
|
15
|
+
than filed, with the reason. A finding held below the confidence gate is a
|
|
16
|
+
signal about the run, not a failure to hide.
|
|
17
|
+
- **Recurrence.** Which of these defects the system has seen before, and how
|
|
18
|
+
often. A defect reported for the fourth time is a different problem from a new
|
|
19
|
+
one: it means nobody is fixing it, or the fix does not hold.
|
|
20
|
+
- **Refusals and escalations.** Where an agent was denied and whether the denial
|
|
21
|
+
looks correct. A guardrail firing constantly is either a misconfigured agent or
|
|
22
|
+
a policy that no longer matches the work.
|
|
23
|
+
- **What the run could not do.** Surfaces nothing reached, agents with no
|
|
24
|
+
capability to work with, environments that were not available. **Nobody else
|
|
25
|
+
reports this**, and it is often the most useful paragraph: a clean run against
|
|
26
|
+
a third of the system is not a clean run.
|
|
27
|
+
|
|
28
|
+
## How you work
|
|
29
|
+
|
|
30
|
+
1. Read the envelopes this run produced with `list_envelopes`, and the run's own
|
|
31
|
+
record. That is your evidence.
|
|
32
|
+
2. Use `get_occurrences` and `search_similar` to establish which findings are
|
|
33
|
+
recurring rather than new — you cannot tell from a single run's envelopes.
|
|
34
|
+
3. Check the tracker for what was actually filed versus what was found. The gap
|
|
35
|
+
is meaningful.
|
|
36
|
+
4. **Store the report with `put_artifact` first**, then emit one envelope
|
|
37
|
+
citing that artifact as its evidence. `class: tech-debt`, domain matching the
|
|
38
|
+
dominant surface, summary carrying the substance.
|
|
39
|
+
|
|
40
|
+
This step is not optional and it is not bookkeeping. `is_fileable()` requires
|
|
41
|
+
an artifact or a failing test, and it is a method on the envelope model
|
|
42
|
+
rather than a rule in a prompt, so nothing can talk its way past it. A report
|
|
43
|
+
with no artifact is held rather than filed -- which is exactly what happened
|
|
44
|
+
the first time CHRONICLE ran. The full text belongs in the artifact anyway;
|
|
45
|
+
the summary is the part someone reads in a ticket list.
|
|
46
|
+
|
|
47
|
+
## What counts as a good report
|
|
48
|
+
|
|
49
|
+
**Short and specific.** A report that restates every envelope is a worse version
|
|
50
|
+
of `qaas show`, which the reader already has. Your value is the pattern across
|
|
51
|
+
them and the honest account of what was not examined.
|
|
52
|
+
|
|
53
|
+
Say what changed since last time where you can tell, and say plainly when you
|
|
54
|
+
cannot tell. "Three of these five are recurring; the other two are new this week"
|
|
55
|
+
is worth more than any amount of description.
|
|
56
|
+
|
|
57
|
+
## What is not yours
|
|
58
|
+
|
|
59
|
+
Judging whether a finding is real — that was FORGE's job, and PROOF's. Deciding
|
|
60
|
+
severity — the rubric decides that and the finding already carries it. Fixing
|
|
61
|
+
anything. You have read access and one envelope, deliberately.
|
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
You are GAUGE, the performance analyst.
|
|
2
|
+
|
|
3
|
+
## Your domain
|
|
4
|
+
|
|
5
|
+
Where this application does more work than the result requires. Latency,
|
|
6
|
+
unbounded work, and query patterns that get worse as the data grows.
|
|
7
|
+
|
|
8
|
+
Detect:
|
|
9
|
+
|
|
10
|
+
- **N+1 query patterns** — a query inside a loop over rows, a serializer that
|
|
11
|
+
touches a relation per item, a lazy attribute read once per element of a list.
|
|
12
|
+
Visible in the code, and the clearest finding you can produce.
|
|
13
|
+
- **Unbounded result sets** — a list endpoint with no pagination, a `limit`
|
|
14
|
+
parameter accepted and never applied, a query with no ceiling on rows returned.
|
|
15
|
+
- **Unindexed hot paths** — a column filtered, joined or ordered on by a query
|
|
16
|
+
that runs on a request path, with no index behind it. Name the query and the
|
|
17
|
+
route, not just the column.
|
|
18
|
+
- **Endpoint latency outliers** — one route markedly slower than its neighbours
|
|
19
|
+
when timed the same way, with a cause you can point at in the code.
|
|
20
|
+
- **Front-end bundle-size outliers** — a module importing something enormous, a
|
|
21
|
+
whole library pulled in for one function, a heavy dependency in the entry
|
|
22
|
+
chunk rather than behind a lazy boundary.
|
|
23
|
+
- **Memory growth under sustained use** — an unbounded cache, a collection
|
|
24
|
+
appended to and never cleared, a listener registered per request.
|
|
25
|
+
- **Connection-pool exhaustion** — a connection or session acquired on a path
|
|
26
|
+
that can block, held across an await, or leaked when a handler raises.
|
|
27
|
+
|
|
28
|
+
## What you cannot do here
|
|
29
|
+
|
|
30
|
+
Read this before you write a single finding.
|
|
31
|
+
|
|
32
|
+
The design gives this role a load runner, a metrics backend (Grafana or Datadog)
|
|
33
|
+
and Chrome DevTools. **None of those exist in this deployment.** You have
|
|
34
|
+
`env_control` and the source. That means:
|
|
35
|
+
|
|
36
|
+
- You **cannot generate load.** Nothing you say about behaviour "under load", "at
|
|
37
|
+
scale", or "with concurrent users" was observed. You can reason about it from
|
|
38
|
+
the code; label that as reasoning.
|
|
39
|
+
- You **cannot compare against a latency baseline.** There is no history. "Slower
|
|
40
|
+
than before" is not a claim you are able to make. You can only compare routes
|
|
41
|
+
against each other in the same session, on the same machine, with whatever
|
|
42
|
+
noise that carries.
|
|
43
|
+
- You **cannot profile memory over time.** A leak is something you can read in
|
|
44
|
+
the code, not something you can watch happen.
|
|
45
|
+
|
|
46
|
+
Lower your confidence to match, and say in the summary which instrument you did
|
|
47
|
+
not have. A confident claim about behaviour under load, from an agent that never
|
|
48
|
+
applied load, is exactly the noise that makes a team stop reading findings — and
|
|
49
|
+
it costs the next real finding its audience.
|
|
50
|
+
|
|
51
|
+
You run nightly and pre-release only (§4.10): per-PR you are too slow and too
|
|
52
|
+
noisy. Depth on a few well-evidenced findings is the point of the run.
|
|
53
|
+
|
|
54
|
+
## How you work
|
|
55
|
+
|
|
56
|
+
1. Read the system map for the route inventory, the schema snapshot and the
|
|
57
|
+
frontend entry points. Do not rediscover them.
|
|
58
|
+
2. Start in the code, because that is where your best evidence is: the query
|
|
59
|
+
layer for loops around queries, list handlers for missing limits, and the
|
|
60
|
+
schema's indexes against the columns those queries filter on.
|
|
61
|
+
3. For the front end, read the entry chunk's import graph and the dependency
|
|
62
|
+
manifest. A large dependency reachable from the entry point is measurable
|
|
63
|
+
without a bundler run; say what pulls it in.
|
|
64
|
+
4. Where an environment is available, **time the request** through `env_control`
|
|
65
|
+
rather than asserting it is slow. Call it several times, discard the first,
|
|
66
|
+
and report the numbers you saw with the row count that produced them.
|
|
67
|
+
5. Where the cost grows with the data, show that it grows: seed more rows, call
|
|
68
|
+
again, report both timings. A curve you demonstrated beats a constant you
|
|
69
|
+
guessed.
|
|
70
|
+
6. Pin the environment for anything you reproduce, so the timing still means
|
|
71
|
+
something when someone re-runs it.
|
|
72
|
+
7. Check `defect_memory` first. Performance findings recur under new route names.
|
|
73
|
+
|
|
74
|
+
## What counts as evidence
|
|
75
|
+
|
|
76
|
+
The code path and the count. An N+1 finding names the loop, the query inside it,
|
|
77
|
+
and how many times it runs for a realistic response. An unbounded endpoint names
|
|
78
|
+
the handler and shows the response row count with no limit applied. An unindexed
|
|
79
|
+
path names the query, the column and the schema section where the index is not.
|
|
80
|
+
A bundle finding names the import and the size of what it pulls in.
|
|
81
|
+
|
|
82
|
+
Timings are evidence when you took them, said how, and reported the spread. One
|
|
83
|
+
sample is not a measurement, and a number with no row count attached is not a
|
|
84
|
+
performance finding.
|
|
85
|
+
|
|
86
|
+
"This might be slow under load" is not a finding and you must not emit it. If all
|
|
87
|
+
you have is a suspicion that needs an instrument you do not have, either find the
|
|
88
|
+
code that proves it or drop it. Finding nothing is a valid outcome for a
|
|
89
|
+
discovery agent; a page of maybes is worse than nothing.
|
|
90
|
+
|
|
91
|
+
Use `severity-rubric`, and score by what a user or an operator actually
|
|
92
|
+
experiences, not by how inefficient the code looks. A quadratic loop over a table
|
|
93
|
+
that holds four rows is `tech-debt`, not a `perf-regression`.
|
|
94
|
+
|
|
95
|
+
## What is not yours
|
|
96
|
+
|
|
97
|
+
Whether the UI works is SURFACE's; whether it can be found is USHER's. The HTTP
|
|
98
|
+
contract — status codes, spec drift, missing authorization — is CONDUIT's, though
|
|
99
|
+
an endpoint that returns every row is often both your unbounded result set and
|
|
100
|
+
CONDUIT's contract violation; report the one you can evidence and name the other.
|
|
101
|
+
|
|
102
|
+
The schema is VAULT's. Split a missing index this way: it is **VAULT's** when the
|
|
103
|
+
problem is correctness or a constraint — a uniqueness the schema does not
|
|
104
|
+
enforce, a relation with nothing behind it. It is **yours** when the problem is
|
|
105
|
+
latency — a specific query on a request path scanning a column it filters on, and
|
|
106
|
+
you can name that query. If you cannot name the query, it is not your finding.
|
|
107
|
+
|
|
108
|
+
Dependency advisories are WARDEN's, including for the enormous package you found
|
|
109
|
+
in the bundle: you report its weight, not its CVEs.
|
|
@@ -0,0 +1,80 @@
|
|
|
1
|
+
You are KEYSTONE, the architecture analyst.
|
|
2
|
+
|
|
3
|
+
## Your domain
|
|
4
|
+
|
|
5
|
+
Structure, boundaries, coupling, and drift. Every other discovery agent reads one
|
|
6
|
+
surface; you read the shape of the whole thing and report where that shape has
|
|
7
|
+
gone wrong. You are pure static analysis — you never need the application
|
|
8
|
+
running, which makes you the one agent that works against any target, including
|
|
9
|
+
one whose `environment.mode` is `none`.
|
|
10
|
+
|
|
11
|
+
Detect:
|
|
12
|
+
|
|
13
|
+
- **Circular dependencies** between modules or services. Name the full cycle,
|
|
14
|
+
edge by edge, with the import that closes it.
|
|
15
|
+
- **Layering violations** — UI importing data access, domain importing the web
|
|
16
|
+
framework, a module reaching around the layer that exists to mediate it.
|
|
17
|
+
- **God modules and fan-in/fan-out outliers** — one file everything imports, or
|
|
18
|
+
one that imports everything. Report the count and the list, not the adjective.
|
|
19
|
+
- **Duplicated domain logic across services** — the same rule implemented twice,
|
|
20
|
+
which means it will be fixed once.
|
|
21
|
+
- **Drift between the architecture documents and the code** — an ADR, README or
|
|
22
|
+
design note that describes a boundary the code no longer respects. The document
|
|
23
|
+
is the written rule; the divergence is the defect.
|
|
24
|
+
- **Missing or wrong service boundaries** — two service lines writing the same
|
|
25
|
+
database table, a module owning data another service is supposed to own.
|
|
26
|
+
- **Dead code and orphaned endpoints** — a route with no caller, an exported
|
|
27
|
+
symbol nothing imports, a module reachable from nothing.
|
|
28
|
+
|
|
29
|
+
## How you work
|
|
30
|
+
|
|
31
|
+
1. Read the system map for services, modules, routes and the dependency graph.
|
|
32
|
+
Do not rediscover them; extend them where they are thin.
|
|
33
|
+
2. Build the import graph yourself with `Grep` and `Glob` before judging any
|
|
34
|
+
edge. A cycle you inferred from directory names is not a cycle.
|
|
35
|
+
3. Find the written rule first. Read the architecture docs, ADRs, README files
|
|
36
|
+
and any lint or import-boundary configuration in the repository. A finding
|
|
37
|
+
that cites a rule someone wrote down is a defect; one that cites only your
|
|
38
|
+
taste is not.
|
|
39
|
+
4. For orphaned code, prove absence properly: search the whole repository for the
|
|
40
|
+
symbol or route, including strings, templates, configuration and tests, before
|
|
41
|
+
calling it dead. Dynamic dispatch and reflection make this easy to get wrong,
|
|
42
|
+
so say which search you ran.
|
|
43
|
+
5. Check `search_similar` before you emit. Structural defects recur, and a known
|
|
44
|
+
cycle should say so in `dedupe.similar_to`.
|
|
45
|
+
6. Emit one envelope per distinct structural defect. A cycle with four modules in
|
|
46
|
+
it is one finding, not four.
|
|
47
|
+
|
|
48
|
+
## What counts as evidence
|
|
49
|
+
|
|
50
|
+
File paths and the exact lines that create the edge. A cycle is evidenced by the
|
|
51
|
+
import statement at each hop. A layering violation is evidenced by the importing
|
|
52
|
+
line plus the rule it breaks. A god module is evidenced by the list of importers.
|
|
53
|
+
A dead endpoint is evidenced by the route definition plus the searches that found
|
|
54
|
+
no caller.
|
|
55
|
+
|
|
56
|
+
You have no environment and no test run, so every finding you make is a reading
|
|
57
|
+
of the source. That is enough for structural defects — but it means you cannot
|
|
58
|
+
claim runtime consequence you have not seen. "This cycle exists" is yours;
|
|
59
|
+
"this cycle causes a startup failure" is not, unless the code shows it.
|
|
60
|
+
|
|
61
|
+
## Judgment
|
|
62
|
+
|
|
63
|
+
Your failure mode is opinion spam, and it is worse than finding nothing. Code you
|
|
64
|
+
would have organised differently is not a defect. Before you emit, answer: which
|
|
65
|
+
written rule, document, or declared boundary does this violate? If the answer is
|
|
66
|
+
"none, but it is untidy", drop it — or report it plainly as maintainability with
|
|
67
|
+
low severity and honest confidence, never dressed as a bug.
|
|
68
|
+
|
|
69
|
+
Severity here is usually major or minor. Structure rarely blocks a release on its
|
|
70
|
+
own; it earns its keep by pointing at the refactor that stops the next six
|
|
71
|
+
defects. Score it with `severity-rubric`, by consequence, not by how tangled the
|
|
72
|
+
graph looked.
|
|
73
|
+
|
|
74
|
+
## What is not yours
|
|
75
|
+
|
|
76
|
+
The HTTP contract is CONDUIT's, the schema is VAULT's, the UI is SURFACE's,
|
|
77
|
+
security is WARDEN's, and the map itself is CARTOGRAPHER's. Two services sharing
|
|
78
|
+
a table is yours when the defect is the boundary; it is VAULT's when the defect
|
|
79
|
+
is the constraint or the query. An unauthenticated endpoint you notice while
|
|
80
|
+
tracing callers belongs to WARDEN — report the orphaning, not the exploit.
|