qaas-python 0.2.3__tar.gz → 0.3.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (166) hide show
  1. {qaas_python-0.2.3 → qaas_python-0.3.1}/CLAUDE.md +5 -3
  2. {qaas_python-0.2.3 → qaas_python-0.3.1}/PKG-INFO +29 -4
  3. {qaas_python-0.2.3 → qaas_python-0.3.1}/README.md +28 -3
  4. {qaas_python-0.2.3 → qaas_python-0.3.1}/pyproject.toml +1 -1
  5. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/conductor.py +25 -0
  6. qaas_python-0.3.1/src/qaas/defaults/config/agents/chronicle.yaml +19 -0
  7. qaas_python-0.3.1/src/qaas/defaults/config/agents/gauge.yaml +26 -0
  8. qaas_python-0.3.1/src/qaas/defaults/config/agents/keystone.yaml +21 -0
  9. qaas_python-0.3.1/src/qaas/defaults/config/agents/pulse.yaml +23 -0
  10. qaas_python-0.3.1/src/qaas/defaults/config/agents/usher.yaml +23 -0
  11. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/system.yaml +7 -3
  12. qaas_python-0.3.1/src/qaas/prompts/CHRONICLE.md +61 -0
  13. qaas_python-0.3.1/src/qaas/prompts/GAUGE.md +109 -0
  14. qaas_python-0.3.1/src/qaas/prompts/KEYSTONE.md +80 -0
  15. qaas_python-0.3.1/src/qaas/prompts/PULSE.md +100 -0
  16. qaas_python-0.3.1/src/qaas/prompts/USHER.md +94 -0
  17. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/registry.py +21 -0
  18. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/target.py +8 -1
  19. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/tasks.py +34 -0
  20. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_conductor.py +39 -2
  21. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_config.py +13 -3
  22. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_registry.py +27 -0
  23. {qaas_python-0.2.3 → qaas_python-0.3.1}/.gitignore +0 -0
  24. {qaas_python-0.2.3 → qaas_python-0.3.1}/ARCHITECTURE.md +0 -0
  25. {qaas_python-0.2.3 → qaas_python-0.3.1}/BUILD_PLAN.md +0 -0
  26. {qaas_python-0.2.3 → qaas_python-0.3.1}/LICENSE +0 -0
  27. {qaas_python-0.2.3 → qaas_python-0.3.1}/config/targets/corvid.yaml +0 -0
  28. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/adapters/__init__.py +0 -0
  29. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/adapters/tracker.py +0 -0
  30. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/adapters/vcs.py +0 -0
  31. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/cli.py +0 -0
  32. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/config.py +0 -0
  33. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/arbiter.yaml +0 -0
  34. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/cartographer.yaml +0 -0
  35. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/clerk.yaml +0 -0
  36. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/conduit.yaml +0 -0
  37. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/forge.yaml +0 -0
  38. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/mender.yaml +0 -0
  39. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/proof.yaml +0 -0
  40. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/surface.yaml +0 -0
  41. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/vault.yaml +0 -0
  42. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/defaults/config/agents/warden.yaml +0 -0
  43. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/discover.py +0 -0
  44. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/envelope.py +0 -0
  45. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/guardrails.py +0 -0
  46. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/__init__.py +0 -0
  47. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/context.py +0 -0
  48. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/contract_diff.py +0 -0
  49. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/defect_memory.py +0 -0
  50. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/env_control.py +0 -0
  51. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/envelope_server.py +0 -0
  52. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/test_runner.py +0 -0
  53. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/tracker.py +0 -0
  54. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/mcp/vcs.py +0 -0
  55. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/paths.py +0 -0
  56. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/.claude-plugin/plugin.json +0 -0
  57. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/a11y-audit/SKILL.md +0 -0
  58. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/adversarial-review/SKILL.md +0 -0
  59. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/api-surface-extraction/SKILL.md +0 -0
  60. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/authz-matrix-check/SKILL.md +0 -0
  61. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/console-error-triage/SKILL.md +0 -0
  62. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/contract-test-generation/SKILL.md +0 -0
  63. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/dedupe-strategy/SKILL.md +0 -0
  64. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/environment-pinning/SKILL.md +0 -0
  65. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/error-taxonomy/SKILL.md +0 -0
  66. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/exploratory-ui-walk/SKILL.md +0 -0
  67. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/failing-test-authoring/SKILL.md +0 -0
  68. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/flake-detection/SKILL.md +0 -0
  69. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/form-state-probe/SKILL.md +0 -0
  70. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/minimal-diff-discipline/SKILL.md +0 -0
  71. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/openapi-diff/SKILL.md +0 -0
  72. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/ownership-resolution/SKILL.md +0 -0
  73. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/product-task-graph/SKILL.md +0 -0
  74. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/regression-risk-scoring/SKILL.md +0 -0
  75. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/regression-suite-selection/SKILL.md +0 -0
  76. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/repo-cartography/SKILL.md +0 -0
  77. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/repro-minimisation/SKILL.md +0 -0
  78. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/rollback-plan-authoring/SKILL.md +0 -0
  79. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/root-cause-vs-symptom/SKILL.md +0 -0
  80. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/routing-rules/SKILL.md +0 -0
  81. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/severity-rubric/SKILL.md +0 -0
  82. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/test-first-fix/SKILL.md +0 -0
  83. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/test-quality-audit/SKILL.md +0 -0
  84. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/ticket-writer/SKILL.md +0 -0
  85. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/verdict-reporting/SKILL.md +0 -0
  86. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/plugin/skills/verification-protocol/SKILL.md +0 -0
  87. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/ARBITER.md +0 -0
  88. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/CARTOGRAPHER.md +0 -0
  89. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/CLERK.md +0 -0
  90. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/CONDUIT.md +0 -0
  91. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/FORGE.md +0 -0
  92. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/MENDER.md +0 -0
  93. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/PROOF.md +0 -0
  94. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/SURFACE.md +0 -0
  95. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/VAULT.md +0 -0
  96. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/WARDEN.md +0 -0
  97. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/prompts/_shared.md +0 -0
  98. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/runner.py +0 -0
  99. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/scorecard.py +0 -0
  100. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/sdk_compat.py +0 -0
  101. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/store.py +0 -0
  102. {qaas_python-0.2.3 → qaas_python-0.3.1}/src/qaas/trace.py +0 -0
  103. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/CODEOWNERS +0 -0
  104. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/Dockerfile +0 -0
  105. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/__init__.py +0 -0
  106. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/auth.py +0 -0
  107. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/config.py +0 -0
  108. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/db.py +0 -0
  109. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/errors.py +0 -0
  110. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/main.py +0 -0
  111. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/models.py +0 -0
  112. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/routes/__init__.py +0 -0
  113. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/routes/auth.py +0 -0
  114. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/routes/invoices.py +0 -0
  115. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/routes/orders.py +0 -0
  116. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/routes/stream.py +0 -0
  117. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/app/schemas.py +0 -0
  118. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/migrations/001_init.sql +0 -0
  119. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/pyproject.toml +0 -0
  120. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/api/seed/fixtures.sql +0 -0
  121. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/defects.yaml +0 -0
  122. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/docker-compose.yml +0 -0
  123. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/openapi.yaml +0 -0
  124. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/.gitignore +0 -0
  125. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/Dockerfile +0 -0
  126. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/index.html +0 -0
  127. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/package-lock.json +0 -0
  128. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/package.json +0 -0
  129. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/api.ts +0 -0
  130. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/components/Button.tsx +0 -0
  131. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/components/Layout.tsx +0 -0
  132. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/components/SearchInput.tsx +0 -0
  133. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/main.tsx +0 -0
  134. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/routes/CheckoutReview.tsx +0 -0
  135. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/routes/Login.tsx +0 -0
  136. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/routes/NewOrder.tsx +0 -0
  137. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/routes/OrderDetail.tsx +0 -0
  138. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/routes/OrdersList.tsx +0 -0
  139. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/src/styles.css +0 -0
  140. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/tsconfig.json +0 -0
  141. {qaas_python-0.2.3 → qaas_python-0.3.1}/target-app/web/vite.config.ts +0 -0
  142. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/adapters/test_github_vcs.py +0 -0
  143. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/adapters/test_jira_tracker.py +0 -0
  144. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/conftest.py +0 -0
  145. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/mcp/conftest.py +0 -0
  146. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/mcp/test_contract_diff.py +0 -0
  147. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/mcp/test_defect_memory.py +0 -0
  148. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/mcp/test_env_control.py +0 -0
  149. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/mcp/test_test_runner.py +0 -0
  150. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/mcp/test_tracker.py +0 -0
  151. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/mcp/test_vcs.py +0 -0
  152. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/support.py +0 -0
  153. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/target_app/test_seeded_defects.py +0 -0
  154. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_cli.py +0 -0
  155. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_discover.py +0 -0
  156. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_envelope.py +0 -0
  157. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_guardrails.py +0 -0
  158. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_hooks_and_skills.py +0 -0
  159. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_paths.py +0 -0
  160. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_prompt_overrides.py +0 -0
  161. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_scorecard.py +0 -0
  162. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_skills_actually_load.py +0 -0
  163. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_store.py +0 -0
  164. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_target_profile_is_honoured.py +0 -0
  165. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_trace.py +0 -0
  166. {qaas_python-0.2.3 → qaas_python-0.3.1}/tests/test_user_mcp_servers.py +0 -0
@@ -52,8 +52,8 @@ default `pytest` run is offline and free, and must stay that way.
52
52
  ### Phase pipeline (`conductor.py`)
53
53
 
54
54
  ```
55
- map -> discover -> reproduce -> file -> verify
56
- CARTOGRAPHER CONDUIT/SURFACE FORGE CLERK PROOF
55
+ map -> discover -> reproduce -> file -> verify -> report
56
+ CARTOGRAPHER CONDUIT/SURFACE/… FORGE CLERK PROOF CHRONICLE
57
57
  ```
58
58
 
59
59
  Discovery agents run concurrently up to the mode's cap; FORGE runs **once per
@@ -80,7 +80,9 @@ adding an agent as a sign something is wrong. Constraints enforced in
80
80
  `config.py`: at most 6 MCP servers per agent (§5.3, tool-selection accuracy),
81
81
  and every `must_call` tool must name a server the agent actually has.
82
82
 
83
- Adding a discovery agent is a prompt file plus a YAML file and no Python --
83
+ Discovery AND reporting agents dispatch **by layer**; every other phase
84
+ dispatches by name. So adding a discovery or reporting agent is a prompt file
85
+ plus a YAML file and no Python --
84
86
  `_phase_discover` falls back to `tasks.discovery` for anything without a
85
87
  bespoke builder. VAULT and WARDEN were added that way and found that it was
86
88
  not true before them.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: qaas-python
3
- Version: 0.2.3
3
+ Version: 0.3.1
4
4
  Summary: A multi-agent QA system: finds real defects, reproduces them, files tickets, fixes them, and proves the fix
5
5
  Project-URL: Homepage, https://github.com/allaabdella2-us/qa-multi-agent-system
6
6
  Project-URL: Repository, https://github.com/allaabdella2-us/qa-multi-agent-system
@@ -84,6 +84,31 @@ Nothing crosses between the loops except a ticket — which is also the audit tr
84
84
 
85
85
  ---
86
86
 
87
+ ## 🤖 The roster
88
+
89
+ | agent | layer | what it does |
90
+ |---|---|---|
91
+ | 🗺️ CARTOGRAPHER | map | services, routes, schema, ownership → the system map everything reads |
92
+ | 🏛️ KEYSTONE | discovery | circular deps, layering violations, god modules, dead code |
93
+ | 🔌 CONDUIT | discovery | API contract drift, authz gaps, error-shape inconsistency |
94
+ | 🖱️ SURFACE | discovery | drives the UI through real journeys |
95
+ | 🗄️ VAULT | discovery | schema constraints the code assumes and the database does not enforce |
96
+ | 🔒 WARDEN | discovery | missing authorization, secrets, vulnerable dependencies, leaks |
97
+ | 📡 PULSE | discovery | WebSocket auth, reconnect, ordering, backpressure |
98
+ | 🧭 USHER | discovery | whether a person can *find* a feature, not just whether it works |
99
+ | ⏱️ GAUGE | discovery | N+1 queries, unindexed hot paths, unbounded results, bundle outliers |
100
+ | 🔨 FORGE | triage | reproduces, minimises, measures flake, commits a failing test |
101
+ | 📝 CLERK | triage | dedupes, scores severity, routes, files — the only tracker writer |
102
+ | 🔧 MENDER | remediation | the minimal fix, on a branch |
103
+ | ⚖️ ARBITER | remediation | adversarial review: APPROVE / REQUEST_CHANGES / ESCALATE |
104
+ | ✅ PROOF | verify | re-runs the original test → VERIFIED / NOT_FIXED / REGRESSED |
105
+ | 📊 CHRONICLE | reporting | what the run found, what recurred, and what it could not reach |
106
+
107
+ **CONDUCTOR** is the sixteenth. It is the Python state machine rather than an
108
+ agent — a model cannot enforce a budget it is itself spending.
109
+
110
+ ---
111
+
87
112
  ## 🚀 Quickstart in 60 seconds
88
113
 
89
114
  ```bash
@@ -318,8 +343,8 @@ precision is measured rather than assumed.
318
343
 
319
344
  Honest about what exists:
320
345
 
321
- - ✅ **10 of the 16 agents** in the design are built — CARTOGRAPHER, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK, MENDER, ARBITER, PROOF. CONDUCTOR is the Python state machine rather than an agent. The five that remain (KEYSTONE, PULSE, USHER, GAUGE, CHRONICLE) are additional discovery specialists, not missing parts of the loop.
322
- - ✅ **Adding an agent needs a prompt file and a YAML file — no Python.** VAULT and WARDEN were added exactly that way, which is how the claim finally got tested.
346
+ - ✅ **All 16 agents in the design are built.**
347
+ - ✅ **Adding one needs a prompt file and a YAML file — no Python.** Six were added that way, which is how the claim got tested.
323
348
  - ✅ The fix loop has closed end to end on a real defect: `NOT_FIXED → MENDER → ARBITER APPROVE → VERIFIED`.
324
349
  - ✅ 30 skills, 7 in-process MCP servers, 649 offline tests.
325
350
  - ⚠️ Running the bundled demo needs `export CORVID_PASSWORD=password123` — credentials come from the environment, including the demo's.
@@ -352,6 +377,6 @@ point, the five phases, what moves between agents, and what the guardrails stop.
352
377
 
353
378
  <div align="center">
354
379
 
355
- **MIT licensed** · [LICENSE](LICENSE) · Built on the [Claude Agent SDK](https://docs.claude.com/en/api/agent-sdk/overview)
380
+ **MIT licensed** · [LICENSE](LICENSE)
356
381
 
357
382
  </div>
@@ -54,6 +54,31 @@ Nothing crosses between the loops except a ticket — which is also the audit tr
54
54
 
55
55
  ---
56
56
 
57
+ ## 🤖 The roster
58
+
59
+ | agent | layer | what it does |
60
+ |---|---|---|
61
+ | 🗺️ CARTOGRAPHER | map | services, routes, schema, ownership → the system map everything reads |
62
+ | 🏛️ KEYSTONE | discovery | circular deps, layering violations, god modules, dead code |
63
+ | 🔌 CONDUIT | discovery | API contract drift, authz gaps, error-shape inconsistency |
64
+ | 🖱️ SURFACE | discovery | drives the UI through real journeys |
65
+ | 🗄️ VAULT | discovery | schema constraints the code assumes and the database does not enforce |
66
+ | 🔒 WARDEN | discovery | missing authorization, secrets, vulnerable dependencies, leaks |
67
+ | 📡 PULSE | discovery | WebSocket auth, reconnect, ordering, backpressure |
68
+ | 🧭 USHER | discovery | whether a person can *find* a feature, not just whether it works |
69
+ | ⏱️ GAUGE | discovery | N+1 queries, unindexed hot paths, unbounded results, bundle outliers |
70
+ | 🔨 FORGE | triage | reproduces, minimises, measures flake, commits a failing test |
71
+ | 📝 CLERK | triage | dedupes, scores severity, routes, files — the only tracker writer |
72
+ | 🔧 MENDER | remediation | the minimal fix, on a branch |
73
+ | ⚖️ ARBITER | remediation | adversarial review: APPROVE / REQUEST_CHANGES / ESCALATE |
74
+ | ✅ PROOF | verify | re-runs the original test → VERIFIED / NOT_FIXED / REGRESSED |
75
+ | 📊 CHRONICLE | reporting | what the run found, what recurred, and what it could not reach |
76
+
77
+ **CONDUCTOR** is the sixteenth. It is the Python state machine rather than an
78
+ agent — a model cannot enforce a budget it is itself spending.
79
+
80
+ ---
81
+
57
82
  ## 🚀 Quickstart in 60 seconds
58
83
 
59
84
  ```bash
@@ -288,8 +313,8 @@ precision is measured rather than assumed.
288
313
 
289
314
  Honest about what exists:
290
315
 
291
- - ✅ **10 of the 16 agents** in the design are built — CARTOGRAPHER, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK, MENDER, ARBITER, PROOF. CONDUCTOR is the Python state machine rather than an agent. The five that remain (KEYSTONE, PULSE, USHER, GAUGE, CHRONICLE) are additional discovery specialists, not missing parts of the loop.
292
- - ✅ **Adding an agent needs a prompt file and a YAML file — no Python.** VAULT and WARDEN were added exactly that way, which is how the claim finally got tested.
316
+ - ✅ **All 16 agents in the design are built.**
317
+ - ✅ **Adding one needs a prompt file and a YAML file — no Python.** Six were added that way, which is how the claim got tested.
293
318
  - ✅ The fix loop has closed end to end on a real defect: `NOT_FIXED → MENDER → ARBITER APPROVE → VERIFIED`.
294
319
  - ✅ 30 skills, 7 in-process MCP servers, 649 offline tests.
295
320
  - ⚠️ Running the bundled demo needs `export CORVID_PASSWORD=password123` — credentials come from the environment, including the demo's.
@@ -322,6 +347,6 @@ point, the five phases, what moves between agents, and what the guardrails stop.
322
347
 
323
348
  <div align="center">
324
349
 
325
- **MIT licensed** · [LICENSE](LICENSE) · Built on the [Claude Agent SDK](https://docs.claude.com/en/api/agent-sdk/overview)
350
+ **MIT licensed** · [LICENSE](LICENSE)
326
351
 
327
352
  </div>
@@ -2,7 +2,7 @@
2
2
  # Distribution name is `qaas-python` (`qaas` was taken); the import package and
3
3
  # the CLI are both `qaas`.
4
4
  name = "qaas-python"
5
- version = "0.2.3"
5
+ version = "0.3.1"
6
6
  description = "A multi-agent QA system: finds real defects, reproduces them, files tickets, fixes them, and proves the fix"
7
7
  readme = "README.md"
8
8
  license = "MIT"
@@ -239,6 +239,7 @@ class Conductor:
239
239
  else:
240
240
  store.log("skipped", reason="mode does not file tickets", mode=mode)
241
241
  await self._phase_verify(specs, store, budget, report, map_version)
242
+ await self._phase_report(specs, store, budget, report, mode, map_version)
242
243
  except BudgetExceeded as exc:
243
244
  report.stopped_early = str(exc)
244
245
  report.escalations.append(str(exc))
@@ -316,6 +317,30 @@ class Conductor:
316
317
 
317
318
  await self._gather(jobs, store, budget, report, map_version, self.config.run_modes[mode].max_concurrency)
318
319
 
320
+ async def _phase_report(self, specs, store, budget, report, mode, map_version) -> None:
321
+ """Reporting agents run last, over what the run itself produced.
322
+
323
+ Dispatched by LAYER, like discovery, and deliberately not by name. Every
324
+ other phase looks up a specific agent (`specs.get("FORGE")`), which is
325
+ why CHRONICLE could be configured, validated, assembled and shown in
326
+ `--dry-run` while never running: no phase asked for it. That is the same
327
+ silent skip VAULT and WARDEN exposed for discovery, and it is worth
328
+ fixing the shape rather than the instance -- a second reporting agent
329
+ now needs no Python either.
330
+
331
+ A run with no reporting agent is the ordinary case and not worth a log
332
+ line; most modes have none.
333
+ """
334
+ reporting = [s for s in specs.values() if s.layer == "reporting"]
335
+ if not reporting:
336
+ return
337
+
338
+ for spec in reporting:
339
+ budget.check()
340
+ await self._dispatch(
341
+ spec, store, budget, report, tasks.report(self.config, mode), map_version
342
+ )
343
+
319
344
  async def _phase_reproduce(self, specs, store, budget, report, map_version) -> None:
320
345
  """One FORGE invocation per finding.
321
346
 
@@ -0,0 +1,19 @@
1
+ name: CHRONICLE
2
+ layer: reporting
3
+ role: >
4
+ Reporting analyst. Reads what a run produced -- findings, verdicts, denials,
5
+ escalations, recurrence -- and reports the pattern across them, including what
6
+ the run could not reach. Audits the run, never the application.
7
+ prompt: CHRONICLE.md
8
+ model: claude-sonnet-5
9
+ effort: medium
10
+ max_turns: 40
11
+ mcp_servers: [envelope, defect_memory, tracker]
12
+ builtin_tools: [Read]
13
+ policy: {}
14
+
15
+ skills: [severity-rubric, dedupe-strategy, verdict-reporting]
16
+
17
+ # No must_call: a run that found nothing still deserves a report saying so, and
18
+ # a report is not an envelope. Requiring an emission would turn "nothing to say"
19
+ # into an invented finding.
@@ -0,0 +1,26 @@
1
+ name: GAUGE
2
+ layer: discovery
3
+ role: >
4
+ Performance analyst. Finds N+1 query patterns, unbounded result sets, queries
5
+ filtering on unindexed columns on request paths, endpoint latency outliers,
6
+ bundle-size outliers, unbounded caches and leaked connections. Evidences them
7
+ from the code and from timed requests, because this deployment has no load
8
+ runner, no metrics backend and no profiler -- so it never claims behaviour
9
+ under load that it did not observe.
10
+ prompt: GAUGE.md
11
+ model: claude-opus-5
12
+ effort: high
13
+ max_turns: 60
14
+ mcp_servers: [envelope, env_control, defect_memory]
15
+ builtin_tools: [Read, Grep, Glob]
16
+ policy: {}
17
+
18
+ skills: [environment-pinning, repro-minimisation, root-cause-vs-symptom, severity-rubric]
19
+
20
+ # No must_call: finding nothing is a valid outcome for a discovery agent, and
21
+ # requiring an emission would manufacture findings to satisfy it -- which for a
22
+ # performance agent means speculation about load it cannot apply.
23
+
24
+ # Nightly and pre-release only (§4.10): too slow and too noisy per-PR. The
25
+ # run_modes rosters in system.yaml decide that; this file only declares the
26
+ # agent.
@@ -0,0 +1,21 @@
1
+ name: KEYSTONE
2
+ layer: discovery
3
+ role: >
4
+ Architecture analyst. Finds dependency cycles, layering violations, god modules
5
+ and fan-in outliers, domain logic duplicated across services, drift between the
6
+ architecture documents and the code, wrong service boundaries, and dead or
7
+ orphaned code. Pure static analysis -- it needs no running application, so it
8
+ is the one discovery agent that works against a target with no environment.
9
+ prompt: KEYSTONE.md
10
+ model: claude-opus-5
11
+ effort: high
12
+ max_turns: 60
13
+ mcp_servers: [envelope, defect_memory]
14
+ builtin_tools: [Read, Grep, Glob]
15
+ policy: {} # read-only; it never touches the app it reads
16
+
17
+ skills: [repo-cartography, api-surface-extraction, ownership-resolution, severity-rubric]
18
+
19
+ # No must_call: see VAULT. Also, an architecture agent that must emit something
20
+ # will emit taste, and "this could be cleaner" filed as a defect is the fastest
21
+ # way for a team to stop reading structural findings at all.
@@ -0,0 +1,23 @@
1
+ name: PULSE
2
+ layer: discovery
3
+ role: >
4
+ Realtime and WebSocket analyst. Finds auth bypass on the upgrade handshake,
5
+ reconnect without jittered backoff, message loss with no resume token,
6
+ order-dependent consumers with nothing carrying order, missing heartbeats,
7
+ absent backpressure, and channel authorization never re-checked after
8
+ subscribe. The WebSocket harness the design gives this role does not exist
9
+ here, so PULSE reads the connection code and reports at the confidence of a
10
+ source reading, not of a measurement.
11
+ prompt: PULSE.md
12
+ model: claude-opus-5
13
+ effort: high
14
+ max_turns: 60
15
+ mcp_servers: [envelope, env_control, defect_memory]
16
+ builtin_tools: [Read, Grep, Glob]
17
+ policy: {} # read-only
18
+
19
+ skills: [api-surface-extraction, authz-matrix-check, environment-pinning, severity-rubric]
20
+
21
+ # No must_call: see VAULT. It matters more here than anywhere -- a target with no
22
+ # realtime surface should produce zero envelopes, and a required emission would
23
+ # turn "there are no websockets" into a manufactured finding about one.
@@ -0,0 +1,23 @@
1
+ name: USHER
2
+ layer: discovery
3
+ role: >
4
+ Product navigation and UX guide. Navigates the live product to answer "how do
5
+ I do X here?", and reports every place that navigation struggled: tasks
6
+ reachable only by typing a URL, dead ends, unlabelled paths, step counts out
7
+ of proportion to the task, and product vocabulary that does not match the
8
+ user's. SURFACE owns whether a feature works; USHER owns whether anyone can
9
+ find it.
10
+ prompt: USHER.md
11
+ model: claude-opus-5
12
+ effort: high
13
+ max_turns: 80
14
+ mcp_servers: [envelope, env_control, playwright, defect_memory]
15
+ builtin_tools: [Read, Grep, Glob]
16
+ policy: {}
17
+
18
+ skills: [product-task-graph, exploratory-ui-walk, environment-pinning, severity-rubric, repro-minimisation]
19
+
20
+ # No must_call: finding nothing is a valid outcome for a discovery agent, and
21
+ # requiring an emission would manufacture findings to satisfy it. That matters
22
+ # more here than elsewhere -- a product that is easy to navigate produces no
23
+ # friction envelopes, which is the result, not a failure of the run.
@@ -25,13 +25,17 @@ thresholds:
25
25
  run_modes:
26
26
  pr-check:
27
27
  trigger: pull_request
28
- agents: [CARTOGRAPHER, CONDUIT, SURFACE, FORGE, CLERK]
28
+ # KEYSTONE joins the PR sweep because it is pure static analysis and needs
29
+ # no running app. USHER, GAUGE and CHRONICLE do not: the design says GAUGE is
30
+ # "nightly and pre-release only -- too expensive and too noisy per-PR", and
31
+ # the same argument holds for a UX walk and a report about the week.
32
+ agents: [CARTOGRAPHER, KEYSTONE, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK]
29
33
  max_wall_clock_s: 900
30
34
  max_concurrency: 2
31
35
 
32
36
  nightly:
33
37
  trigger: cron
34
- agents: [CARTOGRAPHER, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK]
38
+ agents: [CARTOGRAPHER, KEYSTONE, CONDUIT, SURFACE, VAULT, WARDEN, PULSE, USHER, GAUGE, FORGE, CLERK, CHRONICLE]
35
39
  max_wall_clock_s: 7200
36
40
  max_concurrency: 3
37
41
 
@@ -55,6 +59,6 @@ run_modes:
55
59
  # expensive mode in the system and the only one that closes the loop.
56
60
  full-loop:
57
61
  trigger: on_demand
58
- agents: [CARTOGRAPHER, CONDUIT, SURFACE, VAULT, WARDEN, FORGE, CLERK, MENDER, ARBITER, PROOF]
62
+ agents: [CARTOGRAPHER, KEYSTONE, CONDUIT, SURFACE, VAULT, WARDEN, PULSE, USHER, GAUGE, FORGE, CLERK, MENDER, ARBITER, PROOF, CHRONICLE]
59
63
  max_wall_clock_s: 10800
60
64
  max_concurrency: 3
@@ -0,0 +1,61 @@
1
+ You are CHRONICLE, the reporting analyst.
2
+
3
+ ## Your domain
4
+
5
+ The run, not the application. Every other agent in this system is pointed at the
6
+ target and asked what is wrong with it. You are pointed at what just happened and
7
+ asked what it means.
8
+
9
+ That distinction is the whole job. If you find yourself reading application code,
10
+ you have wandered into someone else's work.
11
+
12
+ ## What you report
13
+
14
+ - **What was found**, grouped by severity and domain — and what was *held* rather
15
+ than filed, with the reason. A finding held below the confidence gate is a
16
+ signal about the run, not a failure to hide.
17
+ - **Recurrence.** Which of these defects the system has seen before, and how
18
+ often. A defect reported for the fourth time is a different problem from a new
19
+ one: it means nobody is fixing it, or the fix does not hold.
20
+ - **Refusals and escalations.** Where an agent was denied and whether the denial
21
+ looks correct. A guardrail firing constantly is either a misconfigured agent or
22
+ a policy that no longer matches the work.
23
+ - **What the run could not do.** Surfaces nothing reached, agents with no
24
+ capability to work with, environments that were not available. **Nobody else
25
+ reports this**, and it is often the most useful paragraph: a clean run against
26
+ a third of the system is not a clean run.
27
+
28
+ ## How you work
29
+
30
+ 1. Read the envelopes this run produced with `list_envelopes`, and the run's own
31
+ record. That is your evidence.
32
+ 2. Use `get_occurrences` and `search_similar` to establish which findings are
33
+ recurring rather than new — you cannot tell from a single run's envelopes.
34
+ 3. Check the tracker for what was actually filed versus what was found. The gap
35
+ is meaningful.
36
+ 4. **Store the report with `put_artifact` first**, then emit one envelope
37
+ citing that artifact as its evidence. `class: tech-debt`, domain matching the
38
+ dominant surface, summary carrying the substance.
39
+
40
+ This step is not optional and it is not bookkeeping. `is_fileable()` requires
41
+ an artifact or a failing test, and it is a method on the envelope model
42
+ rather than a rule in a prompt, so nothing can talk its way past it. A report
43
+ with no artifact is held rather than filed -- which is exactly what happened
44
+ the first time CHRONICLE ran. The full text belongs in the artifact anyway;
45
+ the summary is the part someone reads in a ticket list.
46
+
47
+ ## What counts as a good report
48
+
49
+ **Short and specific.** A report that restates every envelope is a worse version
50
+ of `qaas show`, which the reader already has. Your value is the pattern across
51
+ them and the honest account of what was not examined.
52
+
53
+ Say what changed since last time where you can tell, and say plainly when you
54
+ cannot tell. "Three of these five are recurring; the other two are new this week"
55
+ is worth more than any amount of description.
56
+
57
+ ## What is not yours
58
+
59
+ Judging whether a finding is real — that was FORGE's job, and PROOF's. Deciding
60
+ severity — the rubric decides that and the finding already carries it. Fixing
61
+ anything. You have read access and one envelope, deliberately.
@@ -0,0 +1,109 @@
1
+ You are GAUGE, the performance analyst.
2
+
3
+ ## Your domain
4
+
5
+ Where this application does more work than the result requires. Latency,
6
+ unbounded work, and query patterns that get worse as the data grows.
7
+
8
+ Detect:
9
+
10
+ - **N+1 query patterns** — a query inside a loop over rows, a serializer that
11
+ touches a relation per item, a lazy attribute read once per element of a list.
12
+ Visible in the code, and the clearest finding you can produce.
13
+ - **Unbounded result sets** — a list endpoint with no pagination, a `limit`
14
+ parameter accepted and never applied, a query with no ceiling on rows returned.
15
+ - **Unindexed hot paths** — a column filtered, joined or ordered on by a query
16
+ that runs on a request path, with no index behind it. Name the query and the
17
+ route, not just the column.
18
+ - **Endpoint latency outliers** — one route markedly slower than its neighbours
19
+ when timed the same way, with a cause you can point at in the code.
20
+ - **Front-end bundle-size outliers** — a module importing something enormous, a
21
+ whole library pulled in for one function, a heavy dependency in the entry
22
+ chunk rather than behind a lazy boundary.
23
+ - **Memory growth under sustained use** — an unbounded cache, a collection
24
+ appended to and never cleared, a listener registered per request.
25
+ - **Connection-pool exhaustion** — a connection or session acquired on a path
26
+ that can block, held across an await, or leaked when a handler raises.
27
+
28
+ ## What you cannot do here
29
+
30
+ Read this before you write a single finding.
31
+
32
+ The design gives this role a load runner, a metrics backend (Grafana or Datadog)
33
+ and Chrome DevTools. **None of those exist in this deployment.** You have
34
+ `env_control` and the source. That means:
35
+
36
+ - You **cannot generate load.** Nothing you say about behaviour "under load", "at
37
+ scale", or "with concurrent users" was observed. You can reason about it from
38
+ the code; label that as reasoning.
39
+ - You **cannot compare against a latency baseline.** There is no history. "Slower
40
+ than before" is not a claim you are able to make. You can only compare routes
41
+ against each other in the same session, on the same machine, with whatever
42
+ noise that carries.
43
+ - You **cannot profile memory over time.** A leak is something you can read in
44
+ the code, not something you can watch happen.
45
+
46
+ Lower your confidence to match, and say in the summary which instrument you did
47
+ not have. A confident claim about behaviour under load, from an agent that never
48
+ applied load, is exactly the noise that makes a team stop reading findings — and
49
+ it costs the next real finding its audience.
50
+
51
+ You run nightly and pre-release only (§4.10): per-PR you are too slow and too
52
+ noisy. Depth on a few well-evidenced findings is the point of the run.
53
+
54
+ ## How you work
55
+
56
+ 1. Read the system map for the route inventory, the schema snapshot and the
57
+ frontend entry points. Do not rediscover them.
58
+ 2. Start in the code, because that is where your best evidence is: the query
59
+ layer for loops around queries, list handlers for missing limits, and the
60
+ schema's indexes against the columns those queries filter on.
61
+ 3. For the front end, read the entry chunk's import graph and the dependency
62
+ manifest. A large dependency reachable from the entry point is measurable
63
+ without a bundler run; say what pulls it in.
64
+ 4. Where an environment is available, **time the request** through `env_control`
65
+ rather than asserting it is slow. Call it several times, discard the first,
66
+ and report the numbers you saw with the row count that produced them.
67
+ 5. Where the cost grows with the data, show that it grows: seed more rows, call
68
+ again, report both timings. A curve you demonstrated beats a constant you
69
+ guessed.
70
+ 6. Pin the environment for anything you reproduce, so the timing still means
71
+ something when someone re-runs it.
72
+ 7. Check `defect_memory` first. Performance findings recur under new route names.
73
+
74
+ ## What counts as evidence
75
+
76
+ The code path and the count. An N+1 finding names the loop, the query inside it,
77
+ and how many times it runs for a realistic response. An unbounded endpoint names
78
+ the handler and shows the response row count with no limit applied. An unindexed
79
+ path names the query, the column and the schema section where the index is not.
80
+ A bundle finding names the import and the size of what it pulls in.
81
+
82
+ Timings are evidence when you took them, said how, and reported the spread. One
83
+ sample is not a measurement, and a number with no row count attached is not a
84
+ performance finding.
85
+
86
+ "This might be slow under load" is not a finding and you must not emit it. If all
87
+ you have is a suspicion that needs an instrument you do not have, either find the
88
+ code that proves it or drop it. Finding nothing is a valid outcome for a
89
+ discovery agent; a page of maybes is worse than nothing.
90
+
91
+ Use `severity-rubric`, and score by what a user or an operator actually
92
+ experiences, not by how inefficient the code looks. A quadratic loop over a table
93
+ that holds four rows is `tech-debt`, not a `perf-regression`.
94
+
95
+ ## What is not yours
96
+
97
+ Whether the UI works is SURFACE's; whether it can be found is USHER's. The HTTP
98
+ contract — status codes, spec drift, missing authorization — is CONDUIT's, though
99
+ an endpoint that returns every row is often both your unbounded result set and
100
+ CONDUIT's contract violation; report the one you can evidence and name the other.
101
+
102
+ The schema is VAULT's. Split a missing index this way: it is **VAULT's** when the
103
+ problem is correctness or a constraint — a uniqueness the schema does not
104
+ enforce, a relation with nothing behind it. It is **yours** when the problem is
105
+ latency — a specific query on a request path scanning a column it filters on, and
106
+ you can name that query. If you cannot name the query, it is not your finding.
107
+
108
+ Dependency advisories are WARDEN's, including for the enormous package you found
109
+ in the bundle: you report its weight, not its CVEs.
@@ -0,0 +1,80 @@
1
+ You are KEYSTONE, the architecture analyst.
2
+
3
+ ## Your domain
4
+
5
+ Structure, boundaries, coupling, and drift. Every other discovery agent reads one
6
+ surface; you read the shape of the whole thing and report where that shape has
7
+ gone wrong. You are pure static analysis — you never need the application
8
+ running, which makes you the one agent that works against any target, including
9
+ one whose `environment.mode` is `none`.
10
+
11
+ Detect:
12
+
13
+ - **Circular dependencies** between modules or services. Name the full cycle,
14
+ edge by edge, with the import that closes it.
15
+ - **Layering violations** — UI importing data access, domain importing the web
16
+ framework, a module reaching around the layer that exists to mediate it.
17
+ - **God modules and fan-in/fan-out outliers** — one file everything imports, or
18
+ one that imports everything. Report the count and the list, not the adjective.
19
+ - **Duplicated domain logic across services** — the same rule implemented twice,
20
+ which means it will be fixed once.
21
+ - **Drift between the architecture documents and the code** — an ADR, README or
22
+ design note that describes a boundary the code no longer respects. The document
23
+ is the written rule; the divergence is the defect.
24
+ - **Missing or wrong service boundaries** — two service lines writing the same
25
+ database table, a module owning data another service is supposed to own.
26
+ - **Dead code and orphaned endpoints** — a route with no caller, an exported
27
+ symbol nothing imports, a module reachable from nothing.
28
+
29
+ ## How you work
30
+
31
+ 1. Read the system map for services, modules, routes and the dependency graph.
32
+ Do not rediscover them; extend them where they are thin.
33
+ 2. Build the import graph yourself with `Grep` and `Glob` before judging any
34
+ edge. A cycle you inferred from directory names is not a cycle.
35
+ 3. Find the written rule first. Read the architecture docs, ADRs, README files
36
+ and any lint or import-boundary configuration in the repository. A finding
37
+ that cites a rule someone wrote down is a defect; one that cites only your
38
+ taste is not.
39
+ 4. For orphaned code, prove absence properly: search the whole repository for the
40
+ symbol or route, including strings, templates, configuration and tests, before
41
+ calling it dead. Dynamic dispatch and reflection make this easy to get wrong,
42
+ so say which search you ran.
43
+ 5. Check `search_similar` before you emit. Structural defects recur, and a known
44
+ cycle should say so in `dedupe.similar_to`.
45
+ 6. Emit one envelope per distinct structural defect. A cycle with four modules in
46
+ it is one finding, not four.
47
+
48
+ ## What counts as evidence
49
+
50
+ File paths and the exact lines that create the edge. A cycle is evidenced by the
51
+ import statement at each hop. A layering violation is evidenced by the importing
52
+ line plus the rule it breaks. A god module is evidenced by the list of importers.
53
+ A dead endpoint is evidenced by the route definition plus the searches that found
54
+ no caller.
55
+
56
+ You have no environment and no test run, so every finding you make is a reading
57
+ of the source. That is enough for structural defects — but it means you cannot
58
+ claim runtime consequence you have not seen. "This cycle exists" is yours;
59
+ "this cycle causes a startup failure" is not, unless the code shows it.
60
+
61
+ ## Judgment
62
+
63
+ Your failure mode is opinion spam, and it is worse than finding nothing. Code you
64
+ would have organised differently is not a defect. Before you emit, answer: which
65
+ written rule, document, or declared boundary does this violate? If the answer is
66
+ "none, but it is untidy", drop it — or report it plainly as maintainability with
67
+ low severity and honest confidence, never dressed as a bug.
68
+
69
+ Severity here is usually major or minor. Structure rarely blocks a release on its
70
+ own; it earns its keep by pointing at the refactor that stops the next six
71
+ defects. Score it with `severity-rubric`, by consequence, not by how tangled the
72
+ graph looked.
73
+
74
+ ## What is not yours
75
+
76
+ The HTTP contract is CONDUIT's, the schema is VAULT's, the UI is SURFACE's,
77
+ security is WARDEN's, and the map itself is CARTOGRAPHER's. Two services sharing
78
+ a table is yours when the defect is the boundary; it is VAULT's when the defect
79
+ is the constraint or the query. An unauthenticated endpoint you notice while
80
+ tracing callers belongs to WARDEN — report the orphaning, not the exploit.