qaas-python 1.0.0__tar.gz → 2.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {qaas_python-1.0.0 → qaas_python-2.0.0}/ARCHITECTURE.md +50 -50
- {qaas_python-1.0.0 → qaas_python-2.0.0}/BUILD_PLAN.md +47 -47
- qaas_python-2.0.0/CHANGELOG.md +184 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/CLAUDE.md +110 -22
- {qaas_python-1.0.0 → qaas_python-2.0.0}/PKG-INFO +37 -37
- {qaas_python-1.0.0 → qaas_python-2.0.0}/README.md +36 -36
- {qaas_python-1.0.0 → qaas_python-2.0.0}/pyproject.toml +8 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/adapters/tracker.py +96 -10
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/adapters/vcs.py +80 -19
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/cli.py +52 -31
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/config.py +5 -5
- qaas_python-1.0.0/src/qaas/defaults/config/agents/conduit.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/api.yaml +2 -2
- qaas_python-1.0.0/src/qaas/defaults/config/agents/keystone.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/architect.yaml +3 -3
- qaas_python-1.0.0/src/qaas/defaults/config/agents/warden.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/auditor.yaml +3 -3
- qaas_python-1.0.0/src/qaas/defaults/config/agents/surface.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/browser.yaml +2 -2
- qaas_python-1.0.0/src/qaas/defaults/config/agents/vault.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/dba.yaml +2 -2
- qaas_python-1.0.0/src/qaas/defaults/config/agents/mender.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/fixer.yaml +3 -3
- qaas_python-1.0.0/src/qaas/defaults/config/agents/usher.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/guide.yaml +3 -3
- qaas_python-1.0.0/src/qaas/defaults/config/agents/gauge.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/load.yaml +2 -2
- qaas_python-1.0.0/src/qaas/defaults/config/agents/cartographer.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/mapper.yaml +2 -2
- qaas_python-1.0.0/src/qaas/defaults/config/agents/chronicle.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/reporter.yaml +2 -2
- qaas_python-1.0.0/src/qaas/defaults/config/agents/forge.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/reproducer.yaml +3 -3
- qaas_python-1.0.0/src/qaas/defaults/config/agents/arbiter.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/reviewer.yaml +4 -4
- qaas_python-1.0.0/src/qaas/defaults/config/agents/pulse.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/socket.yaml +4 -4
- qaas_python-1.0.0/src/qaas/defaults/config/agents/clerk.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/triage.yaml +2 -2
- qaas_python-1.0.0/src/qaas/defaults/config/agents/proof.yaml → qaas_python-2.0.0/src/qaas/defaults/config/agents/verifier.yaml +2 -2
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/defaults/config/system.yaml +10 -10
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/discover.py +6 -6
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/envelope.py +31 -3
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/envfile.py +9 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/guardrails.py +174 -16
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/context.py +11 -3
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/contract_diff.py +86 -12
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/env_control.py +27 -7
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/envelope_server.py +26 -26
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/test_runner.py +76 -7
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/tracker.py +5 -5
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/vcs.py +29 -34
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/paths.py +2 -2
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/adversarial-review/SKILL.md +5 -5
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/environment-pinning/SKILL.md +1 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/minimal-diff-discipline/SKILL.md +3 -3
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/regression-risk-scoring/SKILL.md +2 -2
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/rollback-plan-authoring/SKILL.md +1 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/root-cause-vs-symptom/SKILL.md +6 -6
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/test-first-fix/SKILL.md +1 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/test-quality-audit/SKILL.md +3 -3
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/ticket-writer/SKILL.md +1 -1
- qaas_python-1.0.0/src/qaas/prompts/CONDUIT.md → qaas_python-2.0.0/src/qaas/prompts/API.md +2 -2
- qaas_python-1.0.0/src/qaas/prompts/KEYSTONE.md → qaas_python-2.0.0/src/qaas/prompts/ARCHITECT.md +5 -5
- qaas_python-1.0.0/src/qaas/prompts/WARDEN.md → qaas_python-2.0.0/src/qaas/prompts/AUDITOR.md +3 -3
- qaas_python-1.0.0/src/qaas/prompts/SURFACE.md → qaas_python-2.0.0/src/qaas/prompts/BROWSER.md +1 -1
- qaas_python-1.0.0/src/qaas/prompts/VAULT.md → qaas_python-2.0.0/src/qaas/prompts/DBA.md +4 -4
- qaas_python-1.0.0/src/qaas/prompts/MENDER.md → qaas_python-2.0.0/src/qaas/prompts/FIXER.md +1 -1
- qaas_python-1.0.0/src/qaas/prompts/USHER.md → qaas_python-2.0.0/src/qaas/prompts/GUIDE.md +6 -6
- qaas_python-1.0.0/src/qaas/prompts/GAUGE.md → qaas_python-2.0.0/src/qaas/prompts/LOAD.md +6 -6
- qaas_python-1.0.0/src/qaas/prompts/CARTOGRAPHER.md → qaas_python-2.0.0/src/qaas/prompts/MAPPER.md +2 -2
- qaas_python-1.0.0/src/qaas/prompts/CHRONICLE.md → qaas_python-2.0.0/src/qaas/prompts/REPORTER.md +3 -3
- qaas_python-1.0.0/src/qaas/prompts/FORGE.md → qaas_python-2.0.0/src/qaas/prompts/REPRODUCER.md +1 -1
- qaas_python-1.0.0/src/qaas/prompts/ARBITER.md → qaas_python-2.0.0/src/qaas/prompts/REVIEWER.md +3 -3
- qaas_python-1.0.0/src/qaas/prompts/PULSE.md → qaas_python-2.0.0/src/qaas/prompts/SOCKET.md +4 -4
- qaas_python-1.0.0/src/qaas/prompts/CLERK.md → qaas_python-2.0.0/src/qaas/prompts/TRIAGE.md +2 -2
- qaas_python-1.0.0/src/qaas/prompts/PROOF.md → qaas_python-2.0.0/src/qaas/prompts/VERIFIER.md +2 -2
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/registry.py +21 -11
- qaas_python-1.0.0/src/qaas/conductor.py → qaas_python-2.0.0/src/qaas/router.py +58 -58
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/runner.py +25 -7
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/scorecard.py +30 -7
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/store.py +57 -24
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/target.py +10 -10
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/tasks.py +13 -13
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/trace.py +3 -3
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/CODEOWNERS +1 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/seed/fixtures.sql +1 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/defects.yaml +6 -6
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/adapters/test_github_vcs.py +107 -14
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/adapters/test_jira_tracker.py +96 -11
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/mcp/conftest.py +6 -6
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/mcp/test_contract_diff.py +124 -29
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/mcp/test_defect_memory.py +5 -5
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/mcp/test_env_control.py +30 -5
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/mcp/test_test_runner.py +67 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/mcp/test_tracker.py +51 -51
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/mcp/test_vcs.py +98 -51
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/target_app/test_seeded_defects.py +1 -1
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_cli.py +3 -3
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_config.py +25 -25
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_envelope.py +5 -5
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_guardrails.py +151 -43
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_hooks_and_skills.py +40 -14
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_paths.py +5 -5
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_prompt_overrides.py +5 -5
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_registry.py +39 -8
- qaas_python-1.0.0/tests/test_conductor.py → qaas_python-2.0.0/tests/test_router.py +105 -105
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_scorecard.py +45 -10
- qaas_python-2.0.0/tests/test_store.py +192 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_trace.py +45 -45
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_user_mcp_servers.py +11 -11
- qaas_python-1.0.0/tests/test_store.py +0 -122
- {qaas_python-1.0.0 → qaas_python-2.0.0}/.gitignore +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/LICENSE +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/config/targets/corvid.yaml +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/adapters/__init__.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/__init__.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/mcp/defect_memory.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/.claude-plugin/plugin.json +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/a11y-audit/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/api-surface-extraction/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/authz-matrix-check/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/console-error-triage/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/contract-test-generation/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/dedupe-strategy/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/error-taxonomy/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/exploratory-ui-walk/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/failing-test-authoring/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/flake-detection/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/form-state-probe/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/openapi-diff/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/ownership-resolution/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/product-task-graph/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/regression-suite-selection/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/repo-cartography/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/repro-minimisation/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/routing-rules/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/severity-rubric/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/verdict-reporting/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/plugin/skills/verification-protocol/SKILL.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/prompts/_shared.md +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/src/qaas/sdk_compat.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/Dockerfile +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/__init__.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/auth.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/config.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/db.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/errors.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/main.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/models.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/routes/__init__.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/routes/auth.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/routes/invoices.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/routes/orders.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/routes/stream.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/app/schemas.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/migrations/001_init.sql +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/api/pyproject.toml +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/docker-compose.yml +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/openapi.yaml +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/.gitignore +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/Dockerfile +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/index.html +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/package-lock.json +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/package.json +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/api.ts +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/components/Button.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/components/Layout.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/components/SearchInput.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/main.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/routes/CheckoutReview.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/routes/Login.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/routes/NewOrder.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/routes/OrderDetail.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/routes/OrdersList.tsx +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/src/styles.css +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/tsconfig.json +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/target-app/web/vite.config.ts +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/conftest.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/support.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_discover.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_envfile.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_skills_actually_load.py +0 -0
- {qaas_python-1.0.0 → qaas_python-2.0.0}/tests/test_target_profile_is_honoured.py +0 -0
|
@@ -64,7 +64,7 @@ Start with `qaas validate`, then `qaas run --dry-run`. Neither calls the API.
|
|
|
64
64
|
```
|
|
65
65
|
src/qaas/
|
|
66
66
|
├── cli.py 866 entry point; every command above
|
|
67
|
-
├──
|
|
67
|
+
├── router.py 452 THE STATE MACHINE — phases, budget, concurrency, loops
|
|
68
68
|
├── runner.py 118 invokes ONE agent, records what it cost
|
|
69
69
|
├── registry.py 275 turns an agent's YAML into SDK options
|
|
70
70
|
├── guardrails.py 413 the permission matrix, enforced in code
|
|
@@ -81,18 +81,18 @@ src/qaas/
|
|
|
81
81
|
└── adapters/ tracker (local | jira), vcs (local | github)
|
|
82
82
|
```
|
|
83
83
|
|
|
84
|
-
If you read three files, read `
|
|
84
|
+
If you read three files, read `router.py`, `envelope.py`, `guardrails.py`.
|
|
85
85
|
Everything else is in service of those.
|
|
86
86
|
|
|
87
87
|
---
|
|
88
88
|
|
|
89
89
|
## 4. The two decisions that explain everything else
|
|
90
90
|
|
|
91
|
-
**
|
|
91
|
+
**ROUTER is Python, not a prompt.** The original design has an orchestrator
|
|
92
92
|
agent. It is implemented as an ordinary state machine instead, because *a model
|
|
93
93
|
cannot enforce a budget it is itself spending*. Phase ordering, concurrency,
|
|
94
94
|
retries, escalation and the budget governor are all plain code in
|
|
95
|
-
`
|
|
95
|
+
`router.py`. This also makes runs reproducible and cheap to unit-test — 547
|
|
96
96
|
tests run offline with no API calls.
|
|
97
97
|
|
|
98
98
|
**Each agent is its own top-level `query()`.** Not SDK subagents of a shared
|
|
@@ -114,34 +114,34 @@ $ qaas run --mode full-loop
|
|
|
114
114
|
cli.py ── loads config/system.yaml + config/agents/*.yaml + the target profile
|
|
115
115
|
│
|
|
116
116
|
▼
|
|
117
|
-
|
|
117
|
+
router.py ── runs five phases in order, in-process
|
|
118
118
|
│
|
|
119
|
-
├── PHASE 1 map
|
|
120
|
-
├── PHASE 2 discover
|
|
121
|
-
├── PHASE 3 reproduce
|
|
122
|
-
├── PHASE 4 file
|
|
123
|
-
└── PHASE 5 verify
|
|
119
|
+
├── PHASE 1 map MAPPER → system-map.json (versioned, pinned)
|
|
120
|
+
├── PHASE 2 discover API, BROWSER, … → DefectEnvelopes [concurrent]
|
|
121
|
+
├── PHASE 3 reproduce REPRODUCER → failing test per finding
|
|
122
|
+
├── PHASE 4 file TRIAGE → tickets
|
|
123
|
+
└── PHASE 5 verify VERIFIER ⇄ FIXER ⇄ REVIEWER [bounded loop]
|
|
124
124
|
```
|
|
125
125
|
|
|
126
|
-
Each phase is a method on `
|
|
126
|
+
Each phase is a method on `Router`: `_phase_map`, `_phase_discover`,
|
|
127
127
|
`_phase_reproduce`, `_phase_file`, `_phase_verify`.
|
|
128
128
|
|
|
129
129
|
Which agents run is **config, not code** — `run_modes` in `config/system.yaml`:
|
|
130
130
|
|
|
131
131
|
```yaml
|
|
132
|
-
pr-check: [
|
|
133
|
-
nightly: [
|
|
134
|
-
fix-cycle: [
|
|
132
|
+
pr-check: [MAPPER, API, BROWSER, REPRODUCER, TRIAGE] $16
|
|
133
|
+
nightly: [MAPPER, API, BROWSER, REPRODUCER, TRIAGE] $40
|
|
134
|
+
fix-cycle: [VERIFIER, FIXER, REVIEWER] $20
|
|
135
135
|
full-loop: all eight $60
|
|
136
136
|
```
|
|
137
137
|
|
|
138
|
-
Note the shape: **
|
|
138
|
+
Note the shape: **REPRODUCER runs once per finding**, in a fresh context each time. So
|
|
139
139
|
cost scales with how much was found, not with how many agents exist.
|
|
140
140
|
|
|
141
141
|
### 5.2 Dispatching one agent
|
|
142
142
|
|
|
143
143
|
```
|
|
144
|
-
|
|
144
|
+
router._dispatch(spec, task)
|
|
145
145
|
│
|
|
146
146
|
├── Budget.check() raises BudgetExceeded → a control working, not an error
|
|
147
147
|
│
|
|
@@ -164,7 +164,7 @@ conductor._dispatch(spec, task)
|
|
|
164
164
|
```
|
|
165
165
|
|
|
166
166
|
Failures are **captured, not raised**. One agent falling over costs the run that
|
|
167
|
-
agent's findings, not the whole run; the
|
|
167
|
+
agent's findings, not the whole run; the router decides whether to retry, skip
|
|
168
168
|
or escalate.
|
|
169
169
|
|
|
170
170
|
### 5.3 The remediation loop (phase 5)
|
|
@@ -175,7 +175,7 @@ explicit bounds:
|
|
|
175
175
|
```
|
|
176
176
|
┌─────────────────────────────────────────────┐
|
|
177
177
|
▼ │
|
|
178
|
-
|
|
178
|
+
VERIFIER ──VERIFIED──▶ done │
|
|
179
179
|
│ │
|
|
180
180
|
├──REGRESSED──▶ escalate to a human │
|
|
181
181
|
│ │
|
|
@@ -183,9 +183,9 @@ explicit bounds:
|
|
|
183
183
|
│ │
|
|
184
184
|
├─ reopens ≥ max_proof_reopens ─▶ escalate │
|
|
185
185
|
▼ │
|
|
186
|
-
|
|
186
|
+
FIXER ──▶ REVIEWER ──APPROVE────────────────────┘
|
|
187
187
|
│
|
|
188
|
-
├─ REQUEST_CHANGES ─▶ back to
|
|
188
|
+
├─ REQUEST_CHANGES ─▶ back to FIXER (capped at 2 trips)
|
|
189
189
|
└─ ESCALATE_TO_HUMAN ─▶ stop
|
|
190
190
|
```
|
|
191
191
|
|
|
@@ -194,7 +194,7 @@ defect cycles until the budget is gone and the run ends with no verdict and no
|
|
|
194
194
|
money left to reach one. Escalating after one reopen costs a human five minutes;
|
|
195
195
|
not escalating costs the whole run.
|
|
196
196
|
|
|
197
|
-
|
|
197
|
+
VERIFIER's verdict is a **typed ledger entry** written through a tool
|
|
198
198
|
(`record_verdict`), never parsed out of the agent's prose.
|
|
199
199
|
|
|
200
200
|
---
|
|
@@ -279,11 +279,11 @@ Playwright is the one exception — a real stdio subprocess, declared in
|
|
|
279
279
|
An agent is **a prompt plus a YAML file**. Nothing else.
|
|
280
280
|
|
|
281
281
|
```
|
|
282
|
-
src/qaas/prompts/
|
|
283
|
-
config/agents/
|
|
282
|
+
src/qaas/prompts/API.md role, standards, what good looks like
|
|
283
|
+
config/agents/api.yaml model, budget, tools, skills, policy
|
|
284
284
|
```
|
|
285
285
|
|
|
286
|
-
Adding an agent should require **no change** to `
|
|
286
|
+
Adding an agent should require **no change** to `router.py`, `runner.py`,
|
|
287
287
|
`registry.py` or `guardrails.py`. A change to those files while adding an agent
|
|
288
288
|
is a sign something is wrong.
|
|
289
289
|
|
|
@@ -305,23 +305,23 @@ that mentions one repo's layout or one app's seeded users works exactly once.
|
|
|
305
305
|
|
|
306
306
|
| agent | layer | does |
|
|
307
307
|
|---|---|---|
|
|
308
|
-
|
|
|
309
|
-
|
|
|
310
|
-
|
|
|
311
|
-
|
|
|
312
|
-
|
|
|
313
|
-
|
|
|
314
|
-
|
|
|
315
|
-
|
|
|
316
|
-
|
|
|
317
|
-
|
|
|
318
|
-
|
|
|
319
|
-
|
|
|
320
|
-
|
|
|
321
|
-
|
|
|
322
|
-
|
|
|
323
|
-
|
|
324
|
-
|
|
308
|
+
| MAPPER | map | services, routes, schema, ownership → `system-map.json` |
|
|
309
|
+
| ARCHITECT | discovery | circular deps, layering violations, god modules, dead code |
|
|
310
|
+
| API | discovery | API contract drift; ships a failing contract test |
|
|
311
|
+
| BROWSER | discovery | drives the UI through real journeys |
|
|
312
|
+
| DBA | discovery | schema constraints the code assumes and the database does not enforce |
|
|
313
|
+
| AUDITOR | discovery | missing authorization, secrets, vulnerable dependencies, leaked internals |
|
|
314
|
+
| SOCKET | discovery | WebSocket auth, reconnect, ordering, backpressure |
|
|
315
|
+
| GUIDE | discovery | whether a person can *find* a feature, not just whether it works |
|
|
316
|
+
| LOAD | discovery | N+1 queries, unindexed hot paths, unbounded results, bundle outliers |
|
|
317
|
+
| REPRODUCER | triage | reproduces, minimises, measures flake, commits a failing test |
|
|
318
|
+
| TRIAGE | triage | dedupes, scores severity, routes, files — the only tracker writer |
|
|
319
|
+
| FIXER | remediation | the minimal fix, on a `fix/*` branch |
|
|
320
|
+
| REVIEWER | remediation | adversarial review: APPROVE / REQUEST_CHANGES / ESCALATE |
|
|
321
|
+
| VERIFIER | verify | re-runs the original test → VERIFIED / NOT_FIXED / REGRESSED |
|
|
322
|
+
| REPORTER | reporting | what the run found, what recurred, and what it could not reach |
|
|
323
|
+
|
|
324
|
+
ROUTER is the sixteenth. It is the Python state machine in `router.py`
|
|
325
325
|
rather than an agent, because a model cannot enforce a budget it is spending.
|
|
326
326
|
|
|
327
327
|
---
|
|
@@ -340,16 +340,16 @@ second layer for calls the allowlist did not auto-approve. Both call one
|
|
|
340
340
|
|
|
341
341
|
What is enforced:
|
|
342
342
|
|
|
343
|
-
- **Path scoping** — writes resolved and checked against `write_paths`.
|
|
344
|
-
|
|
343
|
+
- **Path scoping** — writes resolved and checked against `write_paths`. REPRODUCER and
|
|
344
|
+
FIXER only; everyone else denied.
|
|
345
345
|
- **Branch scoping** — git writes matched against branch patterns
|
|
346
346
|
(`qa/repro/*`, `fix/*`). `main` and force-push denied outright.
|
|
347
347
|
- **Ticket rate limit** — `max_tickets_per_run: 10`; over cap the call is denied
|
|
348
|
-
and the
|
|
348
|
+
and the router escalates instead of filing.
|
|
349
349
|
- **Forbidden path classes** — migrations, auth, payment, secrets, infra, CI stop
|
|
350
350
|
at a human however small the change looks.
|
|
351
351
|
- **Diff budget** — counted per distinct file.
|
|
352
|
-
- **Immutable test** —
|
|
352
|
+
- **Immutable test** — FIXER may not edit the test recorded on the envelope.
|
|
353
353
|
|
|
354
354
|
Denials **return a reason and are logged**; they never kill the turn. The agent
|
|
355
355
|
reads the denial and adapts.
|
|
@@ -366,7 +366,7 @@ drafts. Merging is a human decision.
|
|
|
366
366
|
### Hooks enforce the output contract
|
|
367
367
|
|
|
368
368
|
- **`Stop`** blocks an agent that has not called its `must_call` tools, *while it
|
|
369
|
-
still has a turn to fix it* — the
|
|
369
|
+
still has a turn to fix it* — the router would only find out afterwards. It
|
|
370
370
|
honours `stop_hook_active`; blocking twice burns budget.
|
|
371
371
|
- **`PostToolUse`** tells an agent immediately when an emitted envelope was
|
|
372
372
|
**held** rather than filed, since discovering that at the end is too late to
|
|
@@ -453,8 +453,8 @@ Four tiers, cheapest first:
|
|
|
453
453
|
1. `qaas validate` then `qaas run --mode pr-check --dry-run` — see the machine
|
|
454
454
|
describe itself, for free
|
|
455
455
|
2. `envelope.py` — the contract everything else moves
|
|
456
|
-
3. `
|
|
457
|
-
4. `config/agents/
|
|
456
|
+
3. `router.py::run` — the five phases
|
|
457
|
+
4. `config/agents/api.yaml` + `prompts/API.md` — what an agent *is*
|
|
458
458
|
5. `guardrails.py::check` — the one function both enforcement points call
|
|
459
459
|
6. `.qaas/runs/<id>/ledger.jsonl` from a real run — what actually happened
|
|
460
460
|
|
|
@@ -464,8 +464,8 @@ Four tiers, cheapest first:
|
|
|
464
464
|
|
|
465
465
|
Honesty about what is *not* proven, so nobody inherits a false impression:
|
|
466
466
|
|
|
467
|
-
- **Eight of sixteen agents.**
|
|
468
|
-
|
|
467
|
+
- **Eight of sixteen agents.** ARCHITECT, DBA, SOCKET, GUIDE, AUDITOR, LOAD and
|
|
468
|
+
REPORTER are designed but not built.
|
|
469
469
|
- **The extensibility claim is untested.** "A new agent needs only a prompt and a
|
|
470
470
|
YAML" is the architecture's central promise, and no one has added a ninth agent
|
|
471
471
|
to check it.
|
|
@@ -6,8 +6,8 @@
|
|
|
6
6
|
|
|
7
7
|
This plan turns that design into a running system. Per the answers given:
|
|
8
8
|
|
|
9
|
-
- **Runtime** — a Python service on the **Claude Agent SDK** (`claude-agent-sdk`).
|
|
10
|
-
- **Scope** — the doc's own Phase 1 loop (
|
|
9
|
+
- **Runtime** — a Python service on the **Claude Agent SDK** (`claude-agent-sdk`). ROUTER is a real state machine in code; each agent is a `query()` invocation with a hard tool allowlist.
|
|
10
|
+
- **Scope** — the doc's own Phase 1 loop (MAPPER, BROWSER, API, REPRODUCER, TRIAGE, VERIFIER), genuinely end-to-end, on a framework where agents 7–15 are config rather than code.
|
|
11
11
|
- **Target** — a bundled, deliberately-buggy demo app as the system under test, with a golden defect ledger so precision/recall is measurable, not asserted.
|
|
12
12
|
- **Integrations** — every external system behind an adapter with a local fake. The loop runs offline; real Jira/GitHub is a config swap.
|
|
13
13
|
|
|
@@ -19,9 +19,9 @@ Status: plan approved, no code written yet. The milestone table below is the wor
|
|
|
19
19
|
|
|
20
20
|
## Two design decisions that shape everything
|
|
21
21
|
|
|
22
|
-
**1.
|
|
22
|
+
**1. ROUTER is Python, not a prompt.** §4.1 says "keep its reasoning shallow — routing, not analysis," and §10 makes it the enforcement point for budget and concurrency. A model cannot enforce a budget it is spending. So the run state machine, dispatch, retries, dead-letter queue, and the §8.3 loop breakers are ordinary code. This also makes runs reproducible and cheap to test.
|
|
23
23
|
|
|
24
|
-
**2. Every agent is its own top-level `query()`, not a subagent of a parent.** The SDK's `agents=` parameter nests subagents under one conversation; that blurs the per-agent tool allowlist §5.3 depends on and pools cost into one number. Running each agent as a separate `query()` gives a genuine context boundary (design principle §2), an enforceable per-agent allowlist, and per-agent `total_cost_usd` from its `ResultMessage`. `AgentDefinition`/`agents=` stays available for *intra-agent* fan-out (e.g.
|
|
24
|
+
**2. Every agent is its own top-level `query()`, not a subagent of a parent.** The SDK's `agents=` parameter nests subagents under one conversation; that blurs the per-agent tool allowlist §5.3 depends on and pools cost into one number. Running each agent as a separate `query()` gives a genuine context boundary (design principle §2), an enforceable per-agent allowlist, and per-agent `total_cost_usd` from its `ResultMessage`. `AgentDefinition`/`agents=` stays available for *intra-agent* fan-out (e.g. BROWSER exploring several routes in parallel).
|
|
25
25
|
|
|
26
26
|
---
|
|
27
27
|
|
|
@@ -39,7 +39,7 @@ qa-multi-agent-system/
|
|
|
39
39
|
├── src/qaas/
|
|
40
40
|
│ ├── envelope.py # Pydantic DefectEnvelope v1.0 (§6) — the only inter-agent type
|
|
41
41
|
│ ├── registry.py # AgentSpec (YAML) -> ClaudeAgentOptions
|
|
42
|
-
│ ├──
|
|
42
|
+
│ ├── router.py # run state machine, budget governor, concurrency, dead-letter
|
|
43
43
|
│ ├── runner.py # invoke one agent, stream messages, record cost/usage/artifacts
|
|
44
44
|
│ ├── guardrails.py # can_use_tool + hooks = the §8.1 permission matrix, in code
|
|
45
45
|
│ ├── store.py # run ledger, artifact store, versioned system-map
|
|
@@ -70,22 +70,22 @@ qa-multi-agent-system/
|
|
|
70
70
|
| `contract_diff` | `diff_openapi`, `classify_breaking`, `find_consumers`, `generate_contract_test` | `openapi.yaml` vs. live routes |
|
|
71
71
|
| `tracker` | `create_issue`, `transition`, `link`, `search` | adapter: local JSON tickets or Atlassian |
|
|
72
72
|
|
|
73
|
-
WebSocket Harness MCP is deferred —
|
|
73
|
+
WebSocket Harness MCP is deferred — SOCKET is Phase 2. Off-the-shelf servers (Playwright, Filesystem, GitHub) are declared as stdio configs in `config/agents/*.yaml`, so adding one is a YAML edit.
|
|
74
74
|
|
|
75
75
|
### Guardrails are code, not prompting
|
|
76
76
|
|
|
77
77
|
§8.1's matrix becomes a `Policy` per agent, enforced in `can_use_tool` and a `preToolUse` hook — both of which see the tool name and arguments before execution:
|
|
78
78
|
|
|
79
|
-
- **Path scoping** — Filesystem/Edit/Write calls resolved and checked against the agent's allowed roots.
|
|
79
|
+
- **Path scoping** — Filesystem/Edit/Write calls resolved and checked against the agent's allowed roots. REPRODUCER and FIXER only; everyone else denied.
|
|
80
80
|
- **Branch scoping** — git writes matched against the agent's branch regex (`qa/repro/*`, `fix/*`); `main` and any force-push denied outright.
|
|
81
|
-
- **Ticket rate limit** — `tracker.create_issue` counted per run/project/day; over cap the call is denied with a reason and
|
|
82
|
-
- **Immutable test** —
|
|
81
|
+
- **Ticket rate limit** — `tracker.create_issue` counted per run/project/day; over cap the call is denied with a reason and ROUTER escalates instead of filing (§4.12).
|
|
82
|
+
- **Immutable test** — FIXER (Phase 3) denied any edit to the test path recorded on the envelope (§10).
|
|
83
83
|
|
|
84
84
|
Denials return `PermissionResultDeny` with a message, so the agent gets feedback and adapts rather than dying. Every denial is written to the run ledger — that log is the audit trail the doc asks for.
|
|
85
85
|
|
|
86
86
|
### The envelope is the only contract
|
|
87
87
|
|
|
88
|
-
`envelope.py` is a Pydantic model of §6, and agents cannot emit anything else: their sole write path is `envelope.emit_envelope`, which validates and rejects with field-level errors on failure. Prose never crosses an agent boundary. `confidence < 0.6` routes to a human queue instead of
|
|
88
|
+
`envelope.py` is a Pydantic model of §6, and agents cannot emit anything else: their sole write path is `envelope.emit_envelope`, which validates and rejects with field-level errors on failure. Prose never crosses an agent boundary. `confidence < 0.6` routes to a human queue instead of TRIAGE (§7).
|
|
89
89
|
|
|
90
90
|
---
|
|
91
91
|
|
|
@@ -105,21 +105,21 @@ FastAPI + Postgres + React UI + one WS endpoint. Seed ~14 defects spanning the P
|
|
|
105
105
|
The six servers above plus local tracker/vcs adapters.
|
|
106
106
|
**Verify:** `pytest tests/mcp` — every tool exercised directly, zero API calls. `defect_memory` proven to dedupe two differently-worded reports of the same defect.
|
|
107
107
|
|
|
108
|
-
### M3 — Agent runtime +
|
|
109
|
-
`registry.py`, `runner.py`, `guardrails.py`, then the first real agent.
|
|
110
|
-
**Verify:** `qaas run --agent
|
|
108
|
+
### M3 — Agent runtime + MAPPER ✅ done
|
|
109
|
+
`registry.py`, `runner.py`, `guardrails.py`, then the first real agent. MAPPER goes first because §4.2 is right that everything downstream gets cheaper once the map exists.
|
|
110
|
+
**Verify:** `qaas run --agent MAPPER` writes a versioned `system-map.json` covering the target app's services, routes, schema, and ownership; schema-validated. Guardrail tests assert a write attempt from a read-only agent is denied and logged.
|
|
111
111
|
|
|
112
|
-
### M4 — Discovery:
|
|
113
|
-
|
|
112
|
+
### M4 — Discovery: API + BROWSER ✅ done
|
|
113
|
+
API gets `contract_diff` and ships a failing contract test as evidence (§4.5). BROWSER runs scripted journeys first, then exploratory from the Mapper task graph.
|
|
114
114
|
**Verify:** `qaas run --mode pr-check` emits envelopes; `qaas score` reports how many golden defects in those two domains were found and how many findings were not in the ledger.
|
|
115
115
|
|
|
116
|
-
### M5 — Triage:
|
|
117
|
-
|
|
116
|
+
### M5 — Triage: REPRODUCER + TRIAGE ✅ done
|
|
117
|
+
REPRODUCER reproduces, minimizes, runs N times for flake rate, and commits a failing test to `qa/repro/*`. TRIAGE dedupes, scores against the rubric, resolves owner from the map, routes, and files — the only agent holding tracker write.
|
|
118
118
|
**Verify:** full discovery→triage run produces local tickets with real repro steps and attached failing tests; a second run on the same code files **zero** new tickets and increments occurrence counts instead.
|
|
119
119
|
|
|
120
|
-
### M6 — Close the loop:
|
|
121
|
-
|
|
122
|
-
**Verify:** the acceptance test for the whole system — fix one seeded defect by hand on a branch, run `qaas run --mode fix-cycle --ticket <id>`, and
|
|
120
|
+
### M6 — Close the loop: VERIFIER + ROUTER run modes ✅ done, verified live
|
|
121
|
+
VERIFIER re-runs REPRODUCER's test against a patched build and returns `VERIFIED`/`NOT_FIXED`/`REGRESSED`. ROUTER gains all five discovery run modes, budget governor, concurrency caps, and escalation.
|
|
122
|
+
**Verify:** the acceptance test for the whole system — fix one seeded defect by hand on a branch, run `qaas run --mode fix-cycle --ticket <id>`, and VERIFIER returns `VERIFIED`; revert the fix and it returns `NOT_FIXED`. Then `qaas run --mode nightly && qaas score` prints the §12 scorecard: acceptance rate, duplicate rate, false-positive rate, cost per accepted ticket.
|
|
123
123
|
|
|
124
124
|
### Skills, hooks and loops ✅ done
|
|
125
125
|
|
|
@@ -141,17 +141,17 @@ that nothing counted. `Stop` blocks an agent that skipped its declared
|
|
|
141
141
|
deliverable (`must_call` in its config), honouring `stop_hook_active` so a
|
|
142
142
|
genuinely stuck agent cannot loop the budget away.
|
|
143
143
|
|
|
144
|
-
**Loops** — `_phase_verify` is now a bounded remediation loop:
|
|
145
|
-
→
|
|
146
|
-
`max_mender_arbiter_round_trips` from §8.3.
|
|
147
|
-
entry (`record_verdict`), never parsed from prose. With no
|
|
148
|
-
roster a NOT_FIXED escalates immediately instead of re-running
|
|
144
|
+
**Loops** — `_phase_verify` is now a bounded remediation loop: VERIFIER → NOT_FIXED
|
|
145
|
+
→ FIXER → REVIEWER → VERIFIER, enforcing `max_proof_reopens` and
|
|
146
|
+
`max_mender_arbiter_round_trips` from §8.3. VERIFIER's verdict is a typed ledger
|
|
147
|
+
entry (`record_verdict`), never parsed from prose. With no FIXER in the Phase 1
|
|
148
|
+
roster a NOT_FIXED escalates immediately instead of re-running VERIFIER against
|
|
149
149
|
unchanged code. `qaas sweep` is the cron entry point: run, score, and exit
|
|
150
150
|
non-zero below the §11 precision gate.
|
|
151
151
|
|
|
152
152
|
### M7 — Phase 3: the fix loop (beyond the original plan)
|
|
153
153
|
|
|
154
|
-
|
|
154
|
+
FIXER and REVIEWER, with the §8.2 autonomy envelope enforced in `guardrails.py`
|
|
155
155
|
rather than requested in a prompt: a diff budget counted per distinct file, and
|
|
156
156
|
forbidden path classes (migrations, auth, payment, secrets, infrastructure, CI)
|
|
157
157
|
that stop at a human however small the change looks. `record_review` refuses a
|
|
@@ -161,8 +161,8 @@ review. Two new run modes: `fix-cycle` and `full-loop`.
|
|
|
161
161
|
Merge remains impossible by construction: no merge method exists anywhere in the
|
|
162
162
|
codebase, `gh pr merge` is refused, and pull requests open as drafts.
|
|
163
163
|
|
|
164
|
-
**Verify:**
|
|
165
|
-
|
|
164
|
+
**Verify:** router tests prove the bounded loop (VERIFIER → FIXER → REVIEWER →
|
|
165
|
+
VERIFIER, capped by `max_mender_arbiter_round_trips` and `max_proof_reopens`), and
|
|
166
166
|
guardrail tests prove every forbidden class and the diff budget. Live: a
|
|
167
167
|
`full-loop` run on the demo app that ends in a VERIFIED ticket.
|
|
168
168
|
|
|
@@ -175,9 +175,9 @@ turned up two defects in this system rather than one in the target:
|
|
|
175
175
|
reporting no fixture and no way to pin a reproduction's environment. It
|
|
176
176
|
truncates first now.
|
|
177
177
|
2. `_verify_loop` re-read the branch off the envelope on every pass. That is the
|
|
178
|
-
*repro* branch, written before a fix exists, so
|
|
178
|
+
*repro* branch, written before a fix exists, so VERIFIER was sent back to the
|
|
179
179
|
unfixed branch it had just failed on. VERIFIED was unreachable by
|
|
180
|
-
construction. The loop now reads
|
|
180
|
+
construction. The loop now reads FIXER's branch out of the ledger, scoped to
|
|
181
181
|
the entries one remediation round appended.
|
|
182
182
|
|
|
183
183
|
Both are the point of running the thing live: neither was visible to 528 offline
|
|
@@ -185,10 +185,10 @@ tests, and (2) meant no `full-loop` run could ever have closed.
|
|
|
185
185
|
|
|
186
186
|
**A third defect, found only by re-running.** CORVID-7 then escalated twice with
|
|
187
187
|
a correct one-line fix sitting on the branch, because two rules in this repo
|
|
188
|
-
contradicted each other: `CLAUDE.md` required
|
|
189
|
-
defect's ledger entry in the same commit, and `
|
|
190
|
-
forbade it.
|
|
191
|
-
(`Edit write refused: target-app/defects.yaml is outside
|
|
188
|
+
contradicted each other: `CLAUDE.md` required FIXER to retire the seeded
|
|
189
|
+
defect's ledger entry in the same commit, and `fixer.yaml`'s `write_paths`
|
|
190
|
+
forbade it. FIXER tried and the guardrail refused
|
|
191
|
+
(`Edit write refused: target-app/defects.yaml is outside FIXER's sandbox`).
|
|
192
192
|
Every seeded defect has a ledger entry, so this deadlocked *any* fix to any of
|
|
193
193
|
them. The sandbox was right and the rule was wrong: the ledger is the answer key
|
|
194
194
|
and must stay unwritable by the agents it scores, so the same-commit obligation
|
|
@@ -198,16 +198,16 @@ lesson — never request a change the author is not permitted to make.
|
|
|
198
198
|
**Closed live on CORVID-8**, 2026-09-08, `fix-cycle`, $6.84, zero escalations:
|
|
199
199
|
|
|
200
200
|
```
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
201
|
+
VERIFIER NOT_FIXED defect confirmed present
|
|
202
|
+
FIXER currency: str added to InvoiceOut (1 line of product code)
|
|
203
|
+
REVIEWER APPROVE
|
|
204
|
+
VERIFIER VERIFIED re-verified on FIXER's branch
|
|
205
205
|
```
|
|
206
206
|
|
|
207
|
-
|
|
207
|
+
VERIFIER ran the original failing test 5 of 5 times for flake, then the full 22-test
|
|
208
208
|
`qa/repro` suite, and correctly attributed all 11 failures to other open tickets
|
|
209
209
|
(CORVID-7, CORVID-SEC-6) with source confirmation rather than to the diff. The
|
|
210
|
-
second
|
|
210
|
+
second VERIFIER dispatch is the hop the branch-selection fix above made reachable at
|
|
211
211
|
all. Verified independently afterwards: the diff is one product line and the live
|
|
212
212
|
endpoint returns `currency`.
|
|
213
213
|
|
|
@@ -235,13 +235,13 @@ verdicts, no escalations, no tickets. The audit trail existed; the audit did not
|
|
|
235
235
|
call instead of inventing a 29th kind no reader looks for.
|
|
236
236
|
- **Two provenance gaps closed.** `agent_started` recorded `task_chars=len(task)`
|
|
237
237
|
— the length of the prompt, not the prompt — so the instruction an agent
|
|
238
|
-
actually received was unrecoverable and five
|
|
238
|
+
actually received was unrecoverable and five REPRODUCER dispatches differed only by
|
|
239
239
|
character count; the task now goes to the artifact store with a preview inline.
|
|
240
240
|
And `run_started` recorded no commit, so a run was not pinned to the code it
|
|
241
241
|
examined; it now carries the target's sha, branch and dirty flag.
|
|
242
242
|
|
|
243
243
|
**Verify:** `pytest tests/test_trace.py` — 29 tests over a synthetic ledger and a
|
|
244
|
-
scripted
|
|
244
|
+
scripted router, offline. Includes an AST sweep asserting every literal passed
|
|
245
245
|
to `store.log()` anywhere in `src/qaas/**` is a `LedgerKind` member, so the next
|
|
246
246
|
kind added cannot go missing from the readers.
|
|
247
247
|
|
|
@@ -251,7 +251,7 @@ kind added cannot go missing from the readers.
|
|
|
251
251
|
test", through `ToolContext.repo_root` and `SystemConfig.target_app`. That is
|
|
252
252
|
true of exactly one target — the bundled demo, which happens to sit inside this
|
|
253
253
|
checkout — and false for every other. `ToolContext.target_root` now comes from
|
|
254
|
-
`profile.root_path()`, `SystemConfig.target_app` is gone, and
|
|
254
|
+
`profile.root_path()`, `SystemConfig.target_app` is gone, and FIXER's
|
|
255
255
|
`write_paths` are target-relative (`api/app`, not `target-app/api/app`).
|
|
256
256
|
|
|
257
257
|
`qaas run --repo <path-or-url>` follows from the split: it clones into
|
|
@@ -270,7 +270,7 @@ rather than layered by filename, so the first generated profile hid every
|
|
|
270
270
|
hand-written one.
|
|
271
271
|
|
|
272
272
|
### Adding agents 7–15 afterwards
|
|
273
|
-
A new discovery agent should be a prompt file plus a `config/agents/<NAME>.yaml` naming its allowlist — no changes to
|
|
273
|
+
A new discovery agent should be a prompt file plus a `config/agents/<NAME>.yaml` naming its allowlist — no changes to router, runner, or guardrails. Whether that holds is the real test of M3, so the first Phase-2 agent (DBA) will be added as a smoke test of the extension path before this build is called done.
|
|
274
274
|
|
|
275
275
|
---
|
|
276
276
|
|
|
@@ -278,7 +278,7 @@ A new discovery agent should be a prompt file plus a `config/agents/<NAME>.yaml`
|
|
|
278
278
|
|
|
279
279
|
Every run spends real money, and a nightly sweep is the expensive one. Controls, from the start:
|
|
280
280
|
|
|
281
|
-
- Per-agent and per-run `max_budget_usd` on `ClaudeAgentOptions`;
|
|
281
|
+
- Per-agent and per-run `max_budget_usd` on `ClaudeAgentOptions`; ROUTER aborts the run at the cap and records partial results.
|
|
282
282
|
- `CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH` and `CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS` set via `env` — Opus 5 delegates readily, and an unbounded subagent tree is the fastest way to a surprise bill.
|
|
283
283
|
- `--dry-run` renders each agent's exact options and prompt without calling the API; used in unit tests.
|
|
284
284
|
- A cheap CI profile (lower effort, smaller model for the mechanical agents) separate from the full profile.
|
|
@@ -286,7 +286,7 @@ Every run spends real money, and a nightly sweep is the expensive one. Controls,
|
|
|
286
286
|
|
|
287
287
|
Auth: no `ANTHROPIC_API_KEY` is set here, but the Claude Code CLI (v2.1.263) is installed and authenticated, and the Agent SDK drives it — so it works as-is. `ANTHROPIC_API_KEY` remains the CI path.
|
|
288
288
|
|
|
289
|
-
Model default is `claude-opus-5` for judgment-heavy agents (
|
|
289
|
+
Model default is `claude-opus-5` for judgment-heavy agents (API, BROWSER, REVIEWER later) and a cheaper model for mechanical ones (MAPPER extraction, TRIAGE composition), set per agent in `config/agents/*.yaml`.
|
|
290
290
|
|
|
291
291
|
---
|
|
292
292
|
|
|
@@ -302,5 +302,5 @@ Four tiers, cheapest first:
|
|
|
302
302
|
## Risks
|
|
303
303
|
|
|
304
304
|
- **Hook event names.** The SDK's `HookEvent` literals differ between docs and releases (`preToolUse` vs `PreToolUse`). M3 pins them by introspecting the installed `claude_agent_sdk` rather than trusting the docs, and `can_use_tool` carries the guardrails so a hook-name regression degrades logging, not enforcement.
|
|
305
|
-
- **Exploratory
|
|
305
|
+
- **Exploratory BROWSER is the noisiest agent.** It ships behind the confidence gate and the per-run finding cap from day one; if its false-positive rate is bad in M4, it runs scripted-only until the ledger says otherwise.
|
|
306
306
|
- **Seeded defects are easier than real ones.** The golden ledger measures whether the loop works, not whether it is good at finding hard bugs. Treat 100% recall on `defects.yaml` as a floor, never as evidence the system is ready for a real codebase.
|