adversary-gate 2.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- adversary_gate-2.0.0/PKG-INFO +148 -0
- adversary_gate-2.0.0/README.md +134 -0
- adversary_gate-2.0.0/pyproject.toml +33 -0
- adversary_gate-2.0.0/setup.cfg +4 -0
- adversary_gate-2.0.0/src/adversary_gate.egg-info/PKG-INFO +148 -0
- adversary_gate-2.0.0/src/adversary_gate.egg-info/SOURCES.txt +23 -0
- adversary_gate-2.0.0/src/adversary_gate.egg-info/dependency_links.txt +1 -0
- adversary_gate-2.0.0/src/adversary_gate.egg-info/entry_points.txt +2 -0
- adversary_gate-2.0.0/src/adversary_gate.egg-info/requires.txt +1 -0
- adversary_gate-2.0.0/src/adversary_gate.egg-info/top_level.txt +4 -0
- adversary_gate-2.0.0/src/cli.py +188 -0
- adversary_gate-2.0.0/src/core/circuit_breaker.py +55 -0
- adversary_gate-2.0.0/src/core/contestation.py +84 -0
- adversary_gate-2.0.0/src/core/evidence_log.py +140 -0
- adversary_gate-2.0.0/src/core/exitmap.py +62 -0
- adversary_gate-2.0.0/src/core/gate.py +507 -0
- adversary_gate-2.0.0/src/core/metrics.py +201 -0
- adversary_gate-2.0.0/src/core/path_policy.py +100 -0
- adversary_gate-2.0.0/src/core/quarantine.py +80 -0
- adversary_gate-2.0.0/src/core/types.py +165 -0
- adversary_gate-2.0.0/src/sandbox/runner.py +167 -0
- adversary_gate-2.0.0/src/verifiers/coverage.py +56 -0
- adversary_gate-2.0.0/src/verifiers/stability.py +29 -0
- adversary_gate-2.0.0/src/verifiers/strength.py +42 -0
- adversary_gate-2.0.0/tests/test_gate.py +793 -0
|
@@ -0,0 +1,148 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: adversary-gate
|
|
3
|
+
Version: 2.0.0
|
|
4
|
+
Summary: Evidence-based verification gate for AI coding agents (fail-closed)
|
|
5
|
+
Author: AdversaryGate Authors
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/adversary-gate/adversary-gate
|
|
8
|
+
Project-URL: Documentation, https://github.com/adversary-gate/adversary-gate#readme
|
|
9
|
+
Project-URL: Repository, https://github.com/adversary-gate/adversary-gate.git
|
|
10
|
+
Project-URL: Issues, https://github.com/adversary-gate/adversary-gate/issues
|
|
11
|
+
Requires-Python: >=3.10
|
|
12
|
+
Description-Content-Type: text/markdown
|
|
13
|
+
Requires-Dist: pytest>=8.0.0
|
|
14
|
+
|
|
15
|
+
# AdversaryGate (v2.0.0)
|
|
16
|
+
|
|
17
|
+
> **Uso alto de IA ≠ confiança alta.**
|
|
18
|
+
> O custo de um pipeline com agentes de código não está na inteligência do modelo, está no autoengano do pipeline.
|
|
19
|
+
|
|
20
|
+
An evidence-based fail-closed verification gate for AI coding agents where **uncertainty is a first-class result (`INCONCLUSIVE`)** instead of a silent approval.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## 🎯 The Thesis
|
|
25
|
+
|
|
26
|
+
1. **Uso alto de IA ≠ Confiança alta**: Quanto mais um time depende de agentes no fluxo real de engenharia, mais aparece o custo do *"parece certo"*. Agentes geram código fluente e aparentemente correto, mas pipelines ingênuos que colapsam erros de infraestrutura aprovam patches com testes quebrados ou pulados.
|
|
27
|
+
2. **Troca de Modelo como Sintoma**: Times trocam de modelo (Claude → GPT → Gemini) buscando credibilidade nos Pull Requests. Isso é **falta de verificação determinística, não falta de modelo**. O AdversaryGate executa exatamente o mesmo harness de teste sem invocar LLMs no verificador, tornando os modelos comparáveis empiricamente.
|
|
28
|
+
3. **Perda e Reprocessamento**: O prejuízo financeiro das empresas é concreto: *merged regressions*, tarefas "concluídas" que não estão, rollbacks e horas de code review humano repassando o mesmo PR.
|
|
29
|
+
4. **Menos Autoengano do Pipeline**: O produto não vende "IA mais inteligente". Vende **menos autoengano no pipeline**, medido numericamente pelo **`self_deception_index`**.
|
|
30
|
+
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
## 🔒 The Fail-Closed Model & Double-Filter Rigor
|
|
34
|
+
|
|
35
|
+
A single invariant governs the entire system:
|
|
36
|
+
> *The gate only reports what a healthy test harness actually executed. Unproduced proof is never proof of clean code.*
|
|
37
|
+
|
|
38
|
+
| State | Outcome | Meaning |
|
|
39
|
+
|---|---|---|
|
|
40
|
+
| `Executed & Passed` | `Outcome.VERIFIED` | Test ran to completion and evidence confirms clean execution. |
|
|
41
|
+
| `Executed & Failed` | `Outcome.REFUTED` | Test ran to completion and evidence condemns the patch. |
|
|
42
|
+
| `Harness / Error` | `Outcome.UNVERIFIED` | Collection error, syntax error, missing file, timeout or flaky signal. **Never mergeable.** |
|
|
43
|
+
|
|
44
|
+
### Patch Decision Matrix
|
|
45
|
+
|
|
46
|
+
- **`Decision.MERGE`**: Requires **every** verdict to be `VERIFIED`, `diff_coverage >= 80%`, `suite_strength >= 75%`, and the full repository test suite to pass.
|
|
47
|
+
- **`Decision.BLOCK`**: Triggered if any claim is `REFUTED` or if the full test suite fails (*collateral regression*).
|
|
48
|
+
- **`Decision.INCONCLUSIVE`**: Triggered on `UNVERIFIED` outcomes, open circuit breakers, or weak test suites (`suite_strength < 0.75`). **Never merges.**
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## 📊 Measured Pytest Exit Code Taxonomy
|
|
53
|
+
|
|
54
|
+
| Exit Code | Pytest Meaning | ExecState | Gate Behavior |
|
|
55
|
+
|---|---|---|---|
|
|
56
|
+
| `0` | Tests passed | `PASS` | Evaluated against baseline comparison |
|
|
57
|
+
| `1` | Tests failed | `FAIL` | **The ONLY exit code counted as evidence** |
|
|
58
|
+
| `2` | Collection error (import crash) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
59
|
+
| `3` | Internal error (harness crash) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
60
|
+
| `4` | Usage error (bad path / node ID) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
61
|
+
| `5` | No tests collected | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
62
|
+
| `-1` | Sandbox timeout | `TIMED_OUT` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
63
|
+
|
|
64
|
+
---
|
|
65
|
+
|
|
66
|
+
## 📦 Installation
|
|
67
|
+
|
|
68
|
+
Install via PyPI:
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
pip install adversary-gate
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Or run directly from source:
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
python3 -m cli --help
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
---
|
|
81
|
+
|
|
82
|
+
## 🚀 Quick Start (CLI & GitHub Action)
|
|
83
|
+
|
|
84
|
+
### Command-Line Usage
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
adversary-gate \
|
|
88
|
+
--baseline /path/to/before \
|
|
89
|
+
--patch /path/to/after \
|
|
90
|
+
--test-path tests/test_auth.py \
|
|
91
|
+
--test-id test_token_expiry \
|
|
92
|
+
--coverage-ratio 0.85 \
|
|
93
|
+
--model "${AGENT_MODEL_NAME}" \
|
|
94
|
+
--evidence-log evidence.jsonl \
|
|
95
|
+
--report
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
Exit Codes for CI Integration:
|
|
99
|
+
- `0` — **`MERGE`**: Every claim executed cleanly and cleared coverage & suite strength floors.
|
|
100
|
+
- `1` — **`BLOCK`**: Regressions or collateral suite failures detected.
|
|
101
|
+
- `2` — **`INCONCLUSIVE`**: Infrastructure error, missing test, or weak suite (blocks merge).
|
|
102
|
+
- `3` — **`USAGE_ERROR`**: Bad arguments or malformed JSON.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
### GitHub Action Integration (`action.yml`)
|
|
107
|
+
|
|
108
|
+
Add AdversaryGate to your GitHub Workflow:
|
|
109
|
+
|
|
110
|
+
```yaml
|
|
111
|
+
name: Verification Gate
|
|
112
|
+
|
|
113
|
+
on: [pull_request]
|
|
114
|
+
|
|
115
|
+
jobs:
|
|
116
|
+
verify-agent-patch:
|
|
117
|
+
runs-on: ubuntu-latest
|
|
118
|
+
steps:
|
|
119
|
+
- uses: actions/checkout@v4
|
|
120
|
+
|
|
121
|
+
- name: Run AdversaryGate
|
|
122
|
+
uses: adversary-gate/action@v2
|
|
123
|
+
with:
|
|
124
|
+
baseline: './baseline'
|
|
125
|
+
patch: './patch'
|
|
126
|
+
test-path: 'tests/test_token_expiry.py'
|
|
127
|
+
test-id: 'test_token_expiry'
|
|
128
|
+
coverage-floor: '0.80'
|
|
129
|
+
suite-strength-floor: '0.75'
|
|
130
|
+
model: '${{ matrix.model }}'
|
|
131
|
+
evidence-log: 'evidence.jsonl'
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## 📈 Provenance & `self_deception_index`
|
|
137
|
+
|
|
138
|
+
Every execution logs `ctx_model` and `ctx_commit` into the audit trail. Running patches through the same harness allows `compare_models()` to report verification rates side by side:
|
|
139
|
+
|
|
140
|
+
$$\text{self\_deception\_index} = \frac{\text{unverified\_merges}}{\text{merge\_count}}$$
|
|
141
|
+
|
|
142
|
+
If your product pitch is *"menos autoengano no pipeline"*, this is the dashboard tile that proves it and the metric to watch drop to zero.
|
|
143
|
+
|
|
144
|
+
---
|
|
145
|
+
|
|
146
|
+
## 📄 License
|
|
147
|
+
|
|
148
|
+
Distributed under the [MIT License](LICENSE).
|
|
@@ -0,0 +1,134 @@
|
|
|
1
|
+
# AdversaryGate (v2.0.0)
|
|
2
|
+
|
|
3
|
+
> **Uso alto de IA ≠ confiança alta.**
|
|
4
|
+
> O custo de um pipeline com agentes de código não está na inteligência do modelo, está no autoengano do pipeline.
|
|
5
|
+
|
|
6
|
+
An evidence-based fail-closed verification gate for AI coding agents where **uncertainty is a first-class result (`INCONCLUSIVE`)** instead of a silent approval.
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## 🎯 The Thesis
|
|
11
|
+
|
|
12
|
+
1. **Uso alto de IA ≠ Confiança alta**: Quanto mais um time depende de agentes no fluxo real de engenharia, mais aparece o custo do *"parece certo"*. Agentes geram código fluente e aparentemente correto, mas pipelines ingênuos que colapsam erros de infraestrutura aprovam patches com testes quebrados ou pulados.
|
|
13
|
+
2. **Troca de Modelo como Sintoma**: Times trocam de modelo (Claude → GPT → Gemini) buscando credibilidade nos Pull Requests. Isso é **falta de verificação determinística, não falta de modelo**. O AdversaryGate executa exatamente o mesmo harness de teste sem invocar LLMs no verificador, tornando os modelos comparáveis empiricamente.
|
|
14
|
+
3. **Perda e Reprocessamento**: O prejuízo financeiro das empresas é concreto: *merged regressions*, tarefas "concluídas" que não estão, rollbacks e horas de code review humano repassando o mesmo PR.
|
|
15
|
+
4. **Menos Autoengano do Pipeline**: O produto não vende "IA mais inteligente". Vende **menos autoengano no pipeline**, medido numericamente pelo **`self_deception_index`**.
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## 🔒 The Fail-Closed Model & Double-Filter Rigor
|
|
20
|
+
|
|
21
|
+
A single invariant governs the entire system:
|
|
22
|
+
> *The gate only reports what a healthy test harness actually executed. Unproduced proof is never proof of clean code.*
|
|
23
|
+
|
|
24
|
+
| State | Outcome | Meaning |
|
|
25
|
+
|---|---|---|
|
|
26
|
+
| `Executed & Passed` | `Outcome.VERIFIED` | Test ran to completion and evidence confirms clean execution. |
|
|
27
|
+
| `Executed & Failed` | `Outcome.REFUTED` | Test ran to completion and evidence condemns the patch. |
|
|
28
|
+
| `Harness / Error` | `Outcome.UNVERIFIED` | Collection error, syntax error, missing file, timeout or flaky signal. **Never mergeable.** |
|
|
29
|
+
|
|
30
|
+
### Patch Decision Matrix
|
|
31
|
+
|
|
32
|
+
- **`Decision.MERGE`**: Requires **every** verdict to be `VERIFIED`, `diff_coverage >= 80%`, `suite_strength >= 75%`, and the full repository test suite to pass.
|
|
33
|
+
- **`Decision.BLOCK`**: Triggered if any claim is `REFUTED` or if the full test suite fails (*collateral regression*).
|
|
34
|
+
- **`Decision.INCONCLUSIVE`**: Triggered on `UNVERIFIED` outcomes, open circuit breakers, or weak test suites (`suite_strength < 0.75`). **Never merges.**
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
## 📊 Measured Pytest Exit Code Taxonomy
|
|
39
|
+
|
|
40
|
+
| Exit Code | Pytest Meaning | ExecState | Gate Behavior |
|
|
41
|
+
|---|---|---|---|
|
|
42
|
+
| `0` | Tests passed | `PASS` | Evaluated against baseline comparison |
|
|
43
|
+
| `1` | Tests failed | `FAIL` | **The ONLY exit code counted as evidence** |
|
|
44
|
+
| `2` | Collection error (import crash) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
45
|
+
| `3` | Internal error (harness crash) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
46
|
+
| `4` | Usage error (bad path / node ID) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
47
|
+
| `5` | No tests collected | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
48
|
+
| `-1` | Sandbox timeout | `TIMED_OUT` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## 📦 Installation
|
|
53
|
+
|
|
54
|
+
Install via PyPI:
|
|
55
|
+
|
|
56
|
+
```bash
|
|
57
|
+
pip install adversary-gate
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
Or run directly from source:
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
python3 -m cli --help
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
## 🚀 Quick Start (CLI & GitHub Action)
|
|
69
|
+
|
|
70
|
+
### Command-Line Usage
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
adversary-gate \
|
|
74
|
+
--baseline /path/to/before \
|
|
75
|
+
--patch /path/to/after \
|
|
76
|
+
--test-path tests/test_auth.py \
|
|
77
|
+
--test-id test_token_expiry \
|
|
78
|
+
--coverage-ratio 0.85 \
|
|
79
|
+
--model "${AGENT_MODEL_NAME}" \
|
|
80
|
+
--evidence-log evidence.jsonl \
|
|
81
|
+
--report
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
Exit Codes for CI Integration:
|
|
85
|
+
- `0` — **`MERGE`**: Every claim executed cleanly and cleared coverage & suite strength floors.
|
|
86
|
+
- `1` — **`BLOCK`**: Regressions or collateral suite failures detected.
|
|
87
|
+
- `2` — **`INCONCLUSIVE`**: Infrastructure error, missing test, or weak suite (blocks merge).
|
|
88
|
+
- `3` — **`USAGE_ERROR`**: Bad arguments or malformed JSON.
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
### GitHub Action Integration (`action.yml`)
|
|
93
|
+
|
|
94
|
+
Add AdversaryGate to your GitHub Workflow:
|
|
95
|
+
|
|
96
|
+
```yaml
|
|
97
|
+
name: Verification Gate
|
|
98
|
+
|
|
99
|
+
on: [pull_request]
|
|
100
|
+
|
|
101
|
+
jobs:
|
|
102
|
+
verify-agent-patch:
|
|
103
|
+
runs-on: ubuntu-latest
|
|
104
|
+
steps:
|
|
105
|
+
- uses: actions/checkout@v4
|
|
106
|
+
|
|
107
|
+
- name: Run AdversaryGate
|
|
108
|
+
uses: adversary-gate/action@v2
|
|
109
|
+
with:
|
|
110
|
+
baseline: './baseline'
|
|
111
|
+
patch: './patch'
|
|
112
|
+
test-path: 'tests/test_token_expiry.py'
|
|
113
|
+
test-id: 'test_token_expiry'
|
|
114
|
+
coverage-floor: '0.80'
|
|
115
|
+
suite-strength-floor: '0.75'
|
|
116
|
+
model: '${{ matrix.model }}'
|
|
117
|
+
evidence-log: 'evidence.jsonl'
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
---
|
|
121
|
+
|
|
122
|
+
## 📈 Provenance & `self_deception_index`
|
|
123
|
+
|
|
124
|
+
Every execution logs `ctx_model` and `ctx_commit` into the audit trail. Running patches through the same harness allows `compare_models()` to report verification rates side by side:
|
|
125
|
+
|
|
126
|
+
$$\text{self\_deception\_index} = \frac{\text{unverified\_merges}}{\text{merge\_count}}$$
|
|
127
|
+
|
|
128
|
+
If your product pitch is *"menos autoengano no pipeline"*, this is the dashboard tile that proves it and the metric to watch drop to zero.
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
## 📄 License
|
|
133
|
+
|
|
134
|
+
Distributed under the [MIT License](LICENSE).
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
[build-system]
|
|
2
|
+
requires = ["setuptools>=68"]
|
|
3
|
+
build-backend = "setuptools.build_meta"
|
|
4
|
+
|
|
5
|
+
[project]
|
|
6
|
+
name = "adversary-gate"
|
|
7
|
+
version = "2.0.0"
|
|
8
|
+
description = "Evidence-based verification gate for AI coding agents (fail-closed)"
|
|
9
|
+
readme = "README.md"
|
|
10
|
+
authors = [
|
|
11
|
+
{name = "AdversaryGate Authors"}
|
|
12
|
+
]
|
|
13
|
+
license = "MIT"
|
|
14
|
+
requires-python = ">=3.10"
|
|
15
|
+
dependencies = [
|
|
16
|
+
"pytest>=8.0.0",
|
|
17
|
+
]
|
|
18
|
+
|
|
19
|
+
[project.urls]
|
|
20
|
+
Homepage = "https://github.com/adversary-gate/adversary-gate"
|
|
21
|
+
Documentation = "https://github.com/adversary-gate/adversary-gate#readme"
|
|
22
|
+
Repository = "https://github.com/adversary-gate/adversary-gate.git"
|
|
23
|
+
Issues = "https://github.com/adversary-gate/adversary-gate/issues"
|
|
24
|
+
|
|
25
|
+
[project.scripts]
|
|
26
|
+
adversary-gate = "cli:main"
|
|
27
|
+
|
|
28
|
+
[tool.setuptools]
|
|
29
|
+
package-dir = {"" = "src"}
|
|
30
|
+
py-modules = ["cli"]
|
|
31
|
+
|
|
32
|
+
[tool.setuptools.packages.find]
|
|
33
|
+
where = ["src"]
|
|
@@ -0,0 +1,148 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: adversary-gate
|
|
3
|
+
Version: 2.0.0
|
|
4
|
+
Summary: Evidence-based verification gate for AI coding agents (fail-closed)
|
|
5
|
+
Author: AdversaryGate Authors
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/adversary-gate/adversary-gate
|
|
8
|
+
Project-URL: Documentation, https://github.com/adversary-gate/adversary-gate#readme
|
|
9
|
+
Project-URL: Repository, https://github.com/adversary-gate/adversary-gate.git
|
|
10
|
+
Project-URL: Issues, https://github.com/adversary-gate/adversary-gate/issues
|
|
11
|
+
Requires-Python: >=3.10
|
|
12
|
+
Description-Content-Type: text/markdown
|
|
13
|
+
Requires-Dist: pytest>=8.0.0
|
|
14
|
+
|
|
15
|
+
# AdversaryGate (v2.0.0)
|
|
16
|
+
|
|
17
|
+
> **Uso alto de IA ≠ confiança alta.**
|
|
18
|
+
> O custo de um pipeline com agentes de código não está na inteligência do modelo, está no autoengano do pipeline.
|
|
19
|
+
|
|
20
|
+
An evidence-based fail-closed verification gate for AI coding agents where **uncertainty is a first-class result (`INCONCLUSIVE`)** instead of a silent approval.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## 🎯 The Thesis
|
|
25
|
+
|
|
26
|
+
1. **Uso alto de IA ≠ Confiança alta**: Quanto mais um time depende de agentes no fluxo real de engenharia, mais aparece o custo do *"parece certo"*. Agentes geram código fluente e aparentemente correto, mas pipelines ingênuos que colapsam erros de infraestrutura aprovam patches com testes quebrados ou pulados.
|
|
27
|
+
2. **Troca de Modelo como Sintoma**: Times trocam de modelo (Claude → GPT → Gemini) buscando credibilidade nos Pull Requests. Isso é **falta de verificação determinística, não falta de modelo**. O AdversaryGate executa exatamente o mesmo harness de teste sem invocar LLMs no verificador, tornando os modelos comparáveis empiricamente.
|
|
28
|
+
3. **Perda e Reprocessamento**: O prejuízo financeiro das empresas é concreto: *merged regressions*, tarefas "concluídas" que não estão, rollbacks e horas de code review humano repassando o mesmo PR.
|
|
29
|
+
4. **Menos Autoengano do Pipeline**: O produto não vende "IA mais inteligente". Vende **menos autoengano no pipeline**, medido numericamente pelo **`self_deception_index`**.
|
|
30
|
+
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
## 🔒 The Fail-Closed Model & Double-Filter Rigor
|
|
34
|
+
|
|
35
|
+
A single invariant governs the entire system:
|
|
36
|
+
> *The gate only reports what a healthy test harness actually executed. Unproduced proof is never proof of clean code.*
|
|
37
|
+
|
|
38
|
+
| State | Outcome | Meaning |
|
|
39
|
+
|---|---|---|
|
|
40
|
+
| `Executed & Passed` | `Outcome.VERIFIED` | Test ran to completion and evidence confirms clean execution. |
|
|
41
|
+
| `Executed & Failed` | `Outcome.REFUTED` | Test ran to completion and evidence condemns the patch. |
|
|
42
|
+
| `Harness / Error` | `Outcome.UNVERIFIED` | Collection error, syntax error, missing file, timeout or flaky signal. **Never mergeable.** |
|
|
43
|
+
|
|
44
|
+
### Patch Decision Matrix
|
|
45
|
+
|
|
46
|
+
- **`Decision.MERGE`**: Requires **every** verdict to be `VERIFIED`, `diff_coverage >= 80%`, `suite_strength >= 75%`, and the full repository test suite to pass.
|
|
47
|
+
- **`Decision.BLOCK`**: Triggered if any claim is `REFUTED` or if the full test suite fails (*collateral regression*).
|
|
48
|
+
- **`Decision.INCONCLUSIVE`**: Triggered on `UNVERIFIED` outcomes, open circuit breakers, or weak test suites (`suite_strength < 0.75`). **Never merges.**
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## 📊 Measured Pytest Exit Code Taxonomy
|
|
53
|
+
|
|
54
|
+
| Exit Code | Pytest Meaning | ExecState | Gate Behavior |
|
|
55
|
+
|---|---|---|---|
|
|
56
|
+
| `0` | Tests passed | `PASS` | Evaluated against baseline comparison |
|
|
57
|
+
| `1` | Tests failed | `FAIL` | **The ONLY exit code counted as evidence** |
|
|
58
|
+
| `2` | Collection error (import crash) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
59
|
+
| `3` | Internal error (harness crash) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
60
|
+
| `4` | Usage error (bad path / node ID) | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
61
|
+
| `5` | No tests collected | `UNRUNNABLE` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
62
|
+
| `-1` | Sandbox timeout | `TIMED_OUT` | `UNVERIFIED` $\rightarrow$ `INCONCLUSIVE` |
|
|
63
|
+
|
|
64
|
+
---
|
|
65
|
+
|
|
66
|
+
## 📦 Installation
|
|
67
|
+
|
|
68
|
+
Install via PyPI:
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
pip install adversary-gate
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Or run directly from source:
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
python3 -m cli --help
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
---
|
|
81
|
+
|
|
82
|
+
## 🚀 Quick Start (CLI & GitHub Action)
|
|
83
|
+
|
|
84
|
+
### Command-Line Usage
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
adversary-gate \
|
|
88
|
+
--baseline /path/to/before \
|
|
89
|
+
--patch /path/to/after \
|
|
90
|
+
--test-path tests/test_auth.py \
|
|
91
|
+
--test-id test_token_expiry \
|
|
92
|
+
--coverage-ratio 0.85 \
|
|
93
|
+
--model "${AGENT_MODEL_NAME}" \
|
|
94
|
+
--evidence-log evidence.jsonl \
|
|
95
|
+
--report
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
Exit Codes for CI Integration:
|
|
99
|
+
- `0` — **`MERGE`**: Every claim executed cleanly and cleared coverage & suite strength floors.
|
|
100
|
+
- `1` — **`BLOCK`**: Regressions or collateral suite failures detected.
|
|
101
|
+
- `2` — **`INCONCLUSIVE`**: Infrastructure error, missing test, or weak suite (blocks merge).
|
|
102
|
+
- `3` — **`USAGE_ERROR`**: Bad arguments or malformed JSON.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
### GitHub Action Integration (`action.yml`)
|
|
107
|
+
|
|
108
|
+
Add AdversaryGate to your GitHub Workflow:
|
|
109
|
+
|
|
110
|
+
```yaml
|
|
111
|
+
name: Verification Gate
|
|
112
|
+
|
|
113
|
+
on: [pull_request]
|
|
114
|
+
|
|
115
|
+
jobs:
|
|
116
|
+
verify-agent-patch:
|
|
117
|
+
runs-on: ubuntu-latest
|
|
118
|
+
steps:
|
|
119
|
+
- uses: actions/checkout@v4
|
|
120
|
+
|
|
121
|
+
- name: Run AdversaryGate
|
|
122
|
+
uses: adversary-gate/action@v2
|
|
123
|
+
with:
|
|
124
|
+
baseline: './baseline'
|
|
125
|
+
patch: './patch'
|
|
126
|
+
test-path: 'tests/test_token_expiry.py'
|
|
127
|
+
test-id: 'test_token_expiry'
|
|
128
|
+
coverage-floor: '0.80'
|
|
129
|
+
suite-strength-floor: '0.75'
|
|
130
|
+
model: '${{ matrix.model }}'
|
|
131
|
+
evidence-log: 'evidence.jsonl'
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## 📈 Provenance & `self_deception_index`
|
|
137
|
+
|
|
138
|
+
Every execution logs `ctx_model` and `ctx_commit` into the audit trail. Running patches through the same harness allows `compare_models()` to report verification rates side by side:
|
|
139
|
+
|
|
140
|
+
$$\text{self\_deception\_index} = \frac{\text{unverified\_merges}}{\text{merge\_count}}$$
|
|
141
|
+
|
|
142
|
+
If your product pitch is *"menos autoengano no pipeline"*, this is the dashboard tile that proves it and the metric to watch drop to zero.
|
|
143
|
+
|
|
144
|
+
---
|
|
145
|
+
|
|
146
|
+
## 📄 License
|
|
147
|
+
|
|
148
|
+
Distributed under the [MIT License](LICENSE).
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
README.md
|
|
2
|
+
pyproject.toml
|
|
3
|
+
src/cli.py
|
|
4
|
+
src/adversary_gate.egg-info/PKG-INFO
|
|
5
|
+
src/adversary_gate.egg-info/SOURCES.txt
|
|
6
|
+
src/adversary_gate.egg-info/dependency_links.txt
|
|
7
|
+
src/adversary_gate.egg-info/entry_points.txt
|
|
8
|
+
src/adversary_gate.egg-info/requires.txt
|
|
9
|
+
src/adversary_gate.egg-info/top_level.txt
|
|
10
|
+
src/core/circuit_breaker.py
|
|
11
|
+
src/core/contestation.py
|
|
12
|
+
src/core/evidence_log.py
|
|
13
|
+
src/core/exitmap.py
|
|
14
|
+
src/core/gate.py
|
|
15
|
+
src/core/metrics.py
|
|
16
|
+
src/core/path_policy.py
|
|
17
|
+
src/core/quarantine.py
|
|
18
|
+
src/core/types.py
|
|
19
|
+
src/sandbox/runner.py
|
|
20
|
+
src/verifiers/coverage.py
|
|
21
|
+
src/verifiers/stability.py
|
|
22
|
+
src/verifiers/strength.py
|
|
23
|
+
tests/test_gate.py
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
pytest>=8.0.0
|