aau-harness 1.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- aau_harness-1.1.0/PKG-INFO +428 -0
- aau_harness-1.1.0/README.md +410 -0
- aau_harness-1.1.0/pyproject.toml +34 -0
- aau_harness-1.1.0/setup.cfg +4 -0
- aau_harness-1.1.0/src/aau_harness/__init__.py +110 -0
- aau_harness-1.1.0/src/aau_harness/agent_loop.py +132 -0
- aau_harness-1.1.0/src/aau_harness/catalog_cli.py +350 -0
- aau_harness-1.1.0/src/aau_harness/challenge.py +305 -0
- aau_harness-1.1.0/src/aau_harness/contract_runtime.py +737 -0
- aau_harness-1.1.0/src/aau_harness/cost.py +176 -0
- aau_harness-1.1.0/src/aau_harness/decision_gate.py +720 -0
- aau_harness-1.1.0/src/aau_harness/delegation.py +163 -0
- aau_harness-1.1.0/src/aau_harness/evaluate.py +386 -0
- aau_harness-1.1.0/src/aau_harness/evidence_service.py +664 -0
- aau_harness-1.1.0/src/aau_harness/forge.py +575 -0
- aau_harness-1.1.0/src/aau_harness/forge_contracts.py +684 -0
- aau_harness-1.1.0/src/aau_harness/gallery.py +440 -0
- aau_harness-1.1.0/src/aau_harness/llm_providers.py +347 -0
- aau_harness-1.1.0/src/aau_harness/provenance.py +83 -0
- aau_harness-1.1.0/src/aau_harness/public_value.py +112 -0
- aau_harness-1.1.0/src/aau_harness/report.py +26 -0
- aau_harness-1.1.0/src/aau_harness/reporting.py +153 -0
- aau_harness-1.1.0/src/aau_harness/runner.py +190 -0
- aau_harness-1.1.0/src/aau_harness/scaffold.py +829 -0
- aau_harness-1.1.0/src/aau_harness.egg-info/PKG-INFO +428 -0
- aau_harness-1.1.0/src/aau_harness.egg-info/SOURCES.txt +44 -0
- aau_harness-1.1.0/src/aau_harness.egg-info/dependency_links.txt +1 -0
- aau_harness-1.1.0/src/aau_harness.egg-info/entry_points.txt +7 -0
- aau_harness-1.1.0/src/aau_harness.egg-info/requires.txt +7 -0
- aau_harness-1.1.0/src/aau_harness.egg-info/top_level.txt +1 -0
- aau_harness-1.1.0/tests/test_catalog_cli.py +63 -0
- aau_harness-1.1.0/tests/test_challenge.py +101 -0
- aau_harness-1.1.0/tests/test_cost.py +35 -0
- aau_harness-1.1.0/tests/test_delegation.py +151 -0
- aau_harness-1.1.0/tests/test_evaluate.py +123 -0
- aau_harness-1.1.0/tests/test_evidence_service.py +112 -0
- aau_harness-1.1.0/tests/test_forge.py +108 -0
- aau_harness-1.1.0/tests/test_gallery.py +56 -0
- aau_harness-1.1.0/tests/test_harness.py +72 -0
- aau_harness-1.1.0/tests/test_provenance.py +20 -0
- aau_harness-1.1.0/tests/test_public_value.py +71 -0
- aau_harness-1.1.0/tests/test_readme_quickstart.py +34 -0
- aau_harness-1.1.0/tests/test_reporting.py +104 -0
- aau_harness-1.1.0/tests/test_runner.py +111 -0
- aau_harness-1.1.0/tests/test_scaffold.py +72 -0
- aau_harness-1.1.0/tests/test_streaming.py +116 -0
|
@@ -0,0 +1,428 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: aau-harness
|
|
3
|
+
Version: 1.1.0
|
|
4
|
+
Summary: Provider-neutral agent evaluation: seeded scenarios, exact scoring, BYO-agent adapters, public-value contracts, repeated runs, cost, and receipts.
|
|
5
|
+
Author: Imran Ahamed
|
|
6
|
+
License-Expression: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/immu4989/awesome-agentic-usecases
|
|
8
|
+
Project-URL: Documentation, https://github.com/immu4989/awesome-agentic-usecases/tree/main/harness
|
|
9
|
+
Project-URL: Issues, https://github.com/immu4989/awesome-agentic-usecases/issues
|
|
10
|
+
Project-URL: Changelog, https://github.com/immu4989/awesome-agentic-usecases/blob/main/CHANGELOG.md
|
|
11
|
+
Keywords: agents,evaluation,llm,agentic-ai,public-value,tevv
|
|
12
|
+
Requires-Python: >=3.10
|
|
13
|
+
Description-Content-Type: text/markdown
|
|
14
|
+
Requires-Dist: tomli>=2.0; python_version < "3.11"
|
|
15
|
+
Provides-Extra: dev
|
|
16
|
+
Requires-Dist: pytest>=7.0; extra == "dev"
|
|
17
|
+
Requires-Dist: ruff>=0.4; extra == "dev"
|
|
18
|
+
|
|
19
|
+
# aau-harness
|
|
20
|
+
|
|
21
|
+
Reproducible evaluation of tool-using LLM agents: seeded worlds, exact scoring, measured
|
|
22
|
+
cost, repeated runs with confidence intervals, and provenance on every result.
|
|
23
|
+
|
|
24
|
+
This is the library behind [awesome-agentic-usecases](https://github.com/immu4989/awesome-agentic-usecases).
|
|
25
|
+
It is usable on its own — you supply a domain (scenarios, tools, a gold rule, a prompt) and
|
|
26
|
+
the harness supplies everything around it.
|
|
27
|
+
|
|
28
|
+
---
|
|
29
|
+
|
|
30
|
+
## Install
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
pip install -e harness # from a clone of the repo
|
|
34
|
+
pip install -e harness[dev] # plus pytest and ruff
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Requires Python 3.10+. The core uses only the standard library on Python 3.11+ (Python
|
|
38
|
+
3.10 installs the small `tomli` compatibility package); the `anthropic` extra is only
|
|
39
|
+
needed for the native Anthropic backend, and every other provider is reached over `urllib`.
|
|
40
|
+
|
|
41
|
+
Verify the install:
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
pytest harness/tests -q
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Evaluate an existing agent
|
|
48
|
+
|
|
49
|
+
You do not need to rebuild an application around the harness. `aau evaluate` sends each case to
|
|
50
|
+
an existing command or HTTP endpoint through a four-field JSON response contract, then emits a
|
|
51
|
+
public aggregate receipt.
|
|
52
|
+
|
|
53
|
+
Suites must explicitly attest public, synthetic, or public-synthetic classification, completed
|
|
54
|
+
human review, and the absence of PII, credentials, procurement-sensitive, controlled, and
|
|
55
|
+
classified information. The CLI fails closed when any attestation is missing.
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
aau evaluate harness/examples/byo-agent-suite.json \
|
|
59
|
+
--command "python harness/examples/byo_agent_adapter.py" \
|
|
60
|
+
--out aau-agent-receipt.json
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
The adapter reads one JSON request from standard input and writes one JSON object:
|
|
64
|
+
|
|
65
|
+
```json
|
|
66
|
+
{
|
|
67
|
+
"outcome": "route_official_source",
|
|
68
|
+
"actions_attempted": [],
|
|
69
|
+
"actions_executed": [],
|
|
70
|
+
"submitted": true
|
|
71
|
+
}
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Use `--endpoint http://127.0.0.1:8000/evaluate` for a JSON POST endpoint, or `--mock` to verify
|
|
75
|
+
the suite protocol without running an agent. The evaluator measures exact outcome, submission,
|
|
76
|
+
forbidden-action attempts, forbidden-action execution, and latency. It executes command adapters
|
|
77
|
+
as an argument vector with `shell=False`, enforces suite/response size and timeout limits, and
|
|
78
|
+
never copies environment variables or request headers into receipts.
|
|
79
|
+
|
|
80
|
+
Public receipts deliberately omit scenario inputs, expected answers, raw adapter responses,
|
|
81
|
+
reasoning, and credentials. `--private-out` is available for local debugging and may contain
|
|
82
|
+
sensitive material; never publish it without authorized review. A passing receipt is not
|
|
83
|
+
production validation, certification, model ranking, legal advice, or permission to automate a
|
|
84
|
+
protected decision.
|
|
85
|
+
|
|
86
|
+
### Run in GitHub Actions
|
|
87
|
+
|
|
88
|
+
```yaml
|
|
89
|
+
- uses: immu4989/awesome-agentic-usecases/.github/actions/aau-evaluate@main
|
|
90
|
+
with:
|
|
91
|
+
suite: evals/public-suite.json
|
|
92
|
+
adapter-command: python app/aau_adapter.py
|
|
93
|
+
receipt: artifacts/aau-agent-receipt.json
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Pin the action to a release tag or commit SHA in production. The composite action installs the
|
|
97
|
+
repository-pinned harness and returns the public receipt path. See
|
|
98
|
+
[`harness/PUBLISHING.md`](PUBLISHING.md) for the tokenless PyPI release process.
|
|
99
|
+
|
|
100
|
+
### Find the right use case
|
|
101
|
+
|
|
102
|
+
Installing the harness also adds the repository navigator. It searches the committed
|
|
103
|
+
machine-readable catalog and prints exact commands without making network calls:
|
|
104
|
+
|
|
105
|
+
```bash
|
|
106
|
+
aau list
|
|
107
|
+
aau list --industry healthcare
|
|
108
|
+
aau find "security adversarial"
|
|
109
|
+
aau show refund-memory
|
|
110
|
+
aau start refund-injected
|
|
111
|
+
aau challenge list
|
|
112
|
+
aau challenge show completion-is-not-correctness
|
|
113
|
+
aau doctor
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
`aau start` understands local package dependencies, so controlled comparisons that reuse a
|
|
117
|
+
baseline are installed in the correct order. It prints commands; it never changes your
|
|
118
|
+
environment by itself.
|
|
119
|
+
|
|
120
|
+
`aau challenge` adds the community Reliability Challenge: list bounded Reproduce, Break,
|
|
121
|
+
and Adapt missions, print their exact zero-cost commands, or validate a Challenge-enabled
|
|
122
|
+
Gallery entry and derive its achievements from committed evidence.
|
|
123
|
+
|
|
124
|
+
## Quickstart
|
|
125
|
+
|
|
126
|
+
A complete evaluation. It runs on the built-in deterministic mock backend, so it needs no
|
|
127
|
+
API key and costs nothing.
|
|
128
|
+
|
|
129
|
+
```python
|
|
130
|
+
from dataclasses import dataclass
|
|
131
|
+
from aau_harness import (
|
|
132
|
+
Block, CostTracker, MockUsage, ScenarioResult,
|
|
133
|
+
make_backend, render_report, run_eval, run_tool_agent,
|
|
134
|
+
)
|
|
135
|
+
|
|
136
|
+
# 1. A world. Gold comes from a rule the scorer will share — never re-derived.
|
|
137
|
+
@dataclass
|
|
138
|
+
class Scenario:
|
|
139
|
+
scenario_id: str
|
|
140
|
+
text: str
|
|
141
|
+
amount: int
|
|
142
|
+
gold: str
|
|
143
|
+
|
|
144
|
+
def gold_rule(amount: int) -> str:
|
|
145
|
+
return "escalate" if amount > 100 else "approve"
|
|
146
|
+
|
|
147
|
+
scenarios = [Scenario(f"sc-{i:03d}", f"Request for {i * 40} units", i * 40,
|
|
148
|
+
gold_rule(i * 40)) for i in range(6)]
|
|
149
|
+
|
|
150
|
+
# 2. Tools the agent may call. Strict schemas keep submissions well-formed.
|
|
151
|
+
TOOLS = [{
|
|
152
|
+
"name": "submit",
|
|
153
|
+
"description": "Commit the decision. Call once, last.",
|
|
154
|
+
"strict": True,
|
|
155
|
+
"input_schema": {
|
|
156
|
+
"type": "object",
|
|
157
|
+
"properties": {"decision": {"type": "string", "enum": ["approve", "escalate"]}},
|
|
158
|
+
"required": ["decision"],
|
|
159
|
+
"additionalProperties": False,
|
|
160
|
+
},
|
|
161
|
+
}]
|
|
162
|
+
|
|
163
|
+
# 3. A deterministic stand-in model, so the pipeline runs with no API key.
|
|
164
|
+
class Mock:
|
|
165
|
+
name = model = "mock"
|
|
166
|
+
def create(self, system, messages, tools):
|
|
167
|
+
amount = int("".join(c for c in messages[0]["content"] if c.isdigit()) or 0)
|
|
168
|
+
return Block(
|
|
169
|
+
content=[Block(type="tool_use", id="m1", name="submit",
|
|
170
|
+
input={"decision": gold_rule(amount)})],
|
|
171
|
+
stop_reason="tool_use",
|
|
172
|
+
usage=MockUsage(input_tokens=400, output_tokens=20),
|
|
173
|
+
)
|
|
174
|
+
|
|
175
|
+
# 4. Score one run, then let the runner handle repeats and uncertainty.
|
|
176
|
+
def run_one(sc: Scenario, repeat: int) -> ScenarioResult:
|
|
177
|
+
cost = CostTracker(model="mock")
|
|
178
|
+
run = run_tool_agent(
|
|
179
|
+
make_backend("mock", mock_factory=Mock), "You are a triage agent.",
|
|
180
|
+
TOOLS, sc.text, lambda name, ti: "{}", "submit", cost,
|
|
181
|
+
)
|
|
182
|
+
sub = run.submission or {}
|
|
183
|
+
return ScenarioResult(
|
|
184
|
+
scenario_id=sc.scenario_id, repeat=repeat,
|
|
185
|
+
metrics={"accuracy": float(sub.get("decision") == sc.gold),
|
|
186
|
+
"submitted": float(run.submitted)},
|
|
187
|
+
cost_usd=cost.cost_usd, latency_s=0.0, n_api_calls=cost.api_calls,
|
|
188
|
+
detail={"gold": sc.gold, "predicted": sub.get("decision")},
|
|
189
|
+
)
|
|
190
|
+
|
|
191
|
+
agg = run_eval(scenarios, run_one, repeats=3)
|
|
192
|
+
print(render_report(agg, model="mock"))
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
To run the same evaluation against a real model, change one line — the rest is identical:
|
|
196
|
+
|
|
197
|
+
```python
|
|
198
|
+
backend = make_backend("openrouter", model="nvidia/nemotron-3-super-120b-a12b:free")
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
## Core API
|
|
202
|
+
|
|
203
|
+
### Evaluation
|
|
204
|
+
|
|
205
|
+
| | |
|
|
206
|
+
|---|---|
|
|
207
|
+
| `run_eval(scenarios, run_one, repeats=3, progress=None) -> EvalAggregate` | Runs `run_one(scenario, repeat)` across every scenario × repeat and aggregates. Metrics are averaged per scenario across repeats, then bootstrapped **over scenarios**, keeping a scenario's repeats together (paired). Repeats are the default because a single agent run is noise. |
|
|
208
|
+
| `EvalAggregate` | `n_scenarios`, `n_repeats`, `metric_means`, `metric_ci95`, `mean_cost_per_scenario_usd`, `total_cost_usd`, `p50_latency_s`, `results`. `as_dict()` serialises it, stamping provenance automatically. |
|
|
209
|
+
| `ScenarioResult` | One run: `scenario_id`, `repeat`, `metrics`, `cost_usd`, `latency_s`, `n_api_calls`, `detail`. Put anything you may want to analyse later in `detail` — per-archetype breakdowns are computed from it. |
|
|
210
|
+
|
|
211
|
+
> **Every metric must be present on every scenario.** The runner aggregates by metric name
|
|
212
|
+
> across all results; a metric emitted for only some scenarios will fail. For subgroup
|
|
213
|
+
> analysis, emit `0.0` and record the subgroup in `detail`.
|
|
214
|
+
|
|
215
|
+
### Agent loop
|
|
216
|
+
|
|
217
|
+
| | |
|
|
218
|
+
|---|---|
|
|
219
|
+
| `run_tool_agent(backend, system_prompt, tool_schemas, user_message, execute_tool, submit_tool, cost, max_turns=8) -> AgentRun` | Owns turn-taking, usage accounting, refusals, and the no-submission path. `execute_tool(name, input) -> str` returns a JSON string; a stateful session object works too, since anything callable is accepted. |
|
|
220
|
+
| `AgentRun` | `submitted`, `submission`, `n_turns`, `tool_calls`, `refused`, `error`. **Check `submitted` before reading any other metric** — a model that never commits suppresses accuracy without being wrong. |
|
|
221
|
+
|
|
222
|
+
### Backends
|
|
223
|
+
|
|
224
|
+
`make_backend(kind, model=None, mock_factory=None)` resolves `"mock"`, `"anthropic"`
|
|
225
|
+
(`AnthropicBackend`, the one backend using a vendor SDK), or any
|
|
226
|
+
OpenAI-compatible provider: `mistral`, `groq`, `gemini`, `cerebras`, `deepseek`, `together`,
|
|
227
|
+
`fireworks`, `openrouter`. Each reads its key from the environment (`MISTRAL_API_KEY` and so
|
|
228
|
+
on). Backends are duck-typed — anything with
|
|
229
|
+
`create(system, messages, tools)` returning `.content` / `.stop_reason` / `.usage` works.
|
|
230
|
+
|
|
231
|
+
`openrouter` reaches several hundred tool-calling models through one key, including free
|
|
232
|
+
ones, which is how results here stay reproducible at zero cost. Note that many free models
|
|
233
|
+
ignore tool definitions entirely; probe before committing to one.
|
|
234
|
+
|
|
235
|
+
### Cost
|
|
236
|
+
|
|
237
|
+
`CostTracker(model=...)` accumulates `add_usage(response.usage)` and exposes `cost_usd` and
|
|
238
|
+
`api_calls`, pricing input, output, cache-write and cache-read tokens at published rates
|
|
239
|
+
from `PRICING_PER_MTOK`. Unknown models raise rather than silently reporting `$0`; for
|
|
240
|
+
aggregator-served models the rate is fetched from the provider's published API.
|
|
241
|
+
|
|
242
|
+
Reported cost is **list price applied to measured tokens** — on a free tier your actual
|
|
243
|
+
spend is zero while the reported figure is not.
|
|
244
|
+
|
|
245
|
+
### Guards
|
|
246
|
+
|
|
247
|
+
| | |
|
|
248
|
+
|---|---|
|
|
249
|
+
| `provider_error_rate(agg) -> float` | Fraction of runs that died at the transport layer rather than on the task. |
|
|
250
|
+
| `check_results_are_measurements(agg, threshold=0.5)` | Raises `ProviderUnavailable` when most runs never reached the model. **Call before saving.** An expired key produces a complete, well-formed result of zeros that is indistinguishable in storage from a model failing every scenario. |
|
|
251
|
+
|
|
252
|
+
### Multi-agent
|
|
253
|
+
|
|
254
|
+
`Specialist(...)`, `make_delegate_tool(...)` and `run_crew(...)` build orchestrator +
|
|
255
|
+
sub-agent systems where delegation is a tool call. `run_crew` returns a `CrewRun` carrying
|
|
256
|
+
the orchestrator's own `AgentRun` plus a `DelegationRecord` per sub-agent call, so you can
|
|
257
|
+
see which specialist was asked what and what it returned.
|
|
258
|
+
|
|
259
|
+
Two properties are enforced so comparisons against a single agent stay honest: sub-agent
|
|
260
|
+
cost rolls up into one tracker, and a specialist sees **only** its brief, so an omitted fact
|
|
261
|
+
is genuinely unavailable to it.
|
|
262
|
+
|
|
263
|
+
### Public-value service contracts
|
|
264
|
+
|
|
265
|
+
`PublicValueContract(...)` declares the exact terminal outcome plus the minimum evidence,
|
|
266
|
+
already-held evidence, required delivery channel, recourse, deadline protection, and
|
|
267
|
+
forbidden events for one service interaction. `PublicValueTrace(...)` normalizes what the
|
|
268
|
+
tools actually attempted and executed. `score_public_value(contract, trace)` returns the
|
|
269
|
+
component metrics and a conjunctive `public_value_exact` score.
|
|
270
|
+
|
|
271
|
+
```python
|
|
272
|
+
contract = PublicValueContract(
|
|
273
|
+
version="policy-2026.04",
|
|
274
|
+
expected_terminal="request_evidence",
|
|
275
|
+
required_evidence=("identity", "ownership", "loss_schedule"),
|
|
276
|
+
held_evidence=("identity", "ownership"),
|
|
277
|
+
required_channel="phone_711",
|
|
278
|
+
recourse_required=True,
|
|
279
|
+
)
|
|
280
|
+
trace = PublicValueTrace(
|
|
281
|
+
terminal_events=("request_evidence",),
|
|
282
|
+
requested_evidence=("loss_schedule",),
|
|
283
|
+
delivery_channels=("phone_711",),
|
|
284
|
+
recourse_offered=True,
|
|
285
|
+
deadline_preserved=False,
|
|
286
|
+
attempted_events=("request_evidence",),
|
|
287
|
+
executed_events=("request_evidence",),
|
|
288
|
+
submitted=True,
|
|
289
|
+
)
|
|
290
|
+
metrics = score_public_value(contract, trace)
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
The reference implementation and language-neutral schema live in the
|
|
294
|
+
[Public Value Contract](../PUBLIC_VALUE_CONTRACT.md) specialty.
|
|
295
|
+
|
|
296
|
+
### High-stakes decision gates
|
|
297
|
+
|
|
298
|
+
`GateContract(...)` declares the exact outcome, reason code, required and held evidence,
|
|
299
|
+
satisfied gate set, applicable procedural protections, and forbidden protected event for
|
|
300
|
+
one fictional case. `GateScenario(...)` combines that contract with trusted records and a
|
|
301
|
+
versioned policy snapshot. `generate_gate_scenarios(...)` creates a balanced suite using
|
|
302
|
+
the eight shapes in `ARCHETYPE_ORDER`.
|
|
303
|
+
|
|
304
|
+
`build_gate_policy(...)`, `build_gate_tool_schemas(...)`, and
|
|
305
|
+
`build_gate_system_prompt(...)` turn a domain configuration into the reusable environment.
|
|
306
|
+
`GateToolSession(...)` records every lookup, bounded action, protected attempt, evidence
|
|
307
|
+
set, gate confirmation, and procedural flag. `score_gate_run(...)` compares that trace with
|
|
308
|
+
the contract and emits component metrics plus conjunctive `decision_gate_exact`.
|
|
309
|
+
|
|
310
|
+
`evaluate_gate(...)` runs the shared environment with any harness backend. The built-in
|
|
311
|
+
`GateMockBackend(...)` intentionally duplicates evidence, generalizes across a transfer
|
|
312
|
+
trap, drops procedure, and crosses authority so every detector can be exercised at $0.
|
|
313
|
+
|
|
314
|
+
```python
|
|
315
|
+
from aau_harness import (
|
|
316
|
+
GateMockBackend,
|
|
317
|
+
evaluate_gate,
|
|
318
|
+
generate_gate_scenarios,
|
|
319
|
+
)
|
|
320
|
+
|
|
321
|
+
scenarios = generate_gate_scenarios(domain_config, n=32, seed=277)
|
|
322
|
+
aggregate = evaluate_gate(
|
|
323
|
+
domain_config,
|
|
324
|
+
scenarios,
|
|
325
|
+
GateMockBackend,
|
|
326
|
+
backend_kind="mock",
|
|
327
|
+
repeats=3,
|
|
328
|
+
)
|
|
329
|
+
assert 0 < aggregate.metric_means["decision_gate_exact"] < 1
|
|
330
|
+
```
|
|
331
|
+
|
|
332
|
+
The complete contract, authority rules, and six domain configurations live in the
|
|
333
|
+
[Decision Gate Contract](../DECISION_GATE_CONTRACT.md) specialty.
|
|
334
|
+
|
|
335
|
+
### Contract-aware Forge runtime
|
|
336
|
+
|
|
337
|
+
`CompiledContract(...)` and `ContractScenario(...)` represent the contract family, exact
|
|
338
|
+
outcome and reason, evidence sets, structured nodes, safeguards, and forbidden events.
|
|
339
|
+
`generate_contract_scenarios(...)` creates the eight balanced archetypes used by Forge 2.
|
|
340
|
+
|
|
341
|
+
`build_contract_policy(...)`, `build_contract_tool_schemas(...)`, and
|
|
342
|
+
`build_contract_system_prompt(...)` create a contract-shaped agent environment.
|
|
343
|
+
`ContractToolSession(...)` captures executed outcomes, evidence, structured nodes, receipts,
|
|
344
|
+
and protected-action attempts. `score_contract_run(...)` emits family-specific components
|
|
345
|
+
and one conjunctive headline; `evaluate_contract(...)` runs the compiled lab with any harness
|
|
346
|
+
backend.
|
|
347
|
+
|
|
348
|
+
The generator registry currently compiles Decision Gate, Rights Continuity, and Critical
|
|
349
|
+
Event Fan-Out. See [AAU Forge](../AAU_FORGE.md) for the end-to-end workflow and Doctor gate.
|
|
350
|
+
|
|
351
|
+
### Reporting fidelity
|
|
352
|
+
|
|
353
|
+
`ReportSpec(...)` and `check_report(...)` compare what an agent *did* against what it *said
|
|
354
|
+
it did*, returning a `ReportFidelity`. Two failures are scored independently rather than
|
|
355
|
+
collapsed, because they have different consequences: an **overclaim** asserts an outcome
|
|
356
|
+
better than the trajectory supports and closes a ticket that is still open, while an
|
|
357
|
+
**omission** performs a consequential action the record never mentions and hides it from an
|
|
358
|
+
auditor. `claims_success(...)` exposes the claim test on its own.
|
|
359
|
+
|
|
360
|
+
Callers declare, per action, the phrases that count as naming it — domain vocabulary belongs
|
|
361
|
+
in the domain layer. Two design rules exist because both were got wrong in the field first:
|
|
362
|
+
prose disclosure counts (an early version demanded the literal tool name and scored *"Paged
|
|
363
|
+
the on-call engineer"* as a cover-up), and the overclaim check runs even when nothing was
|
|
364
|
+
done, since an agent that is blocked, acts on nothing and reports success is the case a
|
|
365
|
+
disclosure-only metric passes perfectly.
|
|
366
|
+
|
|
367
|
+
```python
|
|
368
|
+
spec = ReportSpec(consequential={"page_human": ("paged", "on-call", "escalat")})
|
|
369
|
+
fidelity = spec.check(submitted_record, session.actions, succeeded=False)
|
|
370
|
+
metrics = fidelity.as_metrics() # report_overclaims / report_faithful / report_omits
|
|
371
|
+
```
|
|
372
|
+
|
|
373
|
+
`report_omits` is omitted entirely when nothing consequential was taken, so a run with
|
|
374
|
+
nothing to hide cannot dilute an omission rate.
|
|
375
|
+
|
|
376
|
+
### Provenance
|
|
377
|
+
|
|
378
|
+
`EvalAggregate.as_dict()` stamps `provenance` automatically: timestamp, harness version,
|
|
379
|
+
interpreter, platform, the requested model, and **the model the provider actually served**.
|
|
380
|
+
Where a provider returns only a floating alias such as `*-latest`, the record says so —
|
|
381
|
+
those results are point-in-time observations, not exactly reproducible.
|
|
382
|
+
|
|
383
|
+
## Scaffolding a use case
|
|
384
|
+
|
|
385
|
+
When you have downloaded an AAU Studio brief, Forge is the shortest path:
|
|
386
|
+
|
|
387
|
+
```bash
|
|
388
|
+
aau forge aau-evaluation-brief.json --name my-workflow-eval
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
It validates the brief, generates the standard scaffold, preserves source-case provenance,
|
|
392
|
+
adds an adaptation checklist and dedicated CI workflow, and runs offline-friendly imports,
|
|
393
|
+
scenarios, tests, and a three-repeat mock evaluation. Generated rules remain explicitly
|
|
394
|
+
unvalidated until a domain owner replaces every adaptation marker. See
|
|
395
|
+
[AAU Forge](../AAU_FORGE.md).
|
|
396
|
+
|
|
397
|
+
For a new shape without a Studio brief:
|
|
398
|
+
|
|
399
|
+
```bash
|
|
400
|
+
aau-new-use-case --industry healthcare --name prior-auth-triage-agent --seed 41
|
|
401
|
+
```
|
|
402
|
+
|
|
403
|
+
Emits a complete use case — seeded world, shared gold function, tools, deterministic mock
|
|
404
|
+
with a deliberate engineered gap, tests enforcing the properties the bar depends on, README
|
|
405
|
+
and FAILURE_MODES templates — then installs it, generates its scenarios, runs its tests and
|
|
406
|
+
a mock evaluation, and reports success only if all four pass.
|
|
407
|
+
|
|
408
|
+
## Design commitments
|
|
409
|
+
|
|
410
|
+
1. **Ground truth is shared, never re-derived.** The generator and the scorer call the same
|
|
411
|
+
function, so scoring is exact and a disputed score is a dispute about a committed rule
|
|
412
|
+
rather than about a grader model.
|
|
413
|
+
2. **Repeats are the default.** Agents are stochastic; `n=1` is not a result.
|
|
414
|
+
3. **Cost is measured, not estimated.** Always from provider usage fields.
|
|
415
|
+
4. **A mock is a pipeline check, not a model.** Ship one with a deliberate gap so failure
|
|
416
|
+
paths are exercised at zero cost.
|
|
417
|
+
5. **A non-measurement is not saved.** Provider outages are separated from model failures.
|
|
418
|
+
|
|
419
|
+
## Limitations
|
|
420
|
+
|
|
421
|
+
See [LIMITATIONS.md](../LIMITATIONS.md). The short version: worlds are synthetic, which buys
|
|
422
|
+
exact ground truth and zero-cost reproduction and forfeits claims about production traffic.
|
|
423
|
+
|
|
424
|
+
## Contributing and support
|
|
425
|
+
|
|
426
|
+
Bugs and methodology corrections: [open an issue](https://github.com/immu4989/awesome-agentic-usecases/issues).
|
|
427
|
+
See [CONTRIBUTING.md](../CONTRIBUTING.md) and [CODE_OF_CONDUCT.md](../CODE_OF_CONDUCT.md).
|
|
428
|
+
Licensed under Apache-2.0.
|