agent-loss-map 0.18.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (50) hide show
  1. agent_loss_map-0.18.0/CONTRIBUTING.md +173 -0
  2. agent_loss_map-0.18.0/FINDINGS.md +381 -0
  3. agent_loss_map-0.18.0/LICENSE +21 -0
  4. agent_loss_map-0.18.0/MANIFEST.in +8 -0
  5. agent_loss_map-0.18.0/PKG-INFO +283 -0
  6. agent_loss_map-0.18.0/README.md +254 -0
  7. agent_loss_map-0.18.0/agent_loss_map/__init__.py +47 -0
  8. agent_loss_map-0.18.0/agent_loss_map/adapters.py +272 -0
  9. agent_loss_map-0.18.0/agent_loss_map/cli.py +134 -0
  10. agent_loss_map-0.18.0/agent_loss_map/corpus.py +121 -0
  11. agent_loss_map-0.18.0/agent_loss_map/extra_corpus.py +218 -0
  12. agent_loss_map-0.18.0/agent_loss_map/impl_audit.py +325 -0
  13. agent_loss_map-0.18.0/agent_loss_map/report.py +188 -0
  14. agent_loss_map-0.18.0/agent_loss_map/schemas/PROVENANCE.json +14 -0
  15. agent_loss_map-0.18.0/agent_loss_map/schemas/agent-descriptor.schema.json +64 -0
  16. agent_loss_map-0.18.0/agent_loss_map/schemas/capability.schema.json +45 -0
  17. agent_loss_map-0.18.0/agent_loss_map/schemas/envelope.schema.json +159 -0
  18. agent_loss_map-0.18.0/agent_loss_map/schemas/error.schema.json +46 -0
  19. agent_loss_map-0.18.0/agent_loss_map/schemas/handshake.schema.json +79 -0
  20. agent_loss_map-0.18.0/agent_loss_map/schemas/registry.schema.json +86 -0
  21. agent_loss_map-0.18.0/agent_loss_map/schemas.py +44 -0
  22. agent_loss_map-0.18.0/agent_loss_map/spec_audit.py +307 -0
  23. agent_loss_map-0.18.0/agent_loss_map/types.py +94 -0
  24. agent_loss_map-0.18.0/agent_loss_map.egg-info/PKG-INFO +283 -0
  25. agent_loss_map-0.18.0/agent_loss_map.egg-info/SOURCES.txt +48 -0
  26. agent_loss_map-0.18.0/agent_loss_map.egg-info/dependency_links.txt +1 -0
  27. agent_loss_map-0.18.0/agent_loss_map.egg-info/entry_points.txt +2 -0
  28. agent_loss_map-0.18.0/agent_loss_map.egg-info/requires.txt +6 -0
  29. agent_loss_map-0.18.0/agent_loss_map.egg-info/top_level.txt +1 -0
  30. agent_loss_map-0.18.0/docs/LAUNCH.md +193 -0
  31. agent_loss_map-0.18.0/docs/badge.json +6 -0
  32. agent_loss_map-0.18.0/docs/sample-report.json +139 -0
  33. agent_loss_map-0.18.0/docs/sample-report.md +71 -0
  34. agent_loss_map-0.18.0/docs/sample-report.txt +109 -0
  35. agent_loss_map-0.18.0/examples/adoraads-beauty.capture.json +472 -0
  36. agent_loss_map-0.18.0/examples/adoraads-beauty.entry.json +472 -0
  37. agent_loss_map-0.18.0/examples/context7.capture.json +77 -0
  38. agent_loss_map-0.18.0/examples/context7.entry.json +70 -0
  39. agent_loss_map-0.18.0/examples/deepwiki-mcp.capture.json +125 -0
  40. agent_loss_map-0.18.0/examples/deepwiki-mcp.entry.json +111 -0
  41. agent_loss_map-0.18.0/pyproject.toml +54 -0
  42. agent_loss_map-0.18.0/scripts/capture_mcp.py +218 -0
  43. agent_loss_map-0.18.0/scripts/refresh_schemas.py +140 -0
  44. agent_loss_map-0.18.0/scripts/render_report_svg.py +165 -0
  45. agent_loss_map-0.18.0/scripts/verify_corpus_fidelity.py +131 -0
  46. agent_loss_map-0.18.0/setup.cfg +4 -0
  47. agent_loss_map-0.18.0/tests/test_capture_mcp.py +122 -0
  48. agent_loss_map-0.18.0/tests/test_corpus_fidelity.py +118 -0
  49. agent_loss_map-0.18.0/tests/test_extra_corpus.py +213 -0
  50. agent_loss_map-0.18.0/tests/test_interop.py +759 -0
@@ -0,0 +1,173 @@
1
+ # Contributing
2
+
3
+ The value of this harness is coverage. Every framework you add is a framework
4
+ whose interop loss is now measured rather than assumed, and your name is on it.
5
+ Adding one should take about twenty minutes and should not require understanding
6
+ the audit modules.
7
+
8
+ ## Measuring a format, without changing anything
9
+
10
+ A corpus entry is pure data, and the checks are driven entirely off its fields,
11
+ so you can measure a format with a JSON file and no Python at all:
12
+
13
+ ```bash
14
+ cat > my-format.json <<'JSON'
15
+ {
16
+ "framework": "my-framework",
17
+ "agent_id": "my_agent",
18
+ "description": "What this is.",
19
+ "confidence": "documented",
20
+ "provenance": "Where you read the shape, and when",
21
+ "capabilities": [
22
+ { "name": "web.search", "description": "Search the web.",
23
+ "parameters": { "type": "object", "properties": { "query": { "type": "string" } } } }
24
+ ]
25
+ }
26
+ JSON
27
+
28
+ agent-loss-map --corpus my-format.json
29
+ ```
30
+
31
+ `--corpus` also accepts a directory of files, and is repeatable. Your format gets
32
+ its own section in the report, computed by exactly the same checks as the bundled
33
+ entries.
34
+
35
+ Two things the loader refuses rather than guesses:
36
+
37
+ - **`confidence` and `provenance` are required.** A finding with no declared
38
+ provenance looks as trustworthy as a real observation, which is the one thing
39
+ the tier exists to prevent. A misspelt tier is an error, not a downgrade to
40
+ `inferred`.
41
+ - **A typo in a field name is an error.** `capabilties` would otherwise be
42
+ ignored and the entry would quietly under-report.
43
+
44
+ `confidence` accepts short forms (`observed`, `documented`, `inferred`) and is
45
+ normalised to the canonical tier in the report. A bundled framework cannot be
46
+ overridden by a file; those entries are pinned to the upstream commit that was
47
+ read.
48
+
49
+ ## Adding a framework to the corpus
50
+
51
+ This is the version that goes in the package, for a format worth keeping. Open
52
+ `agent_loss_map/corpus.py` and append a `ForeignAgent` to `CORPUS`:
53
+
54
+ ```python
55
+ MY_FRAMEWORK = ForeignAgent(
56
+ framework="my-framework",
57
+ agent_id="my_agent", # must match UACP's ^[a-zA-Z0-9_]{1,128}$
58
+ description="What this agent is for.",
59
+ capabilities=[
60
+ {
61
+ "name": "web.search",
62
+ "description": "Search the web.",
63
+ "parameters": {
64
+ "type": "object",
65
+ "properties": {"query": {"type": "string"}},
66
+ "required": ["query"],
67
+ },
68
+ "outputSchema": {"type": "object", "properties": {}},
69
+ }
70
+ ],
71
+ prose_claims=["What the docs promise a human reader."],
72
+ is_stateful=False, # does a call see prior calls?
73
+ streams_tokens=False, # per-token chunks, or only results?
74
+ supports_multimodal=False, # non-string content parts?
75
+ provenance="https://... — read 2026-09-27",
76
+ confidence="documented-api",
77
+ source_shape="What kind of thing the fields above actually came from.",
78
+ )
79
+ ```
80
+
81
+ Then run it:
82
+
83
+ ```bash
84
+ agent-loss-map
85
+ ```
86
+
87
+ Your framework gets its own `PART 2` section in the report. You did not have to
88
+ write a single check.
89
+
90
+ ### The one rule
91
+
92
+ **Do not invent a serialization.** If your framework has no documented portable
93
+ form for an agent definition, say so in `source_shape` and model the
94
+ constructor or API surface instead. The AutoGen entry does exactly this and is
95
+ therefore the weakest evidence in the corpus; a fabricated schema would be
96
+ stronger-looking and worthless.
97
+
98
+ Every finding is derived from what the source actually declares. That is the
99
+ whole point of the tool, and an entry that guesses undermines every number in
100
+ the report next to it.
101
+
102
+ ### Pick your confidence tier honestly
103
+
104
+ | Tier | Means |
105
+ |---|---|
106
+ | `observed-serialization` | You read a real serialized artifact |
107
+ | `documented-api` | Taken from official documentation of the public API |
108
+ | `inferred` | Your modelling choice, not a documented shape |
109
+
110
+ The CLI prints the tier for every entry, and the report says which findings rest
111
+ on weaker evidence. If yours is `inferred`, the findings are still useful, but
112
+ say so.
113
+
114
+ ### Choosing capability names
115
+
116
+ Two things produce findings before your framework has even been tested, and both
117
+ are worth seeing rather than papering over:
118
+
119
+ - `name` must match `^[a-zA-Z0-9_.]{1,128}$` — use the source's real name. A
120
+ `SCHEMA_VIOLATION` here is a real result, not a harness bug.
121
+ - A name with no dot is a `LOSSY` finding. That is the point.
122
+
123
+ ## Adding a new check
124
+
125
+ Checks live in `adapters.py` (source → UACP), `spec_audit.py` (schema vs the
126
+ prose) or `impl_audit.py` (schema vs the reference implementations).
127
+
128
+ Two obligations come with every check:
129
+
130
+ 1. **A meta-test that proves the check fires.** Mutate a copy of the input so
131
+ the defect is absent, and assert the check goes quiet. A check that silently
132
+ stopped working is worse than no check, because the report still looks
133
+ healthy. The existing `test_meta_*` tests are the pattern to copy.
134
+
135
+ 2. **Evidence on every finding.** `test_every_divergence_carries_evidence`
136
+ enforces a non-empty `detail` and `evidence`. A reader must be able to check
137
+ your work.
138
+
139
+ ### Do not predict what you can observe
140
+
141
+ `SCHEMA_VIOLATION` has exactly one source: `validate_descriptor()`, which runs
142
+ the real JSON Schema validator over the artifact the adapter produced. Do not
143
+ also check a schema pattern with a hand-copied regex. A copied pattern is a
144
+ second source of truth that can drift from the schema it was copied from, and CI
145
+ only pins the schemas. It also counted every such defect twice before this was
146
+ fixed.
147
+
148
+ ### One field, one finding
149
+
150
+ Report a divergent field once, however many witnesses it has. If both reference
151
+ implementations disagree the same way about a field, that is one defect with two
152
+ witnesses: cite both in `evidence` rather than emitting two findings.
153
+ `duplicate_findings()` checks this over the whole report and the suite asserts it
154
+ holds.
155
+
156
+ ## Running the tests
157
+
158
+ ```bash
159
+ pip install -e ".[dev]"
160
+ pytest -q
161
+ ```
162
+
163
+ ## Regenerating the report artifact
164
+
165
+ `docs/sample-report.txt` is a captured run, committed so a reader can see the
166
+ output without installing. After changing a check:
167
+
168
+ ```bash
169
+ agent-loss-map > docs/sample-report.txt
170
+ ```
171
+
172
+ CI re-runs the probe weekly and opens an issue if the corpus or the schemas
173
+ have moved, so the published counts cannot go stale unnoticed.
@@ -0,0 +1,381 @@
1
+ # UACP interop conformance harness
2
+
3
+ Finds where a UACP agent definition and framework-native agent definitions
4
+ actually diverge, and whether UACP's own JSON Schemas agree with UACP's own
5
+ normative text.
6
+
7
+ ## Why this exists
8
+
9
+ UACP's `CONFORMANCE.md` defines L0 to L3 as a feature checklist: does your
10
+ implementation have a stdio transport, does it do rate limiting, does it declare
11
+ a `protocolVersion`. Passing that proves an implementation matches its author's
12
+ reading of the spec. It says nothing about whether two implementations built by
13
+ different people interoperate, which is the premise the whole project rests on.
14
+
15
+ This harness tests the three things the checklist cannot:
16
+
17
+ 1. **Schema versus prose.** Build a message the spec *requires* an
18
+ implementation to accept, then validate it against the shipped schema and
19
+ report the rejection.
20
+ 2. **Schema versus the implementations.** Check that each reference
21
+ implementation can actually represent what the schema declares, and that the
22
+ two implementations model the same envelope. This is the class that found the
23
+ silent credential downgrade.
24
+ 3. **Cross-format divergence, in both directions.** Map a foreign agent
25
+ description into a UACP descriptor, validate the artifact actually produced,
26
+ and separately test whether a UACP capability survives being expressed in a
27
+ framework's native tool format.
28
+
29
+ ## Method
30
+
31
+ Losses are computed, not asserted. `SCHEMA_VIOLATION` findings are observed by
32
+ running the real validator over the artifact the adapter produced. Divergences
33
+ in the reverse direction are found by walking the source JSON Schema and
34
+ collecting keywords outside the target format's supported subset.
35
+
36
+ Every corpus entry declares `confidence`:
37
+
38
+ - `observed-serialization`: a real serialized artifact was read
39
+ - `documented-api`: taken from official documentation of the public API
40
+ - `inferred`: our modelling choice, not a documented shape
41
+
42
+ Findings derived from weaker entries are weaker evidence, and the CLI says so.
43
+
44
+ ## Findings against UACP `main`
45
+
46
+ Against `wippa-uacp@92093793`: **4 schema conflicts, 11 cross-format findings
47
+ (2 blocker, 7 major, 6 minor; 15 in total).**
48
+
49
+ A captured run of exactly these numbers is committed at
50
+ [`docs/sample-report.txt`](docs/sample-report.txt), and CI re-runs the probe
51
+ weekly and opens an issue if the result stops matching it.
52
+
53
+ **UACP is at zero blockers.** Its schemas agree with its normative text, and the
54
+ TypeScript and Python reference implementations model the same envelope. Both
55
+ remaining blockers are loss measured in a *foreign* format:
56
+
57
+ | Blocker | Whose defect |
58
+ |---|---|
59
+ | `csv.analyze.inputSchema.const` | A UACP capability using a JSON Schema keyword an OpenAI function-calling block cannot express |
60
+ | `csv.analyze.inputSchema.oneOf` | As above |
61
+
62
+ **A third blocker here used to be ours.** `$['id']` was a `SCHEMA_VIOLATION`
63
+ against the `openai-function-calling` corpus entry, whose agent id `tool-caller`
64
+ did not match UACP's `^[a-zA-Z0-9_]{1,128}$`. The corpus entry was right and the
65
+ protocol was wrong: the same measurement against live MCP servers that rejected
66
+ hyphens in *capability* names also rejected them in *agent ids*, and two of six
67
+ servers announce a hyphenated `serverInfo.name`. `wippa-uacp` widened the pattern
68
+ to `^[a-zA-Z0-9_\-]{1,128}$` in PR #6. Re-vendoring the schema is what cleared the
69
+ finding — until then it kept reporting as live, which is the failure mode this
70
+ document's provenance section warns about.
71
+
72
+ The last two majors against UACP itself are `capability.async`, whose schema
73
+ default of `true` means a plain request/response consumer waits for a stream
74
+ that never arrives, and the `schemas.$id` / `$ref` finding described below.
75
+
76
+ ## Corrected counts
77
+
78
+ Earlier revisions of this document reported 9 and 13, with 6 blockers and 12
79
+ majors. Those counts were wrong: two defects were being emitted twice. A
80
+ hyphenated agent id produced both a *predicted* `SCHEMA_VIOLATION` from a
81
+ hand-copied regex and an *observed* one from the real validator, and
82
+ `metadata.flowControl` was reported once per reference implementation rather
83
+ than once with both cited as evidence. Both are fixed; `duplicate_findings()`
84
+ now asserts that no two findings in a report share a `(code, field)` pair, and
85
+ the suite includes a guard proving that check can see a duplicate when one is
86
+ injected.
87
+
88
+ ## Fixed upstream
89
+
90
+ The harness was written against its own protocol, so its early reports were
91
+ mostly defects in UACP. Each is fixed upstream and retained here as a
92
+ regression guard, so a regression re-opens it rather than passing silently.
93
+
94
+ | Was | Defect | Fixed in |
95
+ |---|---|---|
96
+ | 2 blockers | `metadata.auth` in the TypeScript model, absent from Python. `from_dict` built metadata from a fixed field list without it, so an authenticated message decoded into a credential-free object with no error raised | `wippa-uacp@b3fe7a1f` |
97
+ | 3 majors | `capability` in the schema's global `required` list, while the spec's message table gives `heartbeat`, `register`, `deregister` and `discover` no target capability. A conforming implementation had to invent a placeholder, and both halves sent the literal `_internal` | `wippa-uacp@b3fe7a1f` |
98
+ | 1 major | `flowControl` documented in SPEC.md under "Flow control" and modelled by both implementations, declared by no schema, so a message carrying it validated against nothing | `wippa-uacp@b3fe7a1f` |
99
+ | 1 blocker | Audit-trail credential persistence in plaintext | `wippa-uacp@e0463b44` |
100
+
101
+ One of those meta-tests had itself gone vacuous and was rewritten. The
102
+ implementation-agreement meta-test mutated the Python model by *adding* `auth`,
103
+ which was correct while the defect was live. Once both halves carried `auth`,
104
+ the mutation became a no-op and the test could no longer fail — a check that had
105
+ quietly stopped working, which is the exact failure this harness exists to
106
+ catch. It now removes the field instead, so reintroducing the silent credential
107
+ downgrade anywhere in the observed field sets fails the suite.
108
+
109
+ ## Resolved in `main` at `7bcc2594`: one decision, three contradictions
110
+
111
+ Section 12 promised that unknown fields "in the envelope or metadata" were
112
+ silently ignored. The schema never did that for the envelope, which sets
113
+ `"additionalProperties": false` at the top level while `$defs/metadata` sets it
114
+ to `true`. Three defects shared that root, and resolving them separately invites
115
+ inconsistent answers, so they were resolved as one decision about what the
116
+ envelope is allowed to mean over time.
117
+
118
+ The schema split was already correct, and the existing extension namespace in
119
+ section 12 already routed growth through `metadata` with `vendor.extension_name`
120
+ keys. So the prose was wrong, not the schema.
121
+
122
+ **Forward compatibility.** Narrowed to `metadata`, which is what the schema
123
+ encoded. The closed envelope is now a stated design decision rather than an
124
+ accident, justified by the typo case: a mistyped `protocolVersion` or `traceId`
125
+ is far likelier than a legitimate new field, and ignoring it would drop routing
126
+ information.
127
+
128
+ **Version pinning.** `"const": "0.1.0"` became the MAJOR-family pattern
129
+ `"^0\.[0-9]+\.[0-9]+$"`. The exact pin made section 12's same-MAJOR
130
+ interoperability rule impossible to honour, because a 0.1.1 message was invalid
131
+ to a 0.1.0 validator. The published schema now tracks the released MAJOR line.
132
+
133
+ **Message type set.** The closed `type` enum became
134
+ `anyOf[known types, naming pattern]`. A MINOR bump may add a message type and
135
+ section 12 requires unrecognised types to be tolerated, so a closed enum made
136
+ both unreachable. This was a third instance of the same contradiction, found
137
+ only by reading section 12 against section 18 together.
138
+
139
+ **The governance cost, stated rather than implied.** No new top-level field can
140
+ ship in a MINOR bump, because an earlier validator will reject it. Growth within
141
+ a MAJOR series goes in `metadata` or as a new message type, and adding a
142
+ top-level field is a MAJOR change that must not be used to avoid adding a
143
+ metadata field. The spec also now notes that this differs from protobuf, which
144
+ preserves and round-trips unknown fields rather than rejecting them.
145
+
146
+ Verified against the validator: 0.1.0, 0.1.1 and 0.9.0 accepted; 1.0.0 and
147
+ `"0.1"` rejected; known and new dot-separated types accepted; `Not A Type`,
148
+ `agent.novel_type` and `9bad` rejected; unknown top-level field rejected;
149
+ unknown metadata field accepted; a `traceid` typo rejected.
150
+
151
+ ## Fixed in `main` at `e0463b44`
152
+
153
+ **Audit-trail credential persistence (was: blocker, `metadata.auth.token`).**
154
+ Section 15.2 said implementations MUST never log token values. Section 16.3
155
+ separately mandated writing "Full metadata JSON" to the audit trail and named no
156
+ redaction hook, and L3 required a "full message log". The reference
157
+ implementation followed 16.3 literally: `log_message` wrote
158
+ `json.dumps(message.metadata.to_dict())` straight into the registry SQLite file.
159
+ `create_redacted_log` already existed in `core/redact.py` and was never called
160
+ from the audit path. So a bearer or API key token, and any secret-looking
161
+ payload key, were persisted in plaintext, in the conformance level named
162
+ "Security".
163
+
164
+ Fixed by routing `log_message` through `create_redacted_log`, adding the
165
+ normative requirement to 16.3, 15.2 and CONFORMANCE L3, correcting the "Full
166
+ trace replay" claim to a redacted replay, and documenting the requirement on
167
+ `metadata.auth` in the schema. Ten new tests read the SQLite file directly,
168
+ bypassing any accessor-layer redaction; seven of them fail against the previous
169
+ implementation. 94 tests pass: 84 pre-existing and unchanged, plus 10 new.
170
+
171
+ The same investigation found a latent crash with no test coverage at all:
172
+ `Message` has no `__post_init__`, so a caller can hand `log_message` a plain dict
173
+ and the old code raised `AttributeError` on a diagnostic path. No existing test
174
+ ever passed metadata to `log_message`, which is why it went unnoticed.
175
+
176
+ ## New blockers, found by comparing code to schema
177
+
178
+ These are invisible from the schemas alone and needed a schema-versus-
179
+ implementation check.
180
+
181
+ **1.** The Python implementation cannot represent `metadata.auth` at all.
182
+ `envelope.schema.json` declares it with `required: ['type','token']`.
183
+ `MessageMetadata` in `python/wippa_uacp/core/types.py` has no such field, and
184
+ `Message.to_dict` and `Message.from_dict` both handle a fixed field list, so a
185
+ conforming envelope from a peer is silently degraded on decode.
186
+
187
+ **2.** The two reference implementations disagree, and the failure mode is silent
188
+ credential downgrade.** `typescript/src/core/types.ts` declares
189
+ `auth?: { type, token }`. The Python one does not. A request authenticated by a
190
+ TypeScript peer is decoded by a Python peer into metadata that has dropped the
191
+ token, with no error raised. The message then relies entirely on section 15.4
192
+ default-deny ACLs to be stopped. Silent downgrade is worse than either rejecting
193
+ the message or honouring it, so 15.2 now requires rejecting a message whose
194
+ `auth` the implementation cannot validate.
195
+
196
+ **3.** `flowControl` is implemented in both languages and declared in neither
197
+ schema.** Both `MessageMetadata` models carry it, and section 13 has a flow
198
+ control section, but `envelope.schema.json` omits it entirely. Anything using
199
+ flow control is outside the specification and will not survive validation against
200
+ a conforming peer.
201
+
202
+
203
+ ### Remaining blockers and majors
204
+
205
+ **4.** The schema set cannot be resolved offline.
206
+ Every `$id` sits on `wippa.dev`, which does not resolve. `agent-descriptor`
207
+ refs `capability.schema.json` *relatively*, which resolves against that `$id` and
208
+ therefore also becomes unreachable. A standards-compliant validator attempts a
209
+ network fetch and fails with `Unretrievable`. This harness had to build a local
210
+ `referencing.Registry` just to validate. Air-gapped builds and offline CI break.
211
+
212
+ **5.** A UACP descriptor cannot name a hyphenated MCP tool.** `capability.name`
213
+ and `agent.id` are `^[a-zA-Z0-9_.]{1,128}$`, which rejects `-`. MCP tool names
214
+ are conventionally hyphenated — `resolve-library-id`, `query-docs`,
215
+ `get-library-docs` — and OpenAI's own function-name grammar
216
+ (`^[a-zA-Z0-9_-]{1,64}$`) accepts all of them. A conforming UACP descriptor has
217
+ to rename the tool, and a renamed tool is a different tool to any client that
218
+ refers to it by name.
219
+
220
+ Measured, not argued: `examples/context7.entry.json` is a live capture of
221
+ Context7 4.1.1 and produces two blockers under released UACP, one per tool.
222
+ DeepWiki's tools are underscored and pass, so this is a property of the name and
223
+ not of the capture path.
224
+
225
+ The two grammars are not subsets of each other, and that asymmetry runs both
226
+ ways — UACP's dot-notation namespace is *also* rejected by OpenAI, so
227
+ `csv.analyze` does not survive a conversion in the other direction either. Any
228
+ format pair with this property forces a rename somewhere, which is the general
229
+ problem this harness exists to make visible.
230
+
231
+ Fixed in `main` at `0a4b166`: `-` added to the character class, additive so no
232
+ valid name became invalid, and `agent.id` deliberately left strict since an agent
233
+ is an identity rather than a namespace. The same change corrected SPEC.md, whose
234
+ stated character set had excluded the dot its own namespace convention requires.
235
+
236
+ Re-measured against the merged spec: Context7 and DeepWiki both go to zero
237
+ blockers, and the tool names survive unrenamed, which was the whole point. The
238
+ other direction is unchanged and is not a UACP defect — `csv.analyze` still does
239
+ not survive conversion to an OpenAI function name, because OpenAI rejects dots.
240
+
241
+ ### Majors
242
+
243
+ **6.** `capability` is mandatory for messages that have none.** It is in the
244
+ global `required` list, so `heartbeat`, `register` and `discover` must all carry
245
+ one, and the pattern requires at least one character. There is no honest value
246
+ to supply. A conforming implementation must invent a placeholder.
247
+
248
+ **7.** `async` defaults to true, inverting the common case.** A capability that
249
+ omits `async` is treated as streaming or deferred, so a plain request/response
250
+ capability must opt out explicitly. A consumer trusting the default waits for a
251
+ `stream.end` that never arrives, which presents as a hang rather than a schema
252
+ error.
253
+
254
+ **8.** Agents with no machine-readable capabilities are a discovery blocker.
255
+ AutoGen's `AssistantAgent` exposes `name`, `description` and a tool list, but no
256
+ structured capability manifest with input and output schemas. A UACP registry
257
+ consumer asking "who can do `csv.analyze`" can only substring-match free text.
258
+ The answer is a guess, not a verifiable claim. This is the capability-overclaim
259
+ problem in its purest form, and it is not fixable by an adapter.
260
+
261
+ **9.** Stateful agents have no UACP representation.** AutoGen documents `run()` as
262
+ mutating internal history and being called with new messages rather than
263
+ complete history. A UACP descriptor has no lifecycle or state field, and the
264
+ envelope is stateless message passing, so two calls that look identical on the
265
+ wire mean different things on either side.
266
+
267
+ **10.** Token streaming and result streaming are different things.** AutoGen emits
268
+ per-token `ModelClientStreamingChunkEvent`s that are part of the agent's own
269
+ message trace. UACP section 13 defines `stream`/`stream.end` as partial *results*
270
+ correlated to a request by `replyTo` and sequence numbers. Mapping one onto the
271
+ other conflates the agent's reasoning trace with the result stream, so a UACP
272
+ consumer would receive internal reasoning it never asked for and could not tell
273
+ the two apart.
274
+
275
+ **11.** Heterogeneous content parts are unrepresentable.** `MultiModalMessage`
276
+ accepts `content=[str, Image]`. The UACP payload is untyped so a list would
277
+ validate, but the spec defines no modality, media type or part ordering, and an
278
+ in-memory `Image` has no wire form. Nothing distinguishes it from an ordinary
279
+ array payload.
280
+
281
+ **12.** A missing output contract has to be fabricated.** UACP requires
282
+ `outputSchema` on every capability. Neither AutoGen's `FunctionTool` nor an
283
+ OpenAI function declaration carries one, so the adapter must invent a permissive
284
+ schema and consumers will believe a guarantee the source never made.
285
+
286
+ **13.** Rich JSON Schema cannot reach an OpenAI-style tool consumer.** A UACP
287
+ capability using `oneOf` or `const` has no representation in an OpenAI
288
+ function-calling `parameters` block, which supports a small fixed subset. The
289
+ capability cannot be exposed without dropping the constraint or flattening it
290
+ into prose. This is the only genuinely *bidirectional* blocker found, and it
291
+ lives in the UACP-to-foreign direction.
292
+
293
+ ### Minors
294
+
295
+ **14.** The OpenAI `strict` flag has no JSON Schema equivalent, so the guarantee
296
+ is dropped in either direction and cannot be re-expressed as a UACP consumer
297
+ check.
298
+
299
+ **15.** Capability names without a namespace prefix (`web_search_func`,
300
+ `get_weather`) must be renamed, so namespace-scoped registry queries will not
301
+ match the original name.
302
+
303
+ **A server whose own name contains a hyphen cannot describe itself.**
304
+
305
+ `agent-descriptor.id` is `^[a-zA-Z0-9_]{1,128}$`, and an MCP server announces itself
306
+ as `serverInfo.name`. Taking that as the agent id is the obvious mapping, and it
307
+ fails for 2 of the 8 live servers measured here — `inside-ads` and
308
+ `adoraads-beauty-gateway` each produce a blocker on their own name.
309
+
310
+ The distinction that matters is that this is *not* the finding 5 defect and not
311
+ fixed by it. There, a tool name had to be renamed, and the fix lets the real name
312
+ through. Here the strictness is deliberate and correct — an agent is an identity,
313
+ not a namespace — and the loss is unavoidable in that direction: the server
314
+ cannot rename itself, since the name is what every existing client asks for.
315
+
316
+ Neither reference implementation would have caught it. Both accept a hyphenated
317
+ `AgentDescriptor.id`; only a schema validator rejects one, so a conformant
318
+ validator and a working implementation disagree about the same object.
319
+
320
+ Captured as `examples/adoraads-beauty.*`, and the capture script deliberately
321
+ does not sanitise the name — doing so would make the probe pass and hide the
322
+ loss, which is the opposite of what it is for.
323
+
324
+ **No cross-pipeline cost ceiling, and a loop between conforming agents is silent.**
325
+
326
+ Rate limiting is specified as a MAY and neither `Bus` implements it. More
327
+ consequentially, nothing bounds a pipeline as a whole: a per-agent ceiling says
328
+ nothing about the sum across every agent on the bus.
329
+
330
+ Measured, not argued. Two agents, each individually conforming and each invoking
331
+ the other's declared capability, ran at roughly 20,000 round trips per second.
332
+ The only bound was the host language's recursion limit — tightening it from 1000
333
+ to 200 scaled the round trips from 124 to 24, proportionally — and the initiating
334
+ caller received an ordinary `response` with nothing indicating the pipeline had
335
+ been cut short.
336
+
337
+ The failure is silent, which is what makes it worth recording. It has the same
338
+ shape as finding 7, where a request/response capability was documented as
339
+ streaming and a consumer observed a hang rather than an error: in both cases the
340
+ caller cannot distinguish a completed run from a truncated one.
341
+
342
+ This cannot be closed by writing better code, because there is no claim to have
343
+ failed — the spec is silent. `wippa-uacp` §15.8 records it as unspecified, and
344
+ the harness reports it as a major `UNVERIFIABLE_CLAIM` that stays reported even
345
+ against a perfect implementation.
346
+
347
+ ## Running it
348
+
349
+ ```bash
350
+ uvx agent-loss-map # human-readable
351
+ uvx agent-loss-map -- --json # machine-readable
352
+ uvx agent-loss-map -- --fail-on-blocker # CI gate
353
+ ```
354
+
355
+ From a checkout, `pip install -e .` and then `agent-loss-map`. Tests need the dev
356
+ extra: `pip install -e ".[dev]" && pytest -q`.
357
+
358
+ ## Adding your own framework
359
+
360
+ Append a `ForeignAgent` to `agent_loss_map/corpus.py` and re-run. You do not have
361
+ to write a check: the divergences for your framework are computed from the
362
+ fields you declare. See [CONTRIBUTING.md](CONTRIBUTING.md), which also covers the
363
+ provenance tiers and the one rule — do not invent a serialization you cannot
364
+ source.
365
+
366
+ ## What would strengthen this
367
+
368
+ - Observed serializations rather than constructor surfaces. The AutoGen entry is
369
+ the weakest evidence in the corpus for exactly this reason, and AutoGen does
370
+ publish a component-serialization mechanism that was not inspected here.
371
+ - The implementation audit records observed field sets as data with provenance,
372
+ rather than parsing the dataclass and TypeScript interface at runtime. Parsing
373
+ them with regex would be more fragile than the value it adds. Re-reading them
374
+ is a manual step, so they can drift.
375
+ - Decisions on the two remaining normative conflicts (`additionalProperties` and
376
+ the `const` version pin), which are design questions rather than defects and
377
+ are deliberately left to the spec owner.
378
+ - LangGraph, CrewAI and OpenAI Assistants entries, which are the formats most
379
+ likely to be lossier than the two here.
380
+ - A behavioural half: run two implementations against each other and compare
381
+ actual message sequences, which is what "interoperate" ultimately means.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 wippa-studios
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,8 @@
1
+ include LICENSE
2
+ include README.md
3
+ include CONTRIBUTING.md
4
+ include FINDINGS.md
5
+ recursive-include agent_loss_map/schemas *.json
6
+ recursive-include docs *.txt *.json *.md
7
+ recursive-include scripts *.py
8
+ recursive-include examples *.json