uacp-interop 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. uacp_interop-0.1.0/CONTRIBUTING.md +131 -0
  2. uacp_interop-0.1.0/FINDINGS.md +265 -0
  3. uacp_interop-0.1.0/LICENSE +21 -0
  4. uacp_interop-0.1.0/MANIFEST.in +6 -0
  5. uacp_interop-0.1.0/PKG-INFO +144 -0
  6. uacp_interop-0.1.0/README.md +115 -0
  7. uacp_interop-0.1.0/docs/sample-report.json +174 -0
  8. uacp_interop-0.1.0/docs/sample-report.txt +117 -0
  9. uacp_interop-0.1.0/pyproject.toml +43 -0
  10. uacp_interop-0.1.0/setup.cfg +4 -0
  11. uacp_interop-0.1.0/tests/test_interop.py +355 -0
  12. uacp_interop-0.1.0/uacp_interop/__init__.py +11 -0
  13. uacp_interop-0.1.0/uacp_interop/adapters.py +224 -0
  14. uacp_interop-0.1.0/uacp_interop/cli.py +107 -0
  15. uacp_interop-0.1.0/uacp_interop/corpus.py +121 -0
  16. uacp_interop-0.1.0/uacp_interop/impl_audit.py +157 -0
  17. uacp_interop-0.1.0/uacp_interop/schemas/PROVENANCE.json +14 -0
  18. uacp_interop-0.1.0/uacp_interop/schemas/agent-descriptor.schema.json +53 -0
  19. uacp_interop-0.1.0/uacp_interop/schemas/capability.schema.json +40 -0
  20. uacp_interop-0.1.0/uacp_interop/schemas/envelope.schema.json +131 -0
  21. uacp_interop-0.1.0/uacp_interop/schemas/error.schema.json +46 -0
  22. uacp_interop-0.1.0/uacp_interop/schemas/handshake.schema.json +79 -0
  23. uacp_interop-0.1.0/uacp_interop/schemas/registry.schema.json +86 -0
  24. uacp_interop-0.1.0/uacp_interop/schemas.py +44 -0
  25. uacp_interop-0.1.0/uacp_interop/spec_audit.py +297 -0
  26. uacp_interop-0.1.0/uacp_interop/types.py +89 -0
  27. uacp_interop-0.1.0/uacp_interop.egg-info/PKG-INFO +144 -0
  28. uacp_interop-0.1.0/uacp_interop.egg-info/SOURCES.txt +30 -0
  29. uacp_interop-0.1.0/uacp_interop.egg-info/dependency_links.txt +1 -0
  30. uacp_interop-0.1.0/uacp_interop.egg-info/entry_points.txt +2 -0
  31. uacp_interop-0.1.0/uacp_interop.egg-info/requires.txt +6 -0
  32. uacp_interop-0.1.0/uacp_interop.egg-info/top_level.txt +1 -0
@@ -0,0 +1,131 @@
1
+ # Contributing
2
+
3
+ The value of this harness is coverage. Every framework you add is a framework
4
+ whose interop loss is now measured rather than assumed, and your name is on it.
5
+ Adding one should take about twenty minutes and should not require understanding
6
+ the audit modules.
7
+
8
+ ## Adding a framework to the corpus
9
+
10
+ Open `uacp_interop/corpus.py` and append a `ForeignAgent` to `CORPUS`:
11
+
12
+ ```python
13
+ MY_FRAMEWORK = ForeignAgent(
14
+ framework="my-framework",
15
+ agent_id="my_agent", # must match UACP's ^[a-zA-Z0-9_]{1,128}$
16
+ description="What this agent is for.",
17
+ capabilities=[
18
+ {
19
+ "name": "web.search",
20
+ "description": "Search the web.",
21
+ "parameters": {
22
+ "type": "object",
23
+ "properties": {"query": {"type": "string"}},
24
+ "required": ["query"],
25
+ },
26
+ "outputSchema": {"type": "object", "properties": {}},
27
+ }
28
+ ],
29
+ prose_claims=["What the docs promise a human reader."],
30
+ is_stateful=False, # does a call see prior calls?
31
+ streams_tokens=False, # per-token chunks, or only results?
32
+ supports_multimodal=False, # non-string content parts?
33
+ provenance="https://... — read 2026-09-27",
34
+ confidence="documented-api",
35
+ source_shape="What kind of thing the fields above actually came from.",
36
+ )
37
+ ```
38
+
39
+ Then run it:
40
+
41
+ ```bash
42
+ uacp-interop
43
+ ```
44
+
45
+ Your framework gets its own `PART 2` section in the report. You did not have to
46
+ write a single check.
47
+
48
+ ### The one rule
49
+
50
+ **Do not invent a serialization.** If your framework has no documented portable
51
+ form for an agent definition, say so in `source_shape` and model the
52
+ constructor or API surface instead. The AutoGen entry does exactly this and is
53
+ therefore the weakest evidence in the corpus; a fabricated schema would be
54
+ stronger-looking and worthless.
55
+
56
+ Every finding is derived from what the source actually declares. That is the
57
+ whole point of the tool, and an entry that guesses undermines every number in
58
+ the report next to it.
59
+
60
+ ### Pick your confidence tier honestly
61
+
62
+ | Tier | Means |
63
+ |---|---|
64
+ | `observed-serialization` | You read a real serialized artifact |
65
+ | `documented-api` | Taken from official documentation of the public API |
66
+ | `inferred` | Your modelling choice, not a documented shape |
67
+
68
+ The CLI prints the tier for every entry, and the report says which findings rest
69
+ on weaker evidence. If yours is `inferred`, the findings are still useful, but
70
+ say so.
71
+
72
+ ### Choosing capability names
73
+
74
+ Two things produce findings before your framework has even been tested, and both
75
+ are worth seeing rather than papering over:
76
+
77
+ - `name` must match `^[a-zA-Z0-9_.]{1,128}$` — use the source's real name. A
78
+ `SCHEMA_VIOLATION` here is a real result, not a harness bug.
79
+ - A name with no dot is a `LOSSY` finding. That is the point.
80
+
81
+ ## Adding a new check
82
+
83
+ Checks live in `adapters.py` (source → UACP), `spec_audit.py` (schema vs the
84
+ prose) or `impl_audit.py` (schema vs the reference implementations).
85
+
86
+ Two obligations come with every check:
87
+
88
+ 1. **A meta-test that proves the check fires.** Mutate a copy of the input so
89
+ the defect is absent, and assert the check goes quiet. A check that silently
90
+ stopped working is worse than no check, because the report still looks
91
+ healthy. The existing `test_meta_*` tests are the pattern to copy.
92
+
93
+ 2. **Evidence on every finding.** `test_every_divergence_carries_evidence`
94
+ enforces a non-empty `detail` and `evidence`. A reader must be able to check
95
+ your work.
96
+
97
+ ### Do not predict what you can observe
98
+
99
+ `SCHEMA_VIOLATION` has exactly one source: `validate_descriptor()`, which runs
100
+ the real JSON Schema validator over the artifact the adapter produced. Do not
101
+ also check a schema pattern with a hand-copied regex. A copied pattern is a
102
+ second source of truth that can drift from the schema it was copied from, and CI
103
+ only pins the schemas. It also counted every such defect twice before this was
104
+ fixed.
105
+
106
+ ### One field, one finding
107
+
108
+ Report a divergent field once, however many witnesses it has. If both reference
109
+ implementations disagree the same way about a field, that is one defect with two
110
+ witnesses: cite both in `evidence` rather than emitting two findings.
111
+ `duplicate_findings()` checks this over the whole report and the suite asserts it
112
+ holds.
113
+
114
+ ## Running the tests
115
+
116
+ ```bash
117
+ pip install -e ".[dev]"
118
+ pytest -q
119
+ ```
120
+
121
+ ## Regenerating the report artifact
122
+
123
+ `docs/sample-report.txt` is a captured run, committed so a reader can see the
124
+ output without installing. After changing a check:
125
+
126
+ ```bash
127
+ uacp-interop > docs/sample-report.txt
128
+ ```
129
+
130
+ CI re-runs the probe weekly and opens an issue if the corpus or the schemas
131
+ have moved, so the published counts cannot go stale unnoticed.
@@ -0,0 +1,265 @@
1
+ # UACP interop conformance harness
2
+
3
+ Finds where a UACP agent definition and framework-native agent definitions
4
+ actually diverge, and whether UACP's own JSON Schemas agree with UACP's own
5
+ normative text.
6
+
7
+ ## Why this exists
8
+
9
+ UACP's `CONFORMANCE.md` defines L0 to L3 as a feature checklist: does your
10
+ implementation have a stdio transport, does it do rate limiting, does it declare
11
+ a `protocolVersion`. Passing that proves an implementation matches its author's
12
+ reading of the spec. It says nothing about whether two implementations built by
13
+ different people interoperate, which is the premise the whole project rests on.
14
+
15
+ This harness tests the three things the checklist cannot:
16
+
17
+ 1. **Schema versus prose.** Build a message the spec *requires* an
18
+ implementation to accept, then validate it against the shipped schema and
19
+ report the rejection.
20
+ 2. **Schema versus the implementations.** Check that each reference
21
+ implementation can actually represent what the schema declares, and that the
22
+ two implementations model the same envelope. This is the class that found the
23
+ silent credential downgrade.
24
+ 3. **Cross-format divergence, in both directions.** Map a foreign agent
25
+ description into a UACP descriptor, validate the artifact actually produced,
26
+ and separately test whether a UACP capability survives being expressed in a
27
+ framework's native tool format.
28
+
29
+ ## Method
30
+
31
+ Losses are computed, not asserted. `SCHEMA_VIOLATION` findings are observed by
32
+ running the real validator over the artifact the adapter produced. Divergences
33
+ in the reverse direction are found by walking the source JSON Schema and
34
+ collecting keywords outside the target format's supported subset.
35
+
36
+ Every corpus entry declares `confidence`:
37
+
38
+ - `observed-serialization`: a real serialized artifact was read
39
+ - `documented-api`: taken from official documentation of the public API
40
+ - `inferred`: our modelling choice, not a documented shape
41
+
42
+ Findings derived from weaker entries are weaker evidence, and the CLI says so.
43
+
44
+ ## Findings against UACP `main`
45
+
46
+ 8 schema conflicts, 12 cross-format findings (5 blocker, 11 major, 4 minor across
47
+ all parts; 20 findings in total). Four previously reported findings are now fixed
48
+ and retained only as regression guards.
49
+
50
+ A captured run of exactly these numbers is committed at
51
+ [`docs/sample-report.txt`](docs/sample-report.txt), and CI re-runs the probe
52
+ weekly and opens an issue if the result stops matching it.
53
+
54
+ Earlier revisions of this document reported 9 and 13, with 6 blockers and 12
55
+ majors. Those counts were wrong: two defects were being emitted twice. A hyphenated
56
+ agent id produced both a *predicted* `SCHEMA_VIOLATION` from a hand-copied regex
57
+ and an *observed* one from the real validator, and `metadata.flowControl` was
58
+ reported once per reference implementation rather than once with both cited as
59
+ evidence. Both are fixed; `duplicate_findings()` now asserts that no two findings
60
+ in a report share a `(code, field)` pair, and the suite includes a guard proving
61
+ that check can see a duplicate when one is injected.
62
+
63
+ ## Resolved in `main` at `7bcc2594`: one decision, three contradictions
64
+
65
+ Section 12 promised that unknown fields "in the envelope or metadata" were
66
+ silently ignored. The schema never did that for the envelope, which sets
67
+ `"additionalProperties": false` at the top level while `$defs/metadata` sets it
68
+ to `true`. Three defects shared that root, and resolving them separately invites
69
+ inconsistent answers, so they were resolved as one decision about what the
70
+ envelope is allowed to mean over time.
71
+
72
+ The schema split was already correct, and the existing extension namespace in
73
+ section 12 already routed growth through `metadata` with `vendor.extension_name`
74
+ keys. So the prose was wrong, not the schema.
75
+
76
+ **Forward compatibility.** Narrowed to `metadata`, which is what the schema
77
+ encoded. The closed envelope is now a stated design decision rather than an
78
+ accident, justified by the typo case: a mistyped `protocolVersion` or `traceId`
79
+ is far likelier than a legitimate new field, and ignoring it would drop routing
80
+ information.
81
+
82
+ **Version pinning.** `"const": "0.1.0"` became the MAJOR-family pattern
83
+ `"^0\.[0-9]+\.[0-9]+$"`. The exact pin made section 12's same-MAJOR
84
+ interoperability rule impossible to honour, because a 0.1.1 message was invalid
85
+ to a 0.1.0 validator. The published schema now tracks the released MAJOR line.
86
+
87
+ **Message type set.** The closed `type` enum became
88
+ `anyOf[known types, naming pattern]`. A MINOR bump may add a message type and
89
+ section 12 requires unrecognised types to be tolerated, so a closed enum made
90
+ both unreachable. This was a third instance of the same contradiction, found
91
+ only by reading section 12 against section 18 together.
92
+
93
+ **The governance cost, stated rather than implied.** No new top-level field can
94
+ ship in a MINOR bump, because an earlier validator will reject it. Growth within
95
+ a MAJOR series goes in `metadata` or as a new message type, and adding a
96
+ top-level field is a MAJOR change that must not be used to avoid adding a
97
+ metadata field. The spec also now notes that this differs from protobuf, which
98
+ preserves and round-trips unknown fields rather than rejecting them.
99
+
100
+ Verified against the validator: 0.1.0, 0.1.1 and 0.9.0 accepted; 1.0.0 and
101
+ `"0.1"` rejected; known and new dot-separated types accepted; `Not A Type`,
102
+ `agent.novel_type` and `9bad` rejected; unknown top-level field rejected;
103
+ unknown metadata field accepted; a `traceid` typo rejected.
104
+
105
+ ## Fixed in `main` at `e0463b44`
106
+
107
+ **Audit-trail credential persistence (was: blocker, `metadata.auth.token`).**
108
+ Section 15.2 said implementations MUST never log token values. Section 16.3
109
+ separately mandated writing "Full metadata JSON" to the audit trail and named no
110
+ redaction hook, and L3 required a "full message log". The reference
111
+ implementation followed 16.3 literally: `log_message` wrote
112
+ `json.dumps(message.metadata.to_dict())` straight into the registry SQLite file.
113
+ `create_redacted_log` already existed in `core/redact.py` and was never called
114
+ from the audit path. So a bearer or API key token, and any secret-looking
115
+ payload key, were persisted in plaintext, in the conformance level named
116
+ "Security".
117
+
118
+ Fixed by routing `log_message` through `create_redacted_log`, adding the
119
+ normative requirement to 16.3, 15.2 and CONFORMANCE L3, correcting the "Full
120
+ trace replay" claim to a redacted replay, and documenting the requirement on
121
+ `metadata.auth` in the schema. Ten new tests read the SQLite file directly,
122
+ bypassing any accessor-layer redaction; seven of them fail against the previous
123
+ implementation. 94 tests pass: 84 pre-existing and unchanged, plus 10 new.
124
+
125
+ The same investigation found a latent crash with no test coverage at all:
126
+ `Message` has no `__post_init__`, so a caller can hand `log_message` a plain dict
127
+ and the old code raised `AttributeError` on a diagnostic path. No existing test
128
+ ever passed metadata to `log_message`, which is why it went unnoticed.
129
+
130
+ ## New blockers, found by comparing code to schema
131
+
132
+ These are invisible from the schemas alone and needed a schema-versus-
133
+ implementation check.
134
+
135
+ **4. The Python implementation cannot represent `metadata.auth` at all.**
136
+ `envelope.schema.json` declares it with `required: ['type','token']`.
137
+ `MessageMetadata` in `python/wippa_uacp/core/types.py` has no such field, and
138
+ `Message.to_dict` and `Message.from_dict` both handle a fixed field list, so a
139
+ conforming envelope from a peer is silently degraded on decode.
140
+
141
+ **5. The two reference implementations disagree, and the failure mode is silent
142
+ credential downgrade.** `typescript/src/core/types.ts` declares
143
+ `auth?: { type, token }`. The Python one does not. A request authenticated by a
144
+ TypeScript peer is decoded by a Python peer into metadata that has dropped the
145
+ token, with no error raised. The message then relies entirely on section 15.4
146
+ default-deny ACLs to be stopped. Silent downgrade is worse than either rejecting
147
+ the message or honouring it, so 15.2 now requires rejecting a message whose
148
+ `auth` the implementation cannot validate.
149
+
150
+ **6. `flowControl` is implemented in both languages and declared in neither
151
+ schema.** Both `MessageMetadata` models carry it, and section 13 has a flow
152
+ control section, but `envelope.schema.json` omits it entirely. Anything using
153
+ flow control is outside the specification and will not survive validation against
154
+ a conforming peer.
155
+
156
+
157
+ ### Remaining blockers and majors
158
+
159
+ **7. The schema set cannot be resolved offline.**
160
+ Every `$id` sits on `wippa.dev`, which does not resolve. `agent-descriptor`
161
+ refs `capability.schema.json` *relatively*, which resolves against that `$id` and
162
+ therefore also becomes unreachable. A standards-compliant validator attempts a
163
+ network fetch and fails with `Unretrievable`. This harness had to build a local
164
+ `referencing.Registry` just to validate. Air-gapped builds and offline CI break.
165
+
166
+ ### Majors
167
+
168
+ **8. `capability` is mandatory for messages that have none.** It is in the
169
+ global `required` list, so `heartbeat`, `register` and `discover` must all carry
170
+ one, and the pattern requires at least one character. There is no honest value
171
+ to supply. A conforming implementation must invent a placeholder.
172
+
173
+ **9. `async` defaults to true, inverting the common case.** A capability that
174
+ omits `async` is treated as streaming or deferred, so a plain request/response
175
+ capability must opt out explicitly. A consumer trusting the default waits for a
176
+ `stream.end` that never arrives, which presents as a hang rather than a schema
177
+ error.
178
+
179
+ **10. Agents with no machine-readable capabilities are a discovery blocker.**
180
+ AutoGen's `AssistantAgent` exposes `name`, `description` and a tool list, but no
181
+ structured capability manifest with input and output schemas. A UACP registry
182
+ consumer asking "who can do `csv.analyze`" can only substring-match free text.
183
+ The answer is a guess, not a verifiable claim. This is the capability-overclaim
184
+ problem in its purest form, and it is not fixable by an adapter.
185
+
186
+ **11. Stateful agents have no UACP representation.** AutoGen documents `run()` as
187
+ mutating internal history and being called with new messages rather than
188
+ complete history. A UACP descriptor has no lifecycle or state field, and the
189
+ envelope is stateless message passing, so two calls that look identical on the
190
+ wire mean different things on either side.
191
+
192
+ **12. Token streaming and result streaming are different things.** AutoGen emits
193
+ per-token `ModelClientStreamingChunkEvent`s that are part of the agent's own
194
+ message trace. UACP section 13 defines `stream`/`stream.end` as partial *results*
195
+ correlated to a request by `replyTo` and sequence numbers. Mapping one onto the
196
+ other conflates the agent's reasoning trace with the result stream, so a UACP
197
+ consumer would receive internal reasoning it never asked for and could not tell
198
+ the two apart.
199
+
200
+ **13. Heterogeneous content parts are unrepresentable.** `MultiModalMessage`
201
+ accepts `content=[str, Image]`. The UACP payload is untyped so a list would
202
+ validate, but the spec defines no modality, media type or part ordering, and an
203
+ in-memory `Image` has no wire form. Nothing distinguishes it from an ordinary
204
+ array payload.
205
+
206
+ **14. A missing output contract has to be fabricated.** UACP requires
207
+ `outputSchema` on every capability. Neither AutoGen's `FunctionTool` nor an
208
+ OpenAI function declaration carries one, so the adapter must invent a permissive
209
+ schema and consumers will believe a guarantee the source never made.
210
+
211
+ **15. Rich JSON Schema cannot reach an OpenAI-style tool consumer.** A UACP
212
+ capability using `oneOf` or `const` has no representation in an OpenAI
213
+ function-calling `parameters` block, which supports a small fixed subset. The
214
+ capability cannot be exposed without dropping the constraint or flattening it
215
+ into prose. This is the only genuinely *bidirectional* blocker found, and it
216
+ lives in the UACP-to-foreign direction.
217
+
218
+ ### Minors
219
+
220
+ **16.** The OpenAI `strict` flag has no JSON Schema equivalent, so the guarantee
221
+ is dropped in either direction and cannot be re-expressed as a UACP consumer
222
+ check.
223
+
224
+ **17.** Capability names without a namespace prefix (`web_search_func`,
225
+ `get_weather`) must be renamed, so namespace-scoped registry queries will not
226
+ match the original name.
227
+
228
+ **18.** Hyphens are legal in most frameworks' identifiers and illegal in UACP
229
+ (`^[a-zA-Z0-9_]{1,128}$`). Observed as a real schema rejection, not predicted.
230
+
231
+ ## Running it
232
+
233
+ ```bash
234
+ uvx uacp-interop # human-readable
235
+ uvx uacp-interop -- --json # machine-readable
236
+ uvx uacp-interop -- --fail-on-blocker # CI gate
237
+ ```
238
+
239
+ From a checkout, `pip install -e .` and then `uacp-interop`. Tests need the dev
240
+ extra: `pip install -e ".[dev]" && pytest -q`.
241
+
242
+ ## Adding your own framework
243
+
244
+ Append a `ForeignAgent` to `uacp_interop/corpus.py` and re-run. You do not have
245
+ to write a check: the divergences for your framework are computed from the
246
+ fields you declare. See [CONTRIBUTING.md](CONTRIBUTING.md), which also covers the
247
+ provenance tiers and the one rule — do not invent a serialization you cannot
248
+ source.
249
+
250
+ ## What would strengthen this
251
+
252
+ - Observed serializations rather than constructor surfaces. The AutoGen entry is
253
+ the weakest evidence in the corpus for exactly this reason, and AutoGen does
254
+ publish a component-serialization mechanism that was not inspected here.
255
+ - The implementation audit records observed field sets as data with provenance,
256
+ rather than parsing the dataclass and TypeScript interface at runtime. Parsing
257
+ them with regex would be more fragile than the value it adds. Re-reading them
258
+ is a manual step, so they can drift.
259
+ - Decisions on the two remaining normative conflicts (`additionalProperties` and
260
+ the `const` version pin), which are design questions rather than defects and
261
+ are deliberately left to the spec owner.
262
+ - LangGraph, CrewAI and OpenAI Assistants entries, which are the formats most
263
+ likely to be lossier than the two here.
264
+ - A behavioural half: run two implementations against each other and compare
265
+ actual message sequences, which is what "interoperate" ultimately means.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 wippa-studios
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,6 @@
1
+ include LICENSE
2
+ include README.md
3
+ include CONTRIBUTING.md
4
+ include FINDINGS.md
5
+ recursive-include uacp_interop/schemas *.json
6
+ recursive-include docs *.txt *.json
@@ -0,0 +1,144 @@
1
+ Metadata-Version: 2.4
2
+ Name: uacp-interop
3
+ Version: 0.1.0
4
+ Summary: Interop conformance probe for UACP: computes where UACP's schemas, prose and reference implementations diverge, and where a UACP agent definition cannot survive a round trip through a framework-native tool format.
5
+ License-Expression: MIT
6
+ Project-URL: Homepage, https://github.com/wippa-studios/uacp-interop
7
+ Project-URL: Repository, https://github.com/wippa-studios/uacp-interop
8
+ Project-URL: Issues, https://github.com/wippa-studios/uacp-interop/issues
9
+ Project-URL: Protocol, https://github.com/wippa-studios/wippa-uacp
10
+ Keywords: uacp,interoperability,conformance,agent,json-schema,protocol
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Topic :: Software Development :: Testing
19
+ Classifier: Topic :: Software Development :: Quality Assurance
20
+ Requires-Python: >=3.10
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Requires-Dist: jsonschema>=4.0
24
+ Provides-Extra: dev
25
+ Requires-Dist: pytest>=7.0; extra == "dev"
26
+ Requires-Dist: build>=1.0; extra == "dev"
27
+ Requires-Dist: twine>=5.0; extra == "dev"
28
+ Dynamic: license-file
29
+
30
+ # uacp-interop
31
+
32
+ Interop conformance probe for [wippa-uacp](https://github.com/wippa-studios/wippa-uacp).
33
+
34
+ UACP's own L0 to L3 conformance levels are a feature checklist. Passing proves an
35
+ implementation matches its author's reading of the spec, and says nothing about
36
+ whether two implementations built by different people interoperate, which is the
37
+ premise the project rests on.
38
+
39
+ This harness tests the three things the checklist cannot:
40
+
41
+ 1. **Schema versus prose.** Build a message the spec *requires* an
42
+ implementation to accept, validate it against the shipped schema, report the
43
+ rejection.
44
+ 2. **Schema versus the implementations.** Check that each reference
45
+ implementation can represent what the schema declares, and that the two
46
+ implementations model the same envelope. This class found a silent credential
47
+ downgrade between the TypeScript and Python implementations.
48
+ 3. **Cross-format divergence, both directions.** Map a framework-native agent
49
+ description into a UACP descriptor, validate the artifact actually produced,
50
+ and separately test whether a UACP capability survives being expressed in a
51
+ framework's native tool format.
52
+
53
+ ## Results
54
+
55
+ See [FINDINGS.md](FINDINGS.md). Against `main` at `7bcc2594`: 8 schema conflicts
56
+ and 12 cross-format findings (5 blocker, 11 major, 4 minor). Four findings have
57
+ since been fixed upstream and are retained only as regression guards, so a
58
+ regression re-opens them rather than passing silently.
59
+
60
+ ## Install
61
+
62
+ ```bash
63
+ uvx uacp-interop # or: pipx run uacp-interop
64
+ ```
65
+
66
+ Nothing to clone and no virtualenv to build. From a checkout:
67
+
68
+ ```bash
69
+ pip install -e . # runtime only
70
+ uacp-interop # the report
71
+ ```
72
+
73
+ ## Quick start
74
+
75
+ ```bash
76
+ uacp-interop
77
+ uacp-interop --json # machine-readable
78
+ uacp-interop --fail-on-blocker # exit 1 if any blocker is found
79
+
80
+ pip install -e ".[dev]" && pytest -q # only to run the tests
81
+ ```
82
+
83
+ A captured run of the current corpus is committed at
84
+ [docs/sample-report.txt](docs/sample-report.txt), so you can see the output
85
+ before installing anything.
86
+
87
+ ## Method
88
+
89
+ Losses are computed, not asserted. `SCHEMA_VIOLATION` findings are observed by
90
+ running the real validator over the artifact an adapter produced, and it is the
91
+ *only* source of them: the adapters do not predict schema violations from
92
+ hand-copied patterns, because a copied pattern is a second source of truth that
93
+ can drift from the schema, and because predicting as well as observing counted
94
+ each such defect twice. Divergences in the reverse direction are found by
95
+ walking the source JSON Schema and collecting keywords outside the target
96
+ format's supported subset.
97
+
98
+ Each field yields at most one finding. When both reference implementations
99
+ disagree the same way about a field, that is one defect with two witnesses and
100
+ both are cited in the evidence line, not two findings. `duplicate_findings()`
101
+ enforces this over the whole report, and the suite asserts it holds for the real
102
+ corpus and that the guard is able to see a duplicate when one is injected.
103
+
104
+ Every corpus entry declares a `confidence`:
105
+
106
+ - `observed-serialization`: a real serialized artifact was read
107
+ - `documented-api`: taken from official documentation of the public API
108
+ - `inferred`: our modelling choice, not a documented shape
109
+
110
+ Findings derived from weaker entries are weaker evidence, and the CLI says so.
111
+ The AutoGen entry currently models a documented constructor surface rather than
112
+ an observed serialization, which makes it the weakest evidence in the corpus.
113
+
114
+ Each check is backed by a meta-test that mutates a copy of the schema and
115
+ confirms the check *fires*, so a check that quietly stopped working fails the
116
+ suite rather than passing silently.
117
+
118
+ ## Vendored schemas
119
+
120
+ `uacp_interop/schemas/` holds a copy of UACP's JSON Schemas so the harness can
121
+ validate offline. A conformance report that silently tests an old protocol is
122
+ worse than no report, so the upstream commit is pinned in
123
+ `uacp_interop/schemas/PROVENANCE.json` and CI fails on drift:
124
+
125
+ ```bash
126
+ python scripts/refresh_schemas.py # refresh and re-pin
127
+ python scripts/refresh_schemas.py --check # verify only, exit 1 on drift
128
+ ```
129
+
130
+ ## Known limits
131
+
132
+ - No behavioural half. Everything here is structural: "these two definitions
133
+ cannot both be true", not "these two implementations produced different
134
+ messages". That is the harder half and the one not yet built.
135
+ - No LangGraph, CrewAI or OpenAI Assistants entries. The LangGraph docs did not
136
+ yield a sourceable serialization format, and an entry that cannot be sourced
137
+ honestly is worse than a missing one.
138
+ - The implementation audit records observed field sets as data with provenance
139
+ rather than parsing the dataclass and TypeScript interface at runtime.
140
+ Re-reading them is a manual step, so they can drift.
141
+
142
+ ## License
143
+
144
+ MIT.