uacp-interop 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- uacp_interop-0.1.0/CONTRIBUTING.md +131 -0
- uacp_interop-0.1.0/FINDINGS.md +265 -0
- uacp_interop-0.1.0/LICENSE +21 -0
- uacp_interop-0.1.0/MANIFEST.in +6 -0
- uacp_interop-0.1.0/PKG-INFO +144 -0
- uacp_interop-0.1.0/README.md +115 -0
- uacp_interop-0.1.0/docs/sample-report.json +174 -0
- uacp_interop-0.1.0/docs/sample-report.txt +117 -0
- uacp_interop-0.1.0/pyproject.toml +43 -0
- uacp_interop-0.1.0/setup.cfg +4 -0
- uacp_interop-0.1.0/tests/test_interop.py +355 -0
- uacp_interop-0.1.0/uacp_interop/__init__.py +11 -0
- uacp_interop-0.1.0/uacp_interop/adapters.py +224 -0
- uacp_interop-0.1.0/uacp_interop/cli.py +107 -0
- uacp_interop-0.1.0/uacp_interop/corpus.py +121 -0
- uacp_interop-0.1.0/uacp_interop/impl_audit.py +157 -0
- uacp_interop-0.1.0/uacp_interop/schemas/PROVENANCE.json +14 -0
- uacp_interop-0.1.0/uacp_interop/schemas/agent-descriptor.schema.json +53 -0
- uacp_interop-0.1.0/uacp_interop/schemas/capability.schema.json +40 -0
- uacp_interop-0.1.0/uacp_interop/schemas/envelope.schema.json +131 -0
- uacp_interop-0.1.0/uacp_interop/schemas/error.schema.json +46 -0
- uacp_interop-0.1.0/uacp_interop/schemas/handshake.schema.json +79 -0
- uacp_interop-0.1.0/uacp_interop/schemas/registry.schema.json +86 -0
- uacp_interop-0.1.0/uacp_interop/schemas.py +44 -0
- uacp_interop-0.1.0/uacp_interop/spec_audit.py +297 -0
- uacp_interop-0.1.0/uacp_interop/types.py +89 -0
- uacp_interop-0.1.0/uacp_interop.egg-info/PKG-INFO +144 -0
- uacp_interop-0.1.0/uacp_interop.egg-info/SOURCES.txt +30 -0
- uacp_interop-0.1.0/uacp_interop.egg-info/dependency_links.txt +1 -0
- uacp_interop-0.1.0/uacp_interop.egg-info/entry_points.txt +2 -0
- uacp_interop-0.1.0/uacp_interop.egg-info/requires.txt +6 -0
- uacp_interop-0.1.0/uacp_interop.egg-info/top_level.txt +1 -0
|
@@ -0,0 +1,131 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
The value of this harness is coverage. Every framework you add is a framework
|
|
4
|
+
whose interop loss is now measured rather than assumed, and your name is on it.
|
|
5
|
+
Adding one should take about twenty minutes and should not require understanding
|
|
6
|
+
the audit modules.
|
|
7
|
+
|
|
8
|
+
## Adding a framework to the corpus
|
|
9
|
+
|
|
10
|
+
Open `uacp_interop/corpus.py` and append a `ForeignAgent` to `CORPUS`:
|
|
11
|
+
|
|
12
|
+
```python
|
|
13
|
+
MY_FRAMEWORK = ForeignAgent(
|
|
14
|
+
framework="my-framework",
|
|
15
|
+
agent_id="my_agent", # must match UACP's ^[a-zA-Z0-9_]{1,128}$
|
|
16
|
+
description="What this agent is for.",
|
|
17
|
+
capabilities=[
|
|
18
|
+
{
|
|
19
|
+
"name": "web.search",
|
|
20
|
+
"description": "Search the web.",
|
|
21
|
+
"parameters": {
|
|
22
|
+
"type": "object",
|
|
23
|
+
"properties": {"query": {"type": "string"}},
|
|
24
|
+
"required": ["query"],
|
|
25
|
+
},
|
|
26
|
+
"outputSchema": {"type": "object", "properties": {}},
|
|
27
|
+
}
|
|
28
|
+
],
|
|
29
|
+
prose_claims=["What the docs promise a human reader."],
|
|
30
|
+
is_stateful=False, # does a call see prior calls?
|
|
31
|
+
streams_tokens=False, # per-token chunks, or only results?
|
|
32
|
+
supports_multimodal=False, # non-string content parts?
|
|
33
|
+
provenance="https://... — read 2026-09-27",
|
|
34
|
+
confidence="documented-api",
|
|
35
|
+
source_shape="What kind of thing the fields above actually came from.",
|
|
36
|
+
)
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Then run it:
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
uacp-interop
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Your framework gets its own `PART 2` section in the report. You did not have to
|
|
46
|
+
write a single check.
|
|
47
|
+
|
|
48
|
+
### The one rule
|
|
49
|
+
|
|
50
|
+
**Do not invent a serialization.** If your framework has no documented portable
|
|
51
|
+
form for an agent definition, say so in `source_shape` and model the
|
|
52
|
+
constructor or API surface instead. The AutoGen entry does exactly this and is
|
|
53
|
+
therefore the weakest evidence in the corpus; a fabricated schema would be
|
|
54
|
+
stronger-looking and worthless.
|
|
55
|
+
|
|
56
|
+
Every finding is derived from what the source actually declares. That is the
|
|
57
|
+
whole point of the tool, and an entry that guesses undermines every number in
|
|
58
|
+
the report next to it.
|
|
59
|
+
|
|
60
|
+
### Pick your confidence tier honestly
|
|
61
|
+
|
|
62
|
+
| Tier | Means |
|
|
63
|
+
|---|---|
|
|
64
|
+
| `observed-serialization` | You read a real serialized artifact |
|
|
65
|
+
| `documented-api` | Taken from official documentation of the public API |
|
|
66
|
+
| `inferred` | Your modelling choice, not a documented shape |
|
|
67
|
+
|
|
68
|
+
The CLI prints the tier for every entry, and the report says which findings rest
|
|
69
|
+
on weaker evidence. If yours is `inferred`, the findings are still useful, but
|
|
70
|
+
say so.
|
|
71
|
+
|
|
72
|
+
### Choosing capability names
|
|
73
|
+
|
|
74
|
+
Two things produce findings before your framework has even been tested, and both
|
|
75
|
+
are worth seeing rather than papering over:
|
|
76
|
+
|
|
77
|
+
- `name` must match `^[a-zA-Z0-9_.]{1,128}$` — use the source's real name. A
|
|
78
|
+
`SCHEMA_VIOLATION` here is a real result, not a harness bug.
|
|
79
|
+
- A name with no dot is a `LOSSY` finding. That is the point.
|
|
80
|
+
|
|
81
|
+
## Adding a new check
|
|
82
|
+
|
|
83
|
+
Checks live in `adapters.py` (source → UACP), `spec_audit.py` (schema vs the
|
|
84
|
+
prose) or `impl_audit.py` (schema vs the reference implementations).
|
|
85
|
+
|
|
86
|
+
Two obligations come with every check:
|
|
87
|
+
|
|
88
|
+
1. **A meta-test that proves the check fires.** Mutate a copy of the input so
|
|
89
|
+
the defect is absent, and assert the check goes quiet. A check that silently
|
|
90
|
+
stopped working is worse than no check, because the report still looks
|
|
91
|
+
healthy. The existing `test_meta_*` tests are the pattern to copy.
|
|
92
|
+
|
|
93
|
+
2. **Evidence on every finding.** `test_every_divergence_carries_evidence`
|
|
94
|
+
enforces a non-empty `detail` and `evidence`. A reader must be able to check
|
|
95
|
+
your work.
|
|
96
|
+
|
|
97
|
+
### Do not predict what you can observe
|
|
98
|
+
|
|
99
|
+
`SCHEMA_VIOLATION` has exactly one source: `validate_descriptor()`, which runs
|
|
100
|
+
the real JSON Schema validator over the artifact the adapter produced. Do not
|
|
101
|
+
also check a schema pattern with a hand-copied regex. A copied pattern is a
|
|
102
|
+
second source of truth that can drift from the schema it was copied from, and CI
|
|
103
|
+
only pins the schemas. It also counted every such defect twice before this was
|
|
104
|
+
fixed.
|
|
105
|
+
|
|
106
|
+
### One field, one finding
|
|
107
|
+
|
|
108
|
+
Report a divergent field once, however many witnesses it has. If both reference
|
|
109
|
+
implementations disagree the same way about a field, that is one defect with two
|
|
110
|
+
witnesses: cite both in `evidence` rather than emitting two findings.
|
|
111
|
+
`duplicate_findings()` checks this over the whole report and the suite asserts it
|
|
112
|
+
holds.
|
|
113
|
+
|
|
114
|
+
## Running the tests
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
pip install -e ".[dev]"
|
|
118
|
+
pytest -q
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
## Regenerating the report artifact
|
|
122
|
+
|
|
123
|
+
`docs/sample-report.txt` is a captured run, committed so a reader can see the
|
|
124
|
+
output without installing. After changing a check:
|
|
125
|
+
|
|
126
|
+
```bash
|
|
127
|
+
uacp-interop > docs/sample-report.txt
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
CI re-runs the probe weekly and opens an issue if the corpus or the schemas
|
|
131
|
+
have moved, so the published counts cannot go stale unnoticed.
|
|
@@ -0,0 +1,265 @@
|
|
|
1
|
+
# UACP interop conformance harness
|
|
2
|
+
|
|
3
|
+
Finds where a UACP agent definition and framework-native agent definitions
|
|
4
|
+
actually diverge, and whether UACP's own JSON Schemas agree with UACP's own
|
|
5
|
+
normative text.
|
|
6
|
+
|
|
7
|
+
## Why this exists
|
|
8
|
+
|
|
9
|
+
UACP's `CONFORMANCE.md` defines L0 to L3 as a feature checklist: does your
|
|
10
|
+
implementation have a stdio transport, does it do rate limiting, does it declare
|
|
11
|
+
a `protocolVersion`. Passing that proves an implementation matches its author's
|
|
12
|
+
reading of the spec. It says nothing about whether two implementations built by
|
|
13
|
+
different people interoperate, which is the premise the whole project rests on.
|
|
14
|
+
|
|
15
|
+
This harness tests the three things the checklist cannot:
|
|
16
|
+
|
|
17
|
+
1. **Schema versus prose.** Build a message the spec *requires* an
|
|
18
|
+
implementation to accept, then validate it against the shipped schema and
|
|
19
|
+
report the rejection.
|
|
20
|
+
2. **Schema versus the implementations.** Check that each reference
|
|
21
|
+
implementation can actually represent what the schema declares, and that the
|
|
22
|
+
two implementations model the same envelope. This is the class that found the
|
|
23
|
+
silent credential downgrade.
|
|
24
|
+
3. **Cross-format divergence, in both directions.** Map a foreign agent
|
|
25
|
+
description into a UACP descriptor, validate the artifact actually produced,
|
|
26
|
+
and separately test whether a UACP capability survives being expressed in a
|
|
27
|
+
framework's native tool format.
|
|
28
|
+
|
|
29
|
+
## Method
|
|
30
|
+
|
|
31
|
+
Losses are computed, not asserted. `SCHEMA_VIOLATION` findings are observed by
|
|
32
|
+
running the real validator over the artifact the adapter produced. Divergences
|
|
33
|
+
in the reverse direction are found by walking the source JSON Schema and
|
|
34
|
+
collecting keywords outside the target format's supported subset.
|
|
35
|
+
|
|
36
|
+
Every corpus entry declares `confidence`:
|
|
37
|
+
|
|
38
|
+
- `observed-serialization`: a real serialized artifact was read
|
|
39
|
+
- `documented-api`: taken from official documentation of the public API
|
|
40
|
+
- `inferred`: our modelling choice, not a documented shape
|
|
41
|
+
|
|
42
|
+
Findings derived from weaker entries are weaker evidence, and the CLI says so.
|
|
43
|
+
|
|
44
|
+
## Findings against UACP `main`
|
|
45
|
+
|
|
46
|
+
8 schema conflicts, 12 cross-format findings (5 blocker, 11 major, 4 minor across
|
|
47
|
+
all parts; 20 findings in total). Four previously reported findings are now fixed
|
|
48
|
+
and retained only as regression guards.
|
|
49
|
+
|
|
50
|
+
A captured run of exactly these numbers is committed at
|
|
51
|
+
[`docs/sample-report.txt`](docs/sample-report.txt), and CI re-runs the probe
|
|
52
|
+
weekly and opens an issue if the result stops matching it.
|
|
53
|
+
|
|
54
|
+
Earlier revisions of this document reported 9 and 13, with 6 blockers and 12
|
|
55
|
+
majors. Those counts were wrong: two defects were being emitted twice. A hyphenated
|
|
56
|
+
agent id produced both a *predicted* `SCHEMA_VIOLATION` from a hand-copied regex
|
|
57
|
+
and an *observed* one from the real validator, and `metadata.flowControl` was
|
|
58
|
+
reported once per reference implementation rather than once with both cited as
|
|
59
|
+
evidence. Both are fixed; `duplicate_findings()` now asserts that no two findings
|
|
60
|
+
in a report share a `(code, field)` pair, and the suite includes a guard proving
|
|
61
|
+
that check can see a duplicate when one is injected.
|
|
62
|
+
|
|
63
|
+
## Resolved in `main` at `7bcc2594`: one decision, three contradictions
|
|
64
|
+
|
|
65
|
+
Section 12 promised that unknown fields "in the envelope or metadata" were
|
|
66
|
+
silently ignored. The schema never did that for the envelope, which sets
|
|
67
|
+
`"additionalProperties": false` at the top level while `$defs/metadata` sets it
|
|
68
|
+
to `true`. Three defects shared that root, and resolving them separately invites
|
|
69
|
+
inconsistent answers, so they were resolved as one decision about what the
|
|
70
|
+
envelope is allowed to mean over time.
|
|
71
|
+
|
|
72
|
+
The schema split was already correct, and the existing extension namespace in
|
|
73
|
+
section 12 already routed growth through `metadata` with `vendor.extension_name`
|
|
74
|
+
keys. So the prose was wrong, not the schema.
|
|
75
|
+
|
|
76
|
+
**Forward compatibility.** Narrowed to `metadata`, which is what the schema
|
|
77
|
+
encoded. The closed envelope is now a stated design decision rather than an
|
|
78
|
+
accident, justified by the typo case: a mistyped `protocolVersion` or `traceId`
|
|
79
|
+
is far likelier than a legitimate new field, and ignoring it would drop routing
|
|
80
|
+
information.
|
|
81
|
+
|
|
82
|
+
**Version pinning.** `"const": "0.1.0"` became the MAJOR-family pattern
|
|
83
|
+
`"^0\.[0-9]+\.[0-9]+$"`. The exact pin made section 12's same-MAJOR
|
|
84
|
+
interoperability rule impossible to honour, because a 0.1.1 message was invalid
|
|
85
|
+
to a 0.1.0 validator. The published schema now tracks the released MAJOR line.
|
|
86
|
+
|
|
87
|
+
**Message type set.** The closed `type` enum became
|
|
88
|
+
`anyOf[known types, naming pattern]`. A MINOR bump may add a message type and
|
|
89
|
+
section 12 requires unrecognised types to be tolerated, so a closed enum made
|
|
90
|
+
both unreachable. This was a third instance of the same contradiction, found
|
|
91
|
+
only by reading section 12 against section 18 together.
|
|
92
|
+
|
|
93
|
+
**The governance cost, stated rather than implied.** No new top-level field can
|
|
94
|
+
ship in a MINOR bump, because an earlier validator will reject it. Growth within
|
|
95
|
+
a MAJOR series goes in `metadata` or as a new message type, and adding a
|
|
96
|
+
top-level field is a MAJOR change that must not be used to avoid adding a
|
|
97
|
+
metadata field. The spec also now notes that this differs from protobuf, which
|
|
98
|
+
preserves and round-trips unknown fields rather than rejecting them.
|
|
99
|
+
|
|
100
|
+
Verified against the validator: 0.1.0, 0.1.1 and 0.9.0 accepted; 1.0.0 and
|
|
101
|
+
`"0.1"` rejected; known and new dot-separated types accepted; `Not A Type`,
|
|
102
|
+
`agent.novel_type` and `9bad` rejected; unknown top-level field rejected;
|
|
103
|
+
unknown metadata field accepted; a `traceid` typo rejected.
|
|
104
|
+
|
|
105
|
+
## Fixed in `main` at `e0463b44`
|
|
106
|
+
|
|
107
|
+
**Audit-trail credential persistence (was: blocker, `metadata.auth.token`).**
|
|
108
|
+
Section 15.2 said implementations MUST never log token values. Section 16.3
|
|
109
|
+
separately mandated writing "Full metadata JSON" to the audit trail and named no
|
|
110
|
+
redaction hook, and L3 required a "full message log". The reference
|
|
111
|
+
implementation followed 16.3 literally: `log_message` wrote
|
|
112
|
+
`json.dumps(message.metadata.to_dict())` straight into the registry SQLite file.
|
|
113
|
+
`create_redacted_log` already existed in `core/redact.py` and was never called
|
|
114
|
+
from the audit path. So a bearer or API key token, and any secret-looking
|
|
115
|
+
payload key, were persisted in plaintext, in the conformance level named
|
|
116
|
+
"Security".
|
|
117
|
+
|
|
118
|
+
Fixed by routing `log_message` through `create_redacted_log`, adding the
|
|
119
|
+
normative requirement to 16.3, 15.2 and CONFORMANCE L3, correcting the "Full
|
|
120
|
+
trace replay" claim to a redacted replay, and documenting the requirement on
|
|
121
|
+
`metadata.auth` in the schema. Ten new tests read the SQLite file directly,
|
|
122
|
+
bypassing any accessor-layer redaction; seven of them fail against the previous
|
|
123
|
+
implementation. 94 tests pass: 84 pre-existing and unchanged, plus 10 new.
|
|
124
|
+
|
|
125
|
+
The same investigation found a latent crash with no test coverage at all:
|
|
126
|
+
`Message` has no `__post_init__`, so a caller can hand `log_message` a plain dict
|
|
127
|
+
and the old code raised `AttributeError` on a diagnostic path. No existing test
|
|
128
|
+
ever passed metadata to `log_message`, which is why it went unnoticed.
|
|
129
|
+
|
|
130
|
+
## New blockers, found by comparing code to schema
|
|
131
|
+
|
|
132
|
+
These are invisible from the schemas alone and needed a schema-versus-
|
|
133
|
+
implementation check.
|
|
134
|
+
|
|
135
|
+
**4. The Python implementation cannot represent `metadata.auth` at all.**
|
|
136
|
+
`envelope.schema.json` declares it with `required: ['type','token']`.
|
|
137
|
+
`MessageMetadata` in `python/wippa_uacp/core/types.py` has no such field, and
|
|
138
|
+
`Message.to_dict` and `Message.from_dict` both handle a fixed field list, so a
|
|
139
|
+
conforming envelope from a peer is silently degraded on decode.
|
|
140
|
+
|
|
141
|
+
**5. The two reference implementations disagree, and the failure mode is silent
|
|
142
|
+
credential downgrade.** `typescript/src/core/types.ts` declares
|
|
143
|
+
`auth?: { type, token }`. The Python one does not. A request authenticated by a
|
|
144
|
+
TypeScript peer is decoded by a Python peer into metadata that has dropped the
|
|
145
|
+
token, with no error raised. The message then relies entirely on section 15.4
|
|
146
|
+
default-deny ACLs to be stopped. Silent downgrade is worse than either rejecting
|
|
147
|
+
the message or honouring it, so 15.2 now requires rejecting a message whose
|
|
148
|
+
`auth` the implementation cannot validate.
|
|
149
|
+
|
|
150
|
+
**6. `flowControl` is implemented in both languages and declared in neither
|
|
151
|
+
schema.** Both `MessageMetadata` models carry it, and section 13 has a flow
|
|
152
|
+
control section, but `envelope.schema.json` omits it entirely. Anything using
|
|
153
|
+
flow control is outside the specification and will not survive validation against
|
|
154
|
+
a conforming peer.
|
|
155
|
+
|
|
156
|
+
|
|
157
|
+
### Remaining blockers and majors
|
|
158
|
+
|
|
159
|
+
**7. The schema set cannot be resolved offline.**
|
|
160
|
+
Every `$id` sits on `wippa.dev`, which does not resolve. `agent-descriptor`
|
|
161
|
+
refs `capability.schema.json` *relatively*, which resolves against that `$id` and
|
|
162
|
+
therefore also becomes unreachable. A standards-compliant validator attempts a
|
|
163
|
+
network fetch and fails with `Unretrievable`. This harness had to build a local
|
|
164
|
+
`referencing.Registry` just to validate. Air-gapped builds and offline CI break.
|
|
165
|
+
|
|
166
|
+
### Majors
|
|
167
|
+
|
|
168
|
+
**8. `capability` is mandatory for messages that have none.** It is in the
|
|
169
|
+
global `required` list, so `heartbeat`, `register` and `discover` must all carry
|
|
170
|
+
one, and the pattern requires at least one character. There is no honest value
|
|
171
|
+
to supply. A conforming implementation must invent a placeholder.
|
|
172
|
+
|
|
173
|
+
**9. `async` defaults to true, inverting the common case.** A capability that
|
|
174
|
+
omits `async` is treated as streaming or deferred, so a plain request/response
|
|
175
|
+
capability must opt out explicitly. A consumer trusting the default waits for a
|
|
176
|
+
`stream.end` that never arrives, which presents as a hang rather than a schema
|
|
177
|
+
error.
|
|
178
|
+
|
|
179
|
+
**10. Agents with no machine-readable capabilities are a discovery blocker.**
|
|
180
|
+
AutoGen's `AssistantAgent` exposes `name`, `description` and a tool list, but no
|
|
181
|
+
structured capability manifest with input and output schemas. A UACP registry
|
|
182
|
+
consumer asking "who can do `csv.analyze`" can only substring-match free text.
|
|
183
|
+
The answer is a guess, not a verifiable claim. This is the capability-overclaim
|
|
184
|
+
problem in its purest form, and it is not fixable by an adapter.
|
|
185
|
+
|
|
186
|
+
**11. Stateful agents have no UACP representation.** AutoGen documents `run()` as
|
|
187
|
+
mutating internal history and being called with new messages rather than
|
|
188
|
+
complete history. A UACP descriptor has no lifecycle or state field, and the
|
|
189
|
+
envelope is stateless message passing, so two calls that look identical on the
|
|
190
|
+
wire mean different things on either side.
|
|
191
|
+
|
|
192
|
+
**12. Token streaming and result streaming are different things.** AutoGen emits
|
|
193
|
+
per-token `ModelClientStreamingChunkEvent`s that are part of the agent's own
|
|
194
|
+
message trace. UACP section 13 defines `stream`/`stream.end` as partial *results*
|
|
195
|
+
correlated to a request by `replyTo` and sequence numbers. Mapping one onto the
|
|
196
|
+
other conflates the agent's reasoning trace with the result stream, so a UACP
|
|
197
|
+
consumer would receive internal reasoning it never asked for and could not tell
|
|
198
|
+
the two apart.
|
|
199
|
+
|
|
200
|
+
**13. Heterogeneous content parts are unrepresentable.** `MultiModalMessage`
|
|
201
|
+
accepts `content=[str, Image]`. The UACP payload is untyped so a list would
|
|
202
|
+
validate, but the spec defines no modality, media type or part ordering, and an
|
|
203
|
+
in-memory `Image` has no wire form. Nothing distinguishes it from an ordinary
|
|
204
|
+
array payload.
|
|
205
|
+
|
|
206
|
+
**14. A missing output contract has to be fabricated.** UACP requires
|
|
207
|
+
`outputSchema` on every capability. Neither AutoGen's `FunctionTool` nor an
|
|
208
|
+
OpenAI function declaration carries one, so the adapter must invent a permissive
|
|
209
|
+
schema and consumers will believe a guarantee the source never made.
|
|
210
|
+
|
|
211
|
+
**15. Rich JSON Schema cannot reach an OpenAI-style tool consumer.** A UACP
|
|
212
|
+
capability using `oneOf` or `const` has no representation in an OpenAI
|
|
213
|
+
function-calling `parameters` block, which supports a small fixed subset. The
|
|
214
|
+
capability cannot be exposed without dropping the constraint or flattening it
|
|
215
|
+
into prose. This is the only genuinely *bidirectional* blocker found, and it
|
|
216
|
+
lives in the UACP-to-foreign direction.
|
|
217
|
+
|
|
218
|
+
### Minors
|
|
219
|
+
|
|
220
|
+
**16.** The OpenAI `strict` flag has no JSON Schema equivalent, so the guarantee
|
|
221
|
+
is dropped in either direction and cannot be re-expressed as a UACP consumer
|
|
222
|
+
check.
|
|
223
|
+
|
|
224
|
+
**17.** Capability names without a namespace prefix (`web_search_func`,
|
|
225
|
+
`get_weather`) must be renamed, so namespace-scoped registry queries will not
|
|
226
|
+
match the original name.
|
|
227
|
+
|
|
228
|
+
**18.** Hyphens are legal in most frameworks' identifiers and illegal in UACP
|
|
229
|
+
(`^[a-zA-Z0-9_]{1,128}$`). Observed as a real schema rejection, not predicted.
|
|
230
|
+
|
|
231
|
+
## Running it
|
|
232
|
+
|
|
233
|
+
```bash
|
|
234
|
+
uvx uacp-interop # human-readable
|
|
235
|
+
uvx uacp-interop -- --json # machine-readable
|
|
236
|
+
uvx uacp-interop -- --fail-on-blocker # CI gate
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
From a checkout, `pip install -e .` and then `uacp-interop`. Tests need the dev
|
|
240
|
+
extra: `pip install -e ".[dev]" && pytest -q`.
|
|
241
|
+
|
|
242
|
+
## Adding your own framework
|
|
243
|
+
|
|
244
|
+
Append a `ForeignAgent` to `uacp_interop/corpus.py` and re-run. You do not have
|
|
245
|
+
to write a check: the divergences for your framework are computed from the
|
|
246
|
+
fields you declare. See [CONTRIBUTING.md](CONTRIBUTING.md), which also covers the
|
|
247
|
+
provenance tiers and the one rule — do not invent a serialization you cannot
|
|
248
|
+
source.
|
|
249
|
+
|
|
250
|
+
## What would strengthen this
|
|
251
|
+
|
|
252
|
+
- Observed serializations rather than constructor surfaces. The AutoGen entry is
|
|
253
|
+
the weakest evidence in the corpus for exactly this reason, and AutoGen does
|
|
254
|
+
publish a component-serialization mechanism that was not inspected here.
|
|
255
|
+
- The implementation audit records observed field sets as data with provenance,
|
|
256
|
+
rather than parsing the dataclass and TypeScript interface at runtime. Parsing
|
|
257
|
+
them with regex would be more fragile than the value it adds. Re-reading them
|
|
258
|
+
is a manual step, so they can drift.
|
|
259
|
+
- Decisions on the two remaining normative conflicts (`additionalProperties` and
|
|
260
|
+
the `const` version pin), which are design questions rather than defects and
|
|
261
|
+
are deliberately left to the spec owner.
|
|
262
|
+
- LangGraph, CrewAI and OpenAI Assistants entries, which are the formats most
|
|
263
|
+
likely to be lossier than the two here.
|
|
264
|
+
- A behavioural half: run two implementations against each other and compare
|
|
265
|
+
actual message sequences, which is what "interoperate" ultimately means.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 wippa-studios
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: uacp-interop
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Interop conformance probe for UACP: computes where UACP's schemas, prose and reference implementations diverge, and where a UACP agent definition cannot survive a round trip through a framework-native tool format.
|
|
5
|
+
License-Expression: MIT
|
|
6
|
+
Project-URL: Homepage, https://github.com/wippa-studios/uacp-interop
|
|
7
|
+
Project-URL: Repository, https://github.com/wippa-studios/uacp-interop
|
|
8
|
+
Project-URL: Issues, https://github.com/wippa-studios/uacp-interop/issues
|
|
9
|
+
Project-URL: Protocol, https://github.com/wippa-studios/wippa-uacp
|
|
10
|
+
Keywords: uacp,interoperability,conformance,agent,json-schema,protocol
|
|
11
|
+
Classifier: Development Status :: 4 - Beta
|
|
12
|
+
Classifier: Intended Audience :: Developers
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Topic :: Software Development :: Testing
|
|
19
|
+
Classifier: Topic :: Software Development :: Quality Assurance
|
|
20
|
+
Requires-Python: >=3.10
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
License-File: LICENSE
|
|
23
|
+
Requires-Dist: jsonschema>=4.0
|
|
24
|
+
Provides-Extra: dev
|
|
25
|
+
Requires-Dist: pytest>=7.0; extra == "dev"
|
|
26
|
+
Requires-Dist: build>=1.0; extra == "dev"
|
|
27
|
+
Requires-Dist: twine>=5.0; extra == "dev"
|
|
28
|
+
Dynamic: license-file
|
|
29
|
+
|
|
30
|
+
# uacp-interop
|
|
31
|
+
|
|
32
|
+
Interop conformance probe for [wippa-uacp](https://github.com/wippa-studios/wippa-uacp).
|
|
33
|
+
|
|
34
|
+
UACP's own L0 to L3 conformance levels are a feature checklist. Passing proves an
|
|
35
|
+
implementation matches its author's reading of the spec, and says nothing about
|
|
36
|
+
whether two implementations built by different people interoperate, which is the
|
|
37
|
+
premise the project rests on.
|
|
38
|
+
|
|
39
|
+
This harness tests the three things the checklist cannot:
|
|
40
|
+
|
|
41
|
+
1. **Schema versus prose.** Build a message the spec *requires* an
|
|
42
|
+
implementation to accept, validate it against the shipped schema, report the
|
|
43
|
+
rejection.
|
|
44
|
+
2. **Schema versus the implementations.** Check that each reference
|
|
45
|
+
implementation can represent what the schema declares, and that the two
|
|
46
|
+
implementations model the same envelope. This class found a silent credential
|
|
47
|
+
downgrade between the TypeScript and Python implementations.
|
|
48
|
+
3. **Cross-format divergence, both directions.** Map a framework-native agent
|
|
49
|
+
description into a UACP descriptor, validate the artifact actually produced,
|
|
50
|
+
and separately test whether a UACP capability survives being expressed in a
|
|
51
|
+
framework's native tool format.
|
|
52
|
+
|
|
53
|
+
## Results
|
|
54
|
+
|
|
55
|
+
See [FINDINGS.md](FINDINGS.md). Against `main` at `7bcc2594`: 8 schema conflicts
|
|
56
|
+
and 12 cross-format findings (5 blocker, 11 major, 4 minor). Four findings have
|
|
57
|
+
since been fixed upstream and are retained only as regression guards, so a
|
|
58
|
+
regression re-opens them rather than passing silently.
|
|
59
|
+
|
|
60
|
+
## Install
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
uvx uacp-interop # or: pipx run uacp-interop
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
Nothing to clone and no virtualenv to build. From a checkout:
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
pip install -e . # runtime only
|
|
70
|
+
uacp-interop # the report
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
## Quick start
|
|
74
|
+
|
|
75
|
+
```bash
|
|
76
|
+
uacp-interop
|
|
77
|
+
uacp-interop --json # machine-readable
|
|
78
|
+
uacp-interop --fail-on-blocker # exit 1 if any blocker is found
|
|
79
|
+
|
|
80
|
+
pip install -e ".[dev]" && pytest -q # only to run the tests
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
A captured run of the current corpus is committed at
|
|
84
|
+
[docs/sample-report.txt](docs/sample-report.txt), so you can see the output
|
|
85
|
+
before installing anything.
|
|
86
|
+
|
|
87
|
+
## Method
|
|
88
|
+
|
|
89
|
+
Losses are computed, not asserted. `SCHEMA_VIOLATION` findings are observed by
|
|
90
|
+
running the real validator over the artifact an adapter produced, and it is the
|
|
91
|
+
*only* source of them: the adapters do not predict schema violations from
|
|
92
|
+
hand-copied patterns, because a copied pattern is a second source of truth that
|
|
93
|
+
can drift from the schema, and because predicting as well as observing counted
|
|
94
|
+
each such defect twice. Divergences in the reverse direction are found by
|
|
95
|
+
walking the source JSON Schema and collecting keywords outside the target
|
|
96
|
+
format's supported subset.
|
|
97
|
+
|
|
98
|
+
Each field yields at most one finding. When both reference implementations
|
|
99
|
+
disagree the same way about a field, that is one defect with two witnesses and
|
|
100
|
+
both are cited in the evidence line, not two findings. `duplicate_findings()`
|
|
101
|
+
enforces this over the whole report, and the suite asserts it holds for the real
|
|
102
|
+
corpus and that the guard is able to see a duplicate when one is injected.
|
|
103
|
+
|
|
104
|
+
Every corpus entry declares a `confidence`:
|
|
105
|
+
|
|
106
|
+
- `observed-serialization`: a real serialized artifact was read
|
|
107
|
+
- `documented-api`: taken from official documentation of the public API
|
|
108
|
+
- `inferred`: our modelling choice, not a documented shape
|
|
109
|
+
|
|
110
|
+
Findings derived from weaker entries are weaker evidence, and the CLI says so.
|
|
111
|
+
The AutoGen entry currently models a documented constructor surface rather than
|
|
112
|
+
an observed serialization, which makes it the weakest evidence in the corpus.
|
|
113
|
+
|
|
114
|
+
Each check is backed by a meta-test that mutates a copy of the schema and
|
|
115
|
+
confirms the check *fires*, so a check that quietly stopped working fails the
|
|
116
|
+
suite rather than passing silently.
|
|
117
|
+
|
|
118
|
+
## Vendored schemas
|
|
119
|
+
|
|
120
|
+
`uacp_interop/schemas/` holds a copy of UACP's JSON Schemas so the harness can
|
|
121
|
+
validate offline. A conformance report that silently tests an old protocol is
|
|
122
|
+
worse than no report, so the upstream commit is pinned in
|
|
123
|
+
`uacp_interop/schemas/PROVENANCE.json` and CI fails on drift:
|
|
124
|
+
|
|
125
|
+
```bash
|
|
126
|
+
python scripts/refresh_schemas.py # refresh and re-pin
|
|
127
|
+
python scripts/refresh_schemas.py --check # verify only, exit 1 on drift
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
## Known limits
|
|
131
|
+
|
|
132
|
+
- No behavioural half. Everything here is structural: "these two definitions
|
|
133
|
+
cannot both be true", not "these two implementations produced different
|
|
134
|
+
messages". That is the harder half and the one not yet built.
|
|
135
|
+
- No LangGraph, CrewAI or OpenAI Assistants entries. The LangGraph docs did not
|
|
136
|
+
yield a sourceable serialization format, and an entry that cannot be sourced
|
|
137
|
+
honestly is worse than a missing one.
|
|
138
|
+
- The implementation audit records observed field sets as data with provenance
|
|
139
|
+
rather than parsing the dataclass and TypeScript interface at runtime.
|
|
140
|
+
Re-reading them is a manual step, so they can drift.
|
|
141
|
+
|
|
142
|
+
## License
|
|
143
|
+
|
|
144
|
+
MIT.
|