@eventcatalog/diff 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/COMPATIBILITY.md +257 -0
- package/LICENSE +31 -0
- package/README.md +210 -0
- package/dist/index.d.mts +244 -0
- package/dist/index.d.ts +244 -0
- package/dist/index.js +851 -0
- package/dist/index.js.map +1 -0
- package/dist/index.mjs +820 -0
- package/dist/index.mjs.map +1 -0
- package/package.json +60 -0
package/COMPATIBILITY.md
ADDED
|
@@ -0,0 +1,257 @@
|
|
|
1
|
+
# Schema compatibility, explained
|
|
2
|
+
|
|
3
|
+
This guide is for anyone who has been told "your change is breaking" and wants to know why, or who wants to change a schema and know in advance whether it is safe. No prior knowledge of schema registries is assumed.
|
|
4
|
+
|
|
5
|
+
It covers JSON Schema, which is the first format `@eventcatalog/diff` understands. The ideas apply to any schema format.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## The one idea behind everything
|
|
10
|
+
|
|
11
|
+
A message has two sides. Something **writes** it, something **reads** it.
|
|
12
|
+
|
|
13
|
+
- A **producer** writes a message using the schema it was built with.
|
|
14
|
+
- A **consumer** reads a message using the schema it was built with.
|
|
15
|
+
|
|
16
|
+
The two sides are usually different services, owned by different teams, deployed at different times. So at any moment there might be messages in flight, in a queue, in a topic, or in a database, written with one version of the schema and read with another.
|
|
17
|
+
|
|
18
|
+
A change is **safe** when the reader's schema accepts everything the writer's schema could have produced.
|
|
19
|
+
|
|
20
|
+
A change is **breaking** when it doesn't.
|
|
21
|
+
|
|
22
|
+
Every rule in this document is that one sentence applied to one keyword.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## The running example
|
|
27
|
+
|
|
28
|
+
Throughout, we use one event.
|
|
29
|
+
|
|
30
|
+
**Orders Service** publishes `OrderCreated`. **Payment Service** subscribes to it.
|
|
31
|
+
|
|
32
|
+
```json
|
|
33
|
+
{
|
|
34
|
+
"type": "object",
|
|
35
|
+
"properties": {
|
|
36
|
+
"orderId": { "type": "string" },
|
|
37
|
+
"customerId": { "type": "string" },
|
|
38
|
+
"status": { "type": "string", "enum": ["pending", "paid"] }
|
|
39
|
+
},
|
|
40
|
+
"required": ["orderId", "customerId"]
|
|
41
|
+
}
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Orders wants to change this schema. Payment is still running on the old one. What happens?
|
|
45
|
+
|
|
46
|
+
---
|
|
47
|
+
|
|
48
|
+
## Two directions
|
|
49
|
+
|
|
50
|
+
Because there are two sides, there are two questions you can ask.
|
|
51
|
+
|
|
52
|
+
### Forward: does the old consumer survive the new messages?
|
|
53
|
+
|
|
54
|
+
Orders ships the new schema. Payment has not been touched. New `OrderCreated` messages start arriving at Payment, which is still validating with the **old** schema.
|
|
55
|
+
|
|
56
|
+
> **Forward compatible** means a consumer on the **old** schema can read messages written with the **new** schema.
|
|
57
|
+
|
|
58
|
+
This is the question most people mean when they ask "will my change break anyone?". In EventCatalog terms: **will this pull request break the consumers in the graph?**
|
|
59
|
+
|
|
60
|
+
### Backward: does the new consumer survive the old messages?
|
|
61
|
+
|
|
62
|
+
Now the other way round. Payment upgrades to the new schema. But there are old messages still in the queue, or Payment replays a week of history from the topic, or reads an event store. Those messages were written with the **old** schema, and Payment is now validating with the **new** one.
|
|
63
|
+
|
|
64
|
+
> **Backward compatible** means a consumer on the **new** schema can read messages written with the **old** schema.
|
|
65
|
+
|
|
66
|
+
This matters whenever old messages can meet a new reader: Kafka replays, event sourcing, dead letter queues, long retention.
|
|
67
|
+
|
|
68
|
+
### A naming trap
|
|
69
|
+
|
|
70
|
+
In the REST API world, "backward compatible" means "old clients keep working". That is the **forward** direction in schema terms. The names come from schema registries like Confluent's, and we use the registry meaning so the tool agrees with the registry. If you only care whether existing consumers keep working, the word you want is **forward**.
|
|
71
|
+
|
|
72
|
+
---
|
|
73
|
+
|
|
74
|
+
## Strategies
|
|
75
|
+
|
|
76
|
+
A strategy tells the diff which questions to ask.
|
|
77
|
+
|
|
78
|
+
| Strategy | Question asked | Use it when |
|
|
79
|
+
| ---------- | ----------------------------------------------------- | ---------------------------------------------------------------------------- |
|
|
80
|
+
| `forward` | Will old consumers survive new messages? | Producers ship first, consumers catch up later. You never replay old data. |
|
|
81
|
+
| `backward` | Will new consumers survive old messages? | Consumers ship first, or consumers replay history. Confluent's default. |
|
|
82
|
+
| `full` | Both of the above | Different teams deploy independently and you want no surprises. **Default.** |
|
|
83
|
+
| `none` | Neither. Report changes, never call anything breaking | You want the report, but you coordinate upgrades yourself. |
|
|
84
|
+
|
|
85
|
+
**Why `full` is the default.** EventCatalog runs the diff on a pull request. The person changing the schema usually owns the producer. The consumers belong to other teams and might deploy tomorrow or next month, and might replay data. A review tool should flag anything that could hurt either of them. If your organisation has settled on one direction, set it explicitly.
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## The rules, with reasons
|
|
90
|
+
|
|
91
|
+
Every change makes the schema accept **more** values, accept **fewer** values, accept a **different** set of values, or accept the **same** values.
|
|
92
|
+
|
|
93
|
+
| | Backward (new reader, old data) | Forward (old reader, new data) |
|
|
94
|
+
| --------------------- | ------------------------------- | ------------------------------ |
|
|
95
|
+
| Accepts **more** | safe | **breaking** |
|
|
96
|
+
| Accepts **fewer** | **breaking** | safe |
|
|
97
|
+
| Accepts **different** | **breaking** | **breaking** |
|
|
98
|
+
| Accepts the **same** | safe | safe |
|
|
99
|
+
|
|
100
|
+
Why? If the new schema accepts more, old data still fits inside it (backward safe), but new data might fall outside the old schema (forward breaking). Accepting fewer is the mirror image.
|
|
101
|
+
|
|
102
|
+
That table is the whole rules engine. The sections below just say which bucket each keyword change lands in.
|
|
103
|
+
|
|
104
|
+
### Properties
|
|
105
|
+
|
|
106
|
+
| Change | Bucket | Why |
|
|
107
|
+
| -------------------------------------------- | --------- | ----------------------------------------------------------------------------------------------------------------- |
|
|
108
|
+
| Add an optional property | same | Old messages don't have it and don't need it. New messages carry an extra field that old readers ignore. |
|
|
109
|
+
| Remove an optional property | same | Nobody was relying on it being there. |
|
|
110
|
+
| Add a required property | fewer | Old messages don't have it, so a new reader rejects them. Old readers ignore the extra field, so forward is fine. |
|
|
111
|
+
| Add a required property that has a `default` | same | The new reader fills the gap with the default when reading old messages. |
|
|
112
|
+
| Remove a required property | more | New messages may omit a field the old reader insists on. |
|
|
113
|
+
| Rename a required property | different | It is a removal and an addition at once. Both directions break. |
|
|
114
|
+
|
|
115
|
+
**Example, forward breaking.** Orders removes `customerId` from `required`. Payment, on the old schema, receives an `OrderCreated` without a `customerId` and rejects it, or crashes reading `undefined`.
|
|
116
|
+
|
|
117
|
+
**Example, backward breaking.** Orders adds `placedAt` to `required`. Payment upgrades and then replays last week's events, none of which have `placedAt`. Every one is rejected.
|
|
118
|
+
|
|
119
|
+
**Fix.** New fields should be optional, or required with a `default`. Removing a field should be done in two steps: make it optional, wait for consumers to stop depending on it, then remove it.
|
|
120
|
+
|
|
121
|
+
### Types
|
|
122
|
+
|
|
123
|
+
| Change | Bucket | Why |
|
|
124
|
+
| ------------------------------------------- | --------- | ------------------------------------------------------------ |
|
|
125
|
+
| `integer` to `number` | more | Every integer is a number. New messages may now carry `2.5`. |
|
|
126
|
+
| `number` to `integer` | fewer | Old messages may have carried `2.5`. |
|
|
127
|
+
| `string` to `["string", "null"]` (nullable) | more | New messages may carry `null`. |
|
|
128
|
+
| `["string", "null"]` to `string` | fewer | Old messages may have carried `null`. |
|
|
129
|
+
| Remove `type` entirely | more | The property now accepts anything. |
|
|
130
|
+
| Add a `type` where there was none | fewer | Old messages could have carried anything. |
|
|
131
|
+
| `string` to `number`, or any other swap | different | Neither side accepts the other's values. |
|
|
132
|
+
|
|
133
|
+
**Example.** Orders changes `orderId` from `string` to `integer`. Payment on the old schema receives `12345` where it expected `"ord_12345"`. Both directions break, because old messages have strings and new messages have integers.
|
|
134
|
+
|
|
135
|
+
### Enums and `const`
|
|
136
|
+
|
|
137
|
+
| Change | Bucket | Why |
|
|
138
|
+
| --------------------------------- | --------- | ------------------------------------------------------------- |
|
|
139
|
+
| Add an enum value | more | New messages may carry a value the old reader has never seen. |
|
|
140
|
+
| Remove an enum value | fewer | Old messages may carry it. |
|
|
141
|
+
| Restrict a free string to an enum | fewer | Old messages could have carried any string. |
|
|
142
|
+
| Lift an enum restriction | more | New messages may carry any string. |
|
|
143
|
+
| Change a `const` | different | The old value is gone and a new one appears. |
|
|
144
|
+
|
|
145
|
+
**Example.** Orders adds `"refunded"` to `status`. Payment on the old schema has a `switch` over `pending` and `paid` with no default branch. The first refunded order hits code that was never written. That is forward breaking, and it is one of the most common real-world breakages.
|
|
146
|
+
|
|
147
|
+
### Constraints
|
|
148
|
+
|
|
149
|
+
`minLength`, `maxLength`, `minimum`, `maximum`, `exclusiveMinimum`, `exclusiveMaximum`, `minItems`, `maxItems`, `minProperties`, `maxProperties`, `pattern`, `format`, `multipleOf`, `uniqueItems`.
|
|
150
|
+
|
|
151
|
+
| Change | Bucket |
|
|
152
|
+
| --------------------------------------------------------------------------------------------- | --------- |
|
|
153
|
+
| Tighten (raise a minimum, lower a maximum, add a pattern, add a format, require unique items) | fewer |
|
|
154
|
+
| Loosen (the reverse of any of the above, or remove the keyword) | more |
|
|
155
|
+
| Change `pattern`, `format` or `multipleOf` to a different value | different |
|
|
156
|
+
|
|
157
|
+
**Why is a changed pattern "different" rather than tighter or looser?** Because working out whether one regular expression accepts everything another one does is not something we can do reliably. Rather than guess, we treat it as breaking in both directions and let a human look.
|
|
158
|
+
|
|
159
|
+
### Additional properties, and closed objects
|
|
160
|
+
|
|
161
|
+
By default a JSON Schema object is **open**: properties not listed in `properties` are allowed. Setting `"additionalProperties": false` makes it **closed**.
|
|
162
|
+
|
|
163
|
+
| Change | Bucket | Why |
|
|
164
|
+
| ---------------------------------------------- | ------ | ------------------------------------------------------ |
|
|
165
|
+
| Open to closed (`additionalProperties: false`) | fewer | Old messages may have carried extra fields. |
|
|
166
|
+
| Closed to open | more | New messages may carry fields the old reader rejects. |
|
|
167
|
+
| Add a schema for additional properties | fewer | Extra fields used to be anything, now they must match. |
|
|
168
|
+
|
|
169
|
+
**Closed objects change the property rules.** The property table above assumes an open object, where a reader shrugs at fields it does not know. A closed reader does not shrug. It rejects the whole message.
|
|
170
|
+
|
|
171
|
+
| Change on a closed object | Bucket | Why |
|
|
172
|
+
| --------------------------- | ------ | ---------------------------------------------------------------------------------------------- |
|
|
173
|
+
| Add an optional property | more | New messages carry a field the **old** closed reader has never heard of. Forward breaks. |
|
|
174
|
+
| Remove an optional property | fewer | Old messages still carry the field, and the **new** closed reader rejects it. Backward breaks. |
|
|
175
|
+
|
|
176
|
+
**Example.** Payment validates `OrderCreated` with `additionalProperties: false`. Orders adds an optional `placedAt`. Under an open schema this is the safest change there is. Under the closed schema, every new message is rejected by Payment. If you use closed objects, adding fields needs a two-step rollout: consumers first, then producers.
|
|
177
|
+
|
|
178
|
+
### Arrays
|
|
179
|
+
|
|
180
|
+
The schema for array items is walked like any other schema, so a change to `items.type` follows the type rules above, with a path like `/properties/lines/items/type`.
|
|
181
|
+
|
|
182
|
+
Tuples (`items` as an array in draft-07, `prefixItems` in 2020-12) describe positions. Constraining a new position accepts fewer; dropping one accepts more.
|
|
183
|
+
|
|
184
|
+
### Composition: `oneOf`, `anyOf`, `allOf`
|
|
185
|
+
|
|
186
|
+
| Change | Bucket | Why |
|
|
187
|
+
| --------------------------------- | ------ | -------------------------------------------------------------------- |
|
|
188
|
+
| Add a `oneOf` / `anyOf` branch | more | New messages may match a shape the old reader has never seen. |
|
|
189
|
+
| Remove a `oneOf` / `anyOf` branch | fewer | Old messages may have matched it. |
|
|
190
|
+
| Add an `allOf` branch | fewer | Every message must now satisfy one more constraint. |
|
|
191
|
+
| Remove an `allOf` branch | more | A constraint is gone. |
|
|
192
|
+
| Change inside a branch | walked | Reported at the branch path, e.g. `/properties/payment/oneOf/0/...`. |
|
|
193
|
+
|
|
194
|
+
**Example.** Orders adds an `apple_pay` branch to `payment.oneOf`. Payment on the old schema receives a payment shape it cannot validate. Forward breaking.
|
|
195
|
+
|
|
196
|
+
### `$ref` and definitions
|
|
197
|
+
|
|
198
|
+
References to `#/definitions/...` and `#/$defs/...` are followed before comparing, including a ref that points at another ref. A change inside a definition is reported **once**, at the definition's path, no matter how many properties use it. Moving an inline object into a definition, or renaming a definition, with identical content is not a change at all.
|
|
199
|
+
|
|
200
|
+
A `$ref` to **another file** cannot be followed, because the diff only sees the one schema. If such a ref changes, it is reported as `keyword.changed` and treated as breaking both ways. If it is unchanged, nothing is reported.
|
|
201
|
+
|
|
202
|
+
### Things that are never breaking
|
|
203
|
+
|
|
204
|
+
Changing `title`, `description`, `examples`, `$comment`, `$id`, `$schema` or `default`. Reordering `properties`, `required` entries, `enum` values or `oneOf` branches. None of these change what values are accepted, so none of them produce any ops.
|
|
205
|
+
|
|
206
|
+
One annotation is reported without being breaking: `"deprecated": true`. Catalogs care about deprecation, so it appears as a `schema.deprecated` op that a UI can show.
|
|
207
|
+
|
|
208
|
+
### Things we refuse to judge
|
|
209
|
+
|
|
210
|
+
`patternProperties`, `propertyNames`, `not`, `if` / `then` / `else`, `dependentRequired`, `dependentSchemas`, `dependencies`, `contains`, `minContains`, `maxContains`, `unevaluatedProperties`, `unevaluatedItems`, `additionalItems`.
|
|
211
|
+
|
|
212
|
+
These keywords can express almost anything, and deciding compatibility for them properly is a research problem. When one of them changes, the diff reports it as `keyword.changed` and treats it as breaking in **both** directions. That is deliberate. A false alarm costs a human a minute. A silent pass costs an outage.
|
|
213
|
+
|
|
214
|
+
---
|
|
215
|
+
|
|
216
|
+
## Reading a diff result
|
|
217
|
+
|
|
218
|
+
For every changed schema the diff gives you:
|
|
219
|
+
|
|
220
|
+
- **`breaking`**: `true`, `false`, or `null` when the index was built without schema content and no verdict was possible.
|
|
221
|
+
- **`direction`**: which side breaks. `forward` means existing consumers are hurt. `backward` means consumers replaying old data are hurt. `both` means both.
|
|
222
|
+
- **`ops`**: every change found, each with a JSON pointer path, a stable `kind`, a human `reason`, and its own `breaking` flag under the chosen strategy.
|
|
223
|
+
|
|
224
|
+
And for every breaking change, an **impact** entry naming the producers and consumers of that message, with their owners, taken from the catalog graph.
|
|
225
|
+
|
|
226
|
+
So "Orders removed `customerId` from `OrderCreated`" comes back as: breaking, direction forward, one op at `/properties/customerId` saying the property is no longer required, and an impact entry saying Payment Service, owned by team-payments, consumes this message.
|
|
227
|
+
|
|
228
|
+
---
|
|
229
|
+
|
|
230
|
+
## Frequently asked questions
|
|
231
|
+
|
|
232
|
+
**I added an optional field and the diff says breaking under `full`. Why?**
|
|
233
|
+
Check the ops. An optional field alone is safe. Usually something else changed in the same commit, often a description tidy-up that also touched an enum, or a field that was made required at the same time.
|
|
234
|
+
|
|
235
|
+
**Adding an enum value is breaking? Everybody does that.**
|
|
236
|
+
Under `forward`, yes, and it is genuinely the cause of real incidents: consumers with an exhaustive switch and no default. Under `backward` it is safe. If your consumers are all written defensively, run with `backward` and the diff will agree with you.
|
|
237
|
+
|
|
238
|
+
**My schema uses `if` / `then`. Every change is flagged.**
|
|
239
|
+
Only changes to the `if` / `then` blocks themselves. Changes elsewhere in the schema are judged normally. If the conditional logic changes, a human should look.
|
|
240
|
+
|
|
241
|
+
**What does a version bump do?**
|
|
242
|
+
If the baseline has `OrderCreated` 1.0.0 and the candidate adds 2.0.0, then 1.0.0 is compared against 2.0.0 and both versions are recorded on the change. A version bump does not make a breaking change safe. It just makes it visible and intentional. Separately, every version that exists on both sides is compared against itself, so editing 1.0.0 in place while 2.0.0 exists is still found.
|
|
243
|
+
|
|
244
|
+
**I renamed the schema file. Will the diff lose track of it?**
|
|
245
|
+
No. Schema files are paired by `id`, then `path`, then the `default` flag, and if exactly one file is left on each side they are paired by elimination. A file that genuinely appears or disappears is reported as an added or removed schema.
|
|
246
|
+
|
|
247
|
+
**My messages use Avro, not JSON Schema.**
|
|
248
|
+
The diff sees that the file changed, because the hash differs, but cannot judge it yet. The change is reported with `breaking: null` and counted in `summary.schemaUnknown`. Nothing marks the diff as breaking, so make sure your policy treats a non-zero unknown count as "a human needs to look". Silently passing an unjudged change is the one thing this tool must never do.
|
|
249
|
+
|
|
250
|
+
**Why doesn't a schema with no consumers still count as breaking?**
|
|
251
|
+
It does. The diff reports facts. The impact entry will show an empty `consumers` list, and the policy layer, for example the CLI, can decide that nobody is hurt and downgrade it to a warning.
|
|
252
|
+
|
|
253
|
+
**Where do these rules come from?**
|
|
254
|
+
The definitions of backward, forward and full match [Confluent Schema Registry](https://docs.confluent.io/platform/current/schema-registry/fundamentals/schema-evolution.html). The per-keyword reasoning follows Confluent's [JSON Schema compatibility notes](https://docs.confluent.io/platform/current/schema-registry/fundamentals/serdes-develop/serdes-json.html), using their lenient reading of the open content model.
|
|
255
|
+
|
|
256
|
+
**A rule looks wrong to me.**
|
|
257
|
+
Open `src/test/json-schema/json-schema.test.ts`, find the strategy you run, copy the closest test, paste your before and after schema, and state the verdict you expect. That is the fastest way to have the conversation.
|
package/LICENSE
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2022-2026 boyney123
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
Note: Source code within the following directories is licensed under the
|
|
26
|
+
EventCatalog Commercial License and is NOT covered by the MIT License above:
|
|
27
|
+
|
|
28
|
+
- `packages/core/eventcatalog/src/enterprise/`
|
|
29
|
+
- `packages/core/src/federation/`
|
|
30
|
+
|
|
31
|
+
See the `LICENSE` file in each directory for the full terms.
|
package/README.md
ADDED
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
# @eventcatalog/diff
|
|
2
|
+
|
|
3
|
+
Compare two EventCatalog indexes and return a single, versioned `ArchitectureDiff` document.
|
|
4
|
+
|
|
5
|
+
This is a pure library. It has no knowledge of git, CI, webhooks or policy. Those are consumers that build on top of the diff document.
|
|
6
|
+
|
|
7
|
+
## Usage
|
|
8
|
+
|
|
9
|
+
```ts
|
|
10
|
+
import createSDK from '@eventcatalog/sdk';
|
|
11
|
+
import { diff } from '@eventcatalog/diff';
|
|
12
|
+
|
|
13
|
+
// Build an index for each side. Include schema content so the diff can
|
|
14
|
+
// compute compatibility verdicts, not just detect that a schema changed.
|
|
15
|
+
const a = await createSDK(baselineDir).buildIndex({ source: 'acme/catalog', commit: 'abc1234', includeSchemaContent: true });
|
|
16
|
+
const b = await createSDK(candidateDir).buildIndex({ source: 'acme/catalog', commit: 'def5678', includeSchemaContent: true });
|
|
17
|
+
|
|
18
|
+
const result = diff(a, b, { strategy: 'backward' });
|
|
19
|
+
|
|
20
|
+
if (result.summary.breaking) {
|
|
21
|
+
// apply your own policy: fail, warn, notify...
|
|
22
|
+
}
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
`a` is the baseline (for example `main`) and `b` is the candidate (for example a PR branch).
|
|
26
|
+
|
|
27
|
+
### Options
|
|
28
|
+
|
|
29
|
+
| Option | Default | What it does |
|
|
30
|
+
| ---------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
31
|
+
| `strategy` | `'full'` | Compatibility strategy used to judge schema changes. See below. |
|
|
32
|
+
| `includeSchemaContent` | `false` | Copy the raw before and after schema text onto each schema change, so a UI can render a side-by-side diff without going back to the index. |
|
|
33
|
+
|
|
34
|
+
## Input: the SDK `Index`
|
|
35
|
+
|
|
36
|
+
The input is the `Index` document returned by `buildIndex()` in `@eventcatalog/sdk`, the same document used for federation. The diff resolves each side with the SDK resolver so pointers such as `version: latest` become concrete edges. The diff engine never touches the filesystem.
|
|
37
|
+
|
|
38
|
+
Schema content is opt in. Build with `includeSchemaContent: true` to embed the raw schema text next to its hash. Without it the diff still reports which schemas changed, but the compatibility verdict is `null`.
|
|
39
|
+
|
|
40
|
+
### Which schemas get compared
|
|
41
|
+
|
|
42
|
+
- **Versions.** Every version of a message that exists on both sides is compared against itself, so a patch to `OrderCreated` 1.0.0 is found even when 2.0.0 exists. When the latest versions differ, the latest on each side are also compared, so a bump from 1.0.0 to 2.0.0 is reported as one change with both versions recorded.
|
|
43
|
+
- **Files.** A message can carry several schema files. They are paired by `id`, then by `path`, then by the `default` flag, and finally, when exactly one file is left unmatched on each side, by elimination. So renaming `schema.json` to `order-created.json` in the same change does not hide what changed inside it. A file that exists on only one side is reported as `change: 'added'` or `change: 'removed'`.
|
|
44
|
+
- **Formats.** `format: 'json-schema'` is compared. A `.json` file with no declared format is compared only when its content uses JSON Schema keywords, so an example payload stored as `schema.json` is not walked as if it were a schema. Everything else gets a `null` verdict.
|
|
45
|
+
|
|
46
|
+
### Unknown verdicts are never silent
|
|
47
|
+
|
|
48
|
+
A schema change with no verdict, because the format is not supported yet, the content was not in the index, the JSON did not parse, or the file was added or removed, has `breaking: null` and is counted in `summary.schemaUnknown`. `summary.breaking` stays `false` for these. A policy layer should treat a non-zero `schemaUnknown` as "a human needs to look", and probably fail the build for it until the format is supported.
|
|
49
|
+
|
|
50
|
+
## Output: `ArchitectureDiff`
|
|
51
|
+
|
|
52
|
+
- `resources` added / removed / changed
|
|
53
|
+
- `edges` added / removed, across every direction the SDK resolves
|
|
54
|
+
- `schemaChanges` with a compatibility verdict, the direction that broke, and the operations that caused it
|
|
55
|
+
- `impact` derived from `sends` / `receives` edges, so callers do not need to walk the graph again
|
|
56
|
+
|
|
57
|
+
### Resources
|
|
58
|
+
|
|
59
|
+
A resource is identified by its `type` and `id`, so an event and a service with the same id are different resources. Because an index carries every version of a resource:
|
|
60
|
+
|
|
61
|
+
- an id only in the candidate is **added**, an id only in the baseline is **removed**, every version listed
|
|
62
|
+
- an id on both sides whose latest version differs is **changed** with `version: { a, b }`, not removed plus added
|
|
63
|
+
- an id on both sides with the same latest version but a different `name`, `owners` or `deprecated` is **changed**, with the differing `fields` named
|
|
64
|
+
- an individual old version that appears or disappears while the id survives is added or removed on its own, so deleting 0.6.0 while 1.0.0 stays is visible
|
|
65
|
+
|
|
66
|
+
Markdown content and schema files are not resource changes. Schemas are covered by `schemaChanges`; markdown is documentation.
|
|
67
|
+
|
|
68
|
+
### Edges
|
|
69
|
+
|
|
70
|
+
Every direction the SDK resolves is compared: `sends`, `receives`, `writesTo`, `readsFrom`, `contains`, `references`, `appliesTo`, `relatesTo`. Edge identity is direction, `via`, from id and to id. Versions are left out on purpose, so bumping a message does not churn every edge that points at it. A consumer on `latest` is resolved before comparing.
|
|
71
|
+
|
|
72
|
+
An edge whose target disappears from the catalog, so that the pointer no longer resolves, is reported as removed. A pointer to something not in the catalog is still reported, with no `type` on that end.
|
|
73
|
+
|
|
74
|
+
### Impact
|
|
75
|
+
|
|
76
|
+
`impact` names who is hurt. Producers and consumers always come from the **baseline** graph: the services that exist today, on the message that is about to change or disappear, with their owners.
|
|
77
|
+
|
|
78
|
+
| Reason | When | Makes the diff breaking |
|
|
79
|
+
| ------------------------ | -------------------------------------------------------------------- | -------------------------- |
|
|
80
|
+
| `schema_breaking_change` | A message schema broke under the strategy | yes |
|
|
81
|
+
| `message_removed` | A message, or one version of it, is gone while services still use it | yes |
|
|
82
|
+
| `consumer_removed` | A service stopped receiving a message that still exists | no, that is its own choice |
|
|
83
|
+
| `producer_removed` | A service stopped sending a message that still exists | no |
|
|
84
|
+
|
|
85
|
+
A removed message nobody used any more is a tidy-up and gets no impact entry.
|
|
86
|
+
|
|
87
|
+
Producers and consumers are whatever declares `sends` or `receives` in the catalog: services, domains and agents all appear, each with its own `type`.
|
|
88
|
+
|
|
89
|
+
```json
|
|
90
|
+
{
|
|
91
|
+
"message": { "type": "event", "id": "order-created", "version": "1.0.0" },
|
|
92
|
+
"reason": "schema_breaking_change",
|
|
93
|
+
"direction": "forward",
|
|
94
|
+
"producers": [{ "type": "service", "id": "orders-service", "version": "3.1.0", "owners": ["team-orders"] }],
|
|
95
|
+
"consumers": [{ "type": "service", "id": "payment-service", "version": "2.0.0", "owners": ["team-payments"] }]
|
|
96
|
+
}
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
`direction` tells a consumer of the diff how to read the lists:
|
|
100
|
+
|
|
101
|
+
| Direction | Who is hurt |
|
|
102
|
+
| ---------- | --------------------------------------------------------------------------------------------- |
|
|
103
|
+
| `forward` | The consumers listed are on the old schema and will fail to read new messages |
|
|
104
|
+
| `backward` | The consumers listed will fail to read old messages once they move to the new schema (replay) |
|
|
105
|
+
| `both` | Both of the above. Only possible under the `full` strategy |
|
|
106
|
+
|
|
107
|
+
Only direct edges to the compared version count, after the SDK has resolved them against the baseline. In practice:
|
|
108
|
+
|
|
109
|
+
- A consumer on `latest`, or with no version at all, resolves to the latest version in the baseline. It is listed when that version is edited in place, and when the message is bumped, because a latest subscriber moves to the new version automatically. It is not listed when an older, non-latest version is patched.
|
|
110
|
+
- A consumer on a semver range such as `^1.0.0` resolves to the highest matching version and is listed when that version changes, even for a major bump. The catalog cannot know whether the producer keeps publishing the old line, so it errs on inclusion.
|
|
111
|
+
- A consumer pinned to an exact older version is not listed for a change to a newer one.
|
|
112
|
+
- A message nobody consumes still gets an entry with an empty `consumers` list, so policy can decide that nobody is affected.
|
|
113
|
+
|
|
114
|
+
The diff reports facts. Whether an impact fails a build, warns, or only matters when the consumer belongs to another team is policy, and belongs in the CLI.
|
|
115
|
+
|
|
116
|
+
## Compatibility strategies
|
|
117
|
+
|
|
118
|
+
> New to schema compatibility? Read [COMPATIBILITY.md](./COMPATIBILITY.md) first. It explains the two directions, every rule, and why, using one running example.
|
|
119
|
+
|
|
120
|
+
The strategy answers one question: **which side upgrades first?** The definitions follow [Confluent Schema Registry](https://docs.confluent.io/platform/current/schema-registry/fundamentals/schema-evolution.html), so they match what teams running Kafka already expect.
|
|
121
|
+
|
|
122
|
+
| Strategy | Meaning | Upgrade order |
|
|
123
|
+
| ---------- | ------------------------------------------------------------------------------- | -------------------------------------------- |
|
|
124
|
+
| `backward` | A consumer on the **new** schema can read messages written with the **old** one | Upgrade consumers first, then producers |
|
|
125
|
+
| `forward` | A consumer on the **old** schema can read messages written with the **new** one | Upgrade producers first, then consumers |
|
|
126
|
+
| `full` | Both of the above | Upgrade producers and consumers in any order |
|
|
127
|
+
| `none` | Compatibility is not checked. Changes are still reported, nothing is breaking | Coordinate the upgrade yourself |
|
|
128
|
+
|
|
129
|
+
**The default is `full`.** Confluent defaults to `backward` because a Kafka consumer may rewind and replay old messages. EventCatalog runs the diff on a pull request where a producer changes an event that other services already consume. Those consumers are on the old schema, which is the `forward` direction, so a `backward` default would pass changes that break every consumer in the graph. Checking both directions is the safe default for a review tool. Set `backward` or `forward` to match the setting on your schema registry.
|
|
130
|
+
|
|
131
|
+
> **Naming warning.** In REST API terms, "backward compatible" means old clients keep working. In schema registry terms that is `forward`. We use the schema registry meaning. If you only want to know whether existing consumers keep working, use `forward`.
|
|
132
|
+
|
|
133
|
+
Every rule below comes from the same test: does the reader's schema accept everything the writer's schema could have produced? A change that makes a schema accept **more** is safe for `backward` and breaking for `forward`. A change that makes it accept **less** is the reverse.
|
|
134
|
+
|
|
135
|
+
### JSON Schema rules
|
|
136
|
+
|
|
137
|
+
Objects are open by default in JSON Schema: unknown properties are accepted. We take Confluent's lenient reading of that: adding a property to an open object is safe, on the basis that producers do not emit undeclared properties. Objects with `additionalProperties: false` are closed, and get their own stricter rows below. A schema-valued `additionalProperties` is treated as open.
|
|
138
|
+
|
|
139
|
+
| Change | `backward` | `forward` | Kind |
|
|
140
|
+
| ------------------------------------------------------------------------------------------------------------------ | ---------- | --------- | ------------------------------------- |
|
|
141
|
+
| Add an optional property to an open object | ok | ok | `property.added` |
|
|
142
|
+
| Add an optional property to a closed object (old readers reject unknown properties) | ok | breaking | `property.added-to-closed-object` |
|
|
143
|
+
| Remove an optional property from an open object | ok | ok | `property.removed` |
|
|
144
|
+
| Remove an optional property from a closed object (new readers reject old messages still carrying it) | breaking | ok | `property.removed-from-closed-object` |
|
|
145
|
+
| Make a property required / add a required property | breaking | ok | `required.added` |
|
|
146
|
+
| Make a property required when it has a `default` | ok | ok | `required.added-with-default` |
|
|
147
|
+
| Make a required property optional / remove a required property | ok | breaking | `required.removed` |
|
|
148
|
+
| Widen a type (`integer` to `number`, `string` to `["string","null"]`, type removed) | ok | breaking | `type.widened` |
|
|
149
|
+
| Narrow a type (`number` to `integer`, nullable to not, type added) | breaking | ok | `type.narrowed` |
|
|
150
|
+
| Change a type outright (`string` to `number`) | breaking | breaking | `type.changed` |
|
|
151
|
+
| Add an enum value | ok | breaking | `enum.value.added` |
|
|
152
|
+
| Remove an enum value | breaking | ok | `enum.value.removed` |
|
|
153
|
+
| Restrict a free value to an enum or `const` | breaking | ok | `enum.added` |
|
|
154
|
+
| Lift an enum restriction | ok | breaking | `enum.removed` |
|
|
155
|
+
| Tighten a constraint (`min*` up, `max*` down, `pattern`/`format`/`multipleOf`/`uniqueItems`/`oneOf`/`allOf` added) | breaking | ok | `constraint.tightened` |
|
|
156
|
+
| Loosen a constraint (the reverse) | ok | breaking | `constraint.loosened` |
|
|
157
|
+
| Change `pattern`, `format` or `multipleOf` to a different value | breaking | breaking | `constraint.changed` |
|
|
158
|
+
| Close an open model (`additionalProperties: false`) | breaking | ok | `additionalProperties.closed` |
|
|
159
|
+
| Open a closed model | ok | breaking | `additionalProperties.opened` |
|
|
160
|
+
| Constrain a new tuple position (`items` array or `prefixItems`) | breaking | ok | `tuple.item.added` |
|
|
161
|
+
| Drop a tuple position | ok | breaking | `tuple.item.removed` |
|
|
162
|
+
| Add a `oneOf` / `anyOf` branch | ok | breaking | `union.branch.added` |
|
|
163
|
+
| Remove a `oneOf` / `anyOf` branch | breaking | ok | `union.branch.removed` |
|
|
164
|
+
| Add an `allOf` branch | breaking | ok | `allOf.branch.added` |
|
|
165
|
+
| Remove an `allOf` branch | ok | breaking | `allOf.branch.removed` |
|
|
166
|
+
| Replace `true` (anything) with a schema, or a schema with `false` | breaking | ok | `schema.restricted` |
|
|
167
|
+
| Replace a schema with `true`, or `false` with a schema | ok | breaking | `schema.relaxed` |
|
|
168
|
+
| Mark a node `deprecated: true` | ok | ok | `schema.deprecated` |
|
|
169
|
+
| Change a keyword we do not reason about (see below) | breaking | breaking | `keyword.changed` |
|
|
170
|
+
|
|
171
|
+
`full` is breaking when either column is breaking. `none` is never breaking.
|
|
172
|
+
|
|
173
|
+
Also handled:
|
|
174
|
+
|
|
175
|
+
- **Nesting.** `properties`, `items`, `additionalProperties` (when it is a schema), and every `oneOf` / `anyOf` / `allOf` branch are walked recursively. Paths are JSON pointers straight to the node, e.g. `/properties/lines/items/properties/sku`.
|
|
176
|
+
- **Local `$ref`.** `#/definitions/...` and `#/$defs/...` are resolved before comparing, following chains of refs. A change inside a definition is reported once, at the definition's own path, no matter how many properties use it. Recursive definitions terminate. Extracting an inline object into a definition, or renaming a definition, with identical content is not a change.
|
|
177
|
+
- **External `$ref`.** A ref to another file (`./common.json#/Address`) cannot be followed. If it changes it is reported as `keyword.changed`, breaking both ways. If it is unchanged it is not a change.
|
|
178
|
+
- **Branch matching.** `oneOf` / `anyOf` / `allOf` branches are matched by content first, so reordering is not a change. Remaining branches are paired by position and walked, so an edit inside one branch is reported at that branch's path.
|
|
179
|
+
- **Annotations.** `title`, `description`, `examples`, `$comment`, `$id`, `$schema` and `default` changes produce no ops. Reordering `properties`, `required` or `enum` entries is not a change. `deprecated: true` is the one annotation that is reported, as a non-breaking `schema.deprecated` op, because catalogs care about it.
|
|
180
|
+
- **`format`.** Treated as a constraint. In draft-07 validators it is, but in 2020-12 it is an annotation unless the validator opts in, so a new `format` on a 2020-12 schema may be flagged more strictly than your validator enforces.
|
|
181
|
+
- **Keywords we do not reason about.** `patternProperties`, `propertyNames`, `not`, `if` / `then` / `else`, `dependentRequired`, `dependentSchemas`, `dependencies`, `contains`, `minContains`, `maxContains`, `unevaluatedProperties`, `unevaluatedItems`, `additionalItems`. A change to any of these is reported as `keyword.changed` and treated as breaking in both directions, because refusing to call it safe is better than a silent pass.
|
|
182
|
+
|
|
183
|
+
The table above is the intent. The source of truth is the rules table in [`src/json-schema/rules.ts`](./src/json-schema/rules.ts) and the scenarios in [`src/test/json-schema/json-schema.test.ts`](./src/test/json-schema/json-schema.test.ts). When they disagree with this README, fix the README.
|
|
184
|
+
|
|
185
|
+
### Other formats
|
|
186
|
+
|
|
187
|
+
Avro, Protobuf and others are not compared yet. A change to their schema file is still reported by hash, with a `null` verdict.
|
|
188
|
+
|
|
189
|
+
## Development
|
|
190
|
+
|
|
191
|
+
```bash
|
|
192
|
+
pnpm --filter @eventcatalog/diff run test
|
|
193
|
+
pnpm --filter @eventcatalog/diff run build
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
### Adding a JSON Schema rule
|
|
197
|
+
|
|
198
|
+
1. Add the change kind to `JsonSchemaChangeKind` in `src/json-schema/types.ts`.
|
|
199
|
+
2. Emit it from the walker in `src/json-schema/compare.ts`.
|
|
200
|
+
3. Add its row to the rules table in `src/json-schema/rules.ts`.
|
|
201
|
+
4. Add a test under each strategy it affects in `src/test/json-schema/json-schema.test.ts`, with the before and after schema written out in full.
|
|
202
|
+
5. Update the table in this README.
|
|
203
|
+
|
|
204
|
+
### Reproducing a user-reported schema problem
|
|
205
|
+
|
|
206
|
+
Open `src/test/json-schema/json-schema.test.ts`, find the strategy the user runs, copy the closest test, paste their before and after schema, and state the verdict you expect. Run the tests. Fix the rule. The test stays as a regression check.
|
|
207
|
+
|
|
208
|
+
### Scenario fixtures
|
|
209
|
+
|
|
210
|
+
End-to-end tests are data driven. Each folder under `src/test/fixtures/scenarios/` holds a baseline index (`a.json`), a candidate index (`b.json`), the expected diff (`expected.json`) and optional `options.json`. Indexes are validated with the SDK's `parseIndex` when loaded, so fixtures cannot drift from what `buildIndex` emits. The scenario runner loads every folder, so adding a case is just adding a folder.
|