setspec 0.2.0__tar.gz → 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. {setspec-0.2.0 → setspec-0.4.0}/.github/workflows/ci.yml +53 -0
  2. {setspec-0.2.0 → setspec-0.4.0}/CHANGELOG.md +131 -0
  3. setspec-0.4.0/PHASE4_ISSUES.md +99 -0
  4. {setspec-0.2.0 → setspec-0.4.0}/PKG-INFO +14 -8
  5. {setspec-0.2.0 → setspec-0.4.0}/README.md +11 -7
  6. {setspec-0.2.0 → setspec-0.4.0}/docs/README.md +1 -0
  7. {setspec-0.2.0 → setspec-0.4.0}/docs/packages/setspec/spec.md +5 -3
  8. setspec-0.4.0/docs/prompts-adoption-checklist.md +86 -0
  9. setspec-0.4.0/docs/schemas.md +195 -0
  10. {setspec-0.2.0 → setspec-0.4.0}/pyproject.toml +7 -0
  11. {setspec-0.2.0 → setspec-0.4.0}/requirements/ci.lock +145 -0
  12. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/__about__.py +1 -1
  13. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/__init__.py +17 -0
  14. setspec-0.4.0/src/setspec/artifacts.py +341 -0
  15. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/benchmark/v1.py +6 -5
  16. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/capability/v1.py +4 -4
  17. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/envelope.py +22 -29
  18. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/goal/v1.py +4 -4
  19. setspec-0.4.0/src/setspec/goldens/benchmark.calibration_report/1.0/full.json +58 -0
  20. setspec-0.4.0/src/setspec/goldens/benchmark.calibration_report/1.0/gate_failed.json +41 -0
  21. setspec-0.4.0/src/setspec/goldens/benchmark.calibration_report/1.0/minimal.json +13 -0
  22. setspec-0.4.0/src/setspec/goldens/benchmark.evidence_bundle/1.0/full.json +218 -0
  23. setspec-0.4.0/src/setspec/goldens/benchmark.evidence_bundle/1.0/minimal.json +4 -0
  24. setspec-0.4.0/src/setspec/goldens/benchmark.evidence_bundle/1.0/unsupported.json +67 -0
  25. setspec-0.4.0/src/setspec/goldens/benchmark.goal_pack/1.0/full.json +69 -0
  26. setspec-0.4.0/src/setspec/goldens/benchmark.goal_pack/1.0/minimal.json +25 -0
  27. setspec-0.4.0/src/setspec/goldens/benchmark.goal_pack/1.0/starter_unforked.json +37 -0
  28. setspec-0.4.0/src/setspec/goldens/benchmark.result/1.0/full.json +208 -0
  29. setspec-0.4.0/src/setspec/goldens/benchmark.result/1.0/minimal.json +52 -0
  30. setspec-0.4.0/src/setspec/goldens/benchmark.result/1.0/unsupported.json +129 -0
  31. setspec-0.4.0/src/setspec/goldens/benchmark.run_summary/1.0/full.json +130 -0
  32. setspec-0.4.0/src/setspec/goldens/benchmark.run_summary/1.0/minimal.json +33 -0
  33. setspec-0.4.0/src/setspec/goldens/benchmark.run_summary/1.0/unsupported.json +80 -0
  34. setspec-0.4.0/src/setspec/goldens/capability.evidence/1.0/full.json +83 -0
  35. setspec-0.4.0/src/setspec/goldens/capability.evidence/1.0/goal.json +104 -0
  36. setspec-0.4.0/src/setspec/goldens/capability.evidence/1.0/minimal.json +25 -0
  37. setspec-0.4.0/src/setspec/goldens/capability.evidence/1.0/unsupported.json +61 -0
  38. setspec-0.4.0/src/setspec/goldens/machine.profile/1.0/full.json +44 -0
  39. setspec-0.4.0/src/setspec/goldens/machine.profile/1.0/minimal.json +9 -0
  40. setspec-0.4.0/src/setspec/goldens/machine.profile/1.0/unsupported.json +27 -0
  41. setspec-0.4.0/src/setspec/goldens/model.identity/1.0/full.json +34 -0
  42. setspec-0.4.0/src/setspec/goldens/model.identity/1.0/minimal.json +7 -0
  43. setspec-0.4.0/src/setspec/goldens/model.identity/1.0/unsupported.json +27 -0
  44. setspec-0.4.0/src/setspec/goldens/prompt.manifest/1.0/full.json +19 -0
  45. setspec-0.4.0/src/setspec/goldens/prompt.manifest/1.0/minimal.json +14 -0
  46. setspec-0.4.0/src/setspec/goldens/prompt.record/1.0/full.json +54 -0
  47. setspec-0.4.0/src/setspec/goldens/prompt.record/1.0/minimal.json +10 -0
  48. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/machine/v1.py +1 -1
  49. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/model/v1.py +9 -6
  50. setspec-0.4.0/src/setspec/prompts.py +940 -0
  51. setspec-0.4.0/src/setspec/schemas/benchmark.calibration_report/1.0.json +264 -0
  52. setspec-0.4.0/src/setspec/schemas/benchmark.evidence_bundle/1.0.json +724 -0
  53. setspec-0.4.0/src/setspec/schemas/benchmark.goal_pack/1.0.json +267 -0
  54. setspec-0.4.0/src/setspec/schemas/benchmark.result/1.0.json +1358 -0
  55. setspec-0.4.0/src/setspec/schemas/benchmark.run_summary/1.0.json +819 -0
  56. setspec-0.4.0/src/setspec/schemas/capability.evidence/1.0.json +693 -0
  57. setspec-0.4.0/src/setspec/schemas/machine.profile/1.0.json +323 -0
  58. setspec-0.4.0/src/setspec/schemas/model.identity/1.0.json +299 -0
  59. setspec-0.4.0/src/setspec/schemas/prompt.manifest/1.0.json +37 -0
  60. setspec-0.4.0/src/setspec/schemas/prompt.record/1.0.json +70 -0
  61. setspec-0.4.0/tests/contract/test_cross_version.py +199 -0
  62. setspec-0.4.0/tests/contract/test_goldens.py +299 -0
  63. setspec-0.4.0/tests/contract/test_schema_snapshots.py +261 -0
  64. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_envelope.py +10 -5
  65. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_payloads_goal.py +3 -3
  66. setspec-0.4.0/tests/unit/test_prompts.py +350 -0
  67. setspec-0.2.0/src/setspec/artifacts.py +0 -4
  68. setspec-0.2.0/src/setspec/goldens/.gitkeep +0 -0
  69. setspec-0.2.0/src/setspec/schemas/.gitkeep +0 -0
  70. setspec-0.2.0/tests/contract/test_cross_version.py +0 -4
  71. setspec-0.2.0/tests/contract/test_goldens.py +0 -4
  72. setspec-0.2.0/tests/contract/test_schema_snapshots.py +0 -4
  73. {setspec-0.2.0 → setspec-0.4.0}/.editorconfig +0 -0
  74. {setspec-0.2.0 → setspec-0.4.0}/.github/workflows/release.yml +0 -0
  75. {setspec-0.2.0 → setspec-0.4.0}/.gitignore +0 -0
  76. {setspec-0.2.0 → setspec-0.4.0}/.importlinter +0 -0
  77. {setspec-0.2.0 → setspec-0.4.0}/.pre-commit-config.yaml +0 -0
  78. {setspec-0.2.0 → setspec-0.4.0}/CONTRIBUTING.md +0 -0
  79. {setspec-0.2.0 → setspec-0.4.0}/LICENSE +0 -0
  80. {setspec-0.2.0 → setspec-0.4.0}/SECURITY.md +0 -0
  81. {setspec-0.2.0 → setspec-0.4.0}/docs/packages/setspec/development-plan.md +0 -0
  82. {setspec-0.2.0 → setspec-0.4.0}/requirements/README.md +0 -0
  83. {setspec-0.2.0 → setspec-0.4.0}/requirements/release.in +0 -0
  84. {setspec-0.2.0 → setspec-0.4.0}/requirements/release.lock +0 -0
  85. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/base.py +0 -0
  86. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/error/v1.py +0 -0
  87. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/errors.py +0 -0
  88. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/event/v1.py +0 -0
  89. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/metrics.py +0 -0
  90. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/provenance.py +0 -0
  91. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/py.typed +0 -0
  92. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/serialization.py +0 -0
  93. {setspec-0.2.0 → setspec-0.4.0}/src/setspec/vocabulary.py +0 -0
  94. {setspec-0.2.0 → setspec-0.4.0}/tests/conftest.py +0 -0
  95. {setspec-0.2.0 → setspec-0.4.0}/tests/contract/test_version_negotiation.py +0 -0
  96. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_base.py +0 -0
  97. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_errors.py +0 -0
  98. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_events.py +0 -0
  99. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_metrics.py +0 -0
  100. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_payloads_benchmark.py +0 -0
  101. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_payloads_capability.py +0 -0
  102. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_serialization.py +0 -0
  103. {setspec-0.2.0 → setspec-0.4.0}/tests/unit/test_vocabulary.py +0 -0
@@ -91,6 +91,44 @@ jobs:
91
91
  - run: pip install . --no-deps
92
92
  - run: pytest -m contract
93
93
 
94
+ # The three published guarantees get their own jobs rather than sharing `contracts`' single
95
+ # red mark, because each answers a different question and the answers are acted on differently:
96
+ # a snapshot diff means "bump the version", a golden failure means "the artifact is wrong", and
97
+ # a cross-version failure means "a consumer at another version just broke".
98
+ schema-snapshot-diff:
99
+ runs-on: ubuntu-latest
100
+ steps:
101
+ - uses: actions/checkout@v4
102
+ - uses: actions/setup-python@v5
103
+ with: { python-version: "3.12" }
104
+ - run: pip install --require-hashes -r requirements/ci.lock
105
+ - run: pip install . --no-deps
106
+ # ADR-0009 rule 7: a schema change without a version bump fails CI. The test regenerates
107
+ # every published schema from its writer model and diffs it against the committed snapshot.
108
+ - run: pytest tests/contract/test_schema_snapshots.py -m contract
109
+
110
+ golden-validation:
111
+ runs-on: ubuntu-latest
112
+ steps:
113
+ - uses: actions/checkout@v4
114
+ - uses: actions/setup-python@v5
115
+ with: { python-version: "3.12" }
116
+ - run: pip install --require-hashes -r requirements/ci.lock
117
+ - run: pip install . --no-deps
118
+ - run: pytest tests/contract/test_goldens.py -m contract
119
+
120
+ cross-version-compatibility:
121
+ runs-on: ubuntu-latest
122
+ steps:
123
+ - uses: actions/checkout@v4
124
+ - uses: actions/setup-python@v5
125
+ with: { python-version: "3.12" }
126
+ - run: pip install --require-hashes -r requirements/ci.lock
127
+ - run: pip install . --no-deps
128
+ # Testing Standards §8.3, proven by the repository that makes the promise: a v1.0 reader
129
+ # accepts a synthetic v1.1 golden without loss and refuses a synthetic v2.0 by major.
130
+ - run: pytest tests/contract/test_cross_version.py -m contract
131
+
94
132
  security:
95
133
  runs-on: ubuntu-latest
96
134
  steps:
@@ -148,3 +186,18 @@ jobs:
148
186
  marker = pathlib.Path(setspec.__file__).parent / 'py.typed'
149
187
  assert marker.is_file(), 'py.typed missing from the installed wheel'
150
188
  "
189
+ # The schemas and goldens are package data (ADR-0009 rule 7), and package data is exactly
190
+ # what a build is most likely to leave out silently: every test in this repository passes
191
+ # against the source tree whether or not the files reach the wheel. This is the only job
192
+ # that would notice, which is why it asserts against the *installed* distribution.
193
+ - name: schemas and goldens ship in the wheel
194
+ run: |
195
+ python -c "
196
+ from setspec import PUBLISHED_SCHEMAS, golden_payloads, json_schema_for
197
+ for schema, versions in PUBLISHED_SCHEMAS.items():
198
+ for version in versions:
199
+ assert json_schema_for(schema, version)['title'] == schema
200
+ goldens = golden_payloads(schema, version)
201
+ assert len(goldens) >= 3, f'{schema} {version} ships {len(goldens)} goldens'
202
+ print(f'{len(PUBLISHED_SCHEMAS)} schemas with goldens loaded from the installed wheel')
203
+ "
@@ -7,7 +7,122 @@ packaging and release standards §3.
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.4.0] — 2026-08-29
11
+
12
+ ### Added
13
+
14
+ - **`setspec.prompts`** (Phase 5): prompt records and prompt packs, moved from
15
+ `freeweight.services.prompts` — `PromptRecord`, `PromptLibrary`, `load_record()`, `load_pack()`,
16
+ `build_manifest()`/`write_manifest()`, and the pack's three hashes
17
+ (`prompt_record_hash`/`prompt_subset_hash`/`pack_hash`, ADR-0028). Rendering runs through a
18
+ sandboxed, loader-less Jinja2 `Environment` with `StrictUndefined` (prompt-management-standards
19
+ §2.1): a referenced-but-unsupplied variable is an error, a template cannot reach the filesystem,
20
+ and dunder attribute access is refused. `load_pack()`'s `root` parameter is required — the
21
+ package carries no opinion about an application's pack layout; an application wanting a default
22
+ location supplies its own thin wrapper (FreeWeight's `freeweight.services.prompts.load_pack`
23
+ keeps its old default this way).
24
+
25
+ FreeWeight's own prompt pack adopts this module in the same change (pulling its P12 forward):
26
+ `freeweight.services.prompts` is now a re-export shim over `setspec.prompts`, verified to
27
+ produce a byte-identical `pack_hash` and per-record hash for every shipped prompt before and
28
+ after the move, with FreeWeight's own gate green afterwards. Two applications reading the same
29
+ pack format from two processes is the scenario ADR-0028's hashes exist for; the FreeWeight move
30
+ is the first proof of it.
31
+
32
+ - **JSON Schema for `prompt.record` and `prompt.manifest`**, committed as package data under
33
+ `setspec/schemas/prompt.record/1.0.json` and `setspec/schemas/prompt.manifest/1.0.json`, with
34
+ golden examples under `setspec/goldens/prompt.{record,manifest}/1.0/{full,minimal}.json`. Both
35
+ are hand-authored rather than generated: `PromptRecord` is a plain dataclass, not a pydantic
36
+ envelope payload — a prompt record is a pack file on disk, never wrapped in a `SchemaEnvelope` —
37
+ so neither schema is registered in `PUBLISHED_SCHEMAS`/reachable through `json_schema_for()`,
38
+ and the goldens are illustrative examples rather than inputs to the envelope contract-test
39
+ matrix in `tests/contract/test_goldens.py`. Every golden round-trips through the real
40
+ `load_pack()` loader (not just schema validation) and its `pack_sha256` is the module's own
41
+ computed hash, not a hand-typed value.
42
+
43
+ ### Changed
44
+
45
+ - Bumped to 0.4.0 for the `setspec.prompts` addition (packaging standards §3: additive,
46
+ non-breaking).
47
+
48
+ ## [0.3.0] — 2026-08-28
49
+
10
50
  ### Added
51
+
52
+ - **The v1.0 freeze.** `DRAFT_SCHEMAS` is now empty. Every registered payload type —
53
+ `model.identity`, `machine.profile`, `benchmark.result`, `benchmark.run_summary`,
54
+ `capability.evidence`, `benchmark.evidence_bundle`, `benchmark.goal_pack` and
55
+ `benchmark.calibration_report` — is a published `1.0` that may no longer be reshaped in place.
56
+ FreeWeight's real output (P6 onward, and P8A–8B for the two goal payloads) demanded no field
57
+ change, so the freeze is a promotion rather than a correction; the one field the pass did add,
58
+ `MetricValueFields.metric_key`, landed before it and is frozen *with* the schema.
59
+
60
+ The set survives its own emptiness deliberately: it is the mechanism a future draft uses, and
61
+ deleting it would leave the next provisional payload type either shipping silently provisional or
62
+ inventing a second way to say so.
63
+
64
+ - **Generated JSON Schema for every published version**, committed as package data under
65
+ `setspec/schemas/<schema>/<version>.json` and reachable through `json_schema_for()`. Generated
66
+ from the **writer** (`extra="forbid"`) half, so a published document says
67
+ `additionalProperties: false` — which is the contract API standards §7 rule 5 states and the one
68
+ a producer's test suite asserts its output against. The reader policy stays where it belongs, in
69
+ the reader: `load_envelope` accepts an unknown minor by comparing majors, never by loosening a
70
+ version's schema to admit fields it cannot describe.
71
+
72
+ Every `$ref` resolves inside the document's own `$defs`. Nothing is fetched (spec §14).
73
+
74
+ - **25 golden payloads**, three or more per version, under
75
+ `setspec/goldens/<schema>/<version>/<name>.json` and reachable through `golden_payloads()` and
76
+ `golden_names()`. Each version ships a `minimal` (only required fields), a `full` (every field
77
+ populated) and, wherever the payload carries a measurement at all, an `unsupported`-heavy example.
78
+ `capability.evidence` adds `goal`, a calibrated `user.noir_tech_voice` record with the whole
79
+ goal-sourced group populated; `benchmark.goal_pack` adds `starter_unforked`; and
80
+ `benchmark.calibration_report` adds `gate_failed` — the outcome `capability.evidence`
81
+ deliberately cannot express, because a goal below its gate emits no evidence record at all.
82
+
83
+ Which versions need an `unsupported` golden is derived from the published schema rather than
84
+ listed in the test: a payload whose document contains no `{"const": "unsupported"}` branch has no
85
+ measurement field to leave unsupported, so `benchmark.goal_pack` and
86
+ `benchmark.calibration_report` are excused by the artifact itself rather than by an exception
87
+ someone maintains.
88
+
89
+ - **`setspec.artifacts`** — `PUBLISHED_SCHEMAS`, `json_schema_for()`, `golden_payloads()`,
90
+ `golden_names()`, `payload_pair()`, `build_json_schema()` and `render_schema_document()`. The
91
+ first four are re-exported from `setspec` directly, matching spec §7's "schema artefacts" group:
92
+ an accessor that *returns* a version's artifacts is not itself versioned.
93
+
94
+ `PUBLISHED_SCHEMAS` records **exact** versions, unlike `SUPPORTED_SCHEMAS`, which records
95
+ supported majors. An artifact is a file that either exists or does not, so "highest minor known"
96
+ is the wrong shape for it — and a contract test asserts the two agree, so a schema that can be
97
+ negotiated but not validated, or validated but not negotiated, fails the build.
98
+
99
+ - **Three contract jobs, one per published guarantee.** `schema-snapshot-diff` regenerates every
100
+ schema from its models and diffs it against the committed snapshot — ADR-0009 rule 7, *a schema
101
+ change without a version bump fails CI*, made mechanical and proven by a deliberate test-only
102
+ mutation. `golden-validation` checks every golden against its writer model, its reader model, the
103
+ published JSON Schema and a canonical round trip. `cross-version-compatibility` builds a synthetic
104
+ `1.1` document from each `full` golden and asserts a `1.0` reader accepts it, preserves the field
105
+ it does not know, and re-exports it intact — then builds a synthetic `2.0` and asserts the refusal
106
+ names both versions. They are separate jobs rather than steps because the three failures are acted
107
+ on differently: bump the version, fix the artifact, or go and warn a consumer.
108
+
109
+ - **`install-check` now asserts the package data reaches the wheel.** Every test in this repository
110
+ passes against the source tree whether or not the schemas and goldens are built into the
111
+ distribution, so this job is the only one that would notice. It loads all eight schemas and their
112
+ goldens out of the *installed* wheel.
113
+
114
+ - **`docs/schemas.md`** — the human-readable catalogue: what ships and where, every payload type
115
+ with its required-field count and its goldens, how to consume the artifacts from a repository that
116
+ shares no code with this one, and a table of every cross-field rule the JSON Schema **cannot**
117
+ express and the models therefore still own.
118
+
119
+ - **`jsonschema` and `types-jsonschema` in the `dev` extra.** Test-only, and justified by the
120
+ assertion they exist for: a golden is validated against the *published document* a non-Python
121
+ consumer receives, not only against the model that generated it. Without a real validator that
122
+ check would be a hand-rolled subset of draft 2020-12 maintained in this repository — the thing
123
+ under test also serving as the thing testing it. Nothing under `src/` imports either, and the base
124
+ install still pulls in no validator. `requirements/ci.lock` regenerated.
125
+
11
126
  - **`metric_key` on `MetricValueFields`.** The model declared value, unit, aggregation, direction,
12
127
  sample count and dispersion — and nothing saying *which metric it is*. Both
13
128
  `BenchmarkResultFields.metrics` and `BenchmarkRunSummaryFields.aggregate_metrics` carry sequences
@@ -58,6 +173,14 @@ packaging and release standards §3.
58
173
 
59
174
  ### Changed
60
175
 
176
+ - The five payload modules' status notes read **frozen (`1.0`)** rather than **draft (`1.0`)**, and
177
+ say what the freeze binds them to: a new optional field is a minor bump, and a removal, rename,
178
+ retype or tightening is a major — never an edit in place.
179
+ - The two tests that asserted the draft state now assert the frozen one
180
+ (`test_no_registered_schema_is_still_draft`, `test_the_schema_is_frozen`). They are the same
181
+ guarantee read from the other side, and leaving them asserting `DRAFT_SCHEMAS == SUPPORTED_SCHEMAS`
182
+ would have made the freeze a change the test suite refused.
183
+
61
184
  - A bare reserved root is now refused by `validate_capability` and reported `False` by
62
185
  `is_known_capability`. `user` is a namespace, not a capability: a payload claiming
63
186
  `capability_id: "user"` has lost the identity that is the entire point of the namespace.
@@ -71,6 +194,7 @@ packaging and release standards §3.
71
194
  declaring `1.2` or later is unaffected. Producers on `1.1` must emit roots that exist at `1.1`.
72
195
 
73
196
  ### Fixed
197
+
74
198
  - **Coverage measured a directory nothing imports.** `[tool.coverage.run] source` named
75
199
  `src/setspec`, the source *path*, while CI installs the built distribution — so the moment the
76
200
  jobs stopped using an editable install, coverage reported **0 %** and failed the 95 % floor for a
@@ -89,6 +213,13 @@ packaging and release standards §3.
89
213
  (`baseaicore`, `modelrack`, `sweatmeter`). Without it, `mypy --strict` in a consuming repository
90
214
  cannot see this package's types at all and treats every import from it as untyped.
91
215
 
216
+ ### Known gaps
217
+
218
+ - `event.envelope` and `error.envelope` (Phase 3) and `prompt.record` / `prompt.manifest`
219
+ (Phase 5) are **not** part of this freeze, because they do not exist yet: a schema is frozen by
220
+ being published, and an unwritten one has nothing to publish. The freeze covers the eight payload
221
+ types that cross an application boundary today, not the eleven ADR-0009 anticipates.
222
+
92
223
  ## [0.2.0] — 2026-08-23
93
224
 
94
225
  Phase 2 of the [development plan](docs/packages/setspec/development-plan.md): provisional
@@ -0,0 +1,99 @@
1
+ # Phase 4 — issues to address
2
+
3
+ Written at the end of Phase 4 (freeze v1.0, publish schemas and goldens; `setspec 0.3.0`). Each
4
+ entry is something a later phase, a docs change, a consumer or a release step has to resolve.
5
+ Nothing here blocks Phase 4's acceptance criteria; everything here would become a defect if it were
6
+ forgotten.
7
+
8
+ ---
9
+
10
+ ## Status — 2026-08-28
11
+
12
+ | # | Issue | Status |
13
+ |---|---|---|
14
+ | 1 | The freeze covers eight payload types, not ADR-0009's eleven | **Open — by design.** Phases 3 and 5 own the other three. |
15
+ | 2 | Goldens are authored inputs, but a SetSpec writer emits every declared key | **Needs a decision.** |
16
+ | 3 | `ContributingMetricFields.metric_key` carries no pattern while `MetricValueFields.metric_key` does | **Needs a decision.** |
17
+ | 4 | Cross-field rules are invisible in the published JSON Schema | **Open — documented.** `docs/schemas.md` §4 lists every one. |
18
+ | 5 | `jsonschema` joined the `dev` extra; `requirements/ci.lock` regenerated | **Closed** — verify `pip-audit` in CI. |
19
+ | 6 | The release is not tagged or published yet | **Owner: you.** Commands below. |
20
+
21
+ ---
22
+
23
+ ## 1. The freeze covers eight payload types, not eleven
24
+
25
+ ADR-0009 lists eleven initial payload types. `event.envelope` and `error.envelope` (Phase 3) and
26
+ `prompt.record` / `prompt.manifest` (Phase 5) are not written, so they are not frozen — a schema is
27
+ frozen by being published, and an unwritten one has nothing to publish. `DRAFT_SCHEMAS` is empty
28
+ **now**; when Phase 3 or 5 lands a new payload type it should re-enter that set until its own
29
+ freeze, which is exactly the mechanism the set survives its emptiness for. The changelog's *Known
30
+ gaps* section says the same. Nothing to do until those phases start, except not to read "empty"
31
+ as "finished".
32
+
33
+ ## 2. Goldens are authored inputs, but a SetSpec writer emits every declared key
34
+
35
+ `dump_envelope(CapabilityEvidenceOut(...))` serialises through `model_dump()`, which writes every
36
+ declared field — a non-goal record carries `goal_hash: null`, `uncalibrated: false`, and so on. The
37
+ `capability.evidence/1.0/full` golden was authored the other way round: it omits the goal group
38
+ entirely, because "fully populated" was read as "every field a *non-goal* record populates". Both
39
+ forms validate under both models, so the contract is not broken, but a producer's structural test
40
+ ("my keys match the full golden's keys") fails against the file and passes against the file *as a
41
+ SetSpec writer would dump it*. FreeWeight's contract test compares against the latter.
42
+
43
+ **Decision needed:** should goldens be committed *as written* (every declared key present, nulls
44
+ included), so that "matches the golden structurally" is a byte-level statement? If yes, regenerate
45
+ the 25 goldens through their `Out` models once and add a contract test that a golden equals its own
46
+ writer-dumped form. If no, say so in `docs/schemas.md` §3 so consumers compare against the dumped
47
+ form, as FreeWeight now does.
48
+
49
+ ## 3. `ContributingMetricFields.metric_key` has no pattern
50
+
51
+ `MetricValueFields.metric_key` is constrained to lower snake case, dot-separable
52
+ (`^[a-z][a-z0-9_]*(\.[a-z][a-z0-9_]*)*$`). `ContributingMetricFields.metric_key` — the key inside
53
+ `capability.evidence.contributing_metrics` — is only `min_length=1`. FreeWeight writes
54
+ `<suite_key>.<metric_key>` there (`native.tool_use.task_success`), `criterion.<key>` for a goal's
55
+ own record and `goal.<slug>.composite_score` for a goal contributing to a shipped capability; all
56
+ three happen to satisfy the stricter pattern. Tightening the field is a **major** change now that
57
+ `1.0` is frozen, so it can only land as `capability.evidence 2.0`, and only if a consumer ever
58
+ needs the guarantee. Recorded so the asymmetry is a decision rather than an accident.
59
+
60
+ ## 4. Cross-field rules are invisible in the published JSON Schema
61
+
62
+ Pydantic renders types, ranges, patterns and required keys; it cannot render a `model_validator`.
63
+ `runtime_profile_hash` agreeing with its profile, `score_method_mix` summing to one,
64
+ `measured_at ≤ computed_at`, the goal group's five coherence rules — every one is enforced by the
65
+ models only. `docs/schemas.md` §4 tables them, and every golden is validated against *both* the
66
+ schema and the model for exactly this reason. A non-Python consumer that needs those rules needs a
67
+ second validation step this package does not ship. Nothing to do unless such a consumer appears.
68
+
69
+ ## 5. `jsonschema` in the `dev` extra
70
+
71
+ Added so the golden contract test validates against the *published* document with a real
72
+ draft-2020-12 validator rather than a hand-rolled subset. Test only — nothing under `src/`
73
+ imports it and the base install pulls in no validator. `requirements/ci.lock` was regenerated with
74
+ `pip-compile --generate-hashes` (jsonschema 4.26.0, jsonschema-specifications, referencing,
75
+ rpds-py, attrs, types-jsonschema). `release.lock` is unchanged. CI's `security` job audits both
76
+ locks; the first run after this lands is the one to watch.
77
+
78
+ ## 6. Tagging and publishing 0.3.0
79
+
80
+ Nothing was tagged or published. `__about__.py` says `0.3.0` and the changelog carries a dated
81
+ `[0.3.0]` section. The release procedure (packaging standards §6), from the SetSpec repository:
82
+
83
+ ```bash
84
+ cd ~/ai/suite/py/SetSpec
85
+ git add -A
86
+ git commit -m "feat(setspec): freeze v1.0, publish JSON Schema and goldens (Phase 4, 0.3.0)"
87
+ git push origin main
88
+ # wait for CI to be green on main, then:
89
+ git tag -a v0.3.0 -m "setspec 0.3.0 — v1.0 contracts frozen; JSON Schema and goldens published"
90
+ git push origin v0.3.0
91
+ # release.yml builds, tests the wheel, publishes via Trusted Publishing and creates the release.
92
+ # Verify:
93
+ python -m venv /tmp/setspec-check && /tmp/setspec-check/bin/pip install setspec==0.3.0 && \
94
+ /tmp/setspec-check/bin/python -c "from setspec import PUBLISHED_SCHEMAS, golden_payloads, SchemaVersion; \
95
+ print(len(PUBLISHED_SCHEMAS), len(golden_payloads('capability.evidence', SchemaVersion(1, 0))))"
96
+ ```
97
+
98
+ FreeWeight's `pyproject.toml` already pins `setspec>=0.3,<0.4`; its `install-check` job cannot
99
+ resolve until this release is on PyPI.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: setspec
3
- Version: 0.2.0
3
+ Version: 0.4.0
4
4
  Summary: Every versioned data contract that crosses an application boundary: benchmark results, capability evidence, event/error envelopes, prompt records.
5
5
  Project-URL: Homepage, https://github.com/JPKell/SetSpec
6
6
  Project-URL: Documentation, https://github.com/JPKell/SetSpec/tree/main/docs
@@ -21,25 +21,31 @@ Requires-Dist: jinja2<4,>=3.1
21
21
  Requires-Dist: pydantic<3,>=2.9
22
22
  Provides-Extra: dev
23
23
  Requires-Dist: import-linter<3,>=2.0; extra == 'dev'
24
+ Requires-Dist: jsonschema<5,>=4.23; extra == 'dev'
24
25
  Requires-Dist: mypy<2,>=1.11; extra == 'dev'
25
26
  Requires-Dist: pytest-cov<6,>=5; extra == 'dev'
26
27
  Requires-Dist: pytest-randomly<4,>=3; extra == 'dev'
27
28
  Requires-Dist: pytest<10,>=9.0.3; extra == 'dev'
28
29
  Requires-Dist: respx<1,>=0.21; extra == 'dev'
29
30
  Requires-Dist: ruff<1,>=0.6; extra == 'dev'
31
+ Requires-Dist: types-jsonschema<5,>=4.23; extra == 'dev'
30
32
  Description-Content-Type: text/markdown
31
33
 
32
34
  # SetSpec
33
35
 
34
36
  Every versioned data contract that crosses an application boundary: benchmark results, capability evidence, event/error envelopes, prompt records.
35
37
 
36
- **Status:** `0.2.0` — Phases 1–2 complete. The envelope, version negotiation and serialization core
37
- are implemented and tested; `model.identity`, `machine.profile`, `benchmark.result`,
38
- `benchmark.run_summary`, `capability.evidence` and `benchmark.evidence_bundle` are registered in
39
- `SUPPORTED_SCHEMAS`, but **draft**: Phase 4 may still reshape a field once FreeWeight has produced
40
- real results against these payloads. Which schemas are still provisional is readable at runtime
41
- from `setspec.DRAFT_SCHEMAS`, not only stated here — freezing one is a deletion from that set.
42
- Event and error envelopes arrive in Phase 3. See the
38
+ **Status:** `0.3.0` — Phases 1–2, 3A and 4 complete, and **the v1.0 contracts are frozen.** Eight
39
+ payload types are published at `1.0` — `model.identity`, `machine.profile`, `benchmark.result`,
40
+ `benchmark.run_summary`, `capability.evidence`, `benchmark.evidence_bundle`,
41
+ `benchmark.goal_pack` and `benchmark.calibration_report` — each with generated JSON Schema and at
42
+ least three golden payloads shipped as package data. `setspec.DRAFT_SCHEMAS` is empty, which is
43
+ where the freeze is readable at runtime rather than only stated here; from now on a new optional
44
+ field is a minor bump and anything else is a major, enforced by a snapshot diff in CI.
45
+
46
+ The [schema catalogue](docs/schemas.md) lists every payload type, its artifacts, and the
47
+ cross-field rules the JSON Schema cannot express. Event and error envelopes (Phase 3) and prompt
48
+ records (Phase 5) are not yet written and are therefore not part of the freeze. See the
43
49
  [development plan](docs/packages/setspec/development-plan.md) for what each phase adds.
44
50
 
45
51
  Part of the **Local AI Suite**.
@@ -2,13 +2,17 @@
2
2
 
3
3
  Every versioned data contract that crosses an application boundary: benchmark results, capability evidence, event/error envelopes, prompt records.
4
4
 
5
- **Status:** `0.2.0` — Phases 1–2 complete. The envelope, version negotiation and serialization core
6
- are implemented and tested; `model.identity`, `machine.profile`, `benchmark.result`,
7
- `benchmark.run_summary`, `capability.evidence` and `benchmark.evidence_bundle` are registered in
8
- `SUPPORTED_SCHEMAS`, but **draft**: Phase 4 may still reshape a field once FreeWeight has produced
9
- real results against these payloads. Which schemas are still provisional is readable at runtime
10
- from `setspec.DRAFT_SCHEMAS`, not only stated here — freezing one is a deletion from that set.
11
- Event and error envelopes arrive in Phase 3. See the
5
+ **Status:** `0.3.0` — Phases 1–2, 3A and 4 complete, and **the v1.0 contracts are frozen.** Eight
6
+ payload types are published at `1.0` — `model.identity`, `machine.profile`, `benchmark.result`,
7
+ `benchmark.run_summary`, `capability.evidence`, `benchmark.evidence_bundle`,
8
+ `benchmark.goal_pack` and `benchmark.calibration_report` — each with generated JSON Schema and at
9
+ least three golden payloads shipped as package data. `setspec.DRAFT_SCHEMAS` is empty, which is
10
+ where the freeze is readable at runtime rather than only stated here; from now on a new optional
11
+ field is a minor bump and anything else is a major, enforced by a snapshot diff in CI.
12
+
13
+ The [schema catalogue](docs/schemas.md) lists every payload type, its artifacts, and the
14
+ cross-field rules the JSON Schema cannot express. Event and error envelopes (Phase 3) and prompt
15
+ records (Phase 5) are not yet written and are therefore not part of the freeze. See the
12
16
  [development plan](docs/packages/setspec/development-plan.md) for what each phase adds.
13
17
 
14
18
  Part of the **Local AI Suite**.
@@ -4,5 +4,6 @@ This directory contains the maintained documentation for SetSpec.
4
4
 
5
5
  ## Documents
6
6
 
7
+ - [Schema catalogue](schemas.md) — every payload type and version, with its artifacts
7
8
  - [Development plan](packages/setspec/development-plan.md)
8
9
  - [Specification](packages/setspec/spec.md)
@@ -124,9 +124,11 @@ RESERVED_ROOTS: frozenset[str] # roots valid ONLY as a specialization; {"
124
124
  validate_capability(capability_id: str) -> CapabilityId
125
125
  is_known_capability(capability_id: str) -> bool
126
126
 
127
- # Schema artefacts
128
- json_schema_for(schema: str, version: SchemaVersion) -> dict
129
- golden_payloads(schema: str, version: SchemaVersion) -> list[dict]
127
+ # Schema artefacts (package data; every published version keeps both for the life of its major)
128
+ PUBLISHED_SCHEMAS: Mapping[str, tuple[SchemaVersion, ...]] # exact versions with committed artefacts
129
+ json_schema_for(schema: str, version: SchemaVersion) -> dict # the committed JSON Schema snapshot
130
+ golden_payloads(schema: str, version: SchemaVersion) -> list[dict] # ≥ 3 per version, by name
131
+ golden_names(schema: str, version: SchemaVersion) -> tuple[str, ...] # the same order as above
130
132
 
131
133
  # Prompts (ADR-0028) — one implementation of a determinism contract, not three
132
134
  class PromptLibrary:
@@ -0,0 +1,86 @@
1
+ # `setspec.prompts` Adoption Checklist
2
+
3
+ For an application that owns a prompt pack today (its own `PromptRecord`/`load_pack`-shaped code,
4
+ or an equivalent) and wants to move onto `setspec.prompts` instead of maintaining a second copy of
5
+ prompt-record loading, hashing and rendering. Written against the first adopter, FreeWeight P12
6
+ (§5 below is that migration, completed); the steps are what the *next* adopter — IdeaPress or
7
+ LoadCoach, whichever writes its first prompt pack second — should follow.
8
+
9
+ ## 1. What to delete outright
10
+
11
+ Anything that duplicates, byte-for-byte in behaviour, what `setspec.prompts` now provides:
12
+ `PromptRecord`/`VariableSpec`/`RenderedPrompt`/`PromptReference`/`PromptLibrary`, `load_record()`,
13
+ `load_pack()`, `build_manifest()`/`write_manifest()`, `prompt_record_hash()`/`prompt_subset_hash()`/
14
+ `pack_hash()`, and the sandboxed `StrictUndefined` Jinja2 environment. If your pack format matches
15
+ prompt-management-standards §2.1–§3 (schema_version/prompt_id/version/template/purpose/metadata,
16
+ a manifest with pack_id/pack_version/prompts/pack_sha256), there should be nothing left to keep.
17
+
18
+ ## 2. The one thing every adopter supplies itself: pack location
19
+
20
+ `load_pack(root: Path, *, override_root: Path | None = None)` takes `root` as a required
21
+ positional argument — the package has no opinion about where an application's pack lives on disk,
22
+ and no default. If your call sites currently call a zero-argument `load_pack()`, write a one-line
23
+ wrapper in your own package that supplies your pack's fixed location and forwards everything else:
24
+
25
+ ```python
26
+ from pathlib import Path
27
+ from setspec.prompts import load_pack as _setspec_load_pack, PromptLibrary
28
+
29
+ PACK_ROOT = Path(__file__).resolve().parent.parent / "prompts"
30
+
31
+ def load_pack(root: Path = PACK_ROOT, *, override_root: Path | None = None) -> PromptLibrary:
32
+ return _setspec_load_pack(root, override_root=override_root)
33
+ ```
34
+
35
+ Do this even if every current call site could be edited to pass `root` explicitly — a re-export
36
+ shim at your old import path means the migration touches one file instead of every call site, and
37
+ is the difference between a multi-file diff and a four-file one (§5).
38
+
39
+ ## 3. Behaviour to re-verify, not assume
40
+
41
+ * **Hashing is over `canonical_json`, not file bytes.** If you ever hashed a record file's raw
42
+ bytes, that is not the same value `prompt_record_hash()` produces; re-indenting a file must not
43
+ change its hash, and canonical-JSON hashing already guarantees that here.
44
+ * **Declared-and-unused is a load-time error, same as undeclared-and-used.** A variable in the
45
+ record's `variables` block that the template never references fails validation — not just the
46
+ reverse. Audit existing records for this before the first `load_pack()` call in CI.
47
+ See [`goldens/prompt.record/1.0/full.json`](../src/setspec/goldens/prompt.record/1.0/full.json)
48
+ for a record where every declared variable — including a non-string one with `min`/`max` — is
49
+ used in the template.
50
+ * **The Jinja2 environment has no loader.** `{% include %}` / `{% extends %}` fail at render time
51
+ (`PromptRenderError`), not load time — a template that used either against your old environment
52
+ needs rewriting, not just re-pointing.
53
+ * **`pack_hash()` is provenance, never a fingerprint input** (ADR-0028 §1). If your evidence
54
+ payloads folded a pack-wide hash into a capability fingerprint, stop — only
55
+ `prompt_subset_hash()` over the *specific prompts a benchmark declares* belongs there.
56
+
57
+ ## 4. Tests that must still pass, unchanged
58
+
59
+ Whatever exercises your prompt pack today — rendering determinism, unknown-variable rejection,
60
+ sandboxing (no filesystem reach, no dunder attribute access from a template) — should still pass
61
+ against `setspec.prompts` with no change to the test's assertions, only to its imports. A test that
62
+ needed a new assertion to pass found a real behaviour difference; treat that as a bug to resolve,
63
+ not a test to relax, before shipping the migration.
64
+
65
+ ## 5. Verification: FreeWeight P12 (completed)
66
+
67
+ FreeWeight adopted `setspec.prompts` in the same change that added it to `setspec` (Phase 5 pulled
68
+ FreeWeight's own P12 forward, rather than leaving `setspec.prompts` unused until a later phase).
69
+ The proof this checklist rests on:
70
+
71
+ 1. **Before**: captured the pack hash and every individual record hash from FreeWeight's original,
72
+ pre-migration `freeweight.services.prompts` — `pack_hash: sha256:b1b0ffd0a5941fee5e0013d2a826732ea02a285b229bdc006ebd6dd25ff4ceb4`
73
+ plus 18 individual record hashes spanning `benchmarks.agent.goal` through `goals.judge.rubric`.
74
+ 2. **After, direct**: loaded the same pack through `setspec.prompts.load_pack()` directly and
75
+ recomputed every hash — identical.
76
+ 3. **After, through the shim**: rewrote `freeweight/services/prompts.py` as a thin re-export over
77
+ `setspec.prompts` (§2's pattern, `PACK_ROOT` default preserved) and recomputed through
78
+ FreeWeight's own public import path — identical again.
79
+ 4. **Full gate, unchanged**: FreeWeight's complete pre-PR gate (ruff, mypy strict over 265 files,
80
+ import-linter, pytest) passed with no test assertions changed — 2297 passed, 28 skipped, 21
81
+ deselected (performance). The only non-shim source edit was 3 import statements in
82
+ `tests/security/test_goal_pack_import.py`, which reached into the module's private
83
+ `_environment()` directly for a sandboxing assertion; every other call site needed no change.
84
+
85
+ Final diff: 4 files changed. The next adopter should expect a diff of similar shape — one shim
86
+ file, a dependency bump, and an edit only where a test reached past the public API.