@topy-ai/maggie 0.7.21 → 0.7.22

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -76,6 +76,15 @@ maggie qa summary --project . --run <run-id>
76
76
  The QA workflow stores secret-free run state under `.maggie/qa-runs/` and
77
77
  keeps project-specific scenarios and evidence outside the npm package.
78
78
 
79
+ Before recording a passing user-facing assertion, lint it against runtime
80
+ evidence; source-only class or markup checks are rejected:
81
+
82
+ ```bash
83
+ maggie qa assertion-audit --project . \
84
+ --assertions .maggie/qa-assertions.json \
85
+ --output docs/qa-assertion-audit.json
86
+ ```
87
+
79
88
  Dashboard and documentation audits are provider-neutral and keep host data
80
89
  behind explicit evidence files:
81
90
 
@@ -123,6 +132,24 @@ variants use market/locale-aware canonical routes and reciprocal hreflang.
123
132
  Release preflight consumes the latest matching scenario QA run when one is
124
133
  configured and blocks incomplete evidence.
125
134
 
135
+ The provider-owned parser also supports safe catalogue and lifecycle review:
136
+
137
+ ```bash
138
+ maggie service catalogue-check "<provider-location-url>" \
139
+ --treatment "Lymphatic drainage massage"
140
+ maggie service sync-report --project .
141
+ maggie service retirement-audit --project . \
142
+ --catalogue .maggie/booking/services.json \
143
+ --evidence .maggie/booking/retirement-evidence.json
144
+ ```
145
+
146
+ `catalogue-check` avoids marketplace noise when the provider exposes an
147
+ authoritative embedded catalogue. `sync-report` shows field and variant
148
+ before/after changes. `retirement-audit` validates sanitized runtime evidence
149
+ for `pending`, `redirect`, `tombstone`, or `gone` endings, including HTTP
150
+ status, booking suppression, noindex/unavailable signals, and sitemap
151
+ exclusion. Response bodies are never stored.
152
+
126
153
  For Google integrations, validate a redacted provider matrix before reporting
127
154
  access. The command fails closed on unknown scopes, missing Ads prerequisites,
128
155
  duplicate provider resources, and unverified edit/publish claims:
@@ -183,7 +210,7 @@ maggie memory ... # confirmed preferences and lessons
183
210
  maggie feedback ... # redact, preview, submit, list
184
211
  maggie qa ... # scenario browser QA, fix/retest, release gate
185
212
  maggie localization ... # plan, validate, review, publish, stale
186
- maggie service ... # import, sync, capability audit, validate
213
+ maggie service ... # import, sync/report, catalogue, lifecycle, validate
187
214
  maggie seo performance ... # sampled PageSpeed/CWV report and baseline
188
215
  maggie seo images ... # inventory, variants, confirmation, validate
189
216
  maggie seo sitemap ... # typed/semantic plan, agent-files, apply, rollback
@@ -325,8 +352,8 @@ artifact schemas.
325
352
  Recommended upgrade sequence for the current release:
326
353
 
327
354
  ```bash
328
- npx @topy-ai/maggie@0.7.21 update --project . --force
329
- npx @topy-ai/maggie@0.7.21 cleanup --project .
355
+ npx @topy-ai/maggie@0.7.22 update --project . --force
356
+ npx @topy-ai/maggie@0.7.22 cleanup --project .
330
357
  ```
331
358
 
332
359
  Maintainers should pass npm credentials through the repository helper, never
@@ -336,6 +363,9 @@ as a command-line argument:
336
363
  node scripts/publish-npm.mjs --maggie-env-file ../.env
337
364
  ```
338
365
 
366
+ The 0.7.22 workflow adds provider-catalogue authority checks, field/variant
367
+ sync reports, separate supply/display states, site-owned slug proposals,
368
+ withdrawal endings, retirement evidence audits, and runtime QA assertion lint.
339
369
  The 0.7.21 workflow adds operation-specific localization quality checks,
340
370
  explicit deployment readiness evidence states, and serialized package assembly.
341
371
  The 0.7.20 workflow completes the MaggieDash Activity and Navigation
package/README.zh-TW.md CHANGED
@@ -8,7 +8,7 @@ Codex、Claude Code 與相容的 coding agents。
8
8
  ## 安裝
9
9
 
10
10
  ```bash
11
- npx @topy-ai/maggie@0.7.21 init --agent all
11
+ npx @topy-ai/maggie@0.7.22 init --agent all
12
12
  npx @topy-ai/maggie doctor --project .
13
13
  ```
14
14
 
@@ -25,6 +25,9 @@ maggie doctor --project . --require-bootstrap --strict
25
25
  deployment、memory、feedback 和 MaggieDash。內容先 draft/review,外部寫入、
26
26
  publish 與 production deployment 需要明確確認。
27
27
 
28
+ 0.7.22 加入 provider catalogue authority check、field/variant sync report、
29
+ supply/display lifecycle state、site-owned slug proposal、withdrawal ending、
30
+ retirement evidence audit,以及 runtime QA assertion lint。
28
31
  0.7.21 增加 localization polish/rewrite 的 operation-specific quality checks、
29
32
  deployment readiness evidence 狀態,以及並發 package assembly 保護。
30
33
  0.7.20 完成 MaggieDash Activity 與 Navigation workspace,以及 dashboard UI v3
@@ -76,6 +79,14 @@ maggie service capability-audit --project . \
76
79
  --catalogue .maggie/booking/services.json \
77
80
  --capabilities-file .maggie/booking/provider-capabilities.json \
78
81
  --fixture .maggie/booking/fixtures/provider.json
82
+
83
+ # Provider catalogue and withdrawal lifecycle review
84
+ maggie service catalogue-check "<provider-location-url>" \
85
+ --treatment "Lymphatic drainage massage"
86
+ maggie service sync-report --project .
87
+ maggie service retirement-audit --project . \
88
+ --catalogue .maggie/booking/services.json \
89
+ --evidence .maggie/booking/retirement-evidence.json
79
90
  ```
80
91
 
81
92
  完整中文說明、19 個 skills 清單和 roadmap:
package/bin/maggie.js CHANGED
@@ -100,6 +100,9 @@ Usage:
100
100
  maggie design status <job-id>
101
101
  maggie service import <provider-url> --project PATH
102
102
  maggie service sync <provider-url> --project PATH
103
+ maggie service catalogue-check <provider-url> --treatment NAME
104
+ maggie service sync-report --project PATH
105
+ maggie service retirement-audit --project PATH --evidence FILE
103
106
  maggie service generate --project PATH
104
107
  maggie service capability-audit --project PATH --catalogue FILE --capabilities-file FILE --fixture FILE
105
108
  maggie service validate --project PATH
@@ -120,7 +123,7 @@ Usage:
120
123
  maggie localization <extract|plan|generate|preview|validate|review|publish|stale|glossary> [options]
121
124
  maggie seo performance|images|sitemap|indexnow|social-cards|head-tags [options] (sitemap supports strict validate and agent-files)
122
125
  maggie feedback <collect|preview|submit|list> [options]
123
- maggie qa <start|record|summary|export> [options]
126
+ maggie qa <start|record|summary|export|assertion-audit> [options]
124
127
  maggie site-audit URL [--crawl] [--access-log FILE] [--require-sitemap-request] [--languages en-GB,es-MX,ja-JP] [--check-hreflang]
125
128
  maggie site-audit URL --crawl --save-baseline FILE --reviewer NAME
126
129
  maggie site-audit URL --crawl --baseline FILE
@@ -57,7 +57,7 @@ The stable model is deliberately small and provider-neutral:
57
57
  - **category**: `level1` and `level2`; a service may belong to more than one
58
58
  provider category through `additional.provider.categories`;
59
59
  - **service**: identity, slug, title, description, active/archived status,
60
- timestamps, and variants;
60
+ timestamps, supply/display state, and variants;
61
61
  - **variant**: identity, display name, duration, price, currency, discount or
62
62
  price range when the provider exposes them;
63
63
  - **booking actions**: booking URL, payment URL, and optional action metadata;
@@ -67,6 +67,44 @@ The stable model is deliberately small and provider-neutral:
67
67
  - **sync**: source, fetched time, content hash, parser version, and change
68
68
  status.
69
69
 
70
+ The portable lifecycle model keeps two owners separate:
71
+
72
+ - `supplyState`: provider sync state, either `live` or `withdrawn`;
73
+ - `displayState`: human publication decision, either `published`, `hidden`, or
74
+ `retired`.
75
+
76
+ Sync may change `supplyState` when the provider adds or removes a service, but
77
+ must preserve `displayState`. A withdrawn/published record is an explicit
78
+ interim state: keep the URL answerable, suppress booking and indexing, and
79
+ open a human retirement decision. The legacy `status` field may remain for
80
+ backwards compatibility, but it must not be used as both provider availability
81
+ and display policy.
82
+
83
+ The `slug` is site-owned. A provider rename may produce a `slugProposal` in a
84
+ sync report, but the stored slug remains unchanged until a person accepts the
85
+ proposal and the host writes the new slug plus its redirect in one transaction.
86
+
87
+ ## Withdrawal and retirement contract
88
+
89
+ Provider removal is not permission to delete a public route. The adapter keeps
90
+ the record with `supplyState: "withdrawn"` and the host chooses a reviewed
91
+ `retirement.ending`:
92
+
93
+ | Ending | Required host response | Required safety signals |
94
+ |---|---|---|
95
+ | `pending` | HTTP 200 interim page | unavailable notice, no booking action, `noindex`, absent from sitemap |
96
+ | `redirect` | HTTP 301 or 308 to a different same-site route | no booking action, absent from sitemap |
97
+ | `tombstone` | HTTP 200 unavailable page | unavailable notice, no booking action, `noindex`, absent from sitemap |
98
+ | `gone` | HTTP 410 | no booking action, absent from sitemap |
99
+
100
+ The host must preserve the old route long enough to apply its chosen ending,
101
+ and must not return a generic 404 as a substitute for the reviewed contract.
102
+ `maggie service retirement-audit` validates sanitized runtime evidence against
103
+ the catalogue and records only service IDs, ending decisions, statuses, and
104
+ safe pass/fail metadata. It does not fetch private routes, store response
105
+ bodies, or invent redirect targets. A person must review redirect destination,
106
+ copy, and accessibility before publication.
107
+
70
108
  The service page must render only canonical fields. `additional` is an
71
109
  explicit extension point for provider-specific or future fields and must be
72
110
  namespaced by concern (`provider`, `location`, `presentation`, `compliance`,
@@ -74,6 +112,21 @@ namespaced by concern (`provider`, `location`, `presentation`, `compliance`,
74
112
  nullable-safe, and must never override canonical fields. Unknown data stays in
75
113
  `additional`; it is not guessed into the stable model.
76
114
 
115
+ ## Source ownership and onboarding precedence
116
+
117
+ When a host onboarding flow collects a value directly from the merchant, that
118
+ explicit value is authoritative for the project. Provider research, imported
119
+ profiles, search results, and inferred metadata may fill an empty field only;
120
+ they must never replace a non-empty merchant entry. Persist the source of each
121
+ value when the host supports it, and show a conflict for human review when two
122
+ non-empty sources disagree. This rule applies to the website URL as well as
123
+ business name, address, phone, category, and booking-provider identity.
124
+
125
+ The shared service-booking parser does not implement a host's onboarding UI.
126
+ Adapters must enforce this precedence before passing project identity into
127
+ provider research or sync. A provider URL is evidence about the provider
128
+ catalogue, not permission to overwrite the merchant's website URL.
129
+
77
130
  Pages can also enter the model without a provider. A normal page conversion
78
131
  creates a `draft` service with `additional.conversion.sourcePage` and no
79
132
  invented price, duration, or booking URL. It becomes `active` only after the
@@ -104,6 +157,14 @@ silent conversion, idempotent add/update/archive synchronisation, source
104
157
  snapshots/checksums, import-run audit records and generated service pages with
105
158
  provider booking CTAs.
106
159
 
160
+ When a provider exposes an authoritative venue-owned catalogue, use that
161
+ adapter extraction for sync decisions and treatment-presence checks. Generic
162
+ JSON-LD, navigation links, and page-wide text can contain marketplace or other
163
+ business entities and are not proof that the connected venue still sells a
164
+ treatment. A read-only catalogue query should call the same parser as import
165
+ and sync so an operator can review a reproducible answer before applying a
166
+ withdrawal.
167
+
107
168
  It must not claim real-time availability, booking creation, cancellation sync
108
169
  or webhooks unless an official Fresha partner/API capability is available for
109
170
  that account. Public booking links remain the safe fallback. The paid Fresha
@@ -74,6 +74,26 @@ the host browser adapter. The adapter owns console, network, authentication,
74
74
  URL, viewport, and screenshot capture; this skill stores only a relative path,
75
75
  URL reference, and hash when a local file exists.
76
76
 
77
+ ## Assert the requirement at runtime
78
+
79
+ Do not use a source class, edited markup fragment, or implementation detail as
80
+ the proof of a user-facing requirement. Describe runtime assertions in a
81
+ project-local manifest and lint it before recording a passing scenario:
82
+
83
+ ```bash
84
+ maggie qa assertion-audit --project . \
85
+ --assertions .maggie/qa-assertions.json \
86
+ --output docs/qa-assertion-audit.json
87
+ ```
88
+
89
+ The `maggie.qa-assertions.v1` manifest requires a route, `transport: "http"`,
90
+ HTTP status evidence, a screenshot/reference, and boolean results for runtime
91
+ checks such as `text-present`, `text-absent`, `meta`, `link-absent`, or
92
+ `redirect`. Source/class/markup-only checks are rejected. The audit stores
93
+ assertion IDs and safe pass/fail metadata, never response bodies, credentials,
94
+ or cookies. A passing assertion audit complements browser evidence; it does
95
+ not replace the host adapter's actual HTTP and screenshot capture.
96
+
77
97
  ## Gate and release evidence
78
98
 
79
99
  The run gate is `fail` when any scenario fails, `blocked` when there is no
@@ -2,7 +2,7 @@
2
2
  name: maggie-service-booking
3
3
  description: Import, synchronise, validate, and design SPA service pages from a booking provider such as Fresha. Use for service catalogues, treatment variants, prices, durations, booking links, payment links, and booking-aware page generation.
4
4
  metadata:
5
- version: 1.1.0
5
+ version: 1.2.0
6
6
  ---
7
7
 
8
8
  # Maggie Service Booking
@@ -49,6 +49,13 @@ passed the provider capability check. A project without Fresha may use Google
49
49
  Calendar as a simple fallback, but it must be recorded as a separate
50
50
  BookingProvider.
51
51
 
52
+ If onboarding also asks the merchant for a website or business identity, the
53
+ merchant's non-empty answer owns that field. Imported provider/profile/search
54
+ values are fallback evidence for empty fields only; a disagreement must become
55
+ a visible review item, never a silent replacement. This skill does not own the
56
+ host onboarding UI, so enforce the rule in the host adapter before starting
57
+ provider research.
58
+
52
59
  ## CLI workflow
53
60
 
54
61
  Run from the project root:
@@ -62,6 +69,8 @@ python3 tools/clis/maggie_service_booking.py sync \
62
69
  "https://www.fresha.com/a/spa-chevy-chase-chevy-chase-4500-north-park-avenue-sbic60h4?pId=512061" \
63
70
  --project .
64
71
 
72
+ python3 tools/clis/maggie_service_booking.py sync-report --project .
73
+
65
74
  python3 tools/clis/maggie_service_booking.py generate --project . --copy-data docs/service-page-copy.json
66
75
  python3 tools/clis/maggie_service_booking.py validate --project .
67
76
 
@@ -116,17 +125,40 @@ python3 tools/clis/maggie_service_booking.py capability-audit \
116
125
  --catalogue .maggie/booking/services.json \
117
126
  --capabilities-file .maggie/booking/provider-capabilities.json \
118
127
  --fixture .maggie/booking/fixtures/provider.json
128
+
129
+ # Answer whether provider-owned treatments are present using the same parser
130
+ # as import/sync; do not search the whole provider page for a name.
131
+ python3 tools/clis/maggie_service_booking.py catalogue-check \
132
+ "https://www.fresha.com/a/your-location" \
133
+ --provider fresha --treatment "Lymphatic drainage massage"
134
+
135
+ python3 tools/clis/maggie_service_booking.py retirement-audit \
136
+ --project . \
137
+ --catalogue .maggie/booking/services.json \
138
+ --evidence .maggie/booking/retirement-evidence.json
119
139
  ```
120
140
 
121
141
  `import` creates the first catalogue. `sync` compares the newly imported
122
142
  catalogue with the previous snapshot and records `added`, `updated`,
123
- `removed`, and `unchanged` services. Removed services are archived in the
124
- manifest before any page is deleted. `generate` is retained as a compatibility
143
+ `removed`, and `unchanged` services. Updated records include field-level
144
+ before/after values, variant-level changes, and any pending slug proposal.
145
+ `sync-report` is read-only and loads a saved `maggie-service-sync-report.v1`
146
+ artifact for review. Removed services are archived in the manifest before any
147
+ page is deleted. `generate` is retained as a compatibility
125
148
  guard and refuses to author public copy. The AI agent and the MaggieDash renderer
126
149
  must consume an approved copy artifact instead. `run` executes parse → persist
127
150
  → validate, then stops before public generation until an AI-authored copy
128
151
  artifact and editorial review are present; it records `.maggie/booking/job.json`.
129
152
  `inspect` is read-only and `status` reports the latest job state.
153
+ `catalogue-check` is read-only and answers one or more treatment-presence
154
+ questions from the same provider-owned extraction used by `import` and `sync`.
155
+ It does not use arbitrary page links or marketplace JSON-LD as evidence when
156
+ the provider exposes an authoritative embedded catalogue.
157
+ `retirement-audit` is also read-only. It consumes a sanitized host-browser/HTTP
158
+ evidence artifact and validates the selected ending (`pending`, `redirect`,
159
+ `tombstone`, or `gone`), HTTP status, booking suppression, unavailable/noindex
160
+ signals, and sitemap exclusion. It never stores response bodies or decides a
161
+ redirect destination for the host.
130
162
  `convert-page` imports a normal page as a draft service without inventing
131
163
  booking facts. `match-pages` reads filesystem pages plus `docs/pages.json` and
132
164
  `docs/page-content.json`, writes candidate evidence to both `.maggie/booking`
@@ -220,6 +252,19 @@ Every active service must have:
220
252
  - a booking URL, and a payment URL only when explicitly supplied;
221
253
  - provider, source URL, `firstSeenAt`, `lastSeenAt`, and sync status.
222
254
 
255
+ The portable catalogue also carries `supplyState` (`live` or `withdrawn`) and
256
+ `displayState` (`published`, `hidden`, or `retired`). The sync owns only
257
+ `supplyState`; it must preserve a person's `displayState` and must not silently
258
+ republish a hidden or retired service. Keep legacy `status` only for
259
+ compatibility, never as both meanings at once.
260
+
261
+ Withdrawn services remain available for an explicit retirement decision. Set
262
+ `retirement.ending` to `pending` while a host still serves a safe interim
263
+ response, or to an approved `redirect`, `tombstone`, or `gone` ending. A
264
+ withdrawn route must suppress booking and sitemap inclusion; the host owns the
265
+ actual route/HTTP implementation and supplies sanitized evidence to
266
+ `retirement-audit` before release.
267
+
223
268
  The canonical shape is documented in
224
269
  [`references/universal-booking-adapter.md`](../../references/universal-booking-adapter.md).
225
270
  Do not invent prices, availability, practitioner claims, ratings, medical
@@ -22,6 +22,9 @@ from urllib.parse import urlsplit, urlunsplit
22
22
  DEFAULT_SCENARIOS = Path(".maggie") / "scenario-manifest.json"
23
23
  VALID_STATUSES = {"pending", "pass", "fail", "blocked", "inconclusive"}
24
24
  VALID_PHASES = {"test", "fix", "retest"}
25
+ ASSERTION_SCHEMA = "maggie.qa-assertions.v1"
26
+ RUNTIME_ASSERTION_TYPES = {"http-status", "final-url", "text-present", "text-absent", "meta", "link-absent", "redirect"}
27
+ SOURCE_ASSERTION_TYPES = {"class-present", "class-absent", "markup-present", "markup-absent", "source-text"}
25
28
  SECRET_RE = re.compile(
26
29
  r"(?i)(bearer\s+|(?:api[_-]?key|token|secret|password|authorization|cookie)\s*[=:]\s*)[^\s,;]+"
27
30
  )
@@ -314,6 +317,86 @@ def export_run(args: argparse.Namespace) -> int:
314
317
  return 0
315
318
 
316
319
 
320
+ def assertion_audit(args: argparse.Namespace) -> int:
321
+ """Ensure QA assertions test the served requirement, not edited source markup."""
322
+ project = project_root(args.project)
323
+ manifest_path = project_root(args.assertions) if Path(args.assertions).is_absolute() else project / args.assertions
324
+ manifest = load_json(manifest_path)
325
+ errors = []
326
+ records = manifest.get("assertions")
327
+ if manifest.get("schemaVersion") != ASSERTION_SCHEMA:
328
+ errors.append(f"manifest schemaVersion must be {ASSERTION_SCHEMA}")
329
+ if not isinstance(records, list) or not records:
330
+ errors.append("assertions must be a non-empty array")
331
+ records = []
332
+ seen = set()
333
+ audited = []
334
+ for index, item in enumerate(records):
335
+ item_errors = []
336
+ if not isinstance(item, dict):
337
+ audited.append({"id": f"assertion-{index + 1}", "passed": False, "errors": ["assertion must be an object"]})
338
+ continue
339
+ identifier = safe_text(item.get("id"), 120) or f"assertion-{index + 1}"
340
+ if identifier in seen:
341
+ item_errors.append("assertion id is duplicated")
342
+ seen.add(identifier)
343
+ if not safe_text(item.get("requirement"), 1000):
344
+ item_errors.append("requirement is required")
345
+ route = safe_text(item.get("route"), 500)
346
+ if not route or not route.startswith("/"):
347
+ item_errors.append("route must be a path on the tested site")
348
+ if item.get("transport") != "http":
349
+ item_errors.append("transport must be http runtime evidence")
350
+ evidence_data = item.get("evidence")
351
+ if not isinstance(evidence_data, dict):
352
+ item_errors.append("evidence must be an object")
353
+ evidence_data = {}
354
+ if not safe_text(evidence_data.get("screenshot"), 500):
355
+ item_errors.append("evidence.screenshot is required")
356
+ if not isinstance(evidence_data.get("httpStatus"), int) or isinstance(evidence_data.get("httpStatus"), bool):
357
+ item_errors.append("evidence.httpStatus must be an integer")
358
+ checks = item.get("checks")
359
+ if not isinstance(checks, list) or not checks:
360
+ item_errors.append("checks must be a non-empty array")
361
+ checks = []
362
+ if not any(isinstance(check, dict) and check.get("type") == "http-status" for check in checks):
363
+ item_errors.append("at least one http-status check is required")
364
+ for check_index, check in enumerate(checks):
365
+ if not isinstance(check, dict):
366
+ item_errors.append(f"checks[{check_index}] must be an object")
367
+ continue
368
+ check_type = check.get("type")
369
+ surface = check.get("surface")
370
+ if check_type in SOURCE_ASSERTION_TYPES or surface == "source":
371
+ item_errors.append(f"checks[{check_index}] is source/markup-only; assert the served route instead")
372
+ if check_type not in RUNTIME_ASSERTION_TYPES:
373
+ item_errors.append(f"checks[{check_index}] has an unsupported runtime assertion type")
374
+ if surface not in {"http", "browser-runtime"}:
375
+ item_errors.append(f"checks[{check_index}] must identify http or browser-runtime evidence")
376
+ if not isinstance(check.get("passed"), bool):
377
+ item_errors.append(f"checks[{check_index}].passed must be boolean evidence")
378
+ if check_type == "http-status" and isinstance(evidence_data.get("httpStatus"), int) and check.get("actual") != evidence_data["httpStatus"]:
379
+ item_errors.append(f"checks[{check_index}] actual status does not match evidence.httpStatus")
380
+ if any(isinstance(check, dict) and check.get("passed") is False for check in checks):
381
+ item_errors.append("one or more runtime checks failed")
382
+ audited.append({"id": identifier, "passed": not item_errors, "errors": item_errors})
383
+ passed = not errors and all(item.get("passed") for item in audited) and bool(audited)
384
+ report = {
385
+ "schemaVersion": "maggie.qa-assertion-audit.v1",
386
+ "assertionSchema": ASSERTION_SCHEMA,
387
+ "checked": len(audited),
388
+ "assertions": audited,
389
+ "errors": errors,
390
+ "passed": passed,
391
+ }
392
+ output = Path(args.output).expanduser()
393
+ if not output.is_absolute():
394
+ output = project / output
395
+ write_json(output, report)
396
+ print(json.dumps({"status": "passed" if passed else "failed", "report": str(output.resolve()), "checked": len(audited), "errors": errors}, ensure_ascii=False, indent=2))
397
+ return 0 if passed else 1
398
+
399
+
317
400
  def parser() -> argparse.ArgumentParser:
318
401
  root = argparse.ArgumentParser(description=__doc__)
319
402
  root.add_argument("--project", default=".")
@@ -350,6 +433,10 @@ def parser() -> argparse.ArgumentParser:
350
433
  command.add_argument("--run", required=True)
351
434
  if name == "summary": command.add_argument("--format", choices=("json", "markdown"), default="json")
352
435
  else: command.add_argument("--output", required=True)
436
+ assertion_parser = commands.add_parser("assertion-audit")
437
+ assertion_parser.add_argument("--project", default=argparse.SUPPRESS)
438
+ assertion_parser.add_argument("--assertions", required=True, help="runtime assertion manifest JSON")
439
+ assertion_parser.add_argument("--output", default=".maggie/qa-assertion-audit.json")
353
440
  return root
354
441
 
355
442
 
@@ -359,6 +446,7 @@ def main() -> int:
359
446
  if args.command == "start": return start(args)
360
447
  if args.command == "record": return record(args)
361
448
  if args.command == "summary": return summary(args)
449
+ if args.command == "assertion-audit": return assertion_audit(args)
362
450
  return export_run(args)
363
451
  except (OSError, ValueError, json.JSONDecodeError) as error:
364
452
  print(f"QA workflow error: {safe_text(error)}", file=sys.stderr)
@@ -17,6 +17,8 @@ from booking_capabilities import audit_matrix # noqa: E402
17
17
  NOW = lambda: datetime.now(timezone.utc).isoformat()
18
18
  MONEY = re.compile(r"(?:£|GBP\s*)\s*([0-9]+(?:[.,][0-9]{1,2})?)", re.I)
19
19
  DURATION = re.compile(r"(\d{2,3})\s*(?:min|mins|minutes?)", re.I)
20
+ SUPPLY_STATES = {"live", "withdrawn"}
21
+ DISPLAY_STATES = {"published", "hidden", "retired"}
20
22
 
21
23
  def slug(value):
22
24
  value = re.sub(r"[^a-z0-9]+", "-", value.lower()).strip("-")
@@ -57,33 +59,37 @@ def flatten(value):
57
59
 
58
60
  def parse(source, provider):
59
61
  raw=fetch(source); parser=Extractor(); parser.feed(raw); records=[]
62
+ embedded_marker = provider == "fresha" and re.search(r'"services"\s*:\s*\[', raw)
60
63
  if provider == "fresha":
61
64
  records.extend(parse_fresha_embedded(raw, source))
62
- for obj in json_objects(parser.scripts):
63
- for item in flatten(obj):
64
- typ=item.get("@type", "")
65
- types=set(typ) if isinstance(typ, list) else {typ}
66
- if not types.intersection({"Service","Product","Offer","ListItem"}) and not any(k in item for k in ("duration","price","providerServiceId")): continue
67
- name=item.get("name") or item.get("title") or (item.get("item") or {}).get("name") if isinstance(item.get("item"),dict) else item.get("name")
68
- if not name or len(str(name)) < 3: continue
69
- offers=item.get("offers") if isinstance(item.get("offers"),list) else [item.get("offers", item)]
70
- variants=[]
71
- for offer in offers:
72
- if not isinstance(offer,dict): offer={}
73
- text=" ".join(str(offer.get(k,"")) for k in ("name","description","duration"))
74
- dm=DURATION.search(text); price=offer.get("price")
75
- if price is None:
76
- match=MONEY.search(text); price=match.group(1).replace(",","") if match else None
77
- currency=(offer.get("priceCurrency") or ("GBP" if price and ("£" in text or provider=="fresha") else None))
78
- if price is None and dm is None: continue
79
- amount=int(round(float(str(price).replace(",",""))*100)) if price is not None else None
80
- provider_variant_id = offer.get("sku") or offer.get("@id") or offer.get("id")
81
- variants.append({"id":str(provider_variant_id or f"{slug(name)}:{dm.group(1) if dm else 'default'}"),"idSource":"provider" if provider_variant_id else "derived","variantSource":"native" if provider_variant_id else "derived-offers","title":offer.get("name") or (f"{dm.group(1)} minutes" if dm else "Standard"),"durationMinutes":int(dm.group(1)) if dm else None,"price":{"amountMinor":amount,"currency":currency,"display":f"£{amount/100:.2f}" if amount is not None and currency=="GBP" else None}})
82
- url=item.get("url") or item.get("sameAs")
83
- if not variants and not url: continue
84
- pid=str(item.get("providerServiceId") or item.get("productID") or item.get("sku") or slug(name))
85
- records.append(make_record(provider, source, pid, str(name), str(item.get("description") or ""), variants, url, parser.links))
86
- if not records:
65
+ # A provider-specific authoritative catalogue is the only service source.
66
+ # Generic JSON-LD may include marketplace entities for another business.
67
+ if not embedded_marker:
68
+ for obj in json_objects(parser.scripts):
69
+ for item in flatten(obj):
70
+ typ=item.get("@type", "")
71
+ types=set(typ) if isinstance(typ, list) else {typ}
72
+ if not types.intersection({"Service","Product","Offer","ListItem"}) and not any(k in item for k in ("providerServiceId","productID","sku")): continue
73
+ name=item.get("name") or item.get("title") or (item.get("item") or {}).get("name") if isinstance(item.get("item"),dict) else item.get("name")
74
+ if not name or len(str(name)) < 3: continue
75
+ offers=item.get("offers") if isinstance(item.get("offers"),list) else [item.get("offers", item)]
76
+ variants=[]
77
+ for offer in offers:
78
+ if not isinstance(offer,dict): offer={}
79
+ text=" ".join(str(offer.get(k,"")) for k in ("name","description","duration"))
80
+ dm=DURATION.search(text); price=offer.get("price")
81
+ if price is None:
82
+ match=MONEY.search(text); price=match.group(1).replace(",","") if match else None
83
+ currency=(offer.get("priceCurrency") or ("GBP" if price and ("£" in text or provider=="fresha") else None))
84
+ if price is None and dm is None: continue
85
+ amount=int(round(float(str(price).replace(",",""))*100)) if price is not None else None
86
+ provider_variant_id = offer.get("sku") or offer.get("@id") or offer.get("id")
87
+ variants.append({"id":str(provider_variant_id or f"{slug(name)}:{dm.group(1) if dm else 'default'}"),"idSource":"provider" if provider_variant_id else "derived","variantSource":"native" if provider_variant_id else "derived-offers","title":offer.get("name") or (f"{dm.group(1)} minutes" if dm else "Standard"),"durationMinutes":int(dm.group(1)) if dm else None,"price":{"amountMinor":amount,"currency":currency,"display":f"£{amount/100:.2f}" if amount is not None and currency=="GBP" else None}})
88
+ url=item.get("url") or item.get("sameAs")
89
+ if not variants and not url: continue
90
+ pid=str(item.get("providerServiceId") or item.get("productID") or item.get("sku") or slug(name))
91
+ records.append(make_record(provider, source, pid, str(name), str(item.get("description") or ""), variants, url, parser.links))
92
+ if not records and not embedded_marker:
87
93
  for _, pieces in parser.headings:
88
94
  name=" ".join(pieces).strip()
89
95
  if len(name)>3: records.append(make_record(provider, source, slug(name), name, "", [], None, parser.links))
@@ -97,10 +103,10 @@ def parse(source, provider):
97
103
  def parse_fresha_embedded(raw, source):
98
104
  """Read the public service catalogue embedded in Fresha's Next data."""
99
105
  records=[]
100
- marker='"services":['; start=raw.find(marker)
101
- if start >= 0:
106
+ marker=re.search(r'"services"\s*:\s*\[', raw); start=marker.start() if marker else -1
107
+ if marker:
102
108
  try:
103
- groups=json.JSONDecoder().raw_decode(raw[start+len('"services":'):])[0]
109
+ groups=json.JSONDecoder().raw_decode(raw[marker.end()-1:])[0]
104
110
  except (ValueError, TypeError): groups=[]
105
111
  for group in groups if isinstance(groups,list) else []:
106
112
  level1=str(group.get("name") or "Uncategorised").strip()
@@ -182,7 +188,7 @@ def make_record(provider, source, pid, name, description, variants, booking, lin
182
188
  if not booking:
183
189
  for href,_ in links:
184
190
  if any(x in href.lower() for x in ("book","appointment","checkout")): booking=urljoin(source,href); break
185
- return {"id":f"{provider}:{pid}","provider":provider,"providerServiceId":pid,"slug":slug(name),"title":name,"description":description.strip(),"sourceUrl":source,"category":{"level1":"Uncategorised","level2":None},"variants":variants,"bookingUrl":booking,"paymentUrl":None,"status":"active","firstSeenAt":NOW(),"lastSeenAt":NOW()}
191
+ return {"id":f"{provider}:{pid}","provider":provider,"providerServiceId":pid,"slug":slug(name),"title":name,"description":description.strip(),"sourceUrl":source,"category":{"level1":"Uncategorised","level2":None},"variants":variants,"bookingUrl":booking,"paymentUrl":None,"status":"active","supplyState":"live","displayState":"published","firstSeenAt":NOW(),"lastSeenAt":NOW()}
186
192
 
187
193
  def root(args): return Path(args.project).resolve()
188
194
  def path(project): return project/".maggie"/"booking"/"services.json"
@@ -199,11 +205,15 @@ def validate(data):
199
205
  if not s.get(key): errors.append(f"{s.get('id','unknown')}: missing {key}")
200
206
  if s.get("id") in seen: errors.append(f"duplicate service id: {s['id']}")
201
207
  seen.add(s.get("id"));
208
+ supply=s.get("supplyState") or ("withdrawn" if s.get("status") in {"archived", "removed"} else "live")
209
+ display=s.get("displayState") or ("retired" if s.get("status") == "archived" else "published")
210
+ if supply not in SUPPLY_STATES: errors.append(f"{s.get('id')}: invalid supplyState")
211
+ if display not in DISPLAY_STATES: errors.append(f"{s.get('id')}: invalid displayState")
202
212
  if any(x.get("slug")==s.get("slug") for x in data.get("services",[]) if x is not s): errors.append(f"duplicate service slug: {s.get('slug')}")
203
213
  pages=s.get("pages",[])
204
214
  if not isinstance(pages,list): errors.append(f"{s.get('id')}: pages must be an array"); pages=[]
205
215
  if sum(p.get("role")=="canonical" for p in pages if isinstance(p,dict))>1: errors.append(f"{s.get('id')}: multiple canonical pages")
206
- if s.get("status")=="active" and not s.get("variants"): errors.append(f"{s.get('id')}: no variants")
216
+ if supply == "live" and not s.get("variants"): errors.append(f"{s.get('id')}: no variants")
207
217
  for v in s.get("variants",[]):
208
218
  if not v.get("durationMinutes") or not v.get("price",{}).get("currency"): errors.append(f"{s.get('id')}: incomplete variant {v.get('id')}")
209
219
  for key in ("bookingUrl","paymentUrl"):
@@ -312,33 +322,172 @@ def cmd_fact_audit(args):
312
322
  print(json.dumps({"status": "passed" if report["passed"] else "failed", "report": str(out), "reviewQueue": str(queue_path), "services": len(services), "variants": len(seen_variants), "sourceUrlsBackfilled": changed, "errors": errors[:20], "errorCount": len(errors), "publish": False}, indent=2, ensure_ascii=False))
313
323
  return 0 if report["passed"] else 1
314
324
 
325
+ def cmd_catalogue_check(args):
326
+ """Answer treatment-presence questions using the import/sync parser."""
327
+ try:
328
+ data = parse(args.source, args.provider)
329
+ except (OSError, ValueError, UnicodeError):
330
+ print(json.dumps({"status": "failed", "error": "source could not be parsed"}, indent=2))
331
+ return 1
332
+ services = data.get("services", [])
333
+ results = []
334
+ for query in args.treatment:
335
+ query_tokens = page_tokens(query)
336
+ matches = [
337
+ {"id": service.get("id"), "title": service.get("title"), "category": service.get("category", {}).get("level1")}
338
+ for service in services
339
+ if query_tokens and query_tokens <= page_tokens(service.get("title", ""))
340
+ ]
341
+ results.append({"treatment": query, "status": "found" if matches else "not-found", "matches": matches})
342
+ passed = all(item["status"] == "found" for item in results)
343
+ print(json.dumps({"status": "passed" if passed else "failed", "provider": args.provider, "serviceCount": len(services), "results": results}, indent=2, ensure_ascii=False))
344
+ return 0 if passed else 1
345
+
346
+ RETIREMENT_SCHEMA = "maggie-service-retirement-evidence.v1"
347
+ RETIREMENT_ENDINGS = {"pending", "redirect", "tombstone", "gone"}
348
+
349
+ def cmd_retirement_audit(args):
350
+ """Validate sanitized runtime evidence for withdrawn service URLs."""
351
+ project=root(args); catalogue_path=Path(args.catalogue); catalogue_path=catalogue_path if catalogue_path.is_absolute() else project/catalogue_path
352
+ evidence_path=Path(args.evidence); evidence_path=evidence_path if evidence_path.is_absolute() else project/evidence_path
353
+ try:
354
+ catalogue=json.loads(catalogue_path.read_text(encoding="utf-8")); evidence=json.loads(evidence_path.read_text(encoding="utf-8"))
355
+ except (OSError, ValueError):
356
+ print(json.dumps({"status":"failed","error":"catalogue or retirement evidence is invalid"},indent=2)); return 1
357
+ services={str(item.get("id")):item for item in catalogue.get("services",[]) if isinstance(item,dict) and item.get("id")}
358
+ rows=evidence.get("services") if isinstance(evidence,dict) else None
359
+ errors=[]; results=[]
360
+ if not isinstance(evidence,dict) or evidence.get("schemaVersion") != RETIREMENT_SCHEMA:
361
+ errors.append(f"evidence schemaVersion must be {RETIREMENT_SCHEMA}")
362
+ if not isinstance(rows,list):
363
+ errors.append("evidence.services must be a non-empty array")
364
+ rows=[]
365
+ if isinstance(rows,list) and not rows: errors.append("evidence.services must be a non-empty array")
366
+ seen=set()
367
+ for index,row in enumerate(rows):
368
+ item_errors=[]
369
+ if not isinstance(row,dict):
370
+ results.append({"index":index,"passed":False,"errors":["evidence row must be an object"]}); continue
371
+ service_id=str(row.get("serviceId") or "")
372
+ if not service_id or service_id in seen: item_errors.append("serviceId is missing or duplicated")
373
+ seen.add(service_id)
374
+ service=services.get(service_id)
375
+ if not service: item_errors.append("service is not in the catalogue")
376
+ supply=(service or {}).get("supplyState") or ("withdrawn" if (service or {}).get("status") in {"archived","removed"} else "live")
377
+ if supply != "withdrawn": item_errors.append("retirement evidence requires a withdrawn service")
378
+ ending=str(row.get("ending") or "")
379
+ if ending not in RETIREMENT_ENDINGS: item_errors.append("ending must be pending, redirect, tombstone, or gone")
380
+ expected_ending=((service or {}).get("retirement") or {}).get("ending")
381
+ if expected_ending and expected_ending != ending: item_errors.append("evidence ending does not match the catalogue decision")
382
+ status=row.get("httpStatus")
383
+ if not isinstance(status,int) or isinstance(status,bool): item_errors.append("httpStatus must be an integer")
384
+ if row.get("inSitemap") is not False: item_errors.append("withdrawn service must be absent from the sitemap")
385
+ if row.get("bookingSuppressed") is not True: item_errors.append("bookingSuppressed must be true")
386
+ if ending in {"pending","tombstone"}:
387
+ if status != 200: item_errors.append(f"{ending} must return HTTP 200")
388
+ if row.get("unavailableNotice") is not True: item_errors.append("unavailableNotice must be true")
389
+ if row.get("noindex") is not True: item_errors.append("noindex must be true")
390
+ elif ending == "redirect":
391
+ if status not in {301,308}: item_errors.append("redirect must return HTTP 301 or 308")
392
+ location=str(row.get("location") or "")
393
+ route=str(row.get("route") or "")
394
+ if not location or not location.startswith("/") or location == route: item_errors.append("redirect needs a different same-site location")
395
+ elif ending == "gone" and status != 410:
396
+ item_errors.append("gone must return HTTP 410")
397
+ results.append({"serviceId":service_id,"ending":ending,"passed":not item_errors,"errors":item_errors})
398
+ passed=not errors and bool(results) and all(item["passed"] for item in results)
399
+ report={"schemaVersion":"maggie-service-retirement-audit.v1","evidenceSchema":RETIREMENT_SCHEMA,"checked":len(results),"services":results,"errors":errors,"passed":passed}
400
+ output=Path(args.output).expanduser(); output=output if output.is_absolute() else project/output; output.parent.mkdir(parents=True,exist_ok=True); output.write_text(json.dumps(report,indent=2,ensure_ascii=False)+"\n",encoding="utf-8")
401
+ print(json.dumps({"status":"passed" if passed else "failed","report":str(output.resolve()),"checked":len(results),"errors":errors},indent=2)); return 0 if passed else 1
402
+
315
403
  def cmd_import(args):
316
404
  project=root(args); data=parse(args.source,args.provider); old=load(project) if path(project).exists() else {"services":[]}; previous={s["id"]:s for s in old.get("services",[])}
317
405
  for service in data.get("services",[]):
318
406
  if service["id"] in previous:
319
- service["pages"]=previous[service["id"]].get("pages",[]); service["additional"]={**previous[service["id"]].get("additional",{}),**service.get("additional",{})}
407
+ old_service=previous[service["id"]]
408
+ preserve_site_fields(service, old_service)
409
+ service["displayState"]=old_service.get("displayState") or ("retired" if old_service.get("status") == "archived" else "published")
410
+ service.setdefault("supplyState", "live"); service.setdefault("displayState", "published")
320
411
  errors=validate(data)
321
412
  if errors: print(json.dumps({"status":"failed","errors":errors},indent=2)); return 1
322
413
  save(project,data); print(json.dumps({"status":"imported","path":str(path(project)),"reviewPath":str(project/"docs"/"services.json"),"services":len(data["services"])},indent=2)); return 0
414
+
415
+ SERVICE_DIFF_FIELDS = ("provider", "providerServiceId", "slug", "title", "description", "category", "bookingUrl", "paymentUrl", "status", "supplyState", "displayState", "retirement")
416
+
417
+ def preserve_site_fields(service, old_service):
418
+ proposal=None
419
+ old_slug=str(old_service.get("slug") or "")
420
+ new_slug=str(service.get("slug") or "")
421
+ if old_slug and new_slug and old_slug != new_slug:
422
+ proposal={"status":"pending","from":old_slug,"to":new_slug,"reason":"provider name implies a different public slug"}
423
+ service["slug"]=old_slug
424
+ service["pages"] = old_service.get("pages", [])
425
+ service["additional"] = {**old_service.get("additional", {}), **service.get("additional", {})}
426
+ return proposal
427
+
428
+ def service_diff(before, after):
429
+ fields=[]
430
+ for field in SERVICE_DIFF_FIELDS:
431
+ old_value = before.get(field) if isinstance(before, dict) else None
432
+ new_value = after.get(field) if isinstance(after, dict) else None
433
+ if old_value != new_value:
434
+ fields.append({"field": field, "before": old_value, "after": new_value})
435
+ old_variants={str(item.get("id")): item for item in (before or {}).get("variants", []) if isinstance(item, dict) and item.get("id")}
436
+ new_variants={str(item.get("id")): item for item in (after or {}).get("variants", []) if isinstance(item, dict) and item.get("id")}
437
+ variants=[]
438
+ for variant_id in sorted(set(old_variants) | set(new_variants)):
439
+ old_variant=old_variants.get(variant_id); new_variant=new_variants.get(variant_id)
440
+ if old_variant is None:
441
+ variants.append({"id": variant_id, "change": "added", "before": None, "after": new_variant})
442
+ elif new_variant is None:
443
+ variants.append({"id": variant_id, "change": "removed", "before": old_variant, "after": None})
444
+ elif old_variant != new_variant:
445
+ variants.append({"id": variant_id, "change": "updated", "before": old_variant, "after": new_variant})
446
+ return {"fields": fields, "variantChanges": variants}
447
+
323
448
  def cmd_sync(args):
324
449
  project=root(args); old=load(project) if path(project).exists() else {"services":[]}; fresh=parse(args.source,args.provider); previous={s["id"]:s for s in old.get("services",[])}; current={s["id"]:s for s in fresh.get("services",[])}; changes=[]
325
- def comparable(service):
326
- value=dict(service)
327
- value.pop("firstSeenAt", None)
328
- value.pop("lastSeenAt", None)
329
- return value
330
450
  for sid, service in current.items():
331
451
  old_service = previous.get(sid)
452
+ slug_proposal=None
332
453
  if old_service:
333
- service["pages"] = old_service.get("pages", [])
334
- service["additional"] = {**old_service.get("additional", {}), **service.get("additional", {})}
335
- changes.append({"id":sid,"change":"added" if sid not in previous else ("updated" if json.dumps(comparable(service),sort_keys=True,ensure_ascii=False)!=json.dumps(comparable(old_service),sort_keys=True,ensure_ascii=False) else "unchanged")})
454
+ slug_proposal=preserve_site_fields(service, old_service)
455
+ # Provider sync owns supply; a person's display decision survives
456
+ # the run and is never reset by a newly fetched row.
457
+ service["displayState"] = old_service.get("displayState") or ("retired" if old_service.get("status") == "archived" else "published")
458
+ service["supplyState"] = "live"
459
+ diff=service_diff(old_service, service)
460
+ change="added" if sid not in previous else ("updated" if diff["fields"] or diff["variantChanges"] else "unchanged")
461
+ if slug_proposal:
462
+ diff["slugProposal"]=slug_proposal
463
+ diff["reviewRequired"]=True
464
+ change="updated"
465
+ changes.append({"id":sid,"change":change,**diff})
336
466
  for sid, service in previous.items():
337
- if sid not in current: archived={**service,"status":"archived","removedAt":NOW()}; current[sid]=archived; changes.append({"id":sid,"change":"removed"})
467
+ if sid not in current:
468
+ archived={**service,"status":"archived","supplyState":"withdrawn","displayState":service.get("displayState") or ("retired" if service.get("status") == "archived" else "published"),"retirement":service.get("retirement") or {"ending":"pending"},"removedAt":NOW()}; current[sid]=archived
469
+ changes.append({"id":sid,"change":"removed","fields":[{"field":"status","before":service.get("status"),"after":"archived"}],"variantChanges":[]})
338
470
  fresh["services"]=list(current.values()); save(project,fresh)
339
- run_id=datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
340
- report={"schemaVersion":"1.0","runId":run_id,"syncedAt":fresh["syncedAt"],"sourceUrl":args.source,"provider":args.provider,"changes":changes,"summary":{"total":len(changes),"added":sum(c["change"]=="added" for c in changes),"updated":sum(c["change"]=="updated" for c in changes),"removed":sum(c["change"]=="removed" for c in changes),"unchanged":sum(c["change"]=="unchanged" for c in changes)},"checkpoint":{"phase":"reconciled","processed":len(changes),"total":len(changes),"resumable":True,"status":"complete"}}
471
+ run_id=datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S%fZ")
472
+ report={"schemaVersion":"maggie-service-sync-report.v1","runId":run_id,"syncedAt":fresh["syncedAt"],"sourceUrl":args.source,"provider":args.provider,"changes":changes,"summary":{"total":len(changes),"added":sum(c["change"]=="added" for c in changes),"updated":sum(c["change"]=="updated" for c in changes),"removed":sum(c["change"]=="removed" for c in changes),"unchanged":sum(c["change"]=="unchanged" for c in changes)},"checkpoint":{"phase":"reconciled","processed":len(changes),"total":len(changes),"resumable":True,"status":"complete"}}
341
473
  sync_dir=path(project).parent/"sync-runs"; sync_dir.mkdir(parents=True,exist_ok=True); (sync_dir/f"{run_id}.json").write_text(json.dumps(report,indent=2,ensure_ascii=False)+"\n",encoding="utf-8"); (path(project).parent/"sync-checkpoint.json").write_text(json.dumps(report["checkpoint"]|{"runId":run_id,"sourceUrl":args.source},indent=2,ensure_ascii=False)+"\n",encoding="utf-8"); (path(project).parent/"last-sync.json").write_text(json.dumps(report,indent=2,ensure_ascii=False)+"\n",encoding="utf-8"); print(json.dumps(report,indent=2,ensure_ascii=False)); return 0 if not validate(fresh) else 1
474
+
475
+ def cmd_sync_report(args):
476
+ project=root(args); directory=path(project).parent/"sync-runs"
477
+ candidates=sorted(directory.glob("*.json")) if directory.exists() else []
478
+ if args.run_id:
479
+ candidates=[directory/f"{args.run_id}.json"]
480
+ if not candidates or not candidates[-1].exists():
481
+ print(json.dumps({"status":"failed","error":"sync report not found"},indent=2)); return 1
482
+ try:
483
+ report=json.loads(candidates[-1].read_text(encoding="utf-8"))
484
+ except (OSError, ValueError):
485
+ print(json.dumps({"status":"failed","error":"sync report is invalid"},indent=2)); return 1
486
+ if report.get("schemaVersion") != "maggie-service-sync-report.v1":
487
+ print(json.dumps({"status":"failed","error":"unsupported sync report schema"},indent=2)); return 1
488
+ if args.output:
489
+ output=Path(args.output).expanduser(); output=output if output.is_absolute() else project/output; output.parent.mkdir(parents=True,exist_ok=True); output.write_text(json.dumps(report,indent=2,ensure_ascii=False)+"\n",encoding="utf-8")
490
+ print(json.dumps(report,indent=2,ensure_ascii=False)); return 0
342
491
  def cmd_generate(args):
343
492
  print(json.dumps({"status":"blocked","error":"Public copy generation is agent-owned","required":"The AI agent must create and review the copy artifact, then the project renderer must consume it. This deterministic CLI never synthesizes public hero, section, CTA, title, or FAQ copy.","copyData":getattr(args, "copy_data", None)},indent=2)); return 1
344
493
  project=root(args); data=load(project); errors=validate(data)
@@ -1004,6 +1153,7 @@ def main():
1004
1153
  for name in ("import","sync","run"):
1005
1154
  q=sub.add_parser(name); q.add_argument("source"); q.add_argument("--project",default="."); q.add_argument("--provider",default="fresha",choices=["fresha"])
1006
1155
  if name == "run": q.add_argument("--copy-data", help="AI-authored service copy artifact; required before public page generation")
1156
+ q=sub.add_parser("sync-report"); q.add_argument("--project",default="."); q.add_argument("--run-id"); q.add_argument("--output")
1007
1157
  q=sub.add_parser("convert-page"); q.add_argument("page"); q.add_argument("--project",default=".")
1008
1158
  for name in ("polish","validate-polish"):
1009
1159
  q=sub.add_parser(name); q.add_argument("--service-id",required=True); q.add_argument("--page",required=True); q.add_argument("--project",default="."); q.add_argument("--rendered",help="rendered HTML file or URL for runtime evidence"); q.add_argument("--copy-data",required=True,help="AI-authored copy JSON with provenance and completed review statuses")
@@ -1013,6 +1163,8 @@ def main():
1013
1163
  q=sub.add_parser("category-editorial-review"); q.add_argument("--project",default="."); q.add_argument("--copy-data",default="docs/category-page-copy.json"); q.add_argument("--category",action="append",help="category to review; defaults to all seven launch categories"); q.add_argument("--reviewer"); q.add_argument("--evidence",action="append",default=[],help="review evidence path or URL; repeatable"); q.add_argument("--apply",action="store_true",help="write explicit editorial approval metadata"); q.add_argument("--confirm",action="store_true",help="confirm that the selected records were actually reviewed by a human")
1014
1164
  q=sub.add_parser("category-hash"); q.add_argument("--project",default="."); q.add_argument("--copy-data",required=True); q.add_argument("--apply",action="store_true",help="backfill deterministic provenance hashes in this generated artifact")
1015
1165
  q=sub.add_parser("fact-audit"); q.add_argument("--project",default="."); q.add_argument("--backfill-source",action="store_true",help="copy the catalogue sourceUrl into records that have no sourceUrl"); q.add_argument("--overrides",help="JSON of human-approved, evidenced descriptions for provider gaps"); q.add_argument("--apply-overrides",action="store_true",help="apply only approved fact overrides to the draft catalogue")
1166
+ q=sub.add_parser("catalogue-check"); q.add_argument("source"); q.add_argument("--project",default="."); q.add_argument("--provider",default="fresha",choices=["fresha"]); q.add_argument("--treatment",action="append",required=True,help="provider treatment name; repeat for multiple checks")
1167
+ q=sub.add_parser("retirement-audit"); q.add_argument("--project",default="."); q.add_argument("--catalogue",default=".maggie/booking/services.json"); q.add_argument("--evidence",required=True,help="sanitized runtime retirement evidence JSON"); q.add_argument("--output",default="docs/service-retirement-audit.json")
1016
1168
  q=sub.add_parser("capability-audit", help="validate provider variant declarations and fixture evidence"); q.add_argument("--project",default="."); q.add_argument("--provider"); q.add_argument("--catalogue",default=".maggie/booking/services.json"); q.add_argument("--capabilities-file",default=".maggie/booking/provider-capabilities.json"); q.add_argument("--fixture",help="sanitized provider fixture JSON"); q.add_argument("--output")
1017
1169
  q=sub.add_parser("category-context"); q.add_argument("--project",default="."); q.add_argument("--output")
1018
1170
  q=sub.add_parser("validate-copy"); q.add_argument("--project",default="."); q.add_argument("--service-id",required=True); q.add_argument("--copy-data",required=True,help="AI-authored service copy JSON")
@@ -1026,6 +1178,7 @@ def main():
1026
1178
  a=p.parse_args(); project=root(a)
1027
1179
  if a.command=="import": return cmd_import(a)
1028
1180
  if a.command=="sync": return cmd_sync(a)
1181
+ if a.command=="sync-report": return cmd_sync_report(a)
1029
1182
  if a.command=="run": return cmd_run(a)
1030
1183
  if a.command=="convert-page": return cmd_convert_page(a)
1031
1184
  if a.command in {"polish","validate-polish"}: return cmd_polish(a)
@@ -1035,6 +1188,8 @@ def main():
1035
1188
  if a.command=="category-editorial-review": return cmd_category_editorial_review(a)
1036
1189
  if a.command=="category-hash": return cmd_category_hash(a)
1037
1190
  if a.command=="fact-audit": return cmd_fact_audit(a)
1191
+ if a.command=="catalogue-check": return cmd_catalogue_check(a)
1192
+ if a.command=="retirement-audit": return cmd_retirement_audit(a)
1038
1193
  if a.command=="capability-audit": return cmd_capability_audit(a)
1039
1194
  if a.command=="category-context": return cmd_category_context(a)
1040
1195
  if a.command=="validate-copy": return cmd_validate_copy(a)
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@topy-ai/maggie",
3
- "version": "0.7.21",
3
+ "version": "0.7.22",
4
4
  "description": "Install and manage Maggie Skills for AI coding agents",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -57,7 +57,7 @@ The stable model is deliberately small and provider-neutral:
57
57
  - **category**: `level1` and `level2`; a service may belong to more than one
58
58
  provider category through `additional.provider.categories`;
59
59
  - **service**: identity, slug, title, description, active/archived status,
60
- timestamps, and variants;
60
+ timestamps, supply/display state, and variants;
61
61
  - **variant**: identity, display name, duration, price, currency, discount or
62
62
  price range when the provider exposes them;
63
63
  - **booking actions**: booking URL, payment URL, and optional action metadata;
@@ -67,6 +67,44 @@ The stable model is deliberately small and provider-neutral:
67
67
  - **sync**: source, fetched time, content hash, parser version, and change
68
68
  status.
69
69
 
70
+ The portable lifecycle model keeps two owners separate:
71
+
72
+ - `supplyState`: provider sync state, either `live` or `withdrawn`;
73
+ - `displayState`: human publication decision, either `published`, `hidden`, or
74
+ `retired`.
75
+
76
+ Sync may change `supplyState` when the provider adds or removes a service, but
77
+ must preserve `displayState`. A withdrawn/published record is an explicit
78
+ interim state: keep the URL answerable, suppress booking and indexing, and
79
+ open a human retirement decision. The legacy `status` field may remain for
80
+ backwards compatibility, but it must not be used as both provider availability
81
+ and display policy.
82
+
83
+ The `slug` is site-owned. A provider rename may produce a `slugProposal` in a
84
+ sync report, but the stored slug remains unchanged until a person accepts the
85
+ proposal and the host writes the new slug plus its redirect in one transaction.
86
+
87
+ ## Withdrawal and retirement contract
88
+
89
+ Provider removal is not permission to delete a public route. The adapter keeps
90
+ the record with `supplyState: "withdrawn"` and the host chooses a reviewed
91
+ `retirement.ending`:
92
+
93
+ | Ending | Required host response | Required safety signals |
94
+ |---|---|---|
95
+ | `pending` | HTTP 200 interim page | unavailable notice, no booking action, `noindex`, absent from sitemap |
96
+ | `redirect` | HTTP 301 or 308 to a different same-site route | no booking action, absent from sitemap |
97
+ | `tombstone` | HTTP 200 unavailable page | unavailable notice, no booking action, `noindex`, absent from sitemap |
98
+ | `gone` | HTTP 410 | no booking action, absent from sitemap |
99
+
100
+ The host must preserve the old route long enough to apply its chosen ending,
101
+ and must not return a generic 404 as a substitute for the reviewed contract.
102
+ `maggie service retirement-audit` validates sanitized runtime evidence against
103
+ the catalogue and records only service IDs, ending decisions, statuses, and
104
+ safe pass/fail metadata. It does not fetch private routes, store response
105
+ bodies, or invent redirect targets. A person must review redirect destination,
106
+ copy, and accessibility before publication.
107
+
70
108
  The service page must render only canonical fields. `additional` is an
71
109
  explicit extension point for provider-specific or future fields and must be
72
110
  namespaced by concern (`provider`, `location`, `presentation`, `compliance`,
@@ -74,6 +112,21 @@ namespaced by concern (`provider`, `location`, `presentation`, `compliance`,
74
112
  nullable-safe, and must never override canonical fields. Unknown data stays in
75
113
  `additional`; it is not guessed into the stable model.
76
114
 
115
+ ## Source ownership and onboarding precedence
116
+
117
+ When a host onboarding flow collects a value directly from the merchant, that
118
+ explicit value is authoritative for the project. Provider research, imported
119
+ profiles, search results, and inferred metadata may fill an empty field only;
120
+ they must never replace a non-empty merchant entry. Persist the source of each
121
+ value when the host supports it, and show a conflict for human review when two
122
+ non-empty sources disagree. This rule applies to the website URL as well as
123
+ business name, address, phone, category, and booking-provider identity.
124
+
125
+ The shared service-booking parser does not implement a host's onboarding UI.
126
+ Adapters must enforce this precedence before passing project identity into
127
+ provider research or sync. A provider URL is evidence about the provider
128
+ catalogue, not permission to overwrite the merchant's website URL.
129
+
77
130
  Pages can also enter the model without a provider. A normal page conversion
78
131
  creates a `draft` service with `additional.conversion.sourcePage` and no
79
132
  invented price, duration, or booking URL. It becomes `active` only after the
@@ -104,6 +157,14 @@ silent conversion, idempotent add/update/archive synchronisation, source
104
157
  snapshots/checksums, import-run audit records and generated service pages with
105
158
  provider booking CTAs.
106
159
 
160
+ When a provider exposes an authoritative venue-owned catalogue, use that
161
+ adapter extraction for sync decisions and treatment-presence checks. Generic
162
+ JSON-LD, navigation links, and page-wide text can contain marketplace or other
163
+ business entities and are not proof that the connected venue still sells a
164
+ treatment. A read-only catalogue query should call the same parser as import
165
+ and sync so an operator can review a reproducible answer before applying a
166
+ withdrawal.
167
+
107
168
  It must not claim real-time availability, booking creation, cancellation sync
108
169
  or webhooks unless an official Fresha partner/API capability is available for
109
170
  that account. Public booking links remain the safe fallback. The paid Fresha