@topy-ai/maggie 0.7.5 → 0.7.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +25 -5
- package/README.zh-TW.md +1 -1
- package/bundled-contracts/maggiedash/agent-content-v1.json +3 -1
- package/bundled-references/localization-extraction-render-prd.md +55 -0
- package/bundled-skills/maggie-content-localization/SKILL.md +1 -1
- package/bundled-skills/maggie-dash/SKILL.md +9 -5
- package/bundled-skills/maggie-feedback/SKILL.md +4 -1
- package/bundled-skills/maggie-seo-geo/SKILL.md +8 -4
- package/bundled-tools/clis/maggie_agent_content.py +72 -12
- package/bundled-tools/clis/maggie_feedback.py +35 -7
- package/bundled-tools/runtime/surface_coverage.py +82 -8
- package/package.json +1 -1
- package/references/localization-extraction-render-prd.md +55 -0
package/README.md
CHANGED
|
@@ -93,7 +93,7 @@ maggie dash install | init | status | migrate | cms ...
|
|
|
93
93
|
maggie dash transition ... # explicit content approval transition
|
|
94
94
|
maggie dash variant ... # service variant create/review/preview/publish
|
|
95
95
|
maggie dash sections ... # validate/catalogue/identity keys
|
|
96
|
-
maggie agent-content write ... # host-authorized content bridge
|
|
96
|
+
maggie agent-content write ... # host-authorized, origin-bound content bridge
|
|
97
97
|
maggie verification coverage ... # changed surface/locale evidence gate
|
|
98
98
|
maggie clone ... # authorized homepage capture
|
|
99
99
|
maggie design ... # authorized interior-page design
|
|
@@ -112,6 +112,18 @@ maggie design icon-inventory --source-dir src --runtime assets/icons.css
|
|
|
112
112
|
maggie api lifecycle | site-audit | ops audit
|
|
113
113
|
```
|
|
114
114
|
|
|
115
|
+
Agent content writes must declare the exact host origin that minted the
|
|
116
|
+
short-lived token; the CLI rejects nested credential keys and cross-origin
|
|
117
|
+
redirects:
|
|
118
|
+
|
|
119
|
+
```bash
|
|
120
|
+
maggie agent-content write \
|
|
121
|
+
--url https://example.test/api/maggie/agent-content.json \
|
|
122
|
+
--allowed-origin https://example.test \
|
|
123
|
+
--token-env MAGGIE_CONTENT_TOKEN \
|
|
124
|
+
--payload .maggie/content-operation.json
|
|
125
|
+
```
|
|
126
|
+
|
|
115
127
|
Run `maggie --help` for the complete syntax. `maggie-emdash` is migration
|
|
116
128
|
history only and is not an installable package skill.
|
|
117
129
|
|
|
@@ -194,8 +206,8 @@ artifact schemas.
|
|
|
194
206
|
Recommended upgrade sequence for the current release:
|
|
195
207
|
|
|
196
208
|
```bash
|
|
197
|
-
npx @topy-ai/maggie@0.7.
|
|
198
|
-
npx @topy-ai/maggie@0.7.
|
|
209
|
+
npx @topy-ai/maggie@0.7.6 update --project . --force
|
|
210
|
+
npx @topy-ai/maggie@0.7.6 cleanup --project .
|
|
199
211
|
```
|
|
200
212
|
|
|
201
213
|
Maintainers should pass npm credentials through the repository helper, never
|
|
@@ -205,7 +217,7 @@ as a command-line argument:
|
|
|
205
217
|
node scripts/publish-npm.mjs --maggie-env-file ../.env
|
|
206
218
|
```
|
|
207
219
|
|
|
208
|
-
The 0.7.
|
|
220
|
+
The 0.7.6 workflow adds the installable MaggieDash admin distribution and
|
|
209
221
|
audited CMS operations (`cms revisions`,
|
|
210
222
|
`trash`, `restore`, `schedule`, `publish-due`, `duplicate`, `redirect`, and
|
|
211
223
|
signed `preview`). It also adds a read-only MaggieDash `--diff`/`--dry-run`
|
|
@@ -216,7 +228,12 @@ adds stable section identities and reusable section presets, a machine-readable
|
|
|
216
228
|
MaggieDash section registry, a short-lived agent content bridge, restart gates
|
|
217
229
|
for out-of-band translation writes, deployment release-runner generation,
|
|
218
230
|
disk-versus-manifest skill inventory checks, and changed-surface/locale
|
|
219
|
-
verification.
|
|
231
|
+
verification. Surface evidence is route-level and rejects duplicate
|
|
232
|
+
observations; every route must include status, canonical, indexability, and
|
|
233
|
+
required metadata checks. Agent content writes require an explicit approved
|
|
234
|
+
origin, same-origin redirects, and recursive credential validation. Feedback
|
|
235
|
+
submissions persist and print only an allowlisted acknowledgement. The release
|
|
236
|
+
retains the existing service matching,
|
|
220
237
|
localization, seed-manifest, lockfile/analytics, sitemap, deployment and
|
|
221
238
|
rollback workflows.
|
|
222
239
|
|
|
@@ -308,6 +325,9 @@ maggie feedback submit .maggie/feedback/<feedback-id>.json \
|
|
|
308
325
|
Submission is never implicit. Project paths, source files, logs, and secrets
|
|
309
326
|
are excluded by default. The hosted endpoint stores the normalized feedback in
|
|
310
327
|
NoBlox for maintainer review; it does not automatically create a GitHub Issue.
|
|
328
|
+
The local submission record and command output keep only `status`, `feedbackId`,
|
|
329
|
+
and `requestId` from the provider acknowledgement; arbitrary response bodies
|
|
330
|
+
are not persisted or printed.
|
|
311
331
|
For public reports, use the repository's
|
|
312
332
|
[GitHub Issue Forms](https://github.com/TOPY-AI-LTD/ai-cmo-skills/issues/new/choose).
|
|
313
333
|
|
package/README.zh-TW.md
CHANGED
|
@@ -3,12 +3,14 @@
|
|
|
3
3
|
"$id": "https://maggiedash.noblox.app/contracts/agent-content-v1.schema.json",
|
|
4
4
|
"title": "MaggieDash agent content bridge v1",
|
|
5
5
|
"type": "object",
|
|
6
|
-
"required": ["method", "authorization", "tokenLifetime", "databaseCredentials"],
|
|
6
|
+
"required": ["method", "authorization", "tokenLifetime", "databaseCredentials", "originPolicy", "payloadSecurity"],
|
|
7
7
|
"properties": {
|
|
8
8
|
"method": {"const": "POST"},
|
|
9
9
|
"authorization": {"const": "Bearer short-lived session token"},
|
|
10
10
|
"tokenLifetime": {"const": "terminal session"},
|
|
11
11
|
"databaseCredentials": {"const": "never exposed to agent"},
|
|
12
|
+
"originPolicy": {"const": "explicit approved origin allowlist; redirects remain same-origin"},
|
|
13
|
+
"payloadSecurity": {"const": "recursive sensitive-key rejection and response redaction"},
|
|
12
14
|
"writes": {"type": "array", "items": {"enum": ["validated content operation", "revision", "audit event"]}}
|
|
13
15
|
},
|
|
14
16
|
"additionalProperties": false
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Localization source extraction and rendered validation
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
`maggie-content-localization` currently accepts a prepared `content.json` and
|
|
6
|
+
validates the resulting job. It does not extract translatable strings from a
|
|
7
|
+
real project, so malformed JSX fragments, multiline text nodes, untranslated
|
|
8
|
+
Latin terms, and script mismatches can pass the record validator while still
|
|
9
|
+
breaking the rendered page.
|
|
10
|
+
|
|
11
|
+
## Proposed stable workflow
|
|
12
|
+
|
|
13
|
+
Add a read-only extractor and a separate rendered validation gate:
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
maggie localization extract --project . --source-dir src \
|
|
17
|
+
--routes-file docs/routes.tsv --output .maggie/localization/source.json
|
|
18
|
+
maggie localization plan --content .maggie/localization/source.json ...
|
|
19
|
+
maggie localization validate .maggie/localization/<job>.json \
|
|
20
|
+
--source .maggie/localization/source.json \
|
|
21
|
+
--render-report .maggie/localization/rendered.json
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
The extractor must use an AST-aware adapter for JSX/TSX/Astro/Vue/Svelte
|
|
25
|
+
where available, with a conservative HTML fallback. It must never rewrite
|
|
26
|
+
source files. Each extracted string records file, line, route, source hash,
|
|
27
|
+
attribute/text-node kind, and an explicit translatable/non-translatable
|
|
28
|
+
decision. Fragments containing braces or JavaScript operators are rejected or
|
|
29
|
+
reported for review, not interpreted as text.
|
|
30
|
+
|
|
31
|
+
The validation report must cover multiline nodes, required metadata, protected
|
|
32
|
+
facts, script expectations, glossary/allowlist exceptions, and browser output.
|
|
33
|
+
The rendered report must identify route, screenshot, console errors, failed
|
|
34
|
+
asset requests, placeholder count, and untranslated required strings. Any
|
|
35
|
+
failure keeps the locale non-indexable and blocks publish.
|
|
36
|
+
|
|
37
|
+
## Acceptance criteria
|
|
38
|
+
|
|
39
|
+
- JSX conditionals and expressions are never emitted as translation strings.
|
|
40
|
+
- Text nodes spanning multiple lines are extracted and tested.
|
|
41
|
+
- Script/Latin-word checks support all registered languages and configurable
|
|
42
|
+
proper-noun/glossary allowlists.
|
|
43
|
+
- Source and translated strings are linked by stable IDs and source revision.
|
|
44
|
+
- A clean extraction plus clean rendered report is required before approval;
|
|
45
|
+
record-only validation is not sufficient.
|
|
46
|
+
- Unit fixtures cover Astro/JSX/Vue/Svelte/HTML, false positives, protected
|
|
47
|
+
facts, multiline nodes, script checks, and route-level browser errors.
|
|
48
|
+
- The feature remains local-first, dependency-adaptable, redacted, and
|
|
49
|
+
backward-compatible with existing `content.json` input.
|
|
50
|
+
|
|
51
|
+
## Out of scope
|
|
52
|
+
|
|
53
|
+
Automatic translation quality scoring, silent source rewriting, and automatic
|
|
54
|
+
publish. Translation generation and owner/Admin approval remain separate
|
|
55
|
+
operations.
|
|
@@ -114,7 +114,7 @@ line numbers, routes, and attribute/text-node kinds while skipping scripts,
|
|
|
114
114
|
styles, SVG, and dynamic expressions. Supplying source and render reports to
|
|
115
115
|
`validate` makes the checks fail closed; the prepared `content.json` flow
|
|
116
116
|
remains backward-compatible without those artifacts. See
|
|
117
|
-
[`docs/localization-extraction-render-prd.md`](../../
|
|
117
|
+
[`docs/localization-extraction-render-prd.md`](../../references/localization-extraction-render-prd.md)
|
|
118
118
|
for the artifact contract and limitations.
|
|
119
119
|
|
|
120
120
|
Use the shared translation index and route predicates from
|
|
@@ -52,14 +52,18 @@ bridge rather than editing release files or receiving database credentials:
|
|
|
52
52
|
```bash
|
|
53
53
|
maggie agent-content write \
|
|
54
54
|
--url https://example.test/api/maggie/agent-content.json \
|
|
55
|
+
--allowed-origin https://example.test \
|
|
55
56
|
--token-env MAGGIE_CONTENT_TOKEN --payload .maggie/content-operation.json
|
|
56
57
|
```
|
|
57
58
|
|
|
58
59
|
The host mints a short-lived session token. The bridge uses `POST` with
|
|
59
|
-
`Authorization: Bearer`,
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
60
|
+
`Authorization: Bearer`, requires the exact approved host origin, rejects
|
|
61
|
+
cross-origin redirects, delegates to the ordinary save/revision/audit path,
|
|
62
|
+
and returns 401/422 for invalid authority or content. Nested credential-like
|
|
63
|
+
payload keys are rejected and response values are recursively redacted. Never
|
|
64
|
+
put database credentials or the short-lived token in the payload or generated
|
|
65
|
+
reports. See
|
|
66
|
+
the [agent content contract](../../bundled-contracts/maggiedash/agent-content-v1.json).
|
|
63
67
|
|
|
64
68
|
## Workflow
|
|
65
69
|
|
|
@@ -161,7 +165,7 @@ the route returns 404 and self/chain redirects are rejected. Preview tokens
|
|
|
161
165
|
are short-lived HMAC-signed tokens; preview responses must use a real preview
|
|
162
166
|
URL with `noindex`, `X-Robots-Tag: noindex`, and `Cache-Control: no-store`.
|
|
163
167
|
Never place the secret in source control or generated reports. See the
|
|
164
|
-
provider-neutral [MaggieDash contract set](../../contracts/maggiedash/README.md).
|
|
168
|
+
provider-neutral [MaggieDash contract set](../../bundled-contracts/maggiedash/README.md).
|
|
165
169
|
|
|
166
170
|
Reusable arrangements are stored without page-specific section IDs and expand
|
|
167
171
|
into ordinary editable sections with fresh IDs on every insert. Use the
|
|
@@ -44,7 +44,10 @@ maggie feedback submit .maggie/feedback/<feedback-id>.json \
|
|
|
44
44
|
|
|
45
45
|
Only HTTPS endpoints are accepted. Project paths, file contents, secrets, and
|
|
46
46
|
logs are excluded by default. `--allow-project-context` is an explicit opt-in;
|
|
47
|
-
even then, send only a short public-safe context note.
|
|
47
|
+
even then, send only a short public-safe context note. After submission, the
|
|
48
|
+
local record and stdout retain only the provider acknowledgement fields
|
|
49
|
+
`status`, `feedbackId`, and `requestId`; arbitrary response bodies are never
|
|
50
|
+
stored or printed.
|
|
48
51
|
|
|
49
52
|
## Lifecycle
|
|
50
53
|
|
|
@@ -30,10 +30,14 @@ maggie verification coverage \
|
|
|
30
30
|
--evidence .maggie/surface-coverage.json
|
|
31
31
|
```
|
|
32
32
|
|
|
33
|
-
The evidence must contain one observation for every declared
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
33
|
+
The evidence must contain one observation for every declared
|
|
34
|
+
surface/locale/route key with a unique `route`, HTTP `status`, `passed: true`,
|
|
35
|
+
`canonical: true`, a boolean `indexable`, and a `metadata` object containing
|
|
36
|
+
every field named by that surface's `requiredMetadata`. Duplicate keys fail;
|
|
37
|
+
the verifier preserves the observation count and reports the duplicate and
|
|
38
|
+
the original failure instead of allowing a later passing observation to hide
|
|
39
|
+
it. Public-page checks alone cannot close an admin or API change, and one
|
|
40
|
+
source-language rendering cannot close a localized change.
|
|
37
41
|
|
|
38
42
|
Crawl reports include `summary.byLocaleTemplate` with total, passed, failed
|
|
39
43
|
and failed URLs per declared locale and template. Regions/scripts remain
|
|
@@ -9,32 +9,94 @@ import os
|
|
|
9
9
|
import sys
|
|
10
10
|
from pathlib import Path
|
|
11
11
|
from urllib.error import HTTPError, URLError
|
|
12
|
-
from urllib.
|
|
12
|
+
from urllib.parse import urlparse
|
|
13
|
+
from urllib.request import HTTPRedirectHandler, Request, build_opener
|
|
13
14
|
|
|
14
15
|
|
|
15
|
-
SENSITIVE_KEYS = {"
|
|
16
|
+
SENSITIVE_KEYS = {"databaseurl", "password", "secret", "token", "apikey", "authorization", "privatekey", "connectionstring", "clientsecret"}
|
|
17
|
+
SENSITIVE_KEY_PARTS = ("databaseurl", "password", "secret", "token", "apikey", "authorization", "privatekey", "connectionstring")
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
def normalized_key(value: object) -> str:
|
|
21
|
+
return "".join(character for character in str(value).lower() if character.isalnum())
|
|
22
|
+
|
|
23
|
+
|
|
24
|
+
def is_sensitive_key(value: object) -> bool:
|
|
25
|
+
key = normalized_key(value)
|
|
26
|
+
return key in SENSITIVE_KEYS or any(part in key for part in SENSITIVE_KEY_PARTS)
|
|
27
|
+
|
|
28
|
+
|
|
29
|
+
def validate_no_credentials(value: object) -> None:
|
|
30
|
+
if isinstance(value, dict):
|
|
31
|
+
for key, child in value.items():
|
|
32
|
+
if is_sensitive_key(key):
|
|
33
|
+
raise ValueError("payload must not contain credentials; use the short-lived token environment variable")
|
|
34
|
+
validate_no_credentials(child)
|
|
35
|
+
elif isinstance(value, list):
|
|
36
|
+
for child in value:
|
|
37
|
+
validate_no_credentials(child)
|
|
16
38
|
|
|
17
39
|
|
|
18
40
|
def validate_payload(value: object) -> dict:
|
|
19
41
|
if not isinstance(value, dict):
|
|
20
42
|
raise ValueError("payload must be a JSON object")
|
|
21
|
-
|
|
22
|
-
if forbidden:
|
|
23
|
-
raise ValueError("payload must not contain credentials; use the short-lived token environment variable")
|
|
43
|
+
validate_no_credentials(value)
|
|
24
44
|
if not value.get("op"):
|
|
25
45
|
raise ValueError("payload.op is required")
|
|
26
46
|
return value
|
|
27
47
|
|
|
28
48
|
|
|
29
|
-
def
|
|
49
|
+
def origin(url: str) -> str:
|
|
50
|
+
parsed = urlparse(url)
|
|
51
|
+
if parsed.scheme not in {"http", "https"} or not parsed.hostname or parsed.username or parsed.password or parsed.query or parsed.fragment:
|
|
52
|
+
raise ValueError("content host must be an absolute HTTP(S) URL without credentials or query parameters")
|
|
53
|
+
try:
|
|
54
|
+
port = parsed.port
|
|
55
|
+
except ValueError as error:
|
|
56
|
+
raise ValueError("content host origin has an invalid port") from error
|
|
57
|
+
default = (parsed.scheme == "http" and port == 80) or (parsed.scheme == "https" and port == 443)
|
|
58
|
+
return f"{parsed.scheme.lower()}://{parsed.hostname.lower()}" + (f":{port}" if port and not default else "")
|
|
59
|
+
|
|
60
|
+
|
|
61
|
+
def approved_origin(value: str) -> str:
|
|
62
|
+
parsed = urlparse(value)
|
|
63
|
+
if parsed.path not in {"", "/"}:
|
|
64
|
+
raise ValueError("approved origin must not contain a path")
|
|
65
|
+
return origin(value)
|
|
66
|
+
|
|
67
|
+
|
|
68
|
+
class SameOriginRedirectHandler(HTTPRedirectHandler):
|
|
69
|
+
def __init__(self, approved_origin: str):
|
|
70
|
+
super().__init__()
|
|
71
|
+
self.approved_origin = approved_origin
|
|
72
|
+
|
|
73
|
+
def redirect_request(self, req, fp, code, msg, headers, newurl): # noqa: N802
|
|
74
|
+
if origin(newurl) != self.approved_origin:
|
|
75
|
+
raise ValueError("content host redirected to a different origin")
|
|
76
|
+
return super().redirect_request(req, fp, code, msg, headers, newurl)
|
|
77
|
+
|
|
78
|
+
|
|
79
|
+
def redact_response(value: object) -> object:
|
|
80
|
+
if isinstance(value, dict):
|
|
81
|
+
return {key: "[REDACTED]" if is_sensitive_key(key) else redact_response(child) for key, child in value.items()}
|
|
82
|
+
if isinstance(value, list):
|
|
83
|
+
return [redact_response(child) for child in value]
|
|
84
|
+
return value
|
|
85
|
+
|
|
86
|
+
|
|
87
|
+
def write(url: str, token_env: str, payload: dict, allowed_origin: str, timeout: int = 30) -> dict:
|
|
30
88
|
token = os.environ.get(token_env, "")
|
|
31
89
|
if not token:
|
|
32
90
|
raise ValueError(f"{token_env} is not set")
|
|
91
|
+
approved = approved_origin(allowed_origin)
|
|
92
|
+
if origin(url) != approved:
|
|
93
|
+
raise ValueError("content host URL must match the approved origin")
|
|
33
94
|
request = Request(url, data=json.dumps(payload, ensure_ascii=False).encode("utf-8"), method="POST", headers={
|
|
34
95
|
"Accept": "application/json", "Content-Type": "application/json", "Authorization": f"Bearer {token}",
|
|
35
96
|
})
|
|
36
97
|
try:
|
|
37
|
-
|
|
98
|
+
opener = build_opener(SameOriginRedirectHandler(approved))
|
|
99
|
+
with opener.open(request, timeout=timeout) as response:
|
|
38
100
|
raw = response.read().decode("utf-8", errors="replace")
|
|
39
101
|
status = response.status
|
|
40
102
|
except HTTPError as error:
|
|
@@ -46,10 +108,7 @@ def write(url: str, token_env: str, payload: dict, timeout: int = 30) -> dict:
|
|
|
46
108
|
body = json.loads(raw) if raw else {}
|
|
47
109
|
except json.JSONDecodeError:
|
|
48
110
|
body = {"error": "host returned non-JSON"}
|
|
49
|
-
|
|
50
|
-
for key in SENSITIVE_KEYS:
|
|
51
|
-
body.pop(key, None)
|
|
52
|
-
return {"status": status, "passed": 200 <= status < 300, "response": body}
|
|
111
|
+
return {"status": status, "passed": 200 <= status < 300, "response": redact_response(body)}
|
|
53
112
|
|
|
54
113
|
|
|
55
114
|
def main() -> int:
|
|
@@ -57,13 +116,14 @@ def main() -> int:
|
|
|
57
116
|
sub = parser.add_subparsers(dest="command", required=True)
|
|
58
117
|
command = sub.add_parser("write", help="POST a validated content operation")
|
|
59
118
|
command.add_argument("--url", required=True)
|
|
119
|
+
command.add_argument("--allowed-origin", required=True, help="exact HTTP(S) origin approved by the host")
|
|
60
120
|
command.add_argument("--token-env", default="MAGGIE_CONTENT_TOKEN")
|
|
61
121
|
command.add_argument("--payload", type=Path, required=True)
|
|
62
122
|
command.add_argument("--timeout", type=int, default=30)
|
|
63
123
|
args = parser.parse_args()
|
|
64
124
|
try:
|
|
65
125
|
payload = validate_payload(json.loads(args.payload.read_text(encoding="utf-8")))
|
|
66
|
-
result = write(args.url, args.token_env, payload, args.timeout)
|
|
126
|
+
result = write(args.url, args.token_env, payload, args.allowed_origin, args.timeout)
|
|
67
127
|
print(json.dumps(result, ensure_ascii=False, indent=2))
|
|
68
128
|
return 0 if result["passed"] else 1
|
|
69
129
|
except (OSError, ValueError, RuntimeError) as error:
|
|
@@ -20,6 +20,9 @@ from urllib.request import Request, urlopen
|
|
|
20
20
|
|
|
21
21
|
DEFAULT_ENDPOINT = "https://feedback.noblox.app/api/feedback"
|
|
22
22
|
SECRET_RE = re.compile(r"(?i)(bearer\s+|(?:api[_-]?key|token|secret|password|npm_config[^=]*|authorization)\s*[=:]\s*)[^\s,;]+")
|
|
23
|
+
VERSION_RE = re.compile(r"\d+\.\d+\.\d+(?:[-+][A-Za-z0-9.-]+)?")
|
|
24
|
+
ACKNOWLEDGEMENT_KEYS = ("status", "feedbackId", "requestId")
|
|
25
|
+
REPO_ROOT = Path(__file__).resolve().parents[2]
|
|
23
26
|
|
|
24
27
|
|
|
25
28
|
def now() -> str:
|
|
@@ -35,17 +38,29 @@ def project_fingerprint(project: Path) -> str:
|
|
|
35
38
|
|
|
36
39
|
|
|
37
40
|
def installed_version(project: Path) -> str:
|
|
38
|
-
"""Resolve the
|
|
41
|
+
"""Resolve the exact runtime version and fail closed on source drift."""
|
|
39
42
|
if os.environ.get("MAGGIE_VERSION"):
|
|
40
|
-
|
|
41
|
-
|
|
43
|
+
value = safe_text(os.environ["MAGGIE_VERSION"])
|
|
44
|
+
return value if VERSION_RE.fullmatch(value) else "unknown"
|
|
45
|
+
install_path = project / ".maggie" / "install.json"
|
|
46
|
+
try:
|
|
47
|
+
value = json.loads(install_path.read_text(encoding="utf-8")).get("version", "")
|
|
48
|
+
if VERSION_RE.fullmatch(str(value)):
|
|
49
|
+
return str(value)
|
|
50
|
+
except (OSError, ValueError, json.JSONDecodeError, AttributeError):
|
|
51
|
+
pass
|
|
52
|
+
|
|
53
|
+
source_versions = []
|
|
54
|
+
for path in (REPO_ROOT / "VERSION", REPO_ROOT / "packages" / "maggie-cli" / "package.json"):
|
|
42
55
|
try:
|
|
43
56
|
value = json.loads(path.read_text(encoding="utf-8")) if path.suffix == ".json" else path.read_text(encoding="utf-8").strip()
|
|
44
57
|
if isinstance(value, dict): value = value.get("version", "")
|
|
45
|
-
if
|
|
46
|
-
|
|
58
|
+
if VERSION_RE.fullmatch(str(value)):
|
|
59
|
+
source_versions.append(str(value))
|
|
47
60
|
except (OSError, ValueError, json.JSONDecodeError):
|
|
48
61
|
continue
|
|
62
|
+
if source_versions and len(set(source_versions)) == 1:
|
|
63
|
+
return source_versions[0]
|
|
49
64
|
return "unknown"
|
|
50
65
|
|
|
51
66
|
|
|
@@ -124,6 +139,18 @@ def to_markdown(data: dict) -> str:
|
|
|
124
139
|
return "\n".join(lines)
|
|
125
140
|
|
|
126
141
|
|
|
142
|
+
def acknowledgement(value: object) -> dict:
|
|
143
|
+
"""Keep only safe, scalar acknowledgement fields from a provider response."""
|
|
144
|
+
if not isinstance(value, dict):
|
|
145
|
+
return {}
|
|
146
|
+
result = {}
|
|
147
|
+
for key in ACKNOWLEDGEMENT_KEYS:
|
|
148
|
+
item = value.get(key)
|
|
149
|
+
if isinstance(item, (str, int, float, bool)):
|
|
150
|
+
result[key] = safe_text(item)
|
|
151
|
+
return result
|
|
152
|
+
|
|
153
|
+
|
|
127
154
|
def read_feedback(path: Path) -> dict:
|
|
128
155
|
return load_json(path)
|
|
129
156
|
|
|
@@ -152,9 +179,10 @@ def submit(args: argparse.Namespace) -> int:
|
|
|
152
179
|
result = json.loads(raw) if raw else {}
|
|
153
180
|
except (HTTPError, URLError, TimeoutError) as error:
|
|
154
181
|
raise RuntimeError(f"feedback submission failed: {error}") from error
|
|
155
|
-
|
|
182
|
+
ack = acknowledgement(result)
|
|
183
|
+
data["submission"] = {"status": "submitted", "endpoint": endpoint, "submittedAt": now(), "acknowledgement": ack}
|
|
156
184
|
path.write_text(json.dumps(data, indent=2, ensure_ascii=False) + "\n", encoding="utf-8")
|
|
157
|
-
print(json.dumps({"feedbackId": data["feedbackId"], "status": "submitted", "
|
|
185
|
+
print(json.dumps({"feedbackId": data["feedbackId"], "status": "submitted", "acknowledgement": ack}, indent=2, ensure_ascii=False))
|
|
158
186
|
return 0
|
|
159
187
|
|
|
160
188
|
|
|
@@ -6,6 +6,7 @@ from typing import Any
|
|
|
6
6
|
|
|
7
7
|
|
|
8
8
|
SCHEMA = "maggie-surface-coverage.v1"
|
|
9
|
+
REQUIRED_EVIDENCE = ("status", "canonical", "indexability", "metadata")
|
|
9
10
|
|
|
10
11
|
|
|
11
12
|
def validate_contract(contract: object) -> dict[str, Any]:
|
|
@@ -23,26 +24,99 @@ def validate_contract(contract: object) -> dict[str, Any]:
|
|
|
23
24
|
if not isinstance(item, dict) or not item.get("name") or not isinstance(item.get("routes"), list) or not item["routes"] or not isinstance(item.get("locales"), list) or not item["locales"]:
|
|
24
25
|
errors.append(f"surfaces[{index}] needs name, routes, and locales")
|
|
25
26
|
continue
|
|
27
|
+
required = item.get("requiredEvidence")
|
|
28
|
+
if not isinstance(required, list) or not required or any(value not in REQUIRED_EVIDENCE for value in required):
|
|
29
|
+
errors.append(f"surfaces[{index}] needs valid requiredEvidence")
|
|
30
|
+
metadata = item.get("requiredMetadata")
|
|
31
|
+
if not isinstance(metadata, list) or any(not isinstance(value, str) or not value for value in metadata):
|
|
32
|
+
errors.append(f"surfaces[{index}] needs requiredMetadata")
|
|
33
|
+
route_values = [str(route) for route in item["routes"]]
|
|
34
|
+
if len(route_values) != len(set(route_values)):
|
|
35
|
+
errors.append(f"duplicate route in surface: {item['name']}")
|
|
36
|
+
locale_values = [str(locale) for locale in item["locales"]]
|
|
37
|
+
if len(locale_values) != len(set(locale_values)):
|
|
38
|
+
errors.append(f"duplicate locale in surface: {item['name']}")
|
|
26
39
|
for locale in item["locales"]:
|
|
27
40
|
key = (str(item["name"]), str(locale))
|
|
28
41
|
if key in seen:
|
|
29
42
|
errors.append(f"duplicate surface/locale: {key[0]}/{key[1]}")
|
|
30
43
|
seen.add(key)
|
|
31
|
-
|
|
44
|
+
expected = [f"{item['name']}/{locale}/{route}" for item in surfaces if isinstance(item, dict) and item.get("name") for locale in item.get("locales", []) for route in item.get("routes", [])]
|
|
45
|
+
return {"schemaVersion": SCHEMA, "passed": not errors, "errors": errors, "expectedPairs": sorted(f"{surface}/{locale}" for surface, locale in seen), "expectedRoutes": sorted(expected)}
|
|
46
|
+
|
|
47
|
+
|
|
48
|
+
def _truthy_check(value: object) -> bool:
|
|
49
|
+
if value is True:
|
|
50
|
+
return True
|
|
51
|
+
return isinstance(value, dict) and value.get("ok") is True
|
|
52
|
+
|
|
53
|
+
|
|
54
|
+
def _key_text(key: tuple[str, str, str]) -> str:
|
|
55
|
+
return f"{key[0]}/{key[1]}|{key[2]}"
|
|
56
|
+
|
|
57
|
+
|
|
58
|
+
def _route_result(item: dict[str, Any], required: list[str], required_metadata: list[str]) -> tuple[bool, list[str]]:
|
|
59
|
+
missing: list[str] = []
|
|
60
|
+
checks = item.get("checks") if isinstance(item.get("checks"), dict) else {}
|
|
61
|
+
if "status" in required:
|
|
62
|
+
try:
|
|
63
|
+
if int(item.get("status")) >= 400:
|
|
64
|
+
missing.append("status")
|
|
65
|
+
except (TypeError, ValueError):
|
|
66
|
+
missing.append("status")
|
|
67
|
+
if "canonical" in required and not _truthy_check(item.get("canonical", checks.get("canonical"))):
|
|
68
|
+
missing.append("canonical")
|
|
69
|
+
indexability = item.get("indexable", item.get("indexability", checks.get("indexability")))
|
|
70
|
+
if "indexability" in required and not isinstance(indexability, bool):
|
|
71
|
+
missing.append("indexability")
|
|
72
|
+
if "metadata" in required:
|
|
73
|
+
metadata = item.get("metadata") if isinstance(item.get("metadata"), dict) else checks.get("metadata")
|
|
74
|
+
if not isinstance(metadata, dict) or any(metadata.get(field) is not True for field in required_metadata):
|
|
75
|
+
missing.append("metadata")
|
|
76
|
+
if item.get("passed") is not True:
|
|
77
|
+
missing.append("passed")
|
|
78
|
+
return not missing, sorted(set(missing))
|
|
32
79
|
|
|
33
80
|
|
|
34
81
|
def compare(contract: dict[str, Any], observations: object) -> dict[str, Any]:
|
|
35
82
|
checked = validate_contract(contract)
|
|
36
83
|
if not checked["passed"]:
|
|
37
84
|
return {**checked, "observedPairs": [], "missing": [], "failed": []}
|
|
38
|
-
expected = {(str(item["name"]), str(locale)) for item in contract["surfaces"] for locale in item["locales"]}
|
|
39
|
-
observed: dict[tuple[str, str], dict[str, Any]] = {}
|
|
85
|
+
expected = {(str(item["name"]), str(locale), str(route)) for item in contract["surfaces"] for locale in item["locales"] for route in item["routes"]}
|
|
86
|
+
observed: dict[tuple[str, str, str], dict[str, Any]] = {}
|
|
87
|
+
duplicate_keys: list[str] = []
|
|
88
|
+
observation_count = 0
|
|
89
|
+
route_results: list[dict[str, Any]] = []
|
|
40
90
|
for item in observations if isinstance(observations, list) else []:
|
|
41
91
|
if not isinstance(item, dict):
|
|
42
92
|
continue
|
|
43
|
-
|
|
93
|
+
observation_count += 1
|
|
94
|
+
key = (str(item.get("surface") or ""), str(item.get("locale") or ""), str(item.get("route") or ""))
|
|
44
95
|
if key in expected:
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
96
|
+
if key in observed:
|
|
97
|
+
duplicate_keys.append(_key_text(key))
|
|
98
|
+
else:
|
|
99
|
+
observed[key] = item
|
|
100
|
+
missing = sorted(_key_text(key) for key in expected - observed.keys())
|
|
101
|
+
failed: list[str] = []
|
|
102
|
+
for key, item in observed.items():
|
|
103
|
+
surface = next(value for value in contract["surfaces"] if str(value["name"]) == key[0])
|
|
104
|
+
passed, missing_evidence = _route_result(item, list(surface["requiredEvidence"]), list(surface["requiredMetadata"]))
|
|
105
|
+
key_text = _key_text(key)
|
|
106
|
+
route_results.append({"route": key_text, "passed": passed, "missingEvidence": missing_evidence})
|
|
107
|
+
if not passed:
|
|
108
|
+
failed.append(key_text)
|
|
109
|
+
failed = sorted(set(failed))
|
|
110
|
+
return {
|
|
111
|
+
"schemaVersion": SCHEMA,
|
|
112
|
+
"passed": not missing and not failed and not duplicate_keys,
|
|
113
|
+
"expectedPairs": sorted(f"{surface}/{locale}" for surface, locale, _route in expected),
|
|
114
|
+
"expectedRoutes": sorted(_key_text(key) for key in expected),
|
|
115
|
+
"observedPairs": sorted(f"{surface}/{locale}" for surface, locale, _route in observed),
|
|
116
|
+
"observedRoutes": sorted(_key_text(key) for key in observed),
|
|
117
|
+
"observationCount": observation_count,
|
|
118
|
+
"duplicates": sorted(set(duplicate_keys)),
|
|
119
|
+
"missing": missing,
|
|
120
|
+
"failed": failed,
|
|
121
|
+
"routeResults": sorted(route_results, key=lambda value: value["route"]),
|
|
122
|
+
}
|
package/package.json
CHANGED
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Localization source extraction and rendered validation
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
`maggie-content-localization` currently accepts a prepared `content.json` and
|
|
6
|
+
validates the resulting job. It does not extract translatable strings from a
|
|
7
|
+
real project, so malformed JSX fragments, multiline text nodes, untranslated
|
|
8
|
+
Latin terms, and script mismatches can pass the record validator while still
|
|
9
|
+
breaking the rendered page.
|
|
10
|
+
|
|
11
|
+
## Proposed stable workflow
|
|
12
|
+
|
|
13
|
+
Add a read-only extractor and a separate rendered validation gate:
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
maggie localization extract --project . --source-dir src \
|
|
17
|
+
--routes-file docs/routes.tsv --output .maggie/localization/source.json
|
|
18
|
+
maggie localization plan --content .maggie/localization/source.json ...
|
|
19
|
+
maggie localization validate .maggie/localization/<job>.json \
|
|
20
|
+
--source .maggie/localization/source.json \
|
|
21
|
+
--render-report .maggie/localization/rendered.json
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
The extractor must use an AST-aware adapter for JSX/TSX/Astro/Vue/Svelte
|
|
25
|
+
where available, with a conservative HTML fallback. It must never rewrite
|
|
26
|
+
source files. Each extracted string records file, line, route, source hash,
|
|
27
|
+
attribute/text-node kind, and an explicit translatable/non-translatable
|
|
28
|
+
decision. Fragments containing braces or JavaScript operators are rejected or
|
|
29
|
+
reported for review, not interpreted as text.
|
|
30
|
+
|
|
31
|
+
The validation report must cover multiline nodes, required metadata, protected
|
|
32
|
+
facts, script expectations, glossary/allowlist exceptions, and browser output.
|
|
33
|
+
The rendered report must identify route, screenshot, console errors, failed
|
|
34
|
+
asset requests, placeholder count, and untranslated required strings. Any
|
|
35
|
+
failure keeps the locale non-indexable and blocks publish.
|
|
36
|
+
|
|
37
|
+
## Acceptance criteria
|
|
38
|
+
|
|
39
|
+
- JSX conditionals and expressions are never emitted as translation strings.
|
|
40
|
+
- Text nodes spanning multiple lines are extracted and tested.
|
|
41
|
+
- Script/Latin-word checks support all registered languages and configurable
|
|
42
|
+
proper-noun/glossary allowlists.
|
|
43
|
+
- Source and translated strings are linked by stable IDs and source revision.
|
|
44
|
+
- A clean extraction plus clean rendered report is required before approval;
|
|
45
|
+
record-only validation is not sufficient.
|
|
46
|
+
- Unit fixtures cover Astro/JSX/Vue/Svelte/HTML, false positives, protected
|
|
47
|
+
facts, multiline nodes, script checks, and route-level browser errors.
|
|
48
|
+
- The feature remains local-first, dependency-adaptable, redacted, and
|
|
49
|
+
backward-compatible with existing `content.json` input.
|
|
50
|
+
|
|
51
|
+
## Out of scope
|
|
52
|
+
|
|
53
|
+
Automatic translation quality scoring, silent source rewriting, and automatic
|
|
54
|
+
publish. Translation generation and owner/Admin approval remain separate
|
|
55
|
+
operations.
|