polytypo 1.2.0 → 1.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +648 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -0,0 +1,1281 @@
|
|
|
1
|
+
# Modes — the L2 contract
|
|
2
|
+
|
|
3
|
+
**Not a rule.** No entry in `spec/rules/order.json`, no locale data, no edits of its own. This
|
|
4
|
+
document specifies how a `html`, `markdown` or `yaml` document is decomposed into processable
|
|
5
|
+
text, how the rule pipeline is applied to it, and how the result is reassembled. It is normative
|
|
6
|
+
for all five runtimes and is parser-agnostic by construction: `parse5`, `nokogiri`, `lxml`,
|
|
7
|
+
`golang.org/x/net/html` and PHP's DOM disagree about almost everything this document does not
|
|
8
|
+
forbid them from doing. `yaml` mode is parser-**free** rather than parser-agnostic, for the
|
|
9
|
+
reason §3.8.1 measures.
|
|
10
|
+
**Spec version:** 1.3.0 (0.1.0 for everything except §3.3's class-membership table rows for
|
|
11
|
+
`nbsp` and `apostrophe`, split in 1.2.0, and §3.8, added in 1.3.0).
|
|
12
|
+
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
## 1. Purpose
|
|
16
|
+
|
|
17
|
+
`text` mode hands the pipeline one string. `html`, `markdown` and `yaml` hand it a document in
|
|
18
|
+
which _most_ of the characters must not be touched at all — markup, code, URLs, keys — and the
|
|
19
|
+
processable text is scattered across dozens of disconnected fragments. Two things then have to
|
|
20
|
+
be decided and neither is obvious: **what the rules see**, and **what comes out**.
|
|
21
|
+
|
|
22
|
+
The second is the easier one and this document answers it absolutely: **the output is the input
|
|
23
|
+
with a set of disjoint substring replacements applied, and nothing else.** The source is
|
|
24
|
+
_located_, never re-emitted. That is the only formulation under which five parsers can agree,
|
|
25
|
+
because it removes serialisation from the contract entirely — and in `yaml` mode it is also what
|
|
26
|
+
makes an entire class of reported corruption unreachable (§3.8.1).
|
|
27
|
+
|
|
28
|
+
The first is the interesting one, it is argued rather than asserted in §3.2, and everything
|
|
29
|
+
else in this document follows from it.
|
|
30
|
+
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
## 2. Locale data consumed
|
|
34
|
+
|
|
35
|
+
**None.** The mode layer is locale-independent. It decides _which characters_ the pipeline
|
|
36
|
+
sees; the pipeline decides what to do with them.
|
|
37
|
+
|
|
38
|
+
---
|
|
39
|
+
|
|
40
|
+
## 3. Algorithm
|
|
41
|
+
|
|
42
|
+
### 3.1 Definitions
|
|
43
|
+
|
|
44
|
+
- A **skipped region** is a maximal span of the source document that the pipeline must never
|
|
45
|
+
see and must never modify.
|
|
46
|
+
- A **processable span** is a maximal span of the source that is not skipped. Each span is
|
|
47
|
+
identified by its offsets in the **original source**, and those offsets are the only handle
|
|
48
|
+
the mode layer keeps.
|
|
49
|
+
- The **span sequence** `S₁ … Sₘ` is the processable spans in document order.
|
|
50
|
+
- `text` mode is the degenerate case: one span covering the whole input, no skipped regions.
|
|
51
|
+
Every statement below holds for it trivially.
|
|
52
|
+
|
|
53
|
+
### 3.2 The span model — the decision everything follows from
|
|
54
|
+
|
|
55
|
+
Three models are available. The choice is not a matter of taste; two of them produce wrong
|
|
56
|
+
output on ordinary documents.
|
|
57
|
+
|
|
58
|
+
**Model A — per-span, independent.** Run the whole pipeline on each span separately. Simple,
|
|
59
|
+
parallelisable, and **wrong**. Consider the entirely ordinary
|
|
60
|
+
|
|
61
|
+
```html
|
|
62
|
+
"He said <em>'hi'</em> loudly"
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
The spans are `"He said `, `'hi'` and ` loudly"`. Processed independently, the two double
|
|
66
|
+
quotes are unmatched in their own spans and stay straight, while `'hi'` pairs _in isolation_ at
|
|
67
|
+
**depth 1** — so it takes the locale's **primary** glyphs and comes out `“hi”` in `en-US`, where
|
|
68
|
+
the correct answer is the secondary pair `‘hi’` nested inside a converted outer quotation. That
|
|
69
|
+
is not a miss, it is visibly wrong output on a construction that appears in every second
|
|
70
|
+
paragraph of edited prose. Model A is rejected.
|
|
71
|
+
|
|
72
|
+
**Model B — naive concatenation.** Concatenate the spans, run the pipeline, redistribute by
|
|
73
|
+
offset. This fixes nesting, and introduces two defects of its own, both from **adjacencies that
|
|
74
|
+
do not exist in the document**:
|
|
75
|
+
|
|
76
|
+
- `"a"<code>x</code>"b"` concatenates to `"a""b"`. The `quotes` pass 1 same-kind adjacency veto
|
|
77
|
+
(`quotes.md` §3.2) sees two identical adjacent marks that are _not_ adjacent in the document,
|
|
78
|
+
vetoes both, and the surviving outer pair then quotes across the whole construction:
|
|
79
|
+
`“a""b”`. Wrong output.
|
|
80
|
+
- `He said <code>x</code> "hi"` concatenates to `He said "hi"` with two spaces that render as
|
|
81
|
+
one and are not adjacent in the source. `spaces` collapses them, deleting a character from a
|
|
82
|
+
text node that had a single space in it.
|
|
83
|
+
|
|
84
|
+
Model B is rejected. Note that both defects are invisible to every per-rule argument, because
|
|
85
|
+
each rule is behaving exactly as specified on the input it was given.
|
|
86
|
+
|
|
87
|
+
**Model C — concatenation with an explicit boundary marker. Adopted.** The spans are
|
|
88
|
+
concatenated with a **boundary marker** between each adjacent pair. The pipeline runs once, over
|
|
89
|
+
the whole marker-separated array. Edits are then redistributed to spans by offset.
|
|
90
|
+
|
|
91
|
+
The marker gives the rules what Model A denies them — the knowledge that `'hi'` sits inside a
|
|
92
|
+
larger quotation — while denying them what Model B wrongly grants: the belief that the last
|
|
93
|
+
character of one span touches the first character of the next.
|
|
94
|
+
|
|
95
|
+
**Markers are negative integers in the code-point array**, not Unicode code points. Every
|
|
96
|
+
runtime's array-of-code-points representation (`number[]`, `[]rune`, `list[int]`, `int[]`) holds
|
|
97
|
+
them without collision, and no input text can contain them. Implementations must not substitute
|
|
98
|
+
real code points — a private-use character or a noncharacter — because that reintroduces the
|
|
99
|
+
possibility of collision with author content and makes the classification table below a lie.
|
|
100
|
+
|
|
101
|
+
**There are two markers, and which one is used is decided by the source text of the gap:**
|
|
102
|
+
|
|
103
|
+
| Marker | Used when | Classified as |
|
|
104
|
+
| ------------------------ | ------------------------------------------------------------------------ | --------------------------------------------------- |
|
|
105
|
+
| **−1** — inline boundary | the skipped region between the two spans contains **no** line terminator | §3.3 |
|
|
106
|
+
| **−2** — line boundary | the skipped region contains **at least one** line terminator | a member of **`BREAK`**, for every rule, everywhere |
|
|
107
|
+
|
|
108
|
+
The test is on the raw source bytes of the gap, so it is decidable without asking the parser
|
|
109
|
+
anything and is identical in five runtimes. `a<em>x</em>b` gives −1; `foo\nbar`, `foo\n\nbar`
|
|
110
|
+
and `<p>a</p>\n<p>b</p>` all give −2.
|
|
111
|
+
|
|
112
|
+
**Why a line boundary is a `BREAK` and not just another opaque marker.** Every rule already has
|
|
113
|
+
correct, specified, fixture-covered behaviour at a line terminator, because `text` mode has real
|
|
114
|
+
ones: `spaces` refuses to touch a run that borders a `BREAK` (`spaces.md` §3.2 step 4), `dashes`
|
|
115
|
+
declines a token with a `BREAK` at `cp[L]`/`cp[R]` (§3.2 step 5), `quotes` counts `BREAK` inside
|
|
116
|
+
`SPACELIKE`, and `nbsp` never inserts against one. Reusing that behaviour is free and makes a
|
|
117
|
+
mode's output on hard-wrapped prose **identical to `text` mode's on the same characters**, which
|
|
118
|
+
is the property a mode that silently converts less than `text` cannot claim. See §7.7 for the
|
|
119
|
+
alternative that was rejected.
|
|
120
|
+
|
|
121
|
+
It also removes a parser dependency that would otherwise have been load-bearing. A Markdown hard
|
|
122
|
+
break is two spaces before a line ending; whether those spaces land inside a span or in a
|
|
123
|
+
structural token is a decision each parser makes differently. With a −2 marker the question does
|
|
124
|
+
not arise: the spaces sit next to a `BREAK`, and `spaces` protects a run bordering a `BREAK` in
|
|
125
|
+
every mode and every runtime.
|
|
126
|
+
|
|
127
|
+
### 3.3 How the marker is classified
|
|
128
|
+
|
|
129
|
+
Every rule already decides from character classes. The **−1 inline marker's** membership is
|
|
130
|
+
fixed here, once, and is normative for all rules present and future.
|
|
131
|
+
|
|
132
|
+
> **The class names below are per-rule, not spec-wide.** Each rule document defines its own
|
|
133
|
+
> `SPACELIKE`, `OPENISH` and `CLOSEISH`, and they differ: `nbsp`'s `SPACELIKE` is the largest
|
|
134
|
+
> (it includes the fixed-width spaces), and its `OPENISH`/`CLOSEISH` are **locale-data-driven**,
|
|
135
|
+
> since they contain the declared quote glyphs. Read each row as "the marker is a member of that
|
|
136
|
+
> rule's class of this name, wherever the rule defines one" — not as a claim that one class of
|
|
137
|
+
> that name exists. Conflating two same-named classes across documents is precisely what
|
|
138
|
+
> produced the `STRIP-BEFORE` divergence recorded in `dashes.md` §3.2 step 9. (The **−2 line marker** is
|
|
139
|
+
simply a member of `BREAK` and of nothing else, so it needs no table: every rule's existing
|
|
140
|
+
`BREAK` handling applies to it unchanged.)
|
|
141
|
+
|
|
142
|
+
| Class family | Marker is a member? |
|
|
143
|
+
| ------------------------------------------------------------------------------------------- | ------------------------------------------- |
|
|
144
|
+
| `SPACELIKE`, `SP`, `NOBREAK`, `OTHER-SPACE`, `BREAK` | **no** |
|
|
145
|
+
| `LETTER`, `DIGIT`, `ALNUM`, `UPPER`, `WORDISH`, `ROMAN` | **no** |
|
|
146
|
+
| `DASH`, `INERT-DASH`, `DASHISH`, `SENTENCE-DASH`, `HY`, `NBHY` | **no** |
|
|
147
|
+
| `DOTLIKE`, `STRIP-BEFORE`, `TERMINAL` | **no** |
|
|
148
|
+
| `STRAIGHT`, `SQ`, `DQ` | **no** |
|
|
149
|
+
| `OPEN-BRACKET`, `CLOSE-BRACKET` | **no** |
|
|
150
|
+
| any literal matching list (`abbreviations`, `beforeUnits`, `hyphen.*`, the trademark table) | **no** — the marker never matches a literal |
|
|
151
|
+
| `OPENISH` **and** `CLOSEISH` (`quotes`, `apostrophe`) | **yes, both** |
|
|
152
|
+
| `CLOSEISH` (`nbsp`) | **yes** |
|
|
153
|
+
| `OPENISH` (`nbsp`), `OPENQUOTE` (`apostrophe`) | **no** |
|
|
154
|
+
|
|
155
|
+
and one exemption:
|
|
156
|
+
|
|
157
|
+
> In `quotes` §3.2's `canOpen` right-test, which rejects a `right` in `CLOSEISH`, the marker
|
|
158
|
+
> receives **the same exemption `STRAIGHT` receives** and does not disqualify.
|
|
159
|
+
|
|
160
|
+
Everywhere else the marker is opaque content: it is "a content character" and nothing more —
|
|
161
|
+
**with one exception, which is normative and which resolves a contradiction between this
|
|
162
|
+
document and `spaces.md`.**
|
|
163
|
+
|
|
164
|
+
> ### Edge tests — the marker behaves as `NONE`
|
|
165
|
+
>
|
|
166
|
+
> Where a rule asks not _"what character is here"_ but _"am I at the edge of the text I am
|
|
167
|
+
> allowed to modify"_, the marker behaves as **`NONE`**, exactly as the end of the array does.
|
|
168
|
+
>
|
|
169
|
+
> There is currently **exactly one such test in the whole spec**: the boundary guard in
|
|
170
|
+
> `spaces.md` §3.2 step 4. Every other test in every other rule asks the first question, and the
|
|
171
|
+
> marker remains opaque content for all of them.
|
|
172
|
+
|
|
173
|
+
**Why the exception exists, and why it is not a licence to add more.** Read without it, this
|
|
174
|
+
document said a span-final run of spaces has content on both sides, so `spaces` may collapse or
|
|
175
|
+
delete it. `spaces.md` step 4 said the opposite in the same breath, and named the reason — a
|
|
176
|
+
trailing run "in `html` mode [is] frequently the only separator between two inline elements".
|
|
177
|
+
Both cannot be true. The contradiction is resolved in favour of `spaces.md` for three reasons:
|
|
178
|
+
|
|
179
|
+
- **`spaces` is the only rule that deletes.** Its boundary guard exists because deletion at an
|
|
180
|
+
edge is irreversible and the rule cannot see past the edge to know whether the character
|
|
181
|
+
mattered. At a span edge it genuinely cannot: there is markup there. Every other rule either
|
|
182
|
+
replaces one-for-one or is already constrained by §3.4, so none of them needs the exception and
|
|
183
|
+
none of them may claim it.
|
|
184
|
+
- **The literal reading is destructive, not merely different.** `<em>mot</em> !` in `fr` has its
|
|
185
|
+
space deleted by `spaces` (`!` is in `STRIP-BEFORE`), after which `nbsp`'s U+202F insertion is
|
|
186
|
+
discarded by the edge-growth rule because the insertion point is now a span edge. Output:
|
|
187
|
+
`<em>mot</em>!` — a character removed and nothing put back — where `text` mode on the same
|
|
188
|
+
prose gives `mot⍹!`. §7.4 calls a mode that diverges from `text` on identical characters
|
|
189
|
+
indefensible, and this would have been the sharpest instance of it. Under the exception the
|
|
190
|
+
space survives, `nbsp` converts it in place (1 → 1, permitted at an edge), and the two modes
|
|
191
|
+
agree.
|
|
192
|
+
- **It makes the round-trip guarantee stronger, never weaker.** The exception only ever causes
|
|
193
|
+
`spaces` to do less: `a <!--x--> b` is returned untouched instead of collapsed. Fewer edits
|
|
194
|
+
to bytes the rule cannot fully see is the direction §4 already points in.
|
|
195
|
+
|
|
196
|
+
**The generalisation, for future rules:** a rule that _deletes_ must treat a span edge as the end
|
|
197
|
+
of the text. A rule that replaces or inserts must not — §3.4 governs it instead. If a rule is
|
|
198
|
+
ever added that deletes, it inherits this clause and must say so here.
|
|
199
|
+
|
|
200
|
+
**Why `nbsp` differs (spec 1.2.0).** `nbsp` reads `CLOSEISH` in one place only, the right-context
|
|
201
|
+
guard of N1/N2 (`nbsp.md` §3.3 step 2). Membership there is what lets the no-break space come back
|
|
202
|
+
after `spaces` deleted the typed one before a mark that ends an inline element:
|
|
203
|
+
`<strong>Label :</strong>` in French. It reads `OPENISH` at the quote-glyph guard (`nbsp.md` §3.3
|
|
204
|
+
step 3), in the left-boundary tests of N3, N7, N9 and N10, and in N3's following-token guard. At
|
|
205
|
+
the quote-glyph guard, membership would lose the narrow space in `<em>non</em> !`. At the other
|
|
206
|
+
tests nobody has measured its effect.
|
|
207
|
+
Up to spec 1.1.0 this row read "yes, both" for `nbsp` too, while every runtime implemented
|
|
208
|
+
"neither". `nbsp.md` §7 item 12 records how the split was decided. `apostrophe`'s `OPENQUOTE`
|
|
209
|
+
(`apostrophe.md` §3.1, spec 1.2.0) excludes the marker for a simpler reason: case 3 already
|
|
210
|
+
accepts it through `CLOSEISH`, so membership would change nothing.
|
|
211
|
+
|
|
212
|
+
**Why dual `OPENISH`/`CLOSEISH` membership plus the exemption.** These three settings are what
|
|
213
|
+
make quotation marks pair correctly across an inline element, and they were derived by working
|
|
214
|
+
the cases, not by analogy:
|
|
215
|
+
|
|
216
|
+
| Input | Concatenation | Result |
|
|
217
|
+
| -------------------------------- | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
218
|
+
| `"<em>hello</em>"` | `"⟦hello⟧"` | the opening mark's `right` is the marker; it is exempted from the `CLOSEISH` rejection, so `canOpen` holds. The closing mark's `left` is the marker — not `NONE`, not `SPACELIKE` — so `canClose` holds. **They pair.** |
|
|
219
|
+
| `"a"<code>x</code>"b"` | `"a"⟦"b"` | the second mark's `right` is the marker, in `CLOSEISH`, so `canClose` holds and `"a"` pairs. The third mark's `left` is the marker, in `OPENISH`, so `canOpen` holds and `"b"` pairs. **Two pairs, correctly.** |
|
|
220
|
+
| `"He said <em>'hi'</em> loudly"` | `"He said ⟦'hi'⟧ loudly"` | outer pair at depth 1 → primary; inner pair enclosed by it at depth 2 → **secondary**. The Model A defect is gone. |
|
|
221
|
+
|
|
222
|
+
Treating the marker as `NONE` instead — the intuitive choice, "a span edge is like the edge of
|
|
223
|
+
the text" — fails the first row: both marks are dropped and nothing converts. Treating it as
|
|
224
|
+
plain opaque content fails the second: the closing mark of `"a"` is not `canClose` because its
|
|
225
|
+
`right` is in none of the accepted classes. Only the dual membership plus the exemption passes
|
|
226
|
+
all three.
|
|
227
|
+
|
|
228
|
+
### 3.4 Every other rule gets the right behaviour for free
|
|
229
|
+
|
|
230
|
+
Because the marker is opaque content and is in none of their classes (apart from the memberships
|
|
231
|
+
in the table of §3.3), the remaining rules decline to work across a boundary **without any
|
|
232
|
+
special-casing**:
|
|
233
|
+
|
|
234
|
+
| Situation | Concatenation | Outcome |
|
|
235
|
+
| ----------------------------- | ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
236
|
+
| `<em>foo </em>- bar` | `foo ⟦- bar` | `dashes`: the code point left of the dash is the marker, not a space, so `lsp = 0` while `rsp = 1` — the symmetry guard declines. Miss, not damage |
|
|
237
|
+
| `He said <code>x</code> "hi"` | `He said ⟦ "hi"` | `spaces`: two runs of length 1 separated by the marker, not one run of 2. No collapse |
|
|
238
|
+
| `5 <em>km</em>` | `5 ⟦km` | `nbsp` N5 requires `cp[a-1]` to be `SP` or `NBSP`; it is the marker. Declines |
|
|
239
|
+
| `из<em>-под</em>` | `из⟦-под` | `hyphen`: no listed form matches across the marker. Declines |
|
|
240
|
+
| `(c<em>)</em>` | `(c⟦)` | `symbols`: the literal `(c)` does not match. Declines |
|
|
241
|
+
|
|
242
|
+
**No rule can produce an edit whose span contains a marker.** Every rule's edit span is a
|
|
243
|
+
contiguous run of characters drawn from classes the marker does not belong to, or a literal
|
|
244
|
+
match the marker cannot participate in. This is what makes redistribution unambiguous, and it
|
|
245
|
+
is a **proof obligation on every future rule**: if a new rule could match across a marker, it
|
|
246
|
+
must state how its edits are redistributed. As a safety net, an implementation that computes an
|
|
247
|
+
edit whose span contains a marker **must discard that edit** and should report it as a bug.
|
|
248
|
+
|
|
249
|
+
**The edge-growth rule.** An edit is described by the span it replaces, `cp[p … q]` in the
|
|
250
|
+
concatenation, and a replacement sequence of length `r`. For an insertion, `q = p - 1` and the
|
|
251
|
+
replaced length is zero. Every edit lies wholly within one processable span `S = cp[s₀ … s₁]`.
|
|
252
|
+
|
|
253
|
+
> **An edit is discarded if it would place code points at an extremity of its span that were
|
|
254
|
+
> not there before.** Formally, with `d = q - p + 1` the replaced length and `w` the replacement:
|
|
255
|
+
>
|
|
256
|
+
> - if `p = s₀` and (`r > d`, **or** `r > 0` and `w[0]` is U+0020 and `cp[p]` is not U+0020)
|
|
257
|
+
> → **discard**;
|
|
258
|
+
> - if `q = s₁` and (`r > d`, **or** `r > 0` and `w[r−1]` is U+0020 and `cp[q]` is not U+0020)
|
|
259
|
+
> → **discard**.
|
|
260
|
+
>
|
|
261
|
+
> An insertion has `d = 0`, so `r > d` always holds and an insertion is discarded exactly when
|
|
262
|
+
> its position coincides with a span edge. That is what this section said before; the rule below
|
|
263
|
+
> generalises it.
|
|
264
|
+
|
|
265
|
+
**The second clause, and why the length test alone was not the rule it claimed to be (spec
|
|
266
|
+
1.3.0).** The sentence above is the rule; `r > d` was an incorrect formalisation of it, and the
|
|
267
|
+
gap is **`r = d`**. `dashes` P3 admits a run of two _or three_ dashes, so a tight `---` is a
|
|
268
|
+
3 → 3 replacement in every `-spaced` locale: the length test sees nothing, and U+0020 lands on
|
|
269
|
+
both extremities of the span anyway. Measured, and both cases are ones this document already
|
|
270
|
+
claims to have closed:
|
|
271
|
+
|
|
272
|
+
| Mode | Input | Locale | Before the second clause | What breaks |
|
|
273
|
+
| ---------- | ----------------- | ------- | ------------------------ | --------------------------------------------------------- |
|
|
274
|
+
| `html` | `a<em>---</em>b` | `de-DE` | `a<em> – </em>b` | the element begins and ends with a space it never held — the harm §7.3 exists to prevent |
|
|
275
|
+
| `markdown` | `x *---* y` | `de-DE` | `x * – * y` | the asterisks are de-flanked, so the span partition of §5 item 2 is not stable |
|
|
276
|
+
| `yaml` | `description: a:---b` | `de-DE` | `description: a: – b` | **the document no longer parses** — `: ` is now a mapping indicator |
|
|
277
|
+
|
|
278
|
+
The second clause tests the **character**, not the length, and it tests only U+0020 because
|
|
279
|
+
U+0020 is the only code point any rule emits whose meaning comes from its position rather than
|
|
280
|
+
from itself (§5 item 2). It leaves every case in the table below unchanged: a one-for-one
|
|
281
|
+
`"` → `“` still applies at an edge, a contraction still applies, and a `-spaced` edit that
|
|
282
|
+
already spans the surrounding spaces still applies, because there a U+0020 replaces a U+0020.
|
|
283
|
+
|
|
284
|
+
**This is a behaviour change to `html` and `markdown`, which shipped at v1.0.0**, not only to the
|
|
285
|
+
mode added in 1.3.0: three dashes at a span extremity are now left alone in a `-spaced` locale,
|
|
286
|
+
exactly as two already were. It needs its own changelog entry, in the terms the rest of the
|
|
287
|
+
public copy uses.
|
|
288
|
+
|
|
289
|
+
**Why it had to be generalised, and what it costs.** The previous formulation discarded
|
|
290
|
+
_insertions_ only. But a rule can grow a span with a **replacement**: `dashes` emits a `-spaced`
|
|
291
|
+
form by replacing the dash token with `U+0020 – U+0020`, which is one code point becoming three.
|
|
292
|
+
When the dash is the whole span, the two new spaces land on the span's edges, and the previous
|
|
293
|
+
filter did not see them because no insertion occurred. Three consequences, all observed:
|
|
294
|
+
|
|
295
|
+
- **html** — `a<em>--</em>b` became `a<em> – </em>b`. The element now begins and ends with a
|
|
296
|
+
space it never contained: the precise harm §7.3 exists to prevent, arriving by a route §7.3
|
|
297
|
+
did not cover.
|
|
298
|
+
- **markdown, span partition** — `*–*` became `* – *`, which de-flanks the asterisks so that on
|
|
299
|
+
the next run they are literal content rather than emphasis delimiters. §5 item 2 requires the
|
|
300
|
+
span partition to be stable; it was not.
|
|
301
|
+
- **markdown, unbounded growth** — `a\n\n–\n\n'` became `a\n\n – \n\n'` and then
|
|
302
|
+
`a\n\n – \n\n'`. The emitted space migrates into the line prefix, which is _outside every
|
|
303
|
+
span_, so the next run starts from a fresh span and emits another one. The document grows
|
|
304
|
+
without bound on every re-processing — which for content re-processed on each save is not a
|
|
305
|
+
cosmetic defect.
|
|
306
|
+
|
|
307
|
+
Both clauses are mechanically checkable from `(p, q, r, s₀, s₁)` and the replacement's own first
|
|
308
|
+
and last code points, with no rule cooperation and no knowledge of Markdown, HTML or YAML syntax.
|
|
309
|
+
Together they distinguish exactly the cases that matter:
|
|
310
|
+
|
|
311
|
+
| Edit | At an edge? | `d` → `r` | Verdict |
|
|
312
|
+
| -------------------------------------- | ----------- | --------- | ----------------------------------------------------------- |
|
|
313
|
+
| `"` → `“` (`quotes`) | yes | 1 → 1 | **applied** — a one-for-one replacement never grows an edge |
|
|
314
|
+
| `--` → `␣–␣` (`dashes`, `-spaced`) | yes | 2 → 3 | **discarded** by the length clause |
|
|
315
|
+
| `---` → `␣–␣` (`dashes` P3, `-spaced`) | yes | 3 → 3 | **discarded** by the character clause — the length clause misses it entirely |
|
|
316
|
+
| `a--b` → `a␣–␣b` (same edit, interior) | no | 2 → 3 | **applied** — growth is only a problem at an extremity |
|
|
317
|
+
| `␣--␣` → `␣–␣` (the edit spans its own spaces) | yes | 4 → 3 | **applied** — a U+0020 replaces a U+0020, so no edge changed |
|
|
318
|
+
| `(c)` → `©` (`symbols`) | yes | 3 → 1 | **applied** — shrinking is always safe |
|
|
319
|
+
| U+0020 → U+00A0 (`nbsp`) | yes | 1 → 1 | **applied** — U+00A0 is not U+0020 |
|
|
320
|
+
| insertion of U+202F (`nbsp` N1/N2) | yes | 0 → 1 | **discarded**, as before |
|
|
321
|
+
|
|
322
|
+
**Deletion at an edge is not restricted**, and does not need to be. A rule can only delete
|
|
323
|
+
U+0020 (`spaces`), and a deletion cannot bring into existence a structural token that was not
|
|
324
|
+
already there — it can only remove a character from inside a span. The whitespace that _is_
|
|
325
|
+
structural in Markdown — a line prefix, a list marker's indentation, the two spaces of a hard
|
|
326
|
+
break — is protected by the −2 line marker (§3.2): it sits beside a `BREAK`, and `spaces`
|
|
327
|
+
refuses to touch a run bordering one, in every mode.
|
|
328
|
+
|
|
329
|
+
`fr` `mot<em>!</em>` therefore still keeps no narrow no-break space, for the reason §7.3 gives,
|
|
330
|
+
and `a<em>--</em>b` is now left alone as well in every `-spaced` locale (§7.9).
|
|
331
|
+
|
|
332
|
+
(The original witness for this defect was `a<em>–</em>b`, with an authored en dash. It is still a
|
|
333
|
+
fixed point, but it no longer demonstrates anything: `dashes` §3.2 step 2a declines every token
|
|
334
|
+
containing an authored U+2013 or U+2014, so that input is now refused a step earlier and never
|
|
335
|
+
reaches the edge-growth rule. `--` is the live carrier.)
|
|
336
|
+
|
|
337
|
+
### 3.5 Applying the result
|
|
338
|
+
|
|
339
|
+
1. Extract the span sequence from the source. Record each span's **source offsets**.
|
|
340
|
+
2. Build the code-point array `S₁ ⌢ [m₁] ⌢ S₂ ⌢ [m₂] ⌢ … ⌢ Sₘ`, where each `mₖ` is **−1 or −2
|
|
341
|
+
by §3.2's test on the raw source of that gap** — not always −1. In `yaml` the −2 is the common
|
|
342
|
+
case, since a block scalar's spans are separated by a line terminator.
|
|
343
|
+
3. Run the pipeline **once**, in `order.json` order, over that array.
|
|
344
|
+
4. Each edit lies wholly within one span (§3.4). Map it back to source offsets.
|
|
345
|
+
5. Emit the **original source bytes**, with those replacements applied and nothing else changed.
|
|
346
|
+
|
|
347
|
+
Step 5 is the round-trip guarantee and is stated in full in §4.
|
|
348
|
+
|
|
349
|
+
### 3.6 Skip list — `html`
|
|
350
|
+
|
|
351
|
+
Skipped, exhaustively and exactly:
|
|
352
|
+
|
|
353
|
+
- the **entire subtree** of `code`, `pre`, `kbd`, `samp`, `var`, `script`, `style`, `textarea`,
|
|
354
|
+
`svg`, `math` — including any nested elements, whatever they are. A `<pre><em>x</em></pre>`
|
|
355
|
+
is skipped whole;
|
|
356
|
+
- **every attribute**, name and value, of every element, without exception;
|
|
357
|
+
- **every well-formed character reference** — `&name;`, `Ӓ`, `—` — as an opaque
|
|
358
|
+
unit. A text node containing one is split into spans around it, so the reference's spelling is
|
|
359
|
+
preserved exactly, which is the only way ` ` does not become a literal U+00A0 and back.
|
|
360
|
+
**A bare `&` that begins no well-formed reference stays inside its span.** No rule emits or
|
|
361
|
+
deletes `&`, so it is in no danger; lifting it out would create a boundary that suppresses
|
|
362
|
+
conversions on both sides of it for nothing — `Tom & Jerry's "book"` would lose the pairing of
|
|
363
|
+
its quotation marks. Well-formedness, not the presence of an ampersand, is what makes a
|
|
364
|
+
skipped region;
|
|
365
|
+
- comments, CDATA, doctype, processing instructions, and any prologue.
|
|
366
|
+
|
|
367
|
+
**Element names compare case-insensitively**, per HTML: `<CODE>`, `<Code>` and `<code>` are all
|
|
368
|
+
skipped. (In MDX, JSX names compare case-sensitively — §3.7.3.)
|
|
369
|
+
|
|
370
|
+
`svg` and `math` are skipped although neither appears in PLAN.md §3.2's list. Neither is prose:
|
|
371
|
+
in MathML a quotation mark, a hyphen and a prime are **operators and identifiers**, where
|
|
372
|
+
substituting a curly glyph or an en dash changes what the expression means, not how it looks.
|
|
373
|
+
The accepted cost is that `<svg><text>` and `<svg><title>` do hold real prose and are now left
|
|
374
|
+
untypeset. That asymmetry is deliberate — widening a skip list later is additive, while
|
|
375
|
+
narrowing one after release breaks every document that depended on the wider behaviour.
|
|
376
|
+
|
|
377
|
+
Everything else is processable, **including unknown and custom elements** (`<my-callout>`,
|
|
378
|
+
`<Foo>`). An unknown element is far more likely to be a wrapper than a code container, and
|
|
379
|
+
guessing from its name is exactly the kind of heuristic that behaves differently in five
|
|
380
|
+
runtimes. **The skip list is closed**: extending it is a spec change, not an implementation
|
|
381
|
+
decision.
|
|
382
|
+
|
|
383
|
+
### 3.7 Skip list — `markdown`
|
|
384
|
+
|
|
385
|
+
#### 3.7.1 The dialect is chosen by the caller, never detected
|
|
386
|
+
|
|
387
|
+
`markdown` is not one language. CommonMark and MDX disagree on ordinary documents: **MDX has no
|
|
388
|
+
indented code blocks and no `<https://…>` autolinks; CommonMark has no `{…}` expressions and no
|
|
389
|
+
JSX.** A document is frequently valid in both and means different things in each.
|
|
390
|
+
|
|
391
|
+
> **`markdown` mode takes a required `dialect` option**, one of:
|
|
392
|
+
>
|
|
393
|
+
> | value | language |
|
|
394
|
+
> | -------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
|
|
395
|
+
> | `"commonmark"` | CommonMark 0.31 **plus GFM** — tables, strikethrough, task lists, autolink literals |
|
|
396
|
+
> | `"mdx"` | MDX 3 — CommonMark plus GFM, **minus** indented code blocks and `<…>` autolinks, **plus** JSX elements and `{…}` expression containers |
|
|
397
|
+
>
|
|
398
|
+
> It has **no default**, and omitting it when `mode` is `"markdown"` throws, exactly as an
|
|
399
|
+
> omitted `locale` does (PLAN.md §5.1). `dialect` is ignored in `text` and `html` modes.
|
|
400
|
+
|
|
401
|
+
**Detection is forbidden**, and that is the important half. A heuristic is available and looks
|
|
402
|
+
reasonable — parse as MDX, and if it succeeds _and_ finds an MDX construct, call it MDX — but it
|
|
403
|
+
is silently wrong in exactly the case that matters. **One `<https://example.com>` autolink makes
|
|
404
|
+
a file invalid MDX**, so the heuristic falls back to CommonMark, and
|
|
405
|
+
`export const meta = {slug: "une-note"}` is then no longer an ESM statement but a **paragraph**:
|
|
406
|
+
it becomes processable text, and `fr` puts a narrow no-break space before its colon. A silent
|
|
407
|
+
false positive inside a machine-read field, in the author's own content format, produced by a
|
|
408
|
+
dialect the caller never chose.
|
|
409
|
+
|
|
410
|
+
The caller always knows which dialect they have — it is the file extension — and the library
|
|
411
|
+
never can. Requiring the option moves an unanswerable question to the party that holds the
|
|
412
|
+
answer, which is the reasoning that already makes `locale` required. A separate mode id for MDX
|
|
413
|
+
was rejected because MDX **is** Markdown with two constructs swapped, so an option on the mode
|
|
414
|
+
it varies says what is true; spec 1.3.0 later added `yaml` as a fourth mode id (§3.8), and the
|
|
415
|
+
two decisions do not conflict — YAML has no mode to hang off, and a `dialect` on nothing is not
|
|
416
|
+
a shape this contract has. **Ratified by the operator as public contract**: the throw carries
|
|
417
|
+
its own code, `POLYTYPO_INVALID_DIALECT`. See §7.8.
|
|
418
|
+
|
|
419
|
+
#### 3.7.2 A document that does not parse in its declared dialect
|
|
420
|
+
|
|
421
|
+
> **`transform` throws, carrying the stable code `POLYTYPO_MALFORMED_INPUT`.** The parser's own
|
|
422
|
+
> error type must never escape. **Ratified by the operator as public contract** — this code and
|
|
423
|
+
> `POLYTYPO_INVALID_DIALECT` are part of the taxonomy, not proposals.
|
|
424
|
+
|
|
425
|
+
**This can only happen in one place, and knowing that is what decides it.** Neither of the other
|
|
426
|
+
two languages can fail: HTML parsing is specified with total error recovery, and **every byte
|
|
427
|
+
sequence is valid CommonMark** — there is no such thing as a CommonMark syntax error. Only
|
|
428
|
+
`dialect: "mdx"` can reject a document, and only because MDX embeds JavaScript: an unterminated
|
|
429
|
+
JSX element, a malformed `{…}` expression, a broken `export`. A document that fails there is one
|
|
430
|
+
the author's own build already refuses.
|
|
431
|
+
|
|
432
|
+
Given that, returning the input untouched is the worse option. It would mean polytypo silently
|
|
433
|
+
succeeding on a file that is broken, hiding a build error behind a typography pass, and — the
|
|
434
|
+
common case — hiding the caller's own mistake of naming the wrong dialect. §3.7.1 forbids
|
|
435
|
+
detection precisely so that a wrong dialect is the caller's error rather than the library's
|
|
436
|
+
guess; swallowing the consequence would give that decision back with none of the information.
|
|
437
|
+
PLAN.md §5.1 fails fast on an unknown locale for the same reason, and this is the same shape.
|
|
438
|
+
|
|
439
|
+
Two constraints on the throw:
|
|
440
|
+
|
|
441
|
+
- **The runtime's parser error must be wrapped, never propagated.** A `VFileMessage`, a
|
|
442
|
+
`Nokogiri::SyntaxError` or a Python exception on the public surface puts a dependency's type in
|
|
443
|
+
the contract and is unreproducible in the other four runtimes. Only the code is contractual
|
|
444
|
+
(ARCHITECTURE.md §4.6); a message and a source position are useful and are not part of it.
|
|
445
|
+
- **This is the only way input can make `transform` throw.** Every other throw is caused by
|
|
446
|
+
options or by locale data. `transform` was previously total on its input, and it no longer is
|
|
447
|
+
— that is a real change in the shape of the public contract, and it is why the new code needs
|
|
448
|
+
sign-off (§7.8) rather than being an implementation detail.
|
|
449
|
+
|
|
450
|
+
#### 3.7.3 What is skipped
|
|
451
|
+
|
|
452
|
+
Skipped, exhaustively:
|
|
453
|
+
|
|
454
|
+
- **frontmatter** — a metadata block at the very start of the document, delimited by `---`
|
|
455
|
+
(YAML) or `+++` (TOML), skipped whole including its delimiters. Without it the closing `---`
|
|
456
|
+
reads as a setext underline, `title: Une note` becomes a paragraph, and `fr` inserts a narrow
|
|
457
|
+
no-break space before the colon of a machine-read field. That is a guaranteed false positive
|
|
458
|
+
on the M4 corpus (PLAN.md §8), where every file opens with frontmatter;
|
|
459
|
+
- fenced code blocks, including the info string and the fences;
|
|
460
|
+
- indented code blocks — **`commonmark` only**; MDX has none;
|
|
461
|
+
- inline code spans, including the backticks;
|
|
462
|
+
- autolinks `<https://…>` — **`commonmark` only**;
|
|
463
|
+
- **link and image destinations and titles** — the `(…)` of `[text](url "title")`, and the
|
|
464
|
+
definition line of a reference link. The **link text** is processable;
|
|
465
|
+
- **GFM constructs**: table pipes, delimiter and alignment rows, task-list checkboxes,
|
|
466
|
+
strikethrough delimiters, and footnote definition labels. GFM is enabled in **both** dialects,
|
|
467
|
+
and saying so is not decoration: §4's promise that table alignment rows survive is empty
|
|
468
|
+
unless tables are recognised at all, and §6's `[P: html, markdown]` marking on `dashes`' URL
|
|
469
|
+
bullet holds **only** because GFM autolink literals make a bare `https://…` a link — in pure
|
|
470
|
+
CommonMark a bare URL is ordinary text and would be typeset;
|
|
471
|
+
- HTML blocks and inline raw HTML, which are handed to the `html` skip list of §3.6 rather than
|
|
472
|
+
processed as markdown. **Inline HTML arrives as isolated tags with Markdown between them**,
|
|
473
|
+
not as a tree, so §3.6's subtree rule must be implemented as a **stack of open skipped
|
|
474
|
+
elements**: push on a start tag whose name is in the skip list, pop on its matching end tag,
|
|
475
|
+
and treat everything as skipped while the stack is non-empty. An unclosed skipped start tag
|
|
476
|
+
skips to the end of the block. Without the stack, `<code>a *b* c</code>` leaks its Markdown
|
|
477
|
+
middle back into a span;
|
|
478
|
+
- **MDX only**: JSX expression containers `{…}` in full, and every JSX attribute. **JSX element
|
|
479
|
+
children are processable** — `<Callout>Some text here</Callout>` is prose the author wrote and
|
|
480
|
+
wants typeset, and MDX is the author's own content format (PLAN.md §8, M4).
|
|
481
|
+
|
|
482
|
+
**Element names compare case-sensitively in JSX and case-insensitively in HTML.** `<Code>` is an
|
|
483
|
+
MDX component and is **not** in the skip list; `<code>`, `<CODE>` and `<Code>` in raw HTML all
|
|
484
|
+
are. This follows the two languages rather than any choice of ours, and an implementation that
|
|
485
|
+
lower-cases JSX names before matching will skip a component's children.
|
|
486
|
+
|
|
487
|
+
Nesting follows the same rule as `html`: a skipped construct is skipped whole, including
|
|
488
|
+
anything that looks processable inside it.
|
|
489
|
+
|
|
490
|
+
### 3.8 Skip list — `yaml`
|
|
491
|
+
|
|
492
|
+
**Spec 1.3.0.** `yaml` is the fourth mode id, and PLAN.md §3.2's "exactly three modes" is amended
|
|
493
|
+
by that decision. Two things about it decide everything below, and neither is true of the other
|
|
494
|
+
two modes: **no parser is used** (§3.8.1), and **the caller names the keys whose values hold
|
|
495
|
+
prose** (§3.8.2). The first is forced by what five ecosystems can report; the second by what YAML
|
|
496
|
+
is.
|
|
497
|
+
|
|
498
|
+
#### 3.8.1 Why a scanner, and why the issue's stated blocker is not one
|
|
499
|
+
|
|
500
|
+
The request (#18) arrived with a data-corruption report attached. A caller with an `openapi.yaml`
|
|
501
|
+
wrote their own splicer over a YAML library, and a block scalar's decoded value carries its
|
|
502
|
+
trailing line terminator only for some chomping indicators; they wrote the value back with the
|
|
503
|
+
wrong assumption and the following key was absorbed into the string. The file stayed
|
|
504
|
+
syntactically valid and became semantically wrong.
|
|
505
|
+
|
|
506
|
+
**That bug class is unreachable here, and it is unreachable by construction rather than by
|
|
507
|
+
care.** It is a property of decode → mutate → **re-encode**, and §1 already forbids the third
|
|
508
|
+
step: the source is located, never serialised. A chomping indicator is a byte of the source that
|
|
509
|
+
no span contains, so nothing this document permits can misread it, rewrite it, or lose it. The
|
|
510
|
+
same sentence disposes of indentation, anchors, tag handles and quoting style.
|
|
511
|
+
|
|
512
|
+
What the mode needs, then, is not a parser but a **locator**: something that reports the source
|
|
513
|
+
offsets of scalar content. The obvious move is to take those offsets from each ecosystem's YAML
|
|
514
|
+
library, and it does not survive contact with the five runtimes. Measured, 2026-09-20:
|
|
515
|
+
|
|
516
|
+
| Runtime | Library | Scalar positions |
|
|
517
|
+
| ------- | -------------------------- | -------------------------------------------------------------------- |
|
|
518
|
+
| JS | `yaml` | start **and** end, exact to the token |
|
|
519
|
+
| Python | PyYAML | start **and** end, exact to the token |
|
|
520
|
+
| Ruby | Psych | start **and** end, as line/column |
|
|
521
|
+
| Go | `gopkg.in/yaml.v3` | **start only** — `Node` carries `Line`/`Column` and no end position |
|
|
522
|
+
| PHP | `symfony/yaml` | **none** — the public surface decodes to values and reports no nodes |
|
|
523
|
+
|
|
524
|
+
Two of five cannot supply what the contract needs, and PHP has no alternative: `ext-yaml` binds
|
|
525
|
+
libyaml's value API and reports no positions either. A parser-backed `yaml` mode is therefore not
|
|
526
|
+
portable, and the §4 cost paragraph — "a runtime whose parser cannot report source offsets cannot
|
|
527
|
+
implement `html` mode conformantly" — would have excluded two runtimes outright rather than
|
|
528
|
+
describing a constraint they could meet.
|
|
529
|
+
|
|
530
|
+
So the locator is **specified here and hand-written per runtime**, like every core rule and like
|
|
531
|
+
locale resolution: a single left-to-right scan over the code-point array, no regular expressions,
|
|
532
|
+
no lookbehind. That is a cost — it is the first format knowledge polytypo owns rather than
|
|
533
|
+
delegates — and it buys the one thing delegation could not: the same spans in five runtimes.
|
|
534
|
+
|
|
535
|
+
#### 3.8.2 Why the caller names the keys, and why no heuristic can
|
|
536
|
+
|
|
537
|
+
> **`yaml` mode takes a required `keys` option: the list of mapping keys whose scalar values are
|
|
538
|
+
> processable.** It has **no default**, and omitting it when `mode` is `"yaml"` throws
|
|
539
|
+
> `POLYTYPO_INVALID_OPTION`, exactly as an omitted `dialect` throws in `markdown` mode. A value
|
|
540
|
+
> that is not a list of strings throws the same code. An **empty list is legal** and yields no
|
|
541
|
+
> spans — "process nothing" is a choice a caller may make, not an error.
|
|
542
|
+
|
|
543
|
+
The first draft of this section had no such option. It processed every string scalar and decided
|
|
544
|
+
prose by a content test — a scalar had to contain a space and a letter. Both the option and this
|
|
545
|
+
paragraph exist because that draft was **measured against this repository's own
|
|
546
|
+
`.github/workflows/`, and it corrupted them**: 37 of 106 spans changed, in `en-US`, `de-DE` and
|
|
547
|
+
`fr` alike. Two of the results:
|
|
548
|
+
|
|
549
|
+
```
|
|
550
|
+
run: |
|
|
551
|
+
if ! git cat-file -e "$SHA"; then → if! git cat-file -e "$SHA"; then
|
|
552
|
+
name: Node ${{ matrix.node }} → name: Node ${{matrix.node}}
|
|
553
|
+
```
|
|
554
|
+
|
|
555
|
+
The first is a shell script that no longer parses. Both passed the content test comfortably —
|
|
556
|
+
they have spaces and letters, because shell and template expressions are written in words.
|
|
557
|
+
|
|
558
|
+
**The reason no better content test exists is structural, and it is the whole argument for the
|
|
559
|
+
option.** `html` and `markdown` are prose formats with islands of code in them, so a closed skip
|
|
560
|
+
list works: the islands are marked by the syntax itself — a `<code>` element, a fence — and
|
|
561
|
+
naming them is a finite job this document can do once. **YAML is the inverse: a data format with
|
|
562
|
+
islands of prose in it**, and nothing in YAML's syntax distinguishes them. `description`,
|
|
563
|
+
`summary` and `title` hold sentences; `run`, `command`, `if`, `image`, `pattern` and `name` hold
|
|
564
|
+
things a machine reads. They are the same construct, spelled the same way, and they differ only
|
|
565
|
+
in what the schema above the YAML means by them — which the caller knows and the library cannot.
|
|
566
|
+
|
|
567
|
+
That is the same shape as the `dialect` decision (§3.7.1) and it is settled the same way:
|
|
568
|
+
**requiring the option moves an unanswerable question to the party that holds the answer.** The
|
|
569
|
+
request itself predicted it — "probably an option for which keys to process, since typesetting a
|
|
570
|
+
`url:` or an `id:` value is not wanted" — and the measurement above is what turned that from a
|
|
571
|
+
reasonable expectation into a requirement.
|
|
572
|
+
|
|
573
|
+
**Matching is exact and deliberately dumb**, so that five hand-written scanners cannot disagree:
|
|
574
|
+
a key matches iff its source text, with trailing U+0020 removed, equals a member of `keys` code
|
|
575
|
+
point for code point. No case folding — ARCHITECTURE.md §4.4 forbids locale-dependent case
|
|
576
|
+
operations and Turkish dotless ı is the standing reason. No paths, no globs, no wildcards, no
|
|
577
|
+
nesting: a key named `description` is processable wherever it occurs and at any depth. A **quoted
|
|
578
|
+
key never matches**, because §3.8.4 step 5 has already declined the line; that is an accepted
|
|
579
|
+
cost, recorded in §7.11.
|
|
580
|
+
|
|
581
|
+
**The option also removes a defect the content test carried**, which is why nothing of it
|
|
582
|
+
survives. A content test is a predicate on the text inside a span, and `spaces` can delete the
|
|
583
|
+
very U+0020 the predicate reads: `ref: ${{ steps.pin.outputs.sha }}` yielded one span on the
|
|
584
|
+
first run, `ref: ${{steps.pin.outputs.sha}}` on the second, and no span at all — §5's obligation
|
|
585
|
+
that the span partition be stable, **falsified on a one-line document**. `keys` is a predicate on
|
|
586
|
+
the key, which no rule can reach, so the partition is fixed by the source and the obligation
|
|
587
|
+
holds again. §5's yaml paragraph records this as the reason the predicate must stay structural.
|
|
588
|
+
|
|
589
|
+
#### 3.8.3 Skip by default — the rule the scan is built around
|
|
590
|
+
|
|
591
|
+
> **A construct the scan does not recognise with certainty yields no spans.** The worst outcome
|
|
592
|
+
> of a gap in the scan is prose left untypeset. It is never a changed byte.
|
|
593
|
+
|
|
594
|
+
This is the inverse of the posture `html` takes, where everything not in a closed skip list is
|
|
595
|
+
processable (§3.6). The asymmetry is deliberate and follows from §3.8.1: an HTML adapter has a
|
|
596
|
+
conforming parser telling it what every construct is, so "process what is not skipped" is a claim
|
|
597
|
+
it can back. The YAML scan has no such oracle, so it may only claim what it has proved. **`yaml`
|
|
598
|
+
mode names what is processable and skips the rest**, and the list below is closed: extending it
|
|
599
|
+
is a spec change.
|
|
600
|
+
|
|
601
|
+
A consequence worth stating: **`transform` never throws on input in `yaml` mode.** There is no
|
|
602
|
+
declared grammar to violate, so §3.7.2's `POLYTYPO_MALFORMED_INPUT` has no `yaml` counterpart. A
|
|
603
|
+
file that is not YAML at all yields few spans or none and comes back byte for byte. The only
|
|
604
|
+
throw this mode adds is the option throw of §3.8.2, which is about the call, not the input.
|
|
605
|
+
|
|
606
|
+
#### 3.8.4 The scan
|
|
607
|
+
|
|
608
|
+
The source is the code-point array of §3.1. A **line** is a maximal run containing no U+000A,
|
|
609
|
+
**and a U+000D immediately before that U+000A is not part of the line** — it is a terminator like
|
|
610
|
+
the U+000A itself, so it lies outside every span and comes back untouched. Line terminators are
|
|
611
|
+
never inside a span, so no rule can create or destroy one. `indent` is the number of leading
|
|
612
|
+
U+0020 on the line and `i` the index of the first code point after them.
|
|
613
|
+
|
|
614
|
+
The carriage-return clause is normative rather than obvious, and it is here because five runtimes
|
|
615
|
+
would otherwise inherit five answers from five standard-library calls. Without it a CRLF file
|
|
616
|
+
diverges from the same bytes with LF: the block header of §3.8.5 reads as `|` followed by U+000D
|
|
617
|
+
and is unrecognised, so the block yields no spans, while a plain scalar carries the carriage
|
|
618
|
+
return **inside** its span and hands it to the rules as content.
|
|
619
|
+
|
|
620
|
+
> **Normative, and the repair of a defect the first draft shipped: a line consumed by a construct
|
|
621
|
+
> is never scanned again.** Step 7 defines the one case — an inline value's continuation lines —
|
|
622
|
+
> and without it a multi-line quoted scalar, a multi-line flow collection and a folded plain
|
|
623
|
+
> scalar all leak their continuation lines back into the scan as if they were mappings.
|
|
624
|
+
|
|
625
|
+
Lines are processed in order.
|
|
626
|
+
|
|
627
|
+
1. The line yields no spans if it is empty or all U+0020, or if it contains U+0009 anywhere. A
|
|
628
|
+
tab makes indentation undecidable, which is what every step below depends on.
|
|
629
|
+
2. The line yields no spans if `cp[i]` is `#` (comment) or `%` (directive).
|
|
630
|
+
3. The line yields no spans if it begins at `i` with `---` or with `...` and the next code point
|
|
631
|
+
is U+0020 or the line ends there. The trailing-content form is declined as well as the bare
|
|
632
|
+
one: `--- key: value` is a node introduced by a document marker, and reading the marker as part
|
|
633
|
+
of a key is how the first draft produced a key of `--- key`.
|
|
634
|
+
4. **Block sequence entries are consumed, not skipped.** While `cp[i]` is `-` and `cp[i+1]` is
|
|
635
|
+
U+0020, advance `i` past the marker and past any U+0020 after it. Repeat — `- - key: v` is
|
|
636
|
+
legal and nests. If `i` reaches the end of the line, it yields no spans.
|
|
637
|
+
5. **Find the key.** Scan `j` upward from `i`.
|
|
638
|
+
- If `cp[j]` is `"`, `'`, `{`, `[`, `&`, `*`, `!` or `#`, the line yields no spans. This test
|
|
639
|
+
applies at **every** `j`, not only at `j = i`: it declines a quoted, flow, anchored, aliased
|
|
640
|
+
or tagged key, and equally a key that merely contains one of those characters anywhere —
|
|
641
|
+
`a!b: value` yields nothing. The wider reading is the normative one, and saying so is what
|
|
642
|
+
stops a second implementer from testing only the first position and getting a different span
|
|
643
|
+
set. The cost is recorded in §7.11.
|
|
644
|
+
- If `cp[j]` is `:` **and** `j` is the last code point of the line or `cp[j+1]` is U+0020, the
|
|
645
|
+
key is `cp[i … j−1]` with trailing U+0020 removed, and the scan continues at step 6.
|
|
646
|
+
- **Otherwise `j` advances by one and the scan continues**, including when `cp[j]` is a colon
|
|
647
|
+
that is not followed by U+0020 or the line end — such a colon is an ordinary character of the
|
|
648
|
+
key, and `a:b: v` has the key `a:b`. Stating the else-branch is not pedantry: without it the
|
|
649
|
+
first draft admitted two readings of that line, which is the exact divergence §3.8.1 says
|
|
650
|
+
the scan exists to prevent.
|
|
651
|
+
- If no such `j` exists, or the key is empty, the line yields no spans.
|
|
652
|
+
6. Let `v` be the index of the first code point after the colon that is not U+0020. **If there is
|
|
653
|
+
none, the value is empty**: the following, more-indented lines are a nested node and are
|
|
654
|
+
scanned on their own. The line yields no spans; go to the next line.
|
|
655
|
+
7. Otherwise the line carries an **inline value**, and the **value run** — every following line
|
|
656
|
+
that is blank or indented more than `indent` — belongs to that value. Those lines are never
|
|
657
|
+
scanned as lines of their own. The only spans they can contribute are the ones step 9 takes
|
|
658
|
+
from a recognised block scalar; every other step below yields no spans for the whole run.
|
|
659
|
+
8. **The key must be listed.** If the key of step 5 is not a member of `keys` (§3.8.2), the line
|
|
660
|
+
and its value run yield no spans. This is the step that makes `run:`, `if:` and `image:`
|
|
661
|
+
unreachable, and it is checked before the value's form so that an unlisted key costs nothing
|
|
662
|
+
to decline.
|
|
663
|
+
9. `cp[v]` selects the form:
|
|
664
|
+
- `&`, `*` or `!` — anchor, alias or tag — and `{` or `[` — a flow collection: no spans.
|
|
665
|
+
- `#`: the value is absent and only a comment follows: no spans.
|
|
666
|
+
- `|` or `>`: §3.8.5, block scalar.
|
|
667
|
+
- `"` or `'`: §3.8.6, quoted scalar.
|
|
668
|
+
- anything else: §3.8.6, plain scalar.
|
|
669
|
+
|
|
670
|
+
#### 3.8.5 Block scalars — one span per content line, and the chomping indicator never enters
|
|
671
|
+
|
|
672
|
+
The header is `cp[v]` followed by **at most one** chomping indicator (`-` or `+`) and **at most
|
|
673
|
+
one** indentation indicator (`1`–`9`), in either order, then optional U+0020 and an optional `#`
|
|
674
|
+
comment, then the end of the line. Any other header is unrecognised and the block yields no
|
|
675
|
+
spans. **The header is never inside a span.**
|
|
676
|
+
|
|
677
|
+
The block's content is the value run of §3.8.4 step 7. Let `bi` be the indentation of its first
|
|
678
|
+
non-blank line, or `indent + n` when the header carried an explicit indentation indicator `n`.
|
|
679
|
+
**One definition, and three conditions that make the whole block yield no spans** — the block is
|
|
680
|
+
consumed either way, never rescanned:
|
|
681
|
+
|
|
682
|
+
- any line of the run contains U+0009;
|
|
683
|
+
- the run's first non-blank line is indented less than `bi`, which can only happen when an
|
|
684
|
+
explicit indicator disagrees with the block as written;
|
|
685
|
+
- any later non-blank line of the run is indented less than `bi`.
|
|
686
|
+
|
|
687
|
+
The last one is the case the first draft left with two readings — `k: |` followed by a line at
|
|
688
|
+
four spaces and then one at two — and it is settled by bailing rather than by choosing, because
|
|
689
|
+
either choice would be a guess about which line the author meant to be content.
|
|
690
|
+
|
|
691
|
+
> **Each non-blank content line contributes exactly one span, `[lineStart + bi, lineEnd)`.** Blank
|
|
692
|
+
> lines contribute none. The indentation is outside every span; so is every line terminator, and
|
|
693
|
+
> so is the run of line terminators at the end of the block that the chomping indicator governs.
|
|
694
|
+
|
|
695
|
+
That last clause is the direct answer to the report in §3.8.1: **the trailing newlines are not in
|
|
696
|
+
any span, so no `yaml`-mode output can add or remove one.** `|`, `|-`, `|+`, `>`, `>-` and `>+`
|
|
697
|
+
are handled identically here because the difference between them lives entirely in bytes the mode
|
|
698
|
+
never touches.
|
|
699
|
+
|
|
700
|
+
Extra indentation beyond `bi` on a content line stays **inside** that line's span, because in a
|
|
701
|
+
literal block it is content. It is safe there because each content line is its own span and the
|
|
702
|
+
gap between two of them contains a line terminator: the marker is a **−2**, a member of `BREAK`,
|
|
703
|
+
and `spaces` refuses to touch a run bordering one (`spaces.md` §3.2 step 4). The −1 marker's
|
|
704
|
+
`NONE` exception of §3.3 is a different mechanism and is not what protects it.
|
|
705
|
+
|
|
706
|
+
#### 3.8.6 Quoted and plain scalars
|
|
707
|
+
|
|
708
|
+
**Quoted.** The scalar must open and close **on the same line**; if the matching quote is not on
|
|
709
|
+
that line the value yields no spans, and its continuation lines are consumed by §3.8.4 step 7
|
|
710
|
+
rather than rescanned. After the closing quote only U+0020 and an optional `#` comment may
|
|
711
|
+
follow. The span is the content between the quotes, exclusive of both.
|
|
712
|
+
|
|
713
|
+
Two bails make source characters and content characters the same thing, which §3.1's offset model
|
|
714
|
+
requires:
|
|
715
|
+
|
|
716
|
+
- a double-quoted scalar whose content contains U+005C yields no spans — `\n`, `\"` and `é`
|
|
717
|
+
are content the source spells with more characters than it has;
|
|
718
|
+
- a single-quoted scalar whose content contains two consecutive U+0027 yields no spans, for the
|
|
719
|
+
same reason.
|
|
720
|
+
|
|
721
|
+
This is the treatment `html` gives a character reference (§3.6), reached from the same constraint
|
|
722
|
+
rather than by analogy. **No colon test applies to a quoted scalar** — quoting neutralises the
|
|
723
|
+
colon, and applying the plain-scalar test here is what made the first draft decline
|
|
724
|
+
`title: "Chapter 1: the beginning"`.
|
|
725
|
+
|
|
726
|
+
**Plain.** The scalar runs to the end of the line, minus a trailing comment — the first `#`
|
|
727
|
+
preceded by U+0020, and that U+0020 with it — and minus any remaining trailing U+0020. Then:
|
|
728
|
+
|
|
729
|
+
- **the continuation bail**: if the value run of §3.8.4 step 7 is not empty, the scalar is a
|
|
730
|
+
multi-line plain scalar and yields no spans. Its line folding is a construct this scan does not
|
|
731
|
+
claim;
|
|
732
|
+
- **compact nesting is not a value**: the scalar yields no spans if it begins with `-` followed by
|
|
733
|
+
U+0020 or the line end (`key: - item` opens a sequence), or if it contains a colon followed by
|
|
734
|
+
U+0020 or the line end (`key: a .:` is a mapping whose key is `a .`). The second test is exact
|
|
735
|
+
rather than conservative — **a plain scalar can never contain a colon in that position** — and
|
|
736
|
+
it runs here, after the comment has been removed, so that `key: prose # note: here` is not
|
|
737
|
+
declined for a colon that is not in the scalar at all.
|
|
738
|
+
|
|
739
|
+
> **`:` and `#` are opaque one-code-point skipped units inside a plain scalar.** The scalar's
|
|
740
|
+
> range is split at each of them, exactly as a character reference splits an HTML text node.
|
|
741
|
+
|
|
742
|
+
Without the split, `yaml` mode corrupts documents, and the two witnesses are ordinary:
|
|
743
|
+
|
|
744
|
+
| Source | Locale | Without the split | Reparsed as |
|
|
745
|
+
| ---------- | ------- | ----------------- | -------------------------------------------------- |
|
|
746
|
+
| `k: a:--b` | `de-DE` | `k: a: – b` | **a parse error** — `: ` is now a mapping indicator |
|
|
747
|
+
| `k: a--#b` | `de-DE` | `k: a – #b` | `k` is `a –` — the rest became a comment |
|
|
748
|
+
|
|
749
|
+
Both come from the same source: `dashes` emits U+0020 in every `-spaced` locale, and **U+0020 is
|
|
750
|
+
the only thing that makes either character structural** — a `:` is a mapping indicator only when
|
|
751
|
+
a U+0020 follows it, a `#` a comment introducer only when a U+0020 precedes it. With the split,
|
|
752
|
+
`:` and `#` lie outside every span, so the dash token sits at a span **extremity**, and §3.4's
|
|
753
|
+
edge-growth rule discards a replacement there that is longer than what it replaced. The `-spaced`
|
|
754
|
+
form is 2 → 3. It is discarded, in `de-DE`, `fr` and `ru` alike.
|
|
755
|
+
|
|
756
|
+
**The em-tight form is 2 → 1, and it is applied** — `k: a:--b` becomes `k: a:—b` in `en-US`,
|
|
757
|
+
which reparses as the same one plain scalar it was. That is not a hole in the argument, it is the
|
|
758
|
+
argument: a contraction cannot emit U+0020, and U+0020 is the whole of what the split protects.
|
|
759
|
+
The length test of §3.4 separates exactly the two, with no knowledge of YAML — which is what
|
|
760
|
+
makes it bind rules not yet written, while a per-rule observation would not.
|
|
761
|
+
|
|
762
|
+
The accepted cost is that a conversion whose replacement would **grow** against a colon or a hash
|
|
763
|
+
is missed: `fr` inserts no narrow no-break space before a colon inside a YAML scalar, and
|
|
764
|
+
`one--two` takes its spaced dash only where the split leaves it interior to a span. That is the
|
|
765
|
+
same trade §7.3 already made for `mot<em>!</em>`, and in the same direction: a miss is visible to
|
|
766
|
+
the author and fixable in the source; a corrupted document is neither.
|
|
767
|
+
|
|
768
|
+
#### 3.8.7 What this was tested against
|
|
769
|
+
|
|
770
|
+
Three bodies of evidence, and the order matters: the first two were run against the **keyless
|
|
771
|
+
draft** and are reported here for what they failed to catch, which is as much a part of this
|
|
772
|
+
section as what they established.
|
|
773
|
+
|
|
774
|
+
**A real corpus, span placement only.** 70 YAML files — GitHub Actions workflows,
|
|
775
|
+
`docker-compose`, Dependabot, issue-form templates, and two OpenAPI documents including the one
|
|
776
|
+
the request came from. The scan's spans were checked against the `yaml` package's own scalar
|
|
777
|
+
ranges: **3135 spans, none overlapping a key, none outside a string value scalar, none whose text
|
|
778
|
+
was not a literal substring of that scalar's decoded value.** That establishes the **geometry** of
|
|
779
|
+
the scan and nothing else. It never ran the pipeline, so every span it approved could still have
|
|
780
|
+
been a shell script — and 37 of them were. Read alone it is exactly the kind of measurement that
|
|
781
|
+
looks like proof and is not.
|
|
782
|
+
|
|
783
|
+
**A bounded exhaustive sweep**, in the form `pipeline-idempotency.md` §6 requires of the rules.
|
|
784
|
+
Payloads of length 1–4 over the alphabet `a`, U+0020, `-`, **`---`**, `:`, `#`, `'`, `"`, `\` and
|
|
785
|
+
`.` in eleven document templates — all three scalar forms, sequence entries, compact nesting,
|
|
786
|
+
nested mappings, an interior block-scalar line, and a quoted scalar left open across a line — each
|
|
787
|
+
transformed in eight locales, reparsed, and compared against the original's **structure and
|
|
788
|
+
non-string leaves**. **977 680 cases, 99 274 of them changed by the transform: no parse failure,
|
|
789
|
+
no structural change, no idempotency failure.**
|
|
790
|
+
|
|
791
|
+
**The sweep is discriminating, and that was checked rather than assumed.** With §3.4's character
|
|
792
|
+
clause reverted, the same run reports **480 structural changes** — `k: a:---a` becoming
|
|
793
|
+
`k: a: – a`, a string turning into a nested mapping — in every locale that emits a spaced dash. A
|
|
794
|
+
measurement that passes both with and without the thing it is supposed to protect proves nothing,
|
|
795
|
+
and two earlier versions of this sweep were in exactly that position.
|
|
796
|
+
|
|
797
|
+
**Three things it could not find**, each worth naming because each cost something.
|
|
798
|
+
|
|
799
|
+
- It did not find the workflow damage, because a shell script inside a block scalar is a string
|
|
800
|
+
before the transform and a string after it. **Structural equality is not integrity**, and a
|
|
801
|
+
sweep shaped like this one cannot be the only guard on a data format.
|
|
802
|
+
- Its first alphabet omitted U+0022, so it could not reach a quoted scalar's own delimiters. The
|
|
803
|
+
multi-line-quoted template and the two characters it needs were added afterwards, and §5 now
|
|
804
|
+
requires them.
|
|
805
|
+
- Its alphabet then held `-` but not `---`, and at a payload length of four no combination of
|
|
806
|
+
single dashes reaches a three-dash run flanked by content. That is why the `r = d` hole in §3.4
|
|
807
|
+
survived two full runs of this sweep at 649 440 cases each, reported clean both times. **`---`
|
|
808
|
+
is in the alphabet as one token for that reason**, and §5's sweep obligation now names it.
|
|
809
|
+
|
|
810
|
+
**A real corpus, end to end.** The scan plus the full pipeline over the same 70 files in eight
|
|
811
|
+
locales, comparing the result against the original byte for byte. This is the measurement that
|
|
812
|
+
produced §3.8.2 and the only one of the three that can see a value damaged without being
|
|
813
|
+
destroyed:
|
|
814
|
+
|
|
815
|
+
| `keys` | Result |
|
|
816
|
+
| ------ | ------- |
|
|
817
|
+
| none — every string scalar processable (the keyless draft) | **37 of 106 spans changed in this repository's own workflows**, including `if !` → `if!` and `"$SHA"` → `“$SHA”` inside a `run:` block |
|
|
818
|
+
| every key that occurs anywhere in the corpus, 575 of them | 560 file/locale cases: no parse failure, no structural change, no idempotency failure — the scan's own safety envelope, at its most exposed |
|
|
819
|
+
| `description`, `summary`, `title` — what a caller would actually pass | **no workflow file changes at all**; 8 files change, and every one of them is prose-bearing: the two OpenAPI documents, four issue-form templates, and a package manifest |
|
|
820
|
+
|
|
821
|
+
The third row is the mode working as specified, and the first is why the option is not optional.
|
|
822
|
+
|
|
823
|
+
---
|
|
824
|
+
|
|
825
|
+
## 4. The round-trip guarantee
|
|
826
|
+
|
|
827
|
+
> **`transform` in `html`, `markdown` and `yaml` mode returns the input source with a set of
|
|
828
|
+
> disjoint substring replacements applied at recorded offsets. No other byte of the input
|
|
829
|
+
> changes, ever.**
|
|
830
|
+
|
|
831
|
+
This is stronger than "byte-identical when no changes are needed" (PLAN.md §3.2) and it implies
|
|
832
|
+
it. It is stated as a prohibition because that is the only form five parsers cannot drift on:
|
|
833
|
+
|
|
834
|
+
- **[P] The document is never serialised.** The parser is used to locate spans and is then
|
|
835
|
+
discarded. An implementation that reconstructs output from a DOM is non-conforming even if
|
|
836
|
+
its output happens to match.
|
|
837
|
+
- **[P] Attribute quoting is preserved** — `class=foo`, `class='foo'`, `class="foo"` all
|
|
838
|
+
survive as written. Attributes are skipped _and_ never re-emitted.
|
|
839
|
+
- **[P] Self-closing and void-element forms are preserved** — `<br>`, `<br/>`, `<br />`.
|
|
840
|
+
- **[P] Character-reference spelling is preserved** — `&` does not become `&`, ` `
|
|
841
|
+
does not become U+00A0, and a bare `&` that a parser would repair stays a bare `&`.
|
|
842
|
+
- **[P] Tag name case, attribute order, duplicate attributes, and whitespace inside tags are
|
|
843
|
+
preserved.**
|
|
844
|
+
- **[P] The prologue, doctype, comments, and trailing whitespace are preserved.**
|
|
845
|
+
- **[P] Markdown is never reformatted** — list markers, emphasis delimiters, heading style,
|
|
846
|
+
table alignment, hard-break spaces and line endings are all outside every span.
|
|
847
|
+
- **[P] Malformed input is not repaired.** Unclosed tags, stray `<`, mismatched nesting: the
|
|
848
|
+
parser's error recovery affects only where spans are found, never what is emitted.
|
|
849
|
+
|
|
850
|
+
- **[P] YAML quoting style, indentation, anchors, tag handles and chomping indicators are
|
|
851
|
+
preserved** — `yaml` mode uses no parser and no emitter, so none of them is ever decoded and
|
|
852
|
+
rewritten (§3.8.1). A block scalar's trailing line terminators lie outside every span and are
|
|
853
|
+
returned exactly as written, whichever chomping indicator governs them.
|
|
854
|
+
|
|
855
|
+
The cost is real and is accepted: a runtime whose parser cannot report source offsets cannot
|
|
856
|
+
implement `html` mode conformantly. That constraint belongs to the implementation, not to this
|
|
857
|
+
contract. It does not bind `yaml` mode, which has no parser to ask — §3.8.1 records why asking
|
|
858
|
+
one was not an option in the first place.
|
|
859
|
+
|
|
860
|
+
---
|
|
861
|
+
|
|
862
|
+
## 5. Idempotency under composition
|
|
863
|
+
|
|
864
|
+
`pipeline-idempotency.md` proves `T(T(x)) = T(x)` for a single text run, on the composition
|
|
865
|
+
obligation **CO**: no rule creates work for an earlier-ordered rule. Modes do not weaken CO, and
|
|
866
|
+
they do not carry it over for free either. Write `M` for the whole mode transform.
|
|
867
|
+
|
|
868
|
+
**`M(M(x)) = M(x)` requires three things**, of which only the first is already proved:
|
|
869
|
+
|
|
870
|
+
1. **The pipeline is idempotent on the concatenated array.** This is exactly
|
|
871
|
+
`pipeline-idempotency.md`, applied to an array that happens to contain markers. The markers
|
|
872
|
+
are inert — no rule matches them, no rule edits them, and they are in none of the classes any
|
|
873
|
+
guard tests except `OPENISH`/`CLOSEISH` for `quotes`, where they behave like a fixed piece of
|
|
874
|
+
punctuation that no rule can change. So every CO discharge in every rule document holds
|
|
875
|
+
verbatim.
|
|
876
|
+
|
|
877
|
+
2. **The span partition is stable.** `M` must produce the same sequence of spans for `M(x)` as
|
|
878
|
+
for `x`. This is a **new obligation, invisible to every per-rule argument**, and it is the
|
|
879
|
+
one a mode adapter can break on its own: if applying edits changed how the document parses,
|
|
880
|
+
the second run would see different spans, the concatenation would differ, and the rules would
|
|
881
|
+
legitimately reach different conclusions.
|
|
882
|
+
|
|
883
|
+
An earlier revision discharged this by claiming that **every code point the rules emit is
|
|
884
|
+
syntactically inert in `html` and `markdown`**. That claim is **false**, and the
|
|
885
|
+
counterexample is the most ordinary character in the set. **U+0020 is not a markup character
|
|
886
|
+
and is not inert**: beside `*` or `_` it decides whether a delimiter run is left- or
|
|
887
|
+
right-flanking, and at the start of a line it is a line prefix, a list continuation or
|
|
888
|
+
indented code. `dashes` emits U+0020 in every `-spaced` locale. The claim was wrong when it
|
|
889
|
+
was written, and it is what hid the defect now repaired in §3.4.
|
|
890
|
+
|
|
891
|
+
The honest discharge has four parts, and the third is **positional rather than alphabetic** —
|
|
892
|
+
the same move as CO-S in `pipeline-idempotency.md` §5.1a, applied to _where_ a character may
|
|
893
|
+
be placed rather than to _which_ characters exist:
|
|
894
|
+
|
|
895
|
+
- **§4 forbids reserialisation**, so markup bytes are untouched by construction. Nothing a
|
|
896
|
+
rule does can move, requote or respell a tag, an entity or a fence.
|
|
897
|
+
- **No rule emits a markup character.** Of the code points the rules can emit —
|
|
898
|
+
U+2018/U+2019/U+201A–U+201F, U+00AB/U+00BB, U+2039/U+203A, U+2013/U+2014, U+2026, U+00A0,
|
|
899
|
+
U+202F, U+2011, U+00D7, U+00B1, U+00A9, U+00AE, U+2122, U+002E, U+0020 — none is `<`, `&`,
|
|
900
|
+
`>`, `*`, `_`, `` ` ``, `[`, `]`, `(`, `)`, `#`, `|` or a line terminator. This part of the
|
|
901
|
+
old claim survives and covers every emitted character **except U+0020**.
|
|
902
|
+
- **U+0020 is made safe by position, not by inertness.** The edge-growth rule (§3.4) means a
|
|
903
|
+
rule can introduce a code point only at a position **strictly interior** to a span. An
|
|
904
|
+
interior position has span content on both sides, so an emitted U+0020 can never become
|
|
905
|
+
adjacent to a delimiter — delimiters are structural tokens, outside every span — and can
|
|
906
|
+
never begin a line, because a line terminator is always a −2 marker and a line start is
|
|
907
|
+
therefore a span edge or outside spans entirely. **The characters that decide Markdown
|
|
908
|
+
structure and the positions a rule may write to are disjoint sets.** That is what makes the
|
|
909
|
+
partition stable; no property of U+0020 itself is relied upon.
|
|
910
|
+
- **Deletion** is only ever of a U+0020, and structural whitespace borders a `BREAK` — a real
|
|
911
|
+
one in `text` mode, a −2 marker in the others — where `spaces` refuses to act
|
|
912
|
+
(`spaces.md` §3.2 step 4). A deletion therefore cannot turn `- item` into `-item`, collapse
|
|
913
|
+
an indented code block, or destroy a hard break.
|
|
914
|
+
|
|
915
|
+
> **Normative, and binding on future rules:** a rule may emit a code point that is
|
|
916
|
+
> syntactically significant in a mode **only** where the edge-growth rule confines it to a
|
|
917
|
+
> span interior. A rule that needs to place such a character at a span edge must specify its
|
|
918
|
+
> own span-stability argument here first, and must not assume a character is inert merely
|
|
919
|
+
> because it is not markup. U+0020 is the standing counterexample.
|
|
920
|
+
|
|
921
|
+
**In `yaml` the same four parts hold, the third needs one addition, and the obligation needs
|
|
922
|
+
a fifth part this mode is the first to require.**
|
|
923
|
+
|
|
924
|
+
The addition to the third: YAML's structural characters are a larger set than Markdown's, but
|
|
925
|
+
only two of them are made structural by an adjacent U+0020 — `:` becomes a mapping indicator
|
|
926
|
+
when a U+0020 follows it, and `#` becomes a comment introducer when a U+0020 precedes it — and
|
|
927
|
+
those are exactly the two that §3.8.6 lifts out of every plain scalar as opaque units. They are
|
|
928
|
+
therefore outside every span, an emitted U+0020 can never become adjacent to one, and the dash
|
|
929
|
+
token that would have emitted it sits at a span extremity, where §3.4's length test discards
|
|
930
|
+
the 2 → 3 replacement that emits U+0020 and admits the 2 → 1 replacement that cannot. In a quoted scalar both are neutralised by the
|
|
931
|
+
quoting; in a block scalar both are literal content. Of the remaining emitted code points, only
|
|
932
|
+
U+002E is structural in YAML **syntax** at all, and only as `...` at the start of a line — a
|
|
933
|
+
position no span contains, since a plain scalar's span begins after its key and its colon, and
|
|
934
|
+
a block scalar's begins after an indentation of at least one U+0020. Deletion is unchanged:
|
|
935
|
+
`spaces` may delete only a U+0020, and a block scalar's indentation borders a −2 marker.
|
|
936
|
+
|
|
937
|
+
> **What that argument does not cover, and what §7.11 accepts.** A plain scalar's **type** is
|
|
938
|
+
> resolved from its whole text, not from any one character, so an edit anywhere inside one can
|
|
939
|
+
> change a value's tag without touching anything syntactic. Measured: `description: 1 .`
|
|
940
|
+
> becomes `description: 1.` in every locale, because `spaces` strips the space before a
|
|
941
|
+
> sentence-final dot — and the value goes from the string `"1 ."` to the number `1`. No
|
|
942
|
+
> enumeration of emitted code points can close this, because the mechanism is not a character
|
|
943
|
+
> but a whole-token resolution; the span-stability obligation of item 2 is unaffected, since
|
|
944
|
+
> the predicate that selects the span is the key. This is a cost of processing plain scalars at
|
|
945
|
+
> all, it applies only to a value a caller has explicitly listed as prose, and it is accepted
|
|
946
|
+
> rather than closed.
|
|
947
|
+
|
|
948
|
+
> **The fifth part, normative and binding on future modes: the predicate that decides whether
|
|
949
|
+
> a span exists must not read anything a rule can change.** In `html` and `markdown` the
|
|
950
|
+
> question never arises, because span selection reads only the document's structure. `yaml`'s
|
|
951
|
+
> first draft decided a plain scalar's processability from **its own text** — it had to contain
|
|
952
|
+
> a space and a letter — and `spaces` can delete that space. One line was enough to falsify the
|
|
953
|
+
> obligation: `ref: ${{ steps.pin.outputs.sha }}` yields one span, comes back as
|
|
954
|
+
> `ref: ${{steps.pin.outputs.sha}}`, and yields none on the next run. `M(M(x)) = M(x)` survived
|
|
955
|
+
> that only by accident — a lost span means fewer edits, and the first run's edits happened to
|
|
956
|
+
> be fixed points — but the obligation itself was false, so the conclusion rested on nothing.
|
|
957
|
+
> §3.8.2's `keys` option is the repair: the predicate reads the **key**, which lies outside
|
|
958
|
+
> every span and which no rule can reach, so the partition is a function of the source alone.
|
|
959
|
+
|
|
960
|
+
A content-dependent predicate is not merely risky here, it is unarguable: item 2's whole
|
|
961
|
+
method is to show that the characters rules may write and the positions they may write to are
|
|
962
|
+
disjoint from what decides structure. A predicate over span content puts the rules' own output
|
|
963
|
+
back inside that decision, and there is no version of the argument that survives it.
|
|
964
|
+
|
|
965
|
+
3. **Redistribution is deterministic.** Guaranteed by §3.4: no edit spans a marker, and any edit
|
|
966
|
+
that would place code points at a span extremity which were not there before — an insertion at
|
|
967
|
+
a boundary, or a replacement longer than what it replaced — is discarded rather than assigned
|
|
968
|
+
to a side. The discard is a pure function of `(p, q, r, s₀, s₁)`, so it happens identically on
|
|
969
|
+
every run: a rule that was declined at an edge once is declined there always, and the output
|
|
970
|
+
contains nothing for the next run to reconsider.
|
|
971
|
+
|
|
972
|
+
Given 1–3, `M(M(x)) = M(x)`.
|
|
973
|
+
|
|
974
|
+
**The testing obligation of `pipeline-idempotency.md` §6 extends to modes**: the bounded
|
|
975
|
+
exhaustive sweep must be run in `html`, `markdown` and `yaml` mode over documents that place
|
|
976
|
+
span boundaries inside the swept string — at minimum a template like `A<em>B</em>C` with the
|
|
977
|
+
sweep alphabet distributed across `A`, `B` and `C`. A sweep that only ever produces one span
|
|
978
|
+
tests none of this document. In `yaml` the sweep alphabet must include `:`, `#`, `-` and **`---`
|
|
979
|
+
as one token**, since those are the characters whose adjacency the argument above turns on and a
|
|
980
|
+
three-dash run is not reachable from single dashes at a bounded payload length; and it must
|
|
981
|
+
include U+0022 and U+005C, without which the sweep cannot reach a quoted scalar's delimiters at
|
|
982
|
+
all. §3.8.7 records what each run did and did not establish — including that a sweep comparing
|
|
983
|
+
structure and types is blind to a string whose content is damaged, and that an alphabet without
|
|
984
|
+
`---` reported clean twice while §3.4's `r = d` hole was open.
|
|
985
|
+
|
|
986
|
+
---
|
|
987
|
+
|
|
988
|
+
## 6. Which §4 claims change
|
|
989
|
+
|
|
990
|
+
Per `pipeline-idempotency.md` §5.2, a **[P]** claim is a promise about `transform`. Several were
|
|
991
|
+
written for `text` mode and are mode-dependent. The notation extends: **[P: html, markdown]**
|
|
992
|
+
means the claim holds in those modes only.
|
|
993
|
+
|
|
994
|
+
- **"Anything inside a skipped region"**, asserted in every rule's §4, is
|
|
995
|
+
**[P: html, markdown, yaml]** and is now _defined_ rather than assumed — §3.6, §3.7 and §3.8
|
|
996
|
+
are what those bullets refer to. In `text` mode there are no skipped regions and the bullet is
|
|
997
|
+
vacuous. In `yaml` the definition runs the other way round (§3.8.2): the skipped region is
|
|
998
|
+
everything the scan did not claim.
|
|
999
|
+
- **`dashes` §4 "URLs, code spans, fenced code, HTML attributes"** — **[P: html, markdown]**.
|
|
1000
|
+
In `text` mode a URL is ordinary text. The rule is nevertheless safe there, but by its own
|
|
1001
|
+
guards (P1 declines the tight hyphens in `a-b`, the cluster guard declines `2026-08-15`), not
|
|
1002
|
+
because anything skipped it. The distinction matters to a reader deciding whether to run
|
|
1003
|
+
`text` mode over a document containing URLs. **`yaml` sits with `text` here, not with the
|
|
1004
|
+
other two**: a URL written inside a listed key's scalar is ordinary text and nothing skips it.
|
|
1005
|
+
A URL that is the whole value of a `url:` key is a different matter and is skipped, but by
|
|
1006
|
+
§3.8.2's `keys` option rather than by any rule of this bullet — the caller did not list that
|
|
1007
|
+
key.
|
|
1008
|
+
- **`symbols` §4 "`(c)` and `(r)` in a call or index position"**, and the whole `0x1F` family —
|
|
1009
|
+
**[P] in all modes**, since they rest on guards S1/M3/M4 rather than on skipping. §7.2 of that
|
|
1010
|
+
document already records that the `(tm)` exemption's exposure is `text`-mode-only, and this is
|
|
1011
|
+
why.
|
|
1012
|
+
- **`ellipsis` §4 "`../`, `./..`"** and **`spaces` §4's path protection** — **[P] in all modes**.
|
|
1013
|
+
They rest on the lone-dot condition (`spaces.md` §3.4), not on skipping.
|
|
1014
|
+
- **`quotes` §4 "Any unbalanced mark"** — **[P] in all modes**, and _strengthened_ by §3.3: a
|
|
1015
|
+
mark that would have been unbalanced within its own span may now pair across an element, which
|
|
1016
|
+
converts more, never less.
|
|
1017
|
+
- **`nbsp` §4 "The start or end of a text unit"** — **[P] in all modes, and strengthened.** A
|
|
1018
|
+
"text unit" is the concatenation, and §3.4 additionally refuses any insertion at a span
|
|
1019
|
+
boundary, so no span ever begins or ends with a character `nbsp` put there. The guarantee is
|
|
1020
|
+
now about element boundaries as well as document boundaries.
|
|
1021
|
+
|
|
1022
|
+
- **`dashes` §4, the `-spaced` forms** — a new **[R]** consequence in `html`/`markdown`, not a
|
|
1023
|
+
change to any existing bullet. A `-spaced` locale converts a dash only where the replacement
|
|
1024
|
+
fits inside a span without touching either edge, so `a<em>--</em>b` is left alone in `de-DE`
|
|
1025
|
+
while `en-US` converts it — only a contraction survives at an edge (§7.9). Nothing in `dashes.md` is false; the rule simply gets fewer chances
|
|
1026
|
+
to fire. §7.9 records it.
|
|
1027
|
+
|
|
1028
|
+
No rule's §4 becomes _false_ in a mode. Two become narrower ([P: html, markdown]); none becomes
|
|
1029
|
+
rule-local.
|
|
1030
|
+
|
|
1031
|
+
---
|
|
1032
|
+
|
|
1033
|
+
## 7. Open questions
|
|
1034
|
+
|
|
1035
|
+
1. _(Settled — both skipped.)_ `svg` and `math` are in §3.6's skip list although neither is in
|
|
1036
|
+
PLAN.md §3.2's. The decisive argument was MathML: a quotation mark, a hyphen and a prime are
|
|
1037
|
+
**operators and identifiers** there, so substitution changes meaning rather than appearance.
|
|
1038
|
+
The accepted cost is `<svg><text>` and `<svg><title>`, which do hold prose and are now left
|
|
1039
|
+
untypeset. Direction matters here — widening a skip list is additive, narrowing one after
|
|
1040
|
+
release breaks documents — so the conservative side was taken deliberately rather than
|
|
1041
|
+
pending evidence.
|
|
1042
|
+
2. **A skipped element between two halves of a word.** `un<code>x</code>believable` yields two
|
|
1043
|
+
spans that no rule joins, which is correct, but it also means `hyphen` and the `nbsp` literal
|
|
1044
|
+
lists silently fail on any word interrupted by markup. Correct and conservative, but worth a
|
|
1045
|
+
fixture so it is a decision rather than a surprise.
|
|
1046
|
+
3. **An insertion at a span boundary is discarded, and the miss is permanent under this
|
|
1047
|
+
contract.** (Since the edge-growth rule of §3.4, this is one case of a general prohibition on
|
|
1048
|
+
writing a code point onto a span edge; the argument below is what established the principle,
|
|
1049
|
+
and it applies unchanged to the replacement case that generalised it.) `fr` `mot<em>!</em>` keeps no narrow no-break space. This is an argued
|
|
1050
|
+
limitation, not an oversight, and the argument is the mirror pair — neither side is
|
|
1051
|
+
universally safe:
|
|
1052
|
+
|
|
1053
|
+
| Input | Spans | Assign to left span | Assign to right span |
|
|
1054
|
+
| --------------- | ---------- | ------------------- | -------------------- |
|
|
1055
|
+
| `mot<em>!</em>` | `mot`, `!` | `mot⍹<em>!</em>` ✓ | `mot<em>⍹!</em>` ✗ |
|
|
1056
|
+
| `<em>mot</em>!` | `mot`, `!` | `<em>mot⍹</em>!` ✗ | `<em>mot</em>⍹!` ✓ |
|
|
1057
|
+
|
|
1058
|
+
Whichever rule is adopted, one of these two ordinary documents ends up with an element that
|
|
1059
|
+
begins or ends with a space that was never in it. That is invisible in a plain rendering and
|
|
1060
|
+
immediately visible the moment the element carries an underline, a border or a background —
|
|
1061
|
+
and the package must never put a character inside an element it did not come from. A miss is
|
|
1062
|
+
visible to the author and fixable in the source; corrupted markup is neither.
|
|
1063
|
+
|
|
1064
|
+
Note that only _insertion_ is affected. `<em>mot</em> !` already has a U+0020 inside the
|
|
1065
|
+
right-hand span, so N2 converts it in place and French spacing works normally; the loss is
|
|
1066
|
+
confined to documents where the space does not exist at all and the punctuation is wrapped.
|
|
1067
|
+
|
|
1068
|
+
**A real fix needs something this contract does not have:** the ability to represent
|
|
1069
|
+
"between two nodes" as a position, so that an inserted character could be emitted as a
|
|
1070
|
+
sibling of both elements rather than a child of either. That is a different document model —
|
|
1071
|
+
the adapter would have to be able to _add_ to the document rather than only replace inside
|
|
1072
|
+
spans, which §4's round-trip prohibition rules out by design. Any future attempt starts by
|
|
1073
|
+
reopening §4, not by revisiting this paragraph.
|
|
1074
|
+
|
|
1075
|
+
4. _(Settled — soft breaks are `BREAK`, not boundaries.)_ A soft line break inside a paragraph
|
|
1076
|
+
is a **−2 line marker** (§3.2), classified as a member of `BREAK` for every rule. The
|
|
1077
|
+
alternative readings were both worse. Treating it as an ordinary **−1 inline boundary** — what
|
|
1078
|
+
an implementation gets by default, since a line ending is a structural token like any other —
|
|
1079
|
+
makes a mode diverge from `text` mode on the same characters: a `"` at the start of a wrapped
|
|
1080
|
+
line has a `BREAK` on its left in `text` (so it can only open) and an opaque content marker in
|
|
1081
|
+
`html`/`markdown` (so it could also close), and `"foo\n"bar` then pairs in one and not the
|
|
1082
|
+
other. A mode that converts _differently_ from `text` on identical prose is indefensible; a
|
|
1083
|
+
mode that converts _less_ is nearly as bad, and hard-wrapped prose is precisely the M4 corpus.
|
|
1084
|
+
Putting the line terminator physically **inside** the span was the other candidate and it
|
|
1085
|
+
fails on structure: a continuation line's block prefix — a blockquote `>`, a list item's
|
|
1086
|
+
indentation — must stay outside, which makes the span non-contiguous and breaks the
|
|
1087
|
+
source-offset model §4 depends on.
|
|
1088
|
+
The −2 marker gets the benefit of both: spans stay contiguous, and every rule's existing,
|
|
1089
|
+
fixture-covered `BREAK` behaviour applies unchanged. It also removes a parser dependency that
|
|
1090
|
+
would otherwise have been load-bearing — whether the two spaces of a Markdown hard break land
|
|
1091
|
+
inside a span or in a structural token is a per-parser decision, and with a −2 marker beside
|
|
1092
|
+
them `spaces` protects them either way.
|
|
1093
|
+
5. **Adjacent text nodes.** Some parsers report `a<!-- c -->b` or a long text run as two adjacent
|
|
1094
|
+
text nodes with no element between them. This document treats every gap between processable
|
|
1095
|
+
spans as a boundary, so such a split would suppress conversions that ought to happen.
|
|
1096
|
+
Implementations **should** coalesce adjacent processable spans separated by nothing in the
|
|
1097
|
+
source; whether that is required, and how it interacts with comments, is not settled here.
|
|
1098
|
+
6. **`markdown` emphasis delimiters inside a span.** `*foo*` is markup, but a naive adapter that
|
|
1099
|
+
walks an `mdast` tree gets `foo` as a text node and the asterisks as structure — so they are
|
|
1100
|
+
outside every span, which is right. An adapter that instead processes raw source lines would
|
|
1101
|
+
have them inside a span and could let a rule edit next to them. The tree-walking approach is
|
|
1102
|
+
assumed throughout; it is not stated as a requirement because it names an implementation
|
|
1103
|
+
strategy, and it probably should be.
|
|
1104
|
+
7. **No mode covers plain-text email or reStructuredText**, and the `rules` option cannot
|
|
1105
|
+
express "skip this region" for a caller with their own format. That is out of scope for v1
|
|
1106
|
+
(PLAN.md §4) and is recorded so the span model is not mistaken for an extension point. _(YAML
|
|
1107
|
+
was in this item until 1.3.0 and is now §3.8; the two formats named here are not, and the
|
|
1108
|
+
argument that moved YAML does not carry over — it turned on a measured corpus and a
|
|
1109
|
+
corruption report, neither of which exists for these.)_
|
|
1110
|
+
|
|
1111
|
+
8. **`dialect` and `POLYTYPO_MALFORMED_INPUT` are new public surface and need operator
|
|
1112
|
+
sign-off.** Both are specified — I did not leave them open — but both change the contract.
|
|
1113
|
+
|
|
1114
|
+
The first is `dialect` (§3.7.1): a required option on `markdown` mode rather than a fourth
|
|
1115
|
+
mode id, because mode ids are public API and PLAN.md §3.2 names exactly three. Two things
|
|
1116
|
+
there are mine to flag rather than to decide: whether the operator prefers the option or a
|
|
1117
|
+
`markdown-mdx` mode id after all; and which code the missing or invalid value raises. I
|
|
1118
|
+
specified a throw with no default, on the fail-fast reasoning that already governs `locale`,
|
|
1119
|
+
but ARCHITECTURE.md §4.6 has no code for it — `POLYTYPO_INVALID_MODE` is the closest fit and
|
|
1120
|
+
is arguably a lie, since the mode is valid and only its dialect is not. A
|
|
1121
|
+
`POLYTYPO_INVALID_DIALECT` code would be honest, and is a contract addition.
|
|
1122
|
+
|
|
1123
|
+
The second is `POLYTYPO_MALFORMED_INPUT` (§3.7.2). It adds a code to the taxonomy **and**
|
|
1124
|
+
makes `transform` non-total on its input for the first time. I judged that right on
|
|
1125
|
+
fail-fast grounds and because the only reachable case is a genuinely broken MDX file, but a
|
|
1126
|
+
caller running polytypo over a large corpus may reasonably want a document-level failure not
|
|
1127
|
+
to abort a batch. If so, the answer is a wrapper in _their_ code, not an option here — this
|
|
1128
|
+
spec has no error-suppression switch and should not acquire one.
|
|
1129
|
+
|
|
1130
|
+
9. **The edge-growth rule costs conversions that `text` mode makes, and only a contraction
|
|
1131
|
+
survives at a span edge.** That sentence is the whole rule, and it follows from §3.4's length
|
|
1132
|
+
test rather than sitting beside it: an edit at an extremity is applied iff `r ≤ d`. So a
|
|
1133
|
+
`dashes` promotion at a span edge converts **only where the locale's parenthetical form is
|
|
1134
|
+
shorter than the input that triggered it**.
|
|
1135
|
+
|
|
1136
|
+
Measured across all nine locales, on `a<em>--</em>b`:
|
|
1137
|
+
|
|
1138
|
+
| Locale | `dash.parenthetical` | At a span edge | Interior (`a<em>x--y</em>b`) |
|
|
1139
|
+
| --- | --- | --- | --- |
|
|
1140
|
+
| `en-US` | `em-tight` | **converts** — `a<em>—</em>b`, since 2 → 1 is a contraction | converts |
|
|
1141
|
+
| `de-CH`, `de-DE`, `en-GB`, `fi`, `sv` | `en-spaced` | unchanged — 2 → 3 grows both edges | converts, `a<em>x – y</em>b` |
|
|
1142
|
+
| `fr`, `ru` | `em-spaced` | unchanged — same reason | converts |
|
|
1143
|
+
| `el` | `none` | unchanged — emits nothing anywhere | unchanged |
|
|
1144
|
+
|
|
1145
|
+
`en-US` is the only locale whose parenthetical form is shorter than the `--` that triggers it,
|
|
1146
|
+
and it is therefore the only one that converts at an edge. The asymmetry is real — the same
|
|
1147
|
+
document is typeset differently in `de-DE` and `en-US` for a reason that has nothing to do
|
|
1148
|
+
with German or American typography — but it applies to a **narrower class** than an earlier
|
|
1149
|
+
revision of this item claimed, and it now has a mechanical explanation rather than a
|
|
1150
|
+
coincidental one.
|
|
1151
|
+
|
|
1152
|
+
That earlier revision was also **factually wrong** in a way worth recording, because the error
|
|
1153
|
+
outlived the thing that caused it. It said `a<em>–</em>b` "keeps its en dash unspaced in
|
|
1154
|
+
`de-DE` … while `en-US` (`em-tight`) converts it happily". Neither half is true any more:
|
|
1155
|
+
`dashes` §3.2 step 2a declines **every** token containing an authored U+2013 or U+2014, so
|
|
1156
|
+
nothing converts that input in any locale. The claim was correct when written and was
|
|
1157
|
+
invalidated by a change in another document; the carrier died and the sentence did not notice.
|
|
1158
|
+
§3.4's worked table had the identical fault and was rebuilt on `--` for the same reason.
|
|
1159
|
+
|
|
1160
|
+
10. **Whether a −2 marker should also end a `quotes` pairing scope** is open. Today a quotation
|
|
1161
|
+
opened in one paragraph and never closed can still pair with a mark in the _next_ paragraph,
|
|
1162
|
+
because the stack is not reset at a `BREAK` — which is exactly the behaviour `text` mode has,
|
|
1163
|
+
so the two agree. It may nonetheless be wrong in both, and `quotes.md` §7.5 already records
|
|
1164
|
+
the multi-paragraph question. If that is ever resolved, it must be resolved for `text` and
|
|
1165
|
+
the modes together, not here.
|
|
1166
|
+
|
|
1167
|
+
11. **`yaml` mode's accepted misses, and one accepted cost.** Every entry but the last is a
|
|
1168
|
+
construct the scan declines to claim, so prose inside it is returned untouched; each could be
|
|
1169
|
+
admitted later without breaking a document that relies on today's behaviour, and none could
|
|
1170
|
+
be withdrawn later without breaking one. They are recorded together because the direction is
|
|
1171
|
+
the argument — and the last entry is here precisely because it is the one that does not share
|
|
1172
|
+
it:
|
|
1173
|
+
|
|
1174
|
+
- **a bare sequence item** — `- Some prose here` yields no spans, because §3.8.4 step 5
|
|
1175
|
+
requires a processable scalar to be the value of a key, and a bare item has none for
|
|
1176
|
+
`keys` to match. `- key: value` is unaffected;
|
|
1177
|
+
- **a key containing `"`, `'`, `{`, `[`, `&`, `*`, `!` or `#` anywhere** — `"description"`,
|
|
1178
|
+
but also `a!b` — declined at step 5 before the key is compared, so a document that quotes
|
|
1179
|
+
its keys gets nothing from this mode;
|
|
1180
|
+
- **a plain scalar whose type is resolved from its whole text** — a listed key's value can
|
|
1181
|
+
change tag without any syntactic character being touched (§5 item 2's note). `1 .` becomes
|
|
1182
|
+
`1.` and stops being a string. This one is not a miss but an accepted cost, and it is the
|
|
1183
|
+
only entry here that is not purely in the widening direction;
|
|
1184
|
+
- **a multi-line plain scalar**, and a quoted scalar or flow collection that does not close
|
|
1185
|
+
on its opening line: all three are consumed by step 7 and yield no spans;
|
|
1186
|
+
- **a flow collection** — `tags: [one, two]` and `{a: b}` — even on one line;
|
|
1187
|
+
- **a quoted scalar containing an escape** — `"say \"hi\""` — where the source spells the
|
|
1188
|
+
content with more characters than it has;
|
|
1189
|
+
- **a block scalar whose indentation is ambiguous**, per §3.8.5's three bail conditions.
|
|
1190
|
+
|
|
1191
|
+
12. **`keys` matches a bare name at any depth, and nothing narrower.** `description` is
|
|
1192
|
+
processable wherever it occurs, which is what makes the option portable — no paths, no
|
|
1193
|
+
globs, no schema. The cost is that a caller who wants `components.schemas.*.description`
|
|
1194
|
+
but not `info.description` cannot say so, and a caller whose document has a machine-read
|
|
1195
|
+
`description` somewhere in it must choose between typesetting that one too and typesetting
|
|
1196
|
+
none of them. A path syntax is the obvious extension and is deliberately not specified here:
|
|
1197
|
+
it is a small language, five runtimes would have to agree on it exactly, and no measured
|
|
1198
|
+
document has yet needed it. If one does, this is where it reopens, and the extension is
|
|
1199
|
+
additive — a caller passing bare names keeps today's behaviour.
|
|
1200
|
+
|
|
1201
|
+
## 8. Fixture coverage strategy (non-normative)
|
|
1202
|
+
|
|
1203
|
+
This section records why `spec/fixtures/` does not carry the full every-locale × four-modes
|
|
1204
|
+
Cartesian product, and what a smaller set must still prove instead. It has **no normative
|
|
1205
|
+
force** — it does not change §§1–7 — and if it drifts out of sync with the actual fixtures, that
|
|
1206
|
+
is a defect in this section, not license to distrust the fixtures.
|
|
1207
|
+
|
|
1208
|
+
**Why the full product is not required.** §3.2–§3.5 establish that a mode adapter's only job is
|
|
1209
|
+
to identify which spans of the source are processable text and which are structural markup:
|
|
1210
|
+
locating skip-list boundaries, computing marker gaps, and reassembling by offset. Once a span is
|
|
1211
|
+
handed off, `runOverSpans` (`src/engine/span-runner.ts`, used by `html` and `markdown`) calls
|
|
1212
|
+
exactly the same per-rule `apply()` functions, against the same locale data, that `text` mode's
|
|
1213
|
+
`runRules` (`src/engine/text-pipeline.ts`) calls — there is one implementation per rule id,
|
|
1214
|
+
shared by every mode, not one per mode. A locale's quote glyphs, dash conventions, or `nbsp`
|
|
1215
|
+
targets are therefore already fully exercised by `text`-mode fixtures; an `html`/`markdown`
|
|
1216
|
+
fixture in the same locale mostly re-tests that same rule logic through an extra layer of span
|
|
1217
|
+
bookkeeping, not something the locale itself changes about it.
|
|
1218
|
+
|
|
1219
|
+
What *does* vary by mode is the span-selection and reassembly machinery, and that machinery is
|
|
1220
|
+
locale-agnostic: the HTML skip list and the CommonMark/MDX skip lists never consult `LocaleData`.
|
|
1221
|
+
A representative sample proves the machinery correct; the full product would mostly multiply
|
|
1222
|
+
proof of the same machinery by locale count, without adding coverage of anything the locale
|
|
1223
|
+
changes.
|
|
1224
|
+
|
|
1225
|
+
**What representative coverage requires instead**, and where it currently lives:
|
|
1226
|
+
|
|
1227
|
+
1. **HTML span selection** — the skip list, character-reference handling, and round-trip
|
|
1228
|
+
guarantee, exercised directly (`tests/modes/html.test.ts`) and via fixtures.
|
|
1229
|
+
2. **CommonMark span selection** — fenced/indented/inline code, autolinks, link destinations and
|
|
1230
|
+
titles, reference-link definitions, and the round-trip guarantee
|
|
1231
|
+
(`tests/modes/markdown.test.ts`).
|
|
1232
|
+
3. **MDX span selection** — expression containers, JSX attributes and children, ESM export
|
|
1233
|
+
blocks, and the JSX-vs-skipped-element case distinction (`tests/modes/markdown.test.ts`).
|
|
1234
|
+
3a. **YAML span selection** — the three scalar forms, the `keys` option (a listed key, an
|
|
1235
|
+
unlisted one, a quoted one, the same name at two depths), block-scalar indentation and
|
|
1236
|
+
chomping, the `:`/`#` split of §3.8.6, step 7's consumption of an inline value's continuation
|
|
1237
|
+
lines, and every bail of §3.8.4. `yaml` carries a heavier fixture burden than the other two
|
|
1238
|
+
for a reason §8's opening argument does not cover: its span selection is **specified rather
|
|
1239
|
+
than delegated** (§3.8.1), so a fixture is the only thing standing between five hand-written
|
|
1240
|
+
scanners and five different answers. The bounded sweep of §3.8.7 is part of this item, not an
|
|
1241
|
+
extra — and so is the end-to-end corpus run, for the reason §3.8.7 gives: a sweep that
|
|
1242
|
+
compares structure cannot see a string being damaged without being destroyed.
|
|
1243
|
+
4. **At least two materially different locale outputs per mode/dialect**, so a fixture is not
|
|
1244
|
+
merely "the same English output with a different `locale` field": `spec/fixtures/fr.json` and
|
|
1245
|
+
`fr-CA.json` carry `html`/`markdown`/`mdx` cases whose guillemets-plus-U+00A0 output is
|
|
1246
|
+
structurally different from `en-US`'s curly quotes, not just a different glyph in the same
|
|
1247
|
+
shape. `yaml` carries `en-US`, `de-DE` and `fr`, and for that mode **one of them must be a
|
|
1248
|
+
`-spaced` locale**: §3.8.6's `:`/`#` split only discriminates where the replacement grows, so
|
|
1249
|
+
an em-tight-only fixture set passes with the split removed.
|
|
1250
|
+
5. **Non-ASCII text and code-point/offset boundaries** — an astral-character (surrogate-pair)
|
|
1251
|
+
preservation case in HTML mode, and genuinely accented non-ASCII prose exercised through the
|
|
1252
|
+
French MDX and CommonMark fixtures and round-trip tests.
|
|
1253
|
+
6. **Byte-identical skipped regions** — the round-trip guarantee ("returns a document/article
|
|
1254
|
+
that needs no changes byte for byte") asserted directly for HTML, CommonMark, MDX and YAML.
|
|
1255
|
+
In `yaml` the guarantee is also what proves §3.8.2: a document the scan does not understand
|
|
1256
|
+
must come back unchanged, so an unrecognised construct is a fixture, not a hope.
|
|
1257
|
+
7. **Every locale-specific transformation independently covered through `text`-mode fixtures** —
|
|
1258
|
+
the conformance suite's own per-locale, per-rule coverage, not this file.
|
|
1259
|
+
|
|
1260
|
+
**Enforcement**, split across the two places that actually check each item — neither file alone
|
|
1261
|
+
covers all eight:
|
|
1262
|
+
|
|
1263
|
+
- `tests/conformance/mode-fixture-strategy.test.ts` reads `spec/fixtures/*.json` directly and
|
|
1264
|
+
protects item 4 (at least two distinct locales carry `html`, `markdown`/`commonmark`,
|
|
1265
|
+
`markdown`/`mdx` and `yaml` fixtures, and at least one `yaml` locale is `-spaced`), the
|
|
1266
|
+
code-point half of item 5 (at least one fixture per mode contains a non-ASCII code point,
|
|
1267
|
+
`yaml` included), and item 7 (every locale has a `text`-mode fixture for
|
|
1268
|
+
*every* canonical rule id in `spec/rules/order.json`, checked per locale/rule id pair, not
|
|
1269
|
+
merely "the locale has a fixture for some rule"). It does not inspect span selection or
|
|
1270
|
+
round-trip behaviour at all.
|
|
1271
|
+
- **In each runtime repository**, `tests/modes/html.test.ts`, `tests/modes/markdown.test.ts` and
|
|
1272
|
+
`tests/modes/yaml.test.ts` protect items 1–3a (HTML, CommonMark, MDX and YAML span selection
|
|
1273
|
+
respectively), the parser-boundary half of item 5 (an astral-character/surrogate-pair
|
|
1274
|
+
preservation case in HTML mode), and item 6 (the round-trip guarantee — "returns a
|
|
1275
|
+
document/article that needs no changes byte for byte" — asserted directly for HTML, CommonMark,
|
|
1276
|
+
MDX and YAML). Those files are not in this repository, which has no engine to run them through,
|
|
1277
|
+
so this bullet is an obligation on every runtime: one that ships `yaml` without them is not
|
|
1278
|
+
conformant for the mode however green its fixtures are.
|
|
1279
|
+
|
|
1280
|
+
If a future edit narrows fixture coverage, or removes a span-selection or round-trip assertion,
|
|
1281
|
+
the corresponding test above fails before this section's claim goes silently stale.
|