polytypo 1.7.0 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +1 -1
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +1 -1
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +1 -1
- data/lib/polytypo/data/fixtures/en-US.json +379 -1
- data/lib/polytypo/data/fixtures/es.json +1 -1
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +1 -1
- data/lib/polytypo/data/fixtures/fr.json +56 -1
- data/lib/polytypo/data/fixtures/it.json +1 -1
- data/lib/polytypo/data/fixtures/locale-resolution.json +1 -1
- data/lib/polytypo/data/fixtures/nl.json +1 -1
- data/lib/polytypo/data/fixtures/pl.json +1 -1
- data/lib/polytypo/data/fixtures/pt-BR.json +1 -1
- data/lib/polytypo/data/fixtures/pt-PT.json +1 -1
- data/lib/polytypo/data/fixtures/ru.json +1 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/tr.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +1 -1
- data/lib/polytypo/data/locales/registry.json +1 -1
- data/lib/polytypo/data/rules/modes.md +207 -20
- data/lib/polytypo/data/rules/order.json +1 -1
- data/lib/polytypo/modes/markdown.rb +238 -64
- data/lib/polytypo/version.rb +1 -1
- metadata +1 -1
|
@@ -7,12 +7,13 @@ for all five runtimes and is parser-agnostic by construction: `parse5`, `nokogir
|
|
|
7
7
|
`golang.org/x/net/html` and PHP's DOM disagree about almost everything this document does not
|
|
8
8
|
forbid them from doing. `yaml` mode is parser-**free** rather than parser-agnostic, for the
|
|
9
9
|
reason §3.8.1 measures.
|
|
10
|
-
**Spec version:** 1.
|
|
10
|
+
**Spec version:** 1.8.0 (0.1.0 for everything except §3.3's class-membership table rows for
|
|
11
11
|
`nbsp` and `apostrophe`, split in 1.2.0, §3.8, added in 1.3.0, §3.3's note on `quotes`'
|
|
12
12
|
span-boundary elision veto reading the marker as a trigger, added in 1.4.0 — which changes no
|
|
13
13
|
row of the table it follows — §3.3's `CLOSEDELIM` entry, added in 1.5.0, and §3.7.4 with the
|
|
14
14
|
amendments it carries to §3.1's definitions, §3.2's Model C, §3.5 step 3, §3.7.3's frontmatter
|
|
15
|
-
bullet, §3.8.6's accepted-cost paragraph, §5, §6 and §7, added in 1.7.0
|
|
15
|
+
bullet, §3.8.6's accepted-cost paragraph, §5, §6 and §7, added in 1.7.0, and §3.7.3a, added in
|
|
16
|
+
1.8.0).
|
|
16
17
|
|
|
17
18
|
---
|
|
18
19
|
|
|
@@ -524,6 +525,149 @@ lower-cases JSX names before matching will skip a component's children.
|
|
|
524
525
|
Nesting follows the same rule as `html`: a skipped construct is skipped whole, including
|
|
525
526
|
anything that looks processable inside it.
|
|
526
527
|
|
|
528
|
+
#### 3.7.3a Where the frontmatter block begins and ends
|
|
529
|
+
|
|
530
|
+
**Spec 1.8.0.** §3.7.3 skips "a metadata block at the very start of the document, delimited by
|
|
531
|
+
`---` (YAML) or `+++` (TOML)". That is prose, and frontmatter is in neither CommonMark nor GFM,
|
|
532
|
+
so until now each runtime reached the construct through its own parser's frontmatter support —
|
|
533
|
+
four extensions, four answers. Measured on the published 1.7.0 line, `en-US`, `commonmark`,
|
|
534
|
+
asking only whether the block is skipped or typeset as prose:
|
|
535
|
+
|
|
536
|
+
| document | JS | Python | Go | Ruby |
|
|
537
|
+
| ------------------------------------------- | ------- | ------- | ----------- | ----------- |
|
|
538
|
+
| `---` / `title: "x - y"` / `---` | skipped | skipped | skipped | skipped |
|
|
539
|
+
| trailing space on the **opening** delimiter | skipped | skipped | **typeset** | **typeset** |
|
|
540
|
+
| trailing space on the **closing** delimiter | skipped | skipped | **typeset** | **typeset** |
|
|
541
|
+
| `...` as the closer | typeset | typeset | typeset | typeset |
|
|
542
|
+
| `---` after a leading blank line | typeset | typeset | typeset | typeset |
|
|
543
|
+
|
|
544
|
+
Two of four turn `date: "2026-09-26"` into a quoted-and-curled string over one trailing space
|
|
545
|
+
that no author typed on purpose, and both were conformant, because no fixture pinned the edge.
|
|
546
|
+
|
|
547
|
+
> **The block's extent is decided by the scan below, not by a parser's frontmatter support.** A
|
|
548
|
+
> runtime whose parser also recognises the construct must produce the same extent as this scan; the
|
|
549
|
+
> scan is what a fixture pins and what a disagreement is measured against.
|
|
550
|
+
>
|
|
551
|
+
> **And the parser is handed the block masked out.** The source given to the Markdown parser is the
|
|
552
|
+
> document with the located block — both delimiter lines included, **and the leading U+FEFF of step
|
|
553
|
+
> 1 if there is one** — replaced by U+0020, line terminators kept as they are, so that nothing
|
|
554
|
+
> inside the block can form or close a construct in the body.
|
|
555
|
+
>
|
|
556
|
+
> **The masked source must be positionally aligned with the original in the unit the runtime maps
|
|
557
|
+
> parser offsets back through.** That is the invariant, and it is not the same as "one U+0020 per
|
|
558
|
+
> code point": a runtime that hands its parser's byte offsets straight through owes byte-length
|
|
559
|
+
> preservation, and masking `😀` to a single space shortens its source by three and shifts every
|
|
560
|
+
> body offset after it. A runtime that converts offsets against the **masked** source before
|
|
561
|
+
> reading them against the original owes code-point alignment only, which one U+0020 per code point
|
|
562
|
+
> gives it — and in a language whose strings are sequences of code points, byte-length preservation
|
|
563
|
+
> cannot even be expressed. Both are conformant; stating it as an index-unit count was not, and the
|
|
564
|
+
> port that indexes code points while its parser counts bytes is what showed it.
|
|
565
|
+
>
|
|
566
|
+
> **And no span may lie inside the block, whatever the parser did with the masked text.** Masking
|
|
567
|
+
> is what makes that true for most parsers and it is not sufficient for all of them: measured,
|
|
568
|
+
> tree-sitter-markdown reads a final all-space line with no terminator as a paragraph, so a
|
|
569
|
+
> document whose closing `---` ends the file comes back with a span over the delimiter itself —
|
|
570
|
+
> the parser did not see the block at all, and step 5's "end of input ends a line" is the clause it
|
|
571
|
+
> does not implement. A runtime whose parser emits such a span **clips it to the part outside the
|
|
572
|
+
> block, and drops it when nothing is left**: the block's own characters must not reach the rules,
|
|
573
|
+
> and body prose past the block must not be lost to a parser's mistake about where the block ended.
|
|
574
|
+
> The mask is there so the parse is not deformed; this rule is there so the spans cannot be wrong
|
|
575
|
+
> even when it is.
|
|
576
|
+
>
|
|
577
|
+
> The mark is masked with the block because leaving it out breaks both: a first line of U+FEFF
|
|
578
|
+
> followed by spaces is not blank, no parser is required to strip the mark, and goldmark does not.
|
|
579
|
+
|
|
580
|
+
Masking is not an implementation note, and the runtime that skipped it is measured. Suppressing a
|
|
581
|
+
span inside the block's range is not enough, because the parser has already read the block's
|
|
582
|
+
characters by then: a fenced-code line inside a metadata value pairs with the body's own fence, and
|
|
583
|
+
the body's code block and its prose swap places. One document, the same call in four runtimes, on
|
|
584
|
+
published 1.7.0:
|
|
585
|
+
|
|
586
|
+
````
|
|
587
|
+
---
|
|
588
|
+
x: |
|
|
589
|
+
```
|
|
590
|
+
---
|
|
591
|
+
|
|
592
|
+
```
|
|
593
|
+
code "q"
|
|
594
|
+
```
|
|
595
|
+
|
|
596
|
+
Body "q".
|
|
597
|
+
````
|
|
598
|
+
|
|
599
|
+
| runtime | result |
|
|
600
|
+
| ---------- | --------------------------------------------------------------------------- |
|
|
601
|
+
| JS, Python | `code "q"` stays straight, `Body “q”` converts — correct |
|
|
602
|
+
| Go | **`code “q”` is typeset inside the code block**, and `Body "q"` is missed |
|
|
603
|
+
| Ruby | correct with this bare-looking fence, wrong the moment the fence has a space |
|
|
604
|
+
|
|
605
|
+
Go reaches that by parsing the whole source and suppressing spans in the block's range, which is
|
|
606
|
+
the obvious way to do it without a frontmatter-aware parser and is wrong for a reason no span-level
|
|
607
|
+
rule can see: the damage is in what the parser concluded, not in which spans were emitted. Masking
|
|
608
|
+
costs one pass over a known range and removes the class.
|
|
609
|
+
|
|
610
|
+
**The scan**, over the source's code-point array, in §3.8.4's terms:
|
|
611
|
+
|
|
612
|
+
1. the document must **begin** with the delimiter — `---` or `+++` at offset 0, no leading blank
|
|
613
|
+
line and no indentation. **A single leading U+FEFF is stepped over first** and is not part of
|
|
614
|
+
the document for this scan. It is a byte-order mark, not content: every editor that writes one
|
|
615
|
+
writes it before the fence, and reading it as content would deny the block to every file some
|
|
616
|
+
Windows editors produce. Measured on 1.7.0, JS and Python already step over it and Go and Ruby
|
|
617
|
+
do not — the same two-against-two split, on the same damaging side, as the trailing space;
|
|
618
|
+
2. the rest of that line must be only U+0020 and U+0009. Anything else and there is no block, and
|
|
619
|
+
what the line then is belongs to the dialect rather than to this scan: `--- yaml` is not a
|
|
620
|
+
thematic break, since a break admits only spaces and tabs after its run;
|
|
621
|
+
3. the **closing line** is the first later line whose first code point begins the same delimiter,
|
|
622
|
+
followed by only U+0020 and U+0009. Indentation disqualifies it exactly as it disqualifies the
|
|
623
|
+
opening line — a closer is not searched for inside a line, it is a line. `...` is not a closer
|
|
624
|
+
in either matter, a delimiter of the other kind is not one either, and a fourth delimiter
|
|
625
|
+
character is not whitespace, so `----` closes nothing;
|
|
626
|
+
4. with no such line there is **no block**, and the opening delimiter is whatever the dialect makes
|
|
627
|
+
of it — a thematic break for `---`, ordinary paragraph text for `+++` — with everything after it
|
|
628
|
+
prose, which is what `en-us-markdown-commonmark-frontmatter-unterminated` already pins;
|
|
629
|
+
5. a **line ends as CommonMark ends one** — at U+000A, at a U+000D that is not followed by
|
|
630
|
+
U+000A, or at the end of input — and the terminator is never part of the line. So a CRLF
|
|
631
|
+
document gives the same block as the same bytes with LF, a file whose last line is `---\r`
|
|
632
|
+
with no final U+000A still closes its block, and a document written with lone U+000D line
|
|
633
|
+
endings has one at all.
|
|
634
|
+
|
|
635
|
+
**This is the Markdown document's line model, and it is deliberately not §3.8.4's.** The block
|
|
636
|
+
is a Markdown construct, so it ends its lines the way the language around it does. The content
|
|
637
|
+
inside it keeps §3.8.4's LF-only model — and the reason is not that YAML agrees, because it
|
|
638
|
+
does not: YAML 1.2.2 §5.4 admits a lone U+000D as a line break too. §3.8.4 is a deliberate
|
|
639
|
+
simplification, taken and measured in 1.3.0, and this section does not widen it, because
|
|
640
|
+
widening the content scan's line model is a change to `yaml` mode for every caller and wants
|
|
641
|
+
its own measurement. A port that harmonises the two has silently made that change. What it
|
|
642
|
+
costs here is recorded in §7.13.
|
|
643
|
+
|
|
644
|
+
All three shapes above are documents whose metadata a stricter reading hands to the rules, and
|
|
645
|
+
the reference runtime already treats all three as blocks — measured on 1.7.0.
|
|
646
|
+
|
|
647
|
+
**Which way to be wrong, and why this way.** A locator errs in one of two directions and they are
|
|
648
|
+
not the same size. Recognising a block that is not one skips text that was prose: a miss, and
|
|
649
|
+
§3.8.3 already accepts misses by the dozen. Failing to recognise a block that is one hands the
|
|
650
|
+
metadata to the rules as prose: `title: Q3 review: what changed` takes a no-break space before
|
|
651
|
+
its colon in `fr` — U+00A0, since `fr` puts `:` in `beforePunctuation` and only `;`, `!` and `?`
|
|
652
|
+
in the narrow list — and a date grows quotation marks. The first is invisible and harmless;
|
|
653
|
+
the second is the damage §3.7.3 exists to prevent. **So the wider reading wins every edge where
|
|
654
|
+
the runtimes disagree**, and the trailing-space clause of steps 2 and 3 is that reading written
|
|
655
|
+
down.
|
|
656
|
+
|
|
657
|
+
The accepted cost is stated rather than hidden: a document whose very first line is a thematic
|
|
658
|
+
break written as `--- `, followed by prose and another `---` line, has **everything up to that
|
|
659
|
+
line** skipped — not one paragraph but the whole first section, heading included, since the scan
|
|
660
|
+
takes the first later delimiter line whatever lies between. It is skipped by JS and Python today,
|
|
661
|
+
has been for every release, and nobody has reported it — while a frontmatter block carrying
|
|
662
|
+
trailing whitespace on its fence is what any editor that trims nothing produces.
|
|
663
|
+
|
|
664
|
+
**What this costs each runtime.** JS and Python already behave this way, so the change there is
|
|
665
|
+
that the extent stops being their parser's opinion and becomes this scan's — which also removes
|
|
666
|
+
the second parse `frontmatterKeys` cost them (polytypo/polytypo#59), since locating the block no
|
|
667
|
+
longer needs one. Go and Ruby change behaviour: both must widen their locator to admit trailing
|
|
668
|
+
U+0020 and U+0009 on either delimiter. Go already owns a hand-written locator for exactly this
|
|
669
|
+
reason and `detectFrontmatter` is where it lives.
|
|
670
|
+
|
|
527
671
|
#### 3.7.4 Frontmatter by named keys — the one opt-out from §3.7.3
|
|
528
672
|
|
|
529
673
|
**Spec 1.7.0.** §3.7.3 skips the frontmatter block whole and its reason for doing so still holds.
|
|
@@ -556,6 +700,35 @@ already recognises, and the option only changes what happens inside it:
|
|
|
556
700
|
terminator to the code point that begins the closing delimiter line. A U+000D before that
|
|
557
701
|
terminator belongs to the terminator, exactly as in §3.8.4, so a CRLF document and the same
|
|
558
702
|
bytes with LF give the same content;
|
|
703
|
+
- **a line of that content containing a U+000D not followed by U+000A yields no spans**, exactly
|
|
704
|
+
as §3.8.4 step 1 already declines a line containing U+0009 and for the same kind of reason. The
|
|
705
|
+
two line models meet here and do not compose: §3.7.3a step 5 finds the block in a lone-U+000D
|
|
706
|
+
document, and §3.8.4's LF-only scan then reads that whole block as one line. Measured, the
|
|
707
|
+
result of letting it through is not merely inert — `title: a "b` and `c" d` on two mapping
|
|
708
|
+
lines pair their marks across the boundary, an unlisted line inside the listed key's scalar
|
|
709
|
+
takes `fr`'s spacing, and the U+000D lands **inside a span**, which §3.8.4 forbids in the same
|
|
710
|
+
breath. Per line rather than per block, because a stray U+000D inside one quoted value is
|
|
711
|
+
something people produce by accident and it should cost that value rather than the whole block:
|
|
712
|
+
a lone-U+000D document is one §3.8.4 line and loses everything, a document with one such value
|
|
713
|
+
loses that line and keeps the rest. **The test runs to the start of the next line, not to the
|
|
714
|
+
end of this one**, because §3.8.4's own splitter treats a trailing U+000D as a terminator even
|
|
715
|
+
without a U+000A after it — so a block whose single line ends in one looks clean if the
|
|
716
|
+
terminator is excluded, and one of the four lone-U+000D fixtures separates the two readings.
|
|
717
|
+
**The decline drops the spans the content scan produced for that line; it does not alter the
|
|
718
|
+
content the scan is given, and a span reaching a declined position is dropped whole rather than
|
|
719
|
+
trimmed.** Three ports reached for the shortcut of substituting the U+000D for another
|
|
720
|
+
character the scan already declines, and it is not equivalent: substitution moves the character
|
|
721
|
+
to a different line and can end a value run, so `title` / `slug` / a content-final U+000D
|
|
722
|
+
converts both keys instead of one, and a stray U+000D on a key whose value run continues onto
|
|
723
|
+
an indented line converts a key that is not a key. Both shapes are pinned, and both agree with
|
|
724
|
+
what `yaml` mode already does with the same characters — measured in two shipped runtimes.
|
|
725
|
+
**Where this rule and `yaml` mode part is on purpose, and it is one shape:** a content line
|
|
726
|
+
whose terminator is a bare U+000D is clean to `yaml` mode, which strips it, and declined here,
|
|
727
|
+
because the window includes it. So `title` / `slug` / a content-final U+000D converts `slug` in
|
|
728
|
+
`yaml` mode and does not here. The wider decline is the direction this section takes everywhere
|
|
729
|
+
else — a miss is invisible, a U+000D inside a span is not — and the cost is conversions lost in
|
|
730
|
+
documents written with line endings from the last century. Widening §3.8.4 instead would be a
|
|
731
|
+
change to `yaml` mode for every caller and wants its own measurement;
|
|
559
732
|
- **both delimiter lines stay outside every span**, as does every line terminator, so no edit
|
|
560
733
|
can reach `---` itself and §3.7.3's setext-underline hazard is unreachable;
|
|
561
734
|
- an **unterminated** block is not a block — §3.7.3 already yields no frontmatter construct
|
|
@@ -567,13 +740,16 @@ already recognises, and the option only changes what happens inside it:
|
|
|
567
740
|
contains one, so specifying a TOML locator would be scope taken on speculation. Recorded as an
|
|
568
741
|
accepted miss in §7.13.
|
|
569
742
|
|
|
570
|
-
**The option adds spans only where §3.7.3's skip removed them.** If
|
|
571
|
-
|
|
572
|
-
anything
|
|
573
|
-
|
|
574
|
-
|
|
575
|
-
|
|
576
|
-
|
|
743
|
+
**The option adds spans only where §3.7.3's skip removed them.** If there is no block — an
|
|
744
|
+
unterminated one, a `---` that is not at the start of the document, an opening line carrying
|
|
745
|
+
anything but whitespace — the text is ordinary prose in the body's own unit and the option
|
|
746
|
+
contributes nothing. That coupling is what makes double processing unreachable: no source position
|
|
747
|
+
can belong to both units — **and since 1.8.0 it is §3.7.3a's mask and its no-span rule that enforce
|
|
748
|
+
it**, not an agreement between two locators. Measured while that mask was still being specified, a
|
|
749
|
+
block whose closer the body's parser did not accept had its content emitted twice, once by each
|
|
750
|
+
unit: `more: b - c` came back as `more: b—cb—c`. **Since 1.8.0 the block's extent is §3.7.3a's
|
|
751
|
+
scan** rather than whatever each parser's frontmatter support decided, so the option no longer
|
|
752
|
+
inherits a variance that was measured in eleven documents out of twenty.
|
|
577
753
|
|
|
578
754
|
**Key matching is §3.8.2's, which means bare names at any depth.** `title` is processable wherever
|
|
579
755
|
it occurs in the block, `seo.title` included — measured: `seo:` then an indented `title:` is
|
|
@@ -1434,18 +1610,29 @@ rule-local.
|
|
|
1434
1610
|
behind the option (polytypo/polytypo#13, #26) asked for TOML, and the 187-file corpus of
|
|
1435
1611
|
§3.7.4.1 contains none. Widening to TOML is additive: a caller passing keys today keeps
|
|
1436
1612
|
today's behaviour.
|
|
1437
|
-
-
|
|
1438
|
-
content range is exact once there is a
|
|
1439
|
-
|
|
1440
|
-
|
|
1441
|
-
|
|
1442
|
-
|
|
1443
|
-
|
|
1444
|
-
|
|
1445
|
-
|
|
1613
|
+
- **~~The block itself is recognised by the mode, not by a scan specified here.~~ Closed in
|
|
1614
|
+
1.8.0 by §3.7.3a.** It was here because §3.7.4's content range is exact once there is a
|
|
1615
|
+
block, while the block itself came from each parser's own frontmatter support — and the
|
|
1616
|
+
measurement that followed found those disagreeing two against two on a trailing space,
|
|
1617
|
+
with the two that declined the block typesetting the metadata. The locator is now
|
|
1618
|
+
specified, which is where it belonged: it governs the skip for every caller, not only
|
|
1619
|
+
those who pass the option.
|
|
1620
|
+
- **A lone-U+000D document has a block, and `frontmatterKeys` yields nothing inside it.**
|
|
1621
|
+
§3.7.3a step 5 finds the block by CommonMark's line model and §3.8.4 reads content by its
|
|
1622
|
+
own, so the block reaches the content scan as a single line. An earlier draft let that
|
|
1623
|
+
stand and called it inert; measuring it showed it was not. On two mapping lines a quotation
|
|
1624
|
+
opened on one paired with a mark on the other, an unlisted line inside the listed key's
|
|
1625
|
+
scalar took `fr`'s spacing, and the U+000D landed inside a span — which §3.8.4 forbids in
|
|
1626
|
+
the same breath. §3.7.4 therefore declines any content line carrying such a U+000D — per
|
|
1627
|
+
line, so that a stray one inside a single quoted value costs that value and not the block.
|
|
1628
|
+
The block is still skipped, so nothing machine-read is typeset either way. A lone-U+000D
|
|
1629
|
+
document is one line to §3.8.4 and therefore loses the whole block, which is the price of
|
|
1630
|
+
not widening that section here.
|
|
1446
1631
|
- **§3.8.6's single-quoted bail costs more here than anywhere it has been measured before.**
|
|
1447
|
-
Of the
|
|
1448
|
-
single-quoted scalar
|
|
1632
|
+
Of the 1858 corpus values a locale would convert across eight locales, **1040 yield no spans
|
|
1633
|
+
when written as a single-quoted scalar** — 130 of 247 in `en-GB` alone, the figure 1.7.0
|
|
1634
|
+
shipped with, and `scripts/miss-census.mjs` in this repository is what reproduces either on
|
|
1635
|
+
demand. An apostrophe inside such a scalar is spelled `''`, and an apostrophe
|
|
1449
1636
|
is exactly what `apostrophe` and `quotes` convert, so the bail falls hardest on the values
|
|
1450
1637
|
the option exists for. Double-quoted, none bail. The corpus as authored is entirely
|
|
1451
1638
|
double-quoted, so its own author never meets this; a caller whose YAML style is single
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"spec": "1.
|
|
2
|
+
"spec": "1.8.0",
|
|
3
3
|
"$comment": "Single source of truth for pipeline order. Rules run in ascending `order`. Disabling a rule removes it from the sequence and never reorders the rest. Rule ids are public API (see docs/ARCHITECTURE.md section 5). \"spec\" here must track spec/VERSION exactly — it is not itself the global version source; scripts/validate-spec.mjs enforces the match.",
|
|
4
4
|
"rules": [
|
|
5
5
|
{
|