makiri 0.11.0.rc2-aarch64-linux → 0.12.0-aarch64-linux
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +47 -0
- data/NOKOGIRI_DIFFERENCES.md +34 -15
- data/lib/makiri/3.2/makiri.so +0 -0
- data/lib/makiri/3.3/makiri.so +0 -0
- data/lib/makiri/3.4/makiri.so +0 -0
- data/lib/makiri/4.0/makiri.so +0 -0
- data/lib/makiri/element.rb +16 -0
- data/lib/makiri/node_path.rb +8 -1
- data/lib/makiri/version.rb +1 -1
- data/script/check_unsafe_boundaries.rb +3 -2
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: '080648651b5e8e38a405373a252b0ded69f1e5a672c6e5031e0f31058b09fa80'
|
|
4
|
+
data.tar.gz: c13624a3a5ed5e94e4d0adc761280e498fbf3e34891ba41d46cfeeed762bcf68
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: ba444c67f4c84cbe0d3a9317ddcfbbf6854daf1c9fa0a020e16e319ac9d8e98e693fa0b245ba8f3abb3bf9e0f7837af77238c7c0d4557d9d268d1b0465bd24cd
|
|
7
|
+
data.tar.gz: c822e0c488221872b85fc6d8390dea3ee1b72def3c35ae160750102a26115047a84ca0f06df2f0b3ee34b7b3dacc011edbb7c56de27cff93e96fb7e9f9132360
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,52 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## [0.12.0] - 2026-10-01
|
|
4
|
+
|
|
5
|
+
Most changes accept what the DOM allows where 0.11.0 raised. The XML
|
|
6
|
+
serializers (`to_xml`, `canonicalize`) raise instead when the tree holds
|
|
7
|
+
something XML cannot write.
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
HTML:
|
|
12
|
+
|
|
13
|
+
* `to_html` / `inner_html` write a `<template>`'s contents, not its own
|
|
14
|
+
children (added with `add_child`), as browsers do.
|
|
15
|
+
* `create_element_ns` in the HTML namespace accepts an upper-case name
|
|
16
|
+
(`BR`, `DIV`), as the DOM does. It makes an unknown element: not void, and
|
|
17
|
+
not matched by type selectors. Importing such an XHTML element from XML
|
|
18
|
+
behaves the same way.
|
|
19
|
+
|
|
20
|
+
XML:
|
|
21
|
+
|
|
22
|
+
* Text, attribute values and comments accept the characters the DOM allows:
|
|
23
|
+
`"\f"`, U+0001, NUL, and `--` in a comment.
|
|
24
|
+
* `set_attribute_ns` accepts names XML cannot write, such as `p:a}b`.
|
|
25
|
+
* `to_xml` and `canonicalize` raise while any of the above is in the tree.
|
|
26
|
+
* `create_cdata` with `]]>` and `create_processing_instruction` with `?>`
|
|
27
|
+
now raise `ArgumentError`.
|
|
28
|
+
* `canonicalize` adds the namespace declarations a name needs instead of
|
|
29
|
+
raising. It still raises where it would have to invent a prefix.
|
|
30
|
+
|
|
31
|
+
HTML to XML (`import_node`):
|
|
32
|
+
|
|
33
|
+
* The copy is in its namespace at once. Before, it was in none until
|
|
34
|
+
inserted.
|
|
35
|
+
* The copy gets no `xmlns` attribute.
|
|
36
|
+
* An attribute XML cannot write (`x-on:click`, `@click`) is copied as it is
|
|
37
|
+
instead of raising.
|
|
38
|
+
* A `<template>` keeps its own children, after its contents.
|
|
39
|
+
|
|
40
|
+
Other:
|
|
41
|
+
|
|
42
|
+
* `Node#path` round-trips when a sibling with a different prefix or case
|
|
43
|
+
would also match the same XPath step.
|
|
44
|
+
|
|
45
|
+
## [0.11.0] - 2026-09-30
|
|
46
|
+
|
|
47
|
+
No code changes since 0.11.0.rc2. Coming from 0.10.x, read the rc2 and rc1
|
|
48
|
+
notes below as well.
|
|
49
|
+
|
|
3
50
|
## [0.11.0.rc2] - 2026-09-30
|
|
4
51
|
|
|
5
52
|
### Added
|
data/NOKOGIRI_DIFFERENCES.md
CHANGED
|
@@ -184,11 +184,11 @@ what browsers do - rather than libxml2. Detailed, test-backed notes live in
|
|
|
184
184
|
that does not fit the name (`set_attribute_ns`, `create_element_ns`)
|
|
185
185
|
raises `Makiri::Error`. Invalid UTF-8 raises `Makiri::Error` for every
|
|
186
186
|
argument, names included.
|
|
187
|
-
* `create_element_ns`
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
187
|
+
* `create_element_ns` keeps the case of an HTML-namespace name, as the DOM
|
|
188
|
+
does: `create_element_ns(XHTML, "BR")` is an unknown element named `BR`,
|
|
189
|
+
not a void `br`, and type selectors do not match it (XPath name tests do,
|
|
190
|
+
folding case on HTML elements as browsers do). Nokogiri has no
|
|
191
|
+
`create_element_ns`.
|
|
192
192
|
* An HTML `<template>` follows the WHATWG content model, which Nokogiri does
|
|
193
193
|
not: its parsed contents live in the separate fragment `Element#content_fragment`
|
|
194
194
|
returns, `template.children` is empty, and `inner_html` / `inner_html=`
|
|
@@ -198,6 +198,11 @@ what browsers do - rather than libxml2. Detailed, test-backed notes live in
|
|
|
198
198
|
ordinary element, with the parsed nodes as its children, so
|
|
199
199
|
`template.inner_html` and `template.children` answer the other way round and
|
|
200
200
|
there is no `content_fragment`.
|
|
201
|
+
* `Makiri::XML` has no template contents: an XHTML `<template>`'s children
|
|
202
|
+
are its children. Crossing into HTML they become its contents, and back
|
|
203
|
+
they become children - what a browser's XML parser, which puts them in the
|
|
204
|
+
contents, and its importNode give for the same document. An XML
|
|
205
|
+
`<template>` whose children should stay children has no way to say so.
|
|
201
206
|
* An HTML document has one root element and no text child, as the DOM requires;
|
|
202
207
|
`doc << element` beside an existing root raises.
|
|
203
208
|
* An insertion the DOM refuses - a child under a text, comment, PI, doctype
|
|
@@ -205,13 +210,20 @@ what browsers do - rather than libxml2. Detailed, test-backed notes live in
|
|
|
205
210
|
`Makiri::Error` in both representations. Nokogiri refuses the same ones with
|
|
206
211
|
`ArgumentError` (or `RuntimeError` for a second XML root).
|
|
207
212
|
* Moving HTML into an XML document (`xml_doc.import_node(html_node)`, or
|
|
208
|
-
inserting one)
|
|
209
|
-
|
|
210
|
-
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
213
|
+
inserting one) copies it as the DOM's clone does: every element in its
|
|
214
|
+
namespace from the start (an imported `<p>` is XHTML before it is
|
|
215
|
+
inserted), every attribute named as it is, and no `xmlns` attribute added -
|
|
216
|
+
`to_xml` and `canonicalize` write the declarations the output needs.
|
|
217
|
+
Nokogiri's copy of an HTML5 `<div>` is in no namespace, and libxml2
|
|
218
|
+
declares `xmlns:svg` on it for an SVG child (in `namespace_definitions`),
|
|
219
|
+
writing the child as `<svg:svg>`. One XML cannot write that way -
|
|
220
|
+
in no namespace with a colon or no XML name, `v-on:click`, `:href`,
|
|
221
|
+
`@click`, `fb:like` - crosses DOM-loose, and `to_xml` refuses the tree
|
|
222
|
+
while it is there; so does an element named with a colon (`<fb:like>`).
|
|
223
|
+
Nokogiri copies them and writes output that is not namespace-well-formed.
|
|
224
|
+
* One exception to the DOM: an HTML attribute named `xml:lang` (in no
|
|
225
|
+
namespace) becomes the XML namespace's `xml:lang`, as the XML reader reads
|
|
226
|
+
that name, so XHTML-style HTML still writes as XML.
|
|
215
227
|
* A known gap, in Lexbor's tag table: an HTML document that already holds a
|
|
216
228
|
parsed element named with a colon (`<x:y>`, one local name) and then
|
|
217
229
|
receives, by `import_node` from another document, a prefixed element
|
|
@@ -284,7 +296,7 @@ what browsers do - rather than libxml2. Detailed, test-backed notes live in
|
|
|
284
296
|
case-sensitivity rule, as browsers do: lower-cased for an HTML element (`LI`
|
|
285
297
|
matches `<li>`), as written for any other (`feGaussianBlur` matches the SVG
|
|
286
298
|
element, `fegaussianblur` does not). An HTML element named in upper case
|
|
287
|
-
(`create_element_ns(XHTML, "
|
|
299
|
+
(`create_element_ns(XHTML, "DIV")`, which keeps its name as the DOM does)
|
|
288
300
|
therefore matches no type selector.
|
|
289
301
|
* `Nokogiri::HTML5` is case-sensitive on HTML elements too, so `LI` does not
|
|
290
302
|
match `<li>` there. `Makiri::XML`'s `#css` is case-sensitive, as XML names
|
|
@@ -321,5 +333,12 @@ what browsers do - rather than libxml2. Detailed, test-backed notes live in
|
|
|
321
333
|
`Makiri::Error`.
|
|
322
334
|
* On re-parse, the HTML tokenizer replaces a U+0000 in text/attributes with
|
|
323
335
|
U+FFFD (WHATWG), so a serialized-then-reparsed round-trip is not byte-identical.
|
|
324
|
-
* `Makiri::XML`
|
|
325
|
-
|
|
336
|
+
* `Makiri::XML` holds the character data the DOM holds, NUL included:
|
|
337
|
+
text or an attribute value with a character XML has no `Char` for (`\f`,
|
|
338
|
+
U+0001, U+0000), a comment with `--`. Nokogiri takes the same, but NUL
|
|
339
|
+
(`ArgumentError`, a Ruby C-string limit). Where they part is the output. `to_xml` / `canonicalize` raise for such a tree; Nokogiri writes
|
|
340
|
+
`\f` in text as U+FFFD (the data changes) and a comment's `--` or an
|
|
341
|
+
attribute value's `\f` as it stands (the output does not parse).
|
|
342
|
+
`create_cdata` refuses `]]>` and `create_processing_instruction` `?>`,
|
|
343
|
+
raising `ArgumentError` as the DOM's factories do; Nokogiri takes the PI
|
|
344
|
+
and writes `<?t a?>b?>`, which re-reads as another tree.
|
data/lib/makiri/3.2/makiri.so
CHANGED
|
Binary file
|
data/lib/makiri/3.3/makiri.so
CHANGED
|
Binary file
|
data/lib/makiri/3.4/makiri.so
CHANGED
|
Binary file
|
data/lib/makiri/4.0/makiri.so
CHANGED
|
Binary file
|
data/lib/makiri/element.rb
CHANGED
|
@@ -22,5 +22,21 @@ module Makiri
|
|
|
22
22
|
def path_step
|
|
23
23
|
name_step("", unprefixed_element_namespace)
|
|
24
24
|
end
|
|
25
|
+
|
|
26
|
+
# See {NodePath#path}. An element step selects by expanded name, not by the
|
|
27
|
+
# string of the test: `*[local-name()='BR' and ...]`, written for `h:BR`,
|
|
28
|
+
# also selects the unprefixed `BR` beside it. A bare name test on an HTML
|
|
29
|
+
# element folds ASCII case, as the engine (and browsers) match it, so `div`
|
|
30
|
+
# also selects the `DIV` createElementNS makes. Compared as strings, each
|
|
31
|
+
# pair was counted apart, and the path found the other one first.
|
|
32
|
+
def path_selects?(other, test)
|
|
33
|
+
return false unless other.element? && other.namespace_uri.to_s == namespace_uri.to_s
|
|
34
|
+
|
|
35
|
+
mine = local_name
|
|
36
|
+
theirs = other.local_name
|
|
37
|
+
return mine == theirs unless test == name && unprefixed_element_namespace
|
|
38
|
+
|
|
39
|
+
mine.downcase(:ascii) == theirs.downcase(:ascii)
|
|
40
|
+
end
|
|
25
41
|
end
|
|
26
42
|
end
|
data/lib/makiri/node_path.rb
CHANGED
|
@@ -48,7 +48,7 @@ module Makiri
|
|
|
48
48
|
parent_node = parent
|
|
49
49
|
return test unless parent_node # a detached top: #path answers "?" anyway
|
|
50
50
|
|
|
51
|
-
siblings = parent_node.children.select { |c| c
|
|
51
|
+
siblings = parent_node.children.select { |c| path_selects?(c, test) }
|
|
52
52
|
return test if siblings.length <= 1
|
|
53
53
|
|
|
54
54
|
"#{test}[#{siblings.index(self) + 1}]"
|
|
@@ -73,6 +73,13 @@ module Makiri
|
|
|
73
73
|
nil
|
|
74
74
|
end
|
|
75
75
|
|
|
76
|
+
# Whether +test+, this node's own step, selects +other+, a sibling: the
|
|
77
|
+
# position counts exactly those. By default the tests are compared, which
|
|
78
|
+
# is the engine's answer for every kind but an element (see Element).
|
|
79
|
+
def path_selects?(other, test)
|
|
80
|
+
other.path_node_test == test
|
|
81
|
+
end
|
|
82
|
+
|
|
76
83
|
# The name test for an element (+axis+ "") or attribute (+axis+ "@"): the
|
|
77
84
|
# bare name when an unprefixed test reaches this node - it is in
|
|
78
85
|
# +plain_namespace+ and its name is a plain NCName - else by expanded name.
|
data/lib/makiri/version.rb
CHANGED
|
@@ -56,9 +56,10 @@ UNSAFE_ISLANDS = {
|
|
|
56
56
|
"lexbor/adapter/arena_bytes.rs" => 14,
|
|
57
57
|
"lexbor/adapter/cross_import.rs" => 2,
|
|
58
58
|
"lexbor/adapter/html/attrs.rs" => 20,
|
|
59
|
-
"lexbor/adapter/html/build.rs" =>
|
|
59
|
+
"lexbor/adapter/html/build.rs" => 19,
|
|
60
60
|
"lexbor/adapter/html/mod.rs" => 60,
|
|
61
61
|
"lexbor/adapter/html/mutate.rs" => 9,
|
|
62
|
+
"lexbor/adapter/html/serialize.rs" => 9,
|
|
62
63
|
"lexbor/adapter/post_parse.rs" => 13,
|
|
63
64
|
"lexbor/adapter/source_loc.rs" => 2,
|
|
64
65
|
"lexbor/adapter/text_index.rs" => 1,
|
|
@@ -461,7 +462,7 @@ end
|
|
|
461
462
|
# is generated, so its signature is Lexbor's. A new hand declaration of a
|
|
462
463
|
# header-declared function belongs in build.rs's allowlist instead; one with no
|
|
463
464
|
# header belongs in build.rs's UNDECLARED_EXPORTS as well as here.
|
|
464
|
-
LEXBOR_HAND_DECLS =
|
|
465
|
+
LEXBOR_HAND_DECLS = 4
|
|
465
466
|
abi_decls = comments_removed(File.binread(File.join(RUST, "lexbor/abi.rs"))).scan(LEXBOR_DECL).length
|
|
466
467
|
if abi_decls != LEXBOR_HAND_DECLS
|
|
467
468
|
errors << "lexbor/abi.rs hand-declares #{abi_decls} Lexbor functions (expected " \
|