leptris 1.9.48-aarch64-linux → 1.9.50-aarch64-linux
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +42 -0
- data/README.adoc +99 -79
- data/lib/leptris/version.rb +1 -1
- data/lib/leptris/xml/element.rb +7 -3
- data/lib/leptris/xml/node.rb +6 -3
- data/lib/leptris/xml/serialization.rb +7 -11
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: d25c7c446cfd4440a669e9a98bae64df75b67af90cd1fa3c0c0094bce573fddf
|
|
4
|
+
data.tar.gz: a8d20211ae6e8fe1155ad1d6e716e95035c70945628f66e859bbdbf881ad342f
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: db0b93273a396fa85a6eaafae23c671d882903df830053ec5eae2e58aea7581cdb24b78f297ee349891ab407ad12abb8742d945a519704664b6a03aa6d9e1707
|
|
7
|
+
data.tar.gz: 6d02abe32d00af709387c18b0551381944805c26c0001170bd65e9408e4f6166f2de5b31c7c273cf5d5b4e945be1609a7d54e206adf601cfe0c27ba85c40264b
|
data/CHANGELOG.md
CHANGED
|
@@ -5,6 +5,48 @@ All notable changes to Leptris will be documented in this file.
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/1.0.0/),
|
|
6
6
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
7
|
|
|
8
|
+
## [1.9.50] - 2026-09-01
|
|
9
|
+
|
|
10
|
+
### Fixed
|
|
11
|
+
|
|
12
|
+
- **Copies keep comment and PI children (leptris-ruby#115)**:
|
|
13
|
+
`leptris_element_copy` silently drops COMMENT and PI children —
|
|
14
|
+
the C child-copy loop reads `/* Skip COMMENT and PI nodes for
|
|
15
|
+
now */` (engine-side, filed as leptris/leptris#696 with the
|
|
16
|
+
line). Blast radius in the binding: `Element#dup/#clone` and the
|
|
17
|
+
element indent-unit path both composed through the C copy. Both
|
|
18
|
+
now round through serialization instead — the serializer
|
|
19
|
+
preserves every child kind, prefixed names, and namespace
|
|
20
|
+
declarations, and re-parsing rebuilds them all; `dup` returns a
|
|
21
|
+
tree in a NEW document (unchanged contract), detached from the
|
|
22
|
+
original (spec-pinned).
|
|
23
|
+
|
|
24
|
+
## [1.9.49] - 2026-09-01
|
|
25
|
+
|
|
26
|
+
### Added
|
|
27
|
+
|
|
28
|
+
- **XSLT 1.0–3.0 pinned through the generic face**: specs
|
|
29
|
+
transform stylesheets of both generations — a 1.0
|
|
30
|
+
count/template sheet and a 3.0 sheet exercising
|
|
31
|
+
`xsl:analyze-string`, `let`, and the `=>` arrow (the XPath 3.1
|
|
32
|
+
core). The engine dispatches on the declared version; the
|
|
33
|
+
binding needed no new surface.
|
|
34
|
+
- **README rewritten to the full current feature set**: the
|
|
35
|
+
document chain (add/remove/mutate document-level PIs),
|
|
36
|
+
`Node#visit`, `Element#inner_html`, the indent unit and display
|
|
37
|
+
form, XSLT with the upstream status of standalone XPath 2/3 and
|
|
38
|
+
XQuery, SAX transports (interest-proportional delivery,
|
|
39
|
+
recorder, reset), pull (prefix events, batch guard), iterparse
|
|
40
|
+
v2, and the refreshed head-to-head table (parse 10–12x, CSS 4x,
|
|
41
|
+
scalar XPath 4.6x, SAX 6x on selective handlers, memory 1.8x).
|
|
42
|
+
|
|
43
|
+
### Meta
|
|
44
|
+
|
|
45
|
+
- Standalone XPath 2.0/3.1 evaluation entries and XQuery are not
|
|
46
|
+
yet engine surface — filed as leptris/leptris#683 with the
|
|
47
|
+
inventory; the binding adopts them lockstep-fashion when they
|
|
48
|
+
land.
|
|
49
|
+
|
|
8
50
|
## [1.9.48] - 2026-09-01
|
|
9
51
|
|
|
10
52
|
### Changed
|
data/README.adoc
CHANGED
|
@@ -8,7 +8,8 @@ image:https://github.com/leptris/leptris-ruby/actions/workflows/build.yml/badge.
|
|
|
8
8
|
A Nokogiri-compatible Ruby binding for
|
|
9
9
|
https://github.com/leptris/leptris[libleptris], a pure-C99 XML 1.0 parser
|
|
10
10
|
with full https://www.w3.org/TR/1999/REC-xpath-19991116/[XPath 1.0],
|
|
11
|
-
XML Namespaces 1.0, SAX, and C14N (1.0 / 1.1 / Exclusive)
|
|
11
|
+
XML Namespaces 1.0, SAX, and C14N (1.0 / 1.1 / Exclusive), plus
|
|
12
|
+
XSLT 1.0–3.0 transforms (the XPath 3.1 core rides inside them).
|
|
12
13
|
|
|
13
14
|
The C DOM is the single source of truth — Ruby objects are thin FFI
|
|
14
15
|
handles over the C pointers, so every Ruby method maps to one FFI call.
|
|
@@ -134,7 +135,18 @@ end
|
|
|
134
135
|
|
|
135
136
|
=== Tree iteration
|
|
136
137
|
|
|
137
|
-
`Node#
|
|
138
|
+
`Node#visit` walks the subtree with ONE C call and C-tracked depth —
|
|
139
|
+
elements yield `(node, entering, depth)` enter/leave pairs, every other
|
|
140
|
+
kind yields once; the leanest full-subtree iteration the binding offers:
|
|
141
|
+
|
|
142
|
+
[source,ruby]
|
|
143
|
+
----
|
|
144
|
+
doc.root.visit do |node, entering, depth|
|
|
145
|
+
puts "#{' ' * depth}#{node.name} #{entering ? 'enter' : 'leave'}"
|
|
146
|
+
end
|
|
147
|
+
----
|
|
148
|
+
|
|
149
|
+
`Node#traverse` walks the subtree in post-order via a single C-side
|
|
138
150
|
callback (one FFI call for the whole traversal, not one per node):
|
|
139
151
|
|
|
140
152
|
[source,ruby]
|
|
@@ -148,6 +160,17 @@ doc.root.traverse do |node|
|
|
|
148
160
|
end
|
|
149
161
|
----
|
|
150
162
|
|
|
163
|
+
=== The document chain
|
|
164
|
+
|
|
165
|
+
`Document#node` is the document-level navigation head (the libxml2
|
|
166
|
+
model): `Document#children` reads `[prolog comments/PIs, the root,
|
|
167
|
+
epilog comments/PIs]` in document order, with whitespace kept per
|
|
168
|
+
libxml2's exact rule. Document-level PIs are first-class:
|
|
169
|
+
`Document#add_pi(target, data)` appends one, `Document#remove_pi(target_or_index)`
|
|
170
|
+
removes by target or index, and `PI#target=`/`PI#data=`/`PI#unlink`
|
|
171
|
+
mutate and detach. `Document#processing_instructions` and
|
|
172
|
+
`Document#comments` remain the memoized readers.
|
|
173
|
+
|
|
151
174
|
== Readonly mode
|
|
152
175
|
|
|
153
176
|
For the dominant parse-query-serialize workload, parse with
|
|
@@ -289,8 +312,38 @@ doc.xpath("//t:title")
|
|
|
289
312
|
== Serialization and canonicalization
|
|
290
313
|
|
|
291
314
|
[horizontal]
|
|
292
|
-
`Document#to_xml(indent: 0, no_decl: false, encoding: nil)` (aliases `#to_s`, `#serialize`) :: serialize the whole document.
|
|
293
|
-
`Element#to_xml(...)` :: serialize a subtree.
|
|
315
|
+
`Document#to_xml(indent: 0, no_decl: false, encoding: nil, indent_text: false)` (aliases `#to_s`, `#serialize`) :: serialize the whole document.
|
|
316
|
+
`Element#to_xml(...)` :: serialize a subtree (also takes `indent_text:` — the unit string).
|
|
317
|
+
|
|
318
|
+
`indent_text` carries two meanings: a STRING is the indent unit with
|
|
319
|
+
Nokogiri's semantics (the unit replaces the default spaces, repeated
|
|
320
|
+
`indent` times per depth level — byte-identical to Nokogiri's output);
|
|
321
|
+
`true` selects the display form, which also indents text and mixed
|
|
322
|
+
content:
|
|
323
|
+
|
|
324
|
+
[source,ruby]
|
|
325
|
+
----
|
|
326
|
+
doc.to_xml(indent: 2, indent_text: "\t") # tab-indented
|
|
327
|
+
doc.to_xml(indent: 2, indent_text: true) # display form (documents only)
|
|
328
|
+
----
|
|
329
|
+
|
|
330
|
+
`Element#inner_html` serializes the children with correct escaping —
|
|
331
|
+
well-formed by construction (a re-parse spec pins it).
|
|
332
|
+
|
|
333
|
+
XSLT 1.0–3.0 transforms run through `Leptris::XML::XSLT`:
|
|
334
|
+
|
|
335
|
+
[source,ruby]
|
|
336
|
+
----
|
|
337
|
+
style = Leptris::XML::XSLT.parse(stylesheet_xml) # or .parse_file
|
|
338
|
+
result = style.apply_to(doc) # => Document
|
|
339
|
+
style.serialize(doc) # => String
|
|
340
|
+
----
|
|
341
|
+
|
|
342
|
+
The engine dispatches on the stylesheet's declared version; the 3.0
|
|
343
|
+
instruction set (grouping, `xsl:accumulator`, `xsl:analyze-string`,
|
|
344
|
+
the XPath 3.1 core) is available inside transforms. Standalone
|
|
345
|
+
XPath 2.0/3.1 evaluation entries and XQuery are
|
|
346
|
+
https://github.com/leptris/leptris/issues/683[tracked upstream].
|
|
294
347
|
`Document#save(path, **opts)` :: serialize to a file.
|
|
295
348
|
`Document#canonicalize(version, inclusive_ns, with_comments:, exclusive:, mode:)` (alias `#c14n`) :: canonical XML.
|
|
296
349
|
`Element#canonicalize(...)` :: subtree canonicalization.
|
|
@@ -351,6 +404,32 @@ to `#read`. The handler callbacks are:
|
|
|
351
404
|
`start_prefix_mapping(prefix, uri)`, `end_prefix_mapping(prefix)` :: namespace events.
|
|
352
405
|
`warning(str)`, `error(msg, line, col)` :: recoverable parser messages.
|
|
353
406
|
|
|
407
|
+
=== Transports: interest-proportional delivery
|
|
408
|
+
|
|
409
|
+
The parser picks its transport by what your handler overrides. Overriding
|
|
410
|
+
one hot kind (say `characters`) attaches only that callback — the engine
|
|
411
|
+
skips C-side emission for the rest entirely. Overriding several rides the
|
|
412
|
+
bulk recorder (one C call stages the whole document, then a lean dispatch
|
|
413
|
+
loop) — measured on a 1.9 MB document: text-only 21 ms, all-events 119 ms
|
|
414
|
+
vs Nokogiri's 130/131 ms. The handler cannot tell the transports apart;
|
|
415
|
+
xmlns declarations ride the attribute pairs and prefix-mapping events
|
|
416
|
+
fire alongside.
|
|
417
|
+
|
|
418
|
+
For raw bulk event streams, `SAX::Recorder.parse(xml, kinds:)` drains
|
|
419
|
+
buffered events per chunk — unwanted kinds cost one array read — and
|
|
420
|
+
`Recorder#reset` reuses one recorder across documents.
|
|
421
|
+
|
|
422
|
+
=== Pull parsing and iterparse
|
|
423
|
+
|
|
424
|
+
`Leptris::XML::Pull` is the StAX-style cursor: `Parser#each` delivers
|
|
425
|
+
`Event` structs with types including `:start_prefix`/`:end_prefix` (the
|
|
426
|
+
default namespace's prefix is `""`); `Parser#each_batch(max)` delivers
|
|
427
|
+
events in bulk with a corruption guard that fails loudly rather than
|
|
428
|
+
delivering garbage. `Leptris::XML::Iterparse.parse(xml, mode:)` yields
|
|
429
|
+
completed subtrees top-level (`:top_level`) or every element post-order
|
|
430
|
+
(`:full_document`) with bounded memory, plus `#namespace_uri` on the last
|
|
431
|
+
yielded element and an `#error` channel for truncated input.
|
|
432
|
+
|
|
354
433
|
== Memory model
|
|
355
434
|
|
|
356
435
|
`Document` is the only object that owns C memory. Everything else
|
|
@@ -436,82 +515,23 @@ Notable differences:
|
|
|
436
515
|
|
|
437
516
|
== Performance
|
|
438
517
|
|
|
439
|
-
|
|
440
|
-
|
|
441
|
-
`benchmark/read_paths.rb` measures the binding's own hot paths
|
|
442
|
-
(readonly read loops, SAX, NodeSet unions, css, parse-query-
|
|
443
|
-
serialize). Machine-relative — compare before/after runs on the same
|
|
444
|
-
machine:
|
|
445
|
-
|
|
446
|
-
bundle exec ruby benchmark/read_paths.rb
|
|
447
|
-
|
|
448
|
-
=== Full-field comparison
|
|
449
|
-
|
|
450
|
-
Measured with metanorma/serialbench (Ruby 3.4.8, leptris 1.6.0,
|
|
451
|
-
libleptris 1.6.0, macOS arm64) — Leptris vs the full field:
|
|
452
|
-
|
|
453
|
-
[cols="2,1,1,1,1",options="header"]
|
|
454
|
-
|===
|
|
455
|
-
|Operation |Leptris |Ox |Nokogiri |vs Ox
|
|
456
|
-
|
|
457
|
-
|parse medium (300 KB)
|
|
458
|
-
|0.97 ms
|
|
459
|
-
|3.00 ms
|
|
460
|
-
|8.57 ms
|
|
461
|
-
|3.1x faster
|
|
462
|
-
|
|
463
|
-
|parse large (4.4 MB)
|
|
464
|
-
|13.51 ms
|
|
465
|
-
|77.06 ms
|
|
466
|
-
|104.22 ms
|
|
467
|
-
|5.7x faster
|
|
468
|
-
|
|
469
|
-
|generate medium
|
|
470
|
-
|1.75 ms
|
|
471
|
-
|14.19 ms
|
|
472
|
-
|23.93 ms
|
|
473
|
-
|8.1x faster
|
|
474
|
-
|
|
475
|
-
|generate large
|
|
476
|
-
|74.57 ms
|
|
477
|
-
|80.91 ms
|
|
478
|
-
|117.16 ms
|
|
479
|
-
|1.1x faster
|
|
480
|
-
|
|
481
|
-
|streaming medium
|
|
482
|
-
|0.73 ms
|
|
483
|
-
|103.82 ms
|
|
484
|
-
|43.74 ms
|
|
485
|
-
|142x faster
|
|
486
|
-
|
|
487
|
-
|streaming large
|
|
488
|
-
|8.52 ms
|
|
489
|
-
|204.14 ms
|
|
490
|
-
|239.14 ms
|
|
491
|
-
|24x faster
|
|
492
|
-
|
|
493
|
-
|Ruby allocations (medium)
|
|
494
|
-
|0.01 MB
|
|
495
|
-
|18.43 MB
|
|
496
|
-
|39.05 MB
|
|
497
|
-
|~1800x less
|
|
498
|
-
|===
|
|
499
|
-
|
|
500
|
-
Leptris takes first place on every XML operation above small-document
|
|
501
|
-
size, including against Ox (the C-extension speed champion). On
|
|
502
|
-
small-document micro-operations (sub-5 µs), C extensions hold an
|
|
503
|
-
inherent edge per call; Leptris's readonly mode closes the
|
|
504
|
-
steady-state gap by memoizing reads behind an immutability guarantee.
|
|
505
|
-
|
|
506
|
-
For XPath-heavy workloads, `Leptris::XML::XPath.compile` evaluates a
|
|
507
|
-
parsed expression repeatedly without re-parsing.
|
|
508
|
-
|
|
509
|
-
Run the local benchmark for numbers on your hardware:
|
|
518
|
+
Head-to-head against published Nokogiri 1.19.4 on a 1.86 MB /
|
|
519
|
+
250k-event document (arm64-darwin, CPU totals, best-of-5):
|
|
510
520
|
|
|
511
|
-
[
|
|
512
|
-
|
|
513
|
-
|
|
514
|
-
|
|
521
|
+
[horizontal]
|
|
522
|
+
DOM parse :: **10–12x faster**
|
|
523
|
+
CSS search :: **4.0–4.8x**
|
|
524
|
+
XPath nodeset :: **2.4–3.0x**
|
|
525
|
+
XPath scalar (`string(//item[1])`) :: **4.6x**
|
|
526
|
+
Serialization :: **2.2–2.3x**
|
|
527
|
+
`at_css` :: **1.4x**
|
|
528
|
+
SAX text-only handler :: **6x** (119 ms all-events vs Nokogiri's 131)
|
|
529
|
+
Memory held :: **17.5 MB/doc vs Nokogiri's 31.3** (1.8x lighter)
|
|
530
|
+
|
|
531
|
+
Two rows sit at the Ruby allocation floor by documented choice: the
|
|
532
|
+
cold full-tree walk (children recursion) remains ~1.5x behind
|
|
533
|
+
Nokogiri's C-extension node creation — `Node#visit` is the wrap-free
|
|
534
|
+
lever when that matters.
|
|
515
535
|
|
|
516
536
|
== Development
|
|
517
537
|
|
data/lib/leptris/version.rb
CHANGED
data/lib/leptris/xml/element.rb
CHANGED
|
@@ -251,10 +251,14 @@ class Leptris::XML::Element < Leptris::XML::Node
|
|
|
251
251
|
self
|
|
252
252
|
end
|
|
253
253
|
|
|
254
|
+
# Deep copy in a NEW document, via serialization round-trip:
|
|
255
|
+
# leptris_element_copy silently drops COMMENT and PI children
|
|
256
|
+
# (leptris/leptris#696, leptris-ruby#115) — the serializer
|
|
257
|
+
# preserves every child kind, prefixed names, and namespace
|
|
258
|
+
# declarations, and re-parsing rebuilds them all.
|
|
254
259
|
def dup
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
Leptris::XML::Node.wrap(copy_ptr, @document)
|
|
260
|
+
ensure_alive!
|
|
261
|
+
Leptris::XML::Document.parse(to_xml).root
|
|
258
262
|
end
|
|
259
263
|
alias_method :clone, :dup
|
|
260
264
|
|
data/lib/leptris/xml/node.rb
CHANGED
|
@@ -369,13 +369,16 @@ class Leptris::XML::Node
|
|
|
369
369
|
result
|
|
370
370
|
end
|
|
371
371
|
|
|
372
|
+
# Deep copy in a NEW document. Rounds through serialization
|
|
373
|
+
# rather than leptris_element_copy: the C copy silently drops
|
|
374
|
+
# COMMENT and PI children (leptris/leptris#696, leptris-ruby#115)
|
|
375
|
+
# — the serializer preserves every child kind, prefixed names,
|
|
376
|
+
# and namespace declarations, and re-parsing rebuilds them all.
|
|
372
377
|
def dup
|
|
373
378
|
ensure_alive!
|
|
374
379
|
elem_ptr = Leptris::XML::FFI.leptris_node_as_element(@c_ptr)
|
|
375
380
|
raise Leptris::XML::Error, "dup is only supported for element nodes" if elem_ptr.null?
|
|
376
|
-
|
|
377
|
-
raise Leptris::XML::Error, "leptris_element_copy failed" if copy_ptr.null?
|
|
378
|
-
Leptris::XML::Node.wrap(copy_ptr, @document)
|
|
381
|
+
Leptris::XML::Document.parse(to_xml).root
|
|
379
382
|
end
|
|
380
383
|
alias_method :clone, :dup
|
|
381
384
|
|
|
@@ -97,19 +97,15 @@ module Leptris::XML::Serialization
|
|
|
97
97
|
|
|
98
98
|
# Element-face indent unit (leptris-ruby#109 residual 2): no
|
|
99
99
|
# element-level ext-serialize entry exists yet, so the unit path
|
|
100
|
-
#
|
|
101
|
-
#
|
|
102
|
-
#
|
|
103
|
-
#
|
|
100
|
+
# rounds the subtree through serialization into a fresh document
|
|
101
|
+
# (leptris_element_copy drops COMMENT/PI children —
|
|
102
|
+
# leptris/leptris#696, leptris-ruby#115) and serializes that
|
|
103
|
+
# document without a declaration — identical output to the
|
|
104
|
+
# element serializer for the standard layout, with the unit and
|
|
105
|
+
# every child kind intact.
|
|
104
106
|
def self.to_xml_element_unit(element, unit, indent: 0)
|
|
105
|
-
document = Leptris::XML::Document.
|
|
107
|
+
document = Leptris::XML::Document.parse(element.to_xml)
|
|
106
108
|
begin
|
|
107
|
-
copy_ptr = Leptris::XML::FFI.leptris_element_copy(
|
|
108
|
-
element.c_ptr, document.c_ptr)
|
|
109
|
-
if copy_ptr.null?
|
|
110
|
-
raise Leptris::XML::Error, "leptris_element_copy failed"
|
|
111
|
-
end
|
|
112
|
-
document.root = Leptris::XML::Node.wrap(copy_ptr, document)
|
|
113
109
|
to_xml_indent_unit(document.c_ptr, unit,
|
|
114
110
|
indent: indent, no_decl: true)
|
|
115
111
|
ensure
|