leptris 1.9.48-aarch64-linux → 1.9.50-aarch64-linux

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 0c77801b701593dafd58576b1df4fab6c79b2ce579c32217a82bf605e964f7f0
4
- data.tar.gz: 681bbd1dfe9e629cb31a017549e0d195b46fcf1c697c788a38a8a73a7161a403
3
+ metadata.gz: d25c7c446cfd4440a669e9a98bae64df75b67af90cd1fa3c0c0094bce573fddf
4
+ data.tar.gz: a8d20211ae6e8fe1155ad1d6e716e95035c70945628f66e859bbdbf881ad342f
5
5
  SHA512:
6
- metadata.gz: e2530d9ba9a342ab1fbfdb7b45e0dcae234cefd91798d7e2150ca5f8b77cad1104ae387bb780677d67ad2db606fbefe64381c0ff869e06dbb3b33f0c8751be9c
7
- data.tar.gz: 3613d6f3e91858ae4c802438cbdcb7933a8bb5d93c6aa0bd546475cd59058508247ae70217c5f420398425e8367b61b3c36e95598fefa78166ed251f18bbc360
6
+ metadata.gz: db0b93273a396fa85a6eaafae23c671d882903df830053ec5eae2e58aea7581cdb24b78f297ee349891ab407ad12abb8742d945a519704664b6a03aa6d9e1707
7
+ data.tar.gz: 6d02abe32d00af709387c18b0551381944805c26c0001170bd65e9408e4f6166f2de5b31c7c273cf5d5b4e945be1609a7d54e206adf601cfe0c27ba85c40264b
data/CHANGELOG.md CHANGED
@@ -5,6 +5,48 @@ All notable changes to Leptris will be documented in this file.
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/1.0.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
 
8
+ ## [1.9.50] - 2026-09-01
9
+
10
+ ### Fixed
11
+
12
+ - **Copies keep comment and PI children (leptris-ruby#115)**:
13
+ `leptris_element_copy` silently drops COMMENT and PI children —
14
+ the C child-copy loop reads `/* Skip COMMENT and PI nodes for
15
+ now */` (engine-side, filed as leptris/leptris#696 with the
16
+ line). Blast radius in the binding: `Element#dup/#clone` and the
17
+ element indent-unit path both composed through the C copy. Both
18
+ now round through serialization instead — the serializer
19
+ preserves every child kind, prefixed names, and namespace
20
+ declarations, and re-parsing rebuilds them all; `dup` returns a
21
+ tree in a NEW document (unchanged contract), detached from the
22
+ original (spec-pinned).
23
+
24
+ ## [1.9.49] - 2026-09-01
25
+
26
+ ### Added
27
+
28
+ - **XSLT 1.0–3.0 pinned through the generic face**: specs
29
+ transform stylesheets of both generations — a 1.0
30
+ count/template sheet and a 3.0 sheet exercising
31
+ `xsl:analyze-string`, `let`, and the `=>` arrow (the XPath 3.1
32
+ core). The engine dispatches on the declared version; the
33
+ binding needed no new surface.
34
+ - **README rewritten to the full current feature set**: the
35
+ document chain (add/remove/mutate document-level PIs),
36
+ `Node#visit`, `Element#inner_html`, the indent unit and display
37
+ form, XSLT with the upstream status of standalone XPath 2/3 and
38
+ XQuery, SAX transports (interest-proportional delivery,
39
+ recorder, reset), pull (prefix events, batch guard), iterparse
40
+ v2, and the refreshed head-to-head table (parse 10–12x, CSS 4x,
41
+ scalar XPath 4.6x, SAX 6x on selective handlers, memory 1.8x).
42
+
43
+ ### Meta
44
+
45
+ - Standalone XPath 2.0/3.1 evaluation entries and XQuery are not
46
+ yet engine surface — filed as leptris/leptris#683 with the
47
+ inventory; the binding adopts them lockstep-fashion when they
48
+ land.
49
+
8
50
  ## [1.9.48] - 2026-09-01
9
51
 
10
52
  ### Changed
data/README.adoc CHANGED
@@ -8,7 +8,8 @@ image:https://github.com/leptris/leptris-ruby/actions/workflows/build.yml/badge.
8
8
  A Nokogiri-compatible Ruby binding for
9
9
  https://github.com/leptris/leptris[libleptris], a pure-C99 XML 1.0 parser
10
10
  with full https://www.w3.org/TR/1999/REC-xpath-19991116/[XPath 1.0],
11
- XML Namespaces 1.0, SAX, and C14N (1.0 / 1.1 / Exclusive).
11
+ XML Namespaces 1.0, SAX, and C14N (1.0 / 1.1 / Exclusive), plus
12
+ XSLT 1.0–3.0 transforms (the XPath 3.1 core rides inside them).
12
13
 
13
14
  The C DOM is the single source of truth — Ruby objects are thin FFI
14
15
  handles over the C pointers, so every Ruby method maps to one FFI call.
@@ -134,7 +135,18 @@ end
134
135
 
135
136
  === Tree iteration
136
137
 
137
- `Node#traverse` walks the subtree in document order via a single C-side
138
+ `Node#visit` walks the subtree with ONE C call and C-tracked depth —
139
+ elements yield `(node, entering, depth)` enter/leave pairs, every other
140
+ kind yields once; the leanest full-subtree iteration the binding offers:
141
+
142
+ [source,ruby]
143
+ ----
144
+ doc.root.visit do |node, entering, depth|
145
+ puts "#{' ' * depth}#{node.name} #{entering ? 'enter' : 'leave'}"
146
+ end
147
+ ----
148
+
149
+ `Node#traverse` walks the subtree in post-order via a single C-side
138
150
  callback (one FFI call for the whole traversal, not one per node):
139
151
 
140
152
  [source,ruby]
@@ -148,6 +160,17 @@ doc.root.traverse do |node|
148
160
  end
149
161
  ----
150
162
 
163
+ === The document chain
164
+
165
+ `Document#node` is the document-level navigation head (the libxml2
166
+ model): `Document#children` reads `[prolog comments/PIs, the root,
167
+ epilog comments/PIs]` in document order, with whitespace kept per
168
+ libxml2's exact rule. Document-level PIs are first-class:
169
+ `Document#add_pi(target, data)` appends one, `Document#remove_pi(target_or_index)`
170
+ removes by target or index, and `PI#target=`/`PI#data=`/`PI#unlink`
171
+ mutate and detach. `Document#processing_instructions` and
172
+ `Document#comments` remain the memoized readers.
173
+
151
174
  == Readonly mode
152
175
 
153
176
  For the dominant parse-query-serialize workload, parse with
@@ -289,8 +312,38 @@ doc.xpath("//t:title")
289
312
  == Serialization and canonicalization
290
313
 
291
314
  [horizontal]
292
- `Document#to_xml(indent: 0, no_decl: false, encoding: nil)` (aliases `#to_s`, `#serialize`) :: serialize the whole document.
293
- `Element#to_xml(...)` :: serialize a subtree.
315
+ `Document#to_xml(indent: 0, no_decl: false, encoding: nil, indent_text: false)` (aliases `#to_s`, `#serialize`) :: serialize the whole document.
316
+ `Element#to_xml(...)` :: serialize a subtree (also takes `indent_text:` — the unit string).
317
+
318
+ `indent_text` carries two meanings: a STRING is the indent unit with
319
+ Nokogiri's semantics (the unit replaces the default spaces, repeated
320
+ `indent` times per depth level — byte-identical to Nokogiri's output);
321
+ `true` selects the display form, which also indents text and mixed
322
+ content:
323
+
324
+ [source,ruby]
325
+ ----
326
+ doc.to_xml(indent: 2, indent_text: "\t") # tab-indented
327
+ doc.to_xml(indent: 2, indent_text: true) # display form (documents only)
328
+ ----
329
+
330
+ `Element#inner_html` serializes the children with correct escaping —
331
+ well-formed by construction (a re-parse spec pins it).
332
+
333
+ XSLT 1.0–3.0 transforms run through `Leptris::XML::XSLT`:
334
+
335
+ [source,ruby]
336
+ ----
337
+ style = Leptris::XML::XSLT.parse(stylesheet_xml) # or .parse_file
338
+ result = style.apply_to(doc) # => Document
339
+ style.serialize(doc) # => String
340
+ ----
341
+
342
+ The engine dispatches on the stylesheet's declared version; the 3.0
343
+ instruction set (grouping, `xsl:accumulator`, `xsl:analyze-string`,
344
+ the XPath 3.1 core) is available inside transforms. Standalone
345
+ XPath 2.0/3.1 evaluation entries and XQuery are
346
+ https://github.com/leptris/leptris/issues/683[tracked upstream].
294
347
  `Document#save(path, **opts)` :: serialize to a file.
295
348
  `Document#canonicalize(version, inclusive_ns, with_comments:, exclusive:, mode:)` (alias `#c14n`) :: canonical XML.
296
349
  `Element#canonicalize(...)` :: subtree canonicalization.
@@ -351,6 +404,32 @@ to `#read`. The handler callbacks are:
351
404
  `start_prefix_mapping(prefix, uri)`, `end_prefix_mapping(prefix)` :: namespace events.
352
405
  `warning(str)`, `error(msg, line, col)` :: recoverable parser messages.
353
406
 
407
+ === Transports: interest-proportional delivery
408
+
409
+ The parser picks its transport by what your handler overrides. Overriding
410
+ one hot kind (say `characters`) attaches only that callback — the engine
411
+ skips C-side emission for the rest entirely. Overriding several rides the
412
+ bulk recorder (one C call stages the whole document, then a lean dispatch
413
+ loop) — measured on a 1.9 MB document: text-only 21 ms, all-events 119 ms
414
+ vs Nokogiri's 130/131 ms. The handler cannot tell the transports apart;
415
+ xmlns declarations ride the attribute pairs and prefix-mapping events
416
+ fire alongside.
417
+
418
+ For raw bulk event streams, `SAX::Recorder.parse(xml, kinds:)` drains
419
+ buffered events per chunk — unwanted kinds cost one array read — and
420
+ `Recorder#reset` reuses one recorder across documents.
421
+
422
+ === Pull parsing and iterparse
423
+
424
+ `Leptris::XML::Pull` is the StAX-style cursor: `Parser#each` delivers
425
+ `Event` structs with types including `:start_prefix`/`:end_prefix` (the
426
+ default namespace's prefix is `""`); `Parser#each_batch(max)` delivers
427
+ events in bulk with a corruption guard that fails loudly rather than
428
+ delivering garbage. `Leptris::XML::Iterparse.parse(xml, mode:)` yields
429
+ completed subtrees top-level (`:top_level`) or every element post-order
430
+ (`:full_document`) with bounded memory, plus `#namespace_uri` on the last
431
+ yielded element and an `#error` channel for truncated input.
432
+
354
433
  == Memory model
355
434
 
356
435
  `Document` is the only object that owns C memory. Everything else
@@ -436,82 +515,23 @@ Notable differences:
436
515
 
437
516
  == Performance
438
517
 
439
- === Internal read-path harness
440
-
441
- `benchmark/read_paths.rb` measures the binding's own hot paths
442
- (readonly read loops, SAX, NodeSet unions, css, parse-query-
443
- serialize). Machine-relative — compare before/after runs on the same
444
- machine:
445
-
446
- bundle exec ruby benchmark/read_paths.rb
447
-
448
- === Full-field comparison
449
-
450
- Measured with metanorma/serialbench (Ruby 3.4.8, leptris 1.6.0,
451
- libleptris 1.6.0, macOS arm64) — Leptris vs the full field:
452
-
453
- [cols="2,1,1,1,1",options="header"]
454
- |===
455
- |Operation |Leptris |Ox |Nokogiri |vs Ox
456
-
457
- |parse medium (300 KB)
458
- |0.97 ms
459
- |3.00 ms
460
- |8.57 ms
461
- |3.1x faster
462
-
463
- |parse large (4.4 MB)
464
- |13.51 ms
465
- |77.06 ms
466
- |104.22 ms
467
- |5.7x faster
468
-
469
- |generate medium
470
- |1.75 ms
471
- |14.19 ms
472
- |23.93 ms
473
- |8.1x faster
474
-
475
- |generate large
476
- |74.57 ms
477
- |80.91 ms
478
- |117.16 ms
479
- |1.1x faster
480
-
481
- |streaming medium
482
- |0.73 ms
483
- |103.82 ms
484
- |43.74 ms
485
- |142x faster
486
-
487
- |streaming large
488
- |8.52 ms
489
- |204.14 ms
490
- |239.14 ms
491
- |24x faster
492
-
493
- |Ruby allocations (medium)
494
- |0.01 MB
495
- |18.43 MB
496
- |39.05 MB
497
- |~1800x less
498
- |===
499
-
500
- Leptris takes first place on every XML operation above small-document
501
- size, including against Ox (the C-extension speed champion). On
502
- small-document micro-operations (sub-5 µs), C extensions hold an
503
- inherent edge per call; Leptris's readonly mode closes the
504
- steady-state gap by memoizing reads behind an immutability guarantee.
505
-
506
- For XPath-heavy workloads, `Leptris::XML::XPath.compile` evaluates a
507
- parsed expression repeatedly without re-parsing.
508
-
509
- Run the local benchmark for numbers on your hardware:
518
+ Head-to-head against published Nokogiri 1.19.4 on a 1.86 MB /
519
+ 250k-event document (arm64-darwin, CPU totals, best-of-5):
510
520
 
511
- [source,shell]
512
- ----
513
- bundle exec ruby benchmark/leptris_vs_nokogiri.rb
514
- ----
521
+ [horizontal]
522
+ DOM parse :: **10–12x faster**
523
+ CSS search :: **4.0–4.8x**
524
+ XPath nodeset :: **2.4–3.0x**
525
+ XPath scalar (`string(//item[1])`) :: **4.6x**
526
+ Serialization :: **2.2–2.3x**
527
+ `at_css` :: **1.4x**
528
+ SAX text-only handler :: **6x** (119 ms all-events vs Nokogiri's 131)
529
+ Memory held :: **17.5 MB/doc vs Nokogiri's 31.3** (1.8x lighter)
530
+
531
+ Two rows sit at the Ruby allocation floor by documented choice: the
532
+ cold full-tree walk (children recursion) remains ~1.5x behind
533
+ Nokogiri's C-extension node creation — `Node#visit` is the wrap-free
534
+ lever when that matters.
515
535
 
516
536
  == Development
517
537
 
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Leptris
4
- VERSION = "1.9.48"
4
+ VERSION = "1.9.50"
5
5
  end
@@ -251,10 +251,14 @@ class Leptris::XML::Element < Leptris::XML::Node
251
251
  self
252
252
  end
253
253
 
254
+ # Deep copy in a NEW document, via serialization round-trip:
255
+ # leptris_element_copy silently drops COMMENT and PI children
256
+ # (leptris/leptris#696, leptris-ruby#115) — the serializer
257
+ # preserves every child kind, prefixed names, and namespace
258
+ # declarations, and re-parsing rebuilds them all.
254
259
  def dup
255
- copy_ptr = Leptris::XML::FFI.leptris_element_copy(@c_ptr, @document.c_ptr)
256
- raise Leptris::XML::Error, "leptris_element_copy failed" if copy_ptr.null?
257
- Leptris::XML::Node.wrap(copy_ptr, @document)
260
+ ensure_alive!
261
+ Leptris::XML::Document.parse(to_xml).root
258
262
  end
259
263
  alias_method :clone, :dup
260
264
 
@@ -369,13 +369,16 @@ class Leptris::XML::Node
369
369
  result
370
370
  end
371
371
 
372
+ # Deep copy in a NEW document. Rounds through serialization
373
+ # rather than leptris_element_copy: the C copy silently drops
374
+ # COMMENT and PI children (leptris/leptris#696, leptris-ruby#115)
375
+ # — the serializer preserves every child kind, prefixed names,
376
+ # and namespace declarations, and re-parsing rebuilds them all.
372
377
  def dup
373
378
  ensure_alive!
374
379
  elem_ptr = Leptris::XML::FFI.leptris_node_as_element(@c_ptr)
375
380
  raise Leptris::XML::Error, "dup is only supported for element nodes" if elem_ptr.null?
376
- copy_ptr = Leptris::XML::FFI.leptris_element_copy(elem_ptr, @document.c_ptr)
377
- raise Leptris::XML::Error, "leptris_element_copy failed" if copy_ptr.null?
378
- Leptris::XML::Node.wrap(copy_ptr, @document)
381
+ Leptris::XML::Document.parse(to_xml).root
379
382
  end
380
383
  alias_method :clone, :dup
381
384
 
@@ -97,19 +97,15 @@ module Leptris::XML::Serialization
97
97
 
98
98
  # Element-face indent unit (leptris-ruby#109 residual 2): no
99
99
  # element-level ext-serialize entry exists yet, so the unit path
100
- # copies the element into a fresh document (C-side copy, one
101
- # pool allocation), serializes that document without a
102
- # declaration, and returns the subtree — identical output to the
103
- # element serializer for the standard layout, with the unit.
100
+ # rounds the subtree through serialization into a fresh document
101
+ # (leptris_element_copy drops COMMENT/PI children —
102
+ # leptris/leptris#696, leptris-ruby#115) and serializes that
103
+ # document without a declaration — identical output to the
104
+ # element serializer for the standard layout, with the unit and
105
+ # every child kind intact.
104
106
  def self.to_xml_element_unit(element, unit, indent: 0)
105
- document = Leptris::XML::Document.create
107
+ document = Leptris::XML::Document.parse(element.to_xml)
106
108
  begin
107
- copy_ptr = Leptris::XML::FFI.leptris_element_copy(
108
- element.c_ptr, document.c_ptr)
109
- if copy_ptr.null?
110
- raise Leptris::XML::Error, "leptris_element_copy failed"
111
- end
112
- document.root = Leptris::XML::Node.wrap(copy_ptr, document)
113
109
  to_xml_indent_unit(document.c_ptr, unit,
114
110
  indent: indent, no_decl: true)
115
111
  ensure
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: leptris
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.9.48
4
+ version: 1.9.50
5
5
  platform: aarch64-linux
6
6
  authors:
7
7
  - Ribose Inc.