wp2txt 2.3.3 → 2.3.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 051b5501ecf2aad35ca972af6527c5afbce88222c17e9c93dc76f806b91a8e9d
4
- data.tar.gz: 02efabb7102da40c144a3c8204f564a8a40dfa1e77a54813e40bd6740693a264
3
+ metadata.gz: 3972afe52897df60befbb96bed211e4fd05e33cce380ed46ced223637aed6954
4
+ data.tar.gz: 6674325567d9b21cfb8819110827a16327aaee2e7e0e509115953e3b97229b74
5
5
  SHA512:
6
- metadata.gz: d89ff3e231200c36f647284fee43565021c69c73ac8d9f23ca42b58008d2d0d3ac5598fe63aa3736a45ce933622dd28f649f721635a63e782d638f444602d08b
7
- data.tar.gz: c70fccdb332688f805e77dbefd291563045e2f53c07e7cc994dbff0a7f8858dded28f18bce4bcb1296e7594f12ee54eea12a4573beee04a8944d4e39b7d35620
6
+ metadata.gz: b7e36ac5cc6ab08699f3765768d9629b8cf193d1c134fcaeb9925228739b960bbc71c9ac5056c0bda5add9d1110cf993220b0a5a508e29caa9529eac84830278
7
+ data.tar.gz: abba849021598f7003a6fce9798517b2f497486853ce65522f687ea767eea1a1b91b87046c715f90d7908bff228cc441beec6cf4907c4885adeddd9967335165
data/CHANGELOG.md CHANGED
@@ -8,6 +8,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
8
8
  Entries were rewritten in August 2026 to describe what changed for people using
9
9
  wp2txt, rather than how it was implemented. The changes themselves are unaltered.
10
10
 
11
+ ## [2.3.4] - 2026-09-30
12
+
13
+ - **Records are no longer duplicated in `--no-turbo` output or in large `extract_corpus` runs**: the last record written before each batch could appear once more for every worker process — up to eight identical lines in a row. The default mode for `.bz2` dumps was not affected. If you extracted more than 200 articles with `extract_corpus`, or used `--no-turbo`, check earlier output for repeated lines; the repeats are byte-for-byte identical, so removing exact duplicate lines restores it
14
+ - **`--num-procs` is now honoured**: any value smaller than the number of CPU cores was silently replaced with the automatic choice. The value you give is used, and one outside 1 to the core count is adjusted with a warning
15
+ - **`--articles` no longer fails on large dumps**: extracting specific articles could stop with `Invalid argument` (seen on macOS) when the last requested article sat early in a multi-gigabyte dump
16
+ - **JSON records include `page_id` and `revision_id`**: the dump's own IDs for the page and its revision, right after `title`, in every mode — so each record can be traced to the exact version it came from. From Ruby, `StreamProcessor#each_page(with_ids: true)` yields them as a third value; plain `each_page` is unchanged
17
+ - **Corrupt input now stops processing instead of being cleaned silently**: invalid UTF-8 raises `Wp2txt::EncodingError` naming where it was found. Official dumps are valid UTF-8, so this only triggers on damaged files
18
+ - **Articles that fail to render are reported**: the default mode skipped them without a word; each one is now named on stderr with the reason
19
+
11
20
  ## [2.3.3] - 2026-09-08
12
21
 
13
22
  - **Short searches no longer report a false zero**: in Japanese, Chinese, and Korean indexes a phrase of one or two characters could match nothing and come back as `0 matches`, which reads exactly like a term that is genuinely absent from the dump. Such a search now fails with an explicit error instead. If you recorded a zero result for a short term with an earlier version, re-check it
data/README.md CHANGED
@@ -172,13 +172,16 @@ CATEGORIES: Category1, Category2, Category3
172
172
  Each line contains one JSON object:
173
173
 
174
174
  ```json
175
- {"title": "Article Title", "categories": ["Cat1", "Cat2"], "text": "...", "redirect": null}
175
+ {"title": "Article Title", "page_id": 12345, "revision_id": 67890, "categories": ["Cat1", "Cat2"], "text": "...", "redirect": null}
176
176
  ```
177
177
 
178
+ `page_id` and `revision_id` are the dump's own `<id>` values for the page and its
179
+ revision, so a record can be traced back to the exact version of the article it came from.
180
+
178
181
  For redirect articles:
179
182
 
180
183
  ```json
181
- {"title": "NYC", "categories": [], "text": "", "redirect": "New York City"}
184
+ {"title": "NYC", "page_id": 23456, "revision_id": 78901, "categories": [], "text": "", "redirect": "New York City"}
182
185
  ```
183
186
 
184
187
  ## Cache Management
data/README_ja.md CHANGED
@@ -156,13 +156,16 @@ CATEGORIES: カテゴリ1, カテゴリ2, カテゴリ3
156
156
  各行に1つのJSONオブジェクト:
157
157
 
158
158
  ```json
159
- {"title": "記事タイトル", "categories": ["カテゴリ1", "カテゴリ2"], "text": "...", "redirect": null}
159
+ {"title": "記事タイトル", "page_id": 12345, "revision_id": 67890, "categories": ["カテゴリ1", "カテゴリ2"], "text": "...", "redirect": null}
160
160
  ```
161
161
 
162
+ `page_id` と `revision_id` はダンプに記録されたページと版の `<id>` です。各レコードが
163
+ どの記事のどの版から取り出されたかを後から辿れます。
164
+
162
165
  リダイレクト記事の場合:
163
166
 
164
167
  ```json
165
- {"title": "NYC", "categories": [], "text": "", "redirect": "New York City"}
168
+ {"title": "NYC", "page_id": 23456, "revision_id": 78901, "categories": [], "text": "", "redirect": "New York City"}
166
169
  ```
167
170
 
168
171
  ## キャッシュ管理
data/bin/wp2txt CHANGED
@@ -45,14 +45,14 @@ class WpApp
45
45
  def calculate_num_processes(opts)
46
46
  optimal = Wp2txt::MemoryMonitor.optimal_processes
47
47
 
48
- if opts[:num_procs]
49
- # User specified a value - use it if reasonable
50
- requested = opts[:num_procs].to_i
51
- max_allowed = Etc.nprocessors
52
- [requested, max_allowed, 1].max == requested ? requested : optimal
53
- else
54
- optimal
55
- end.tap { |n| n = 1 if n < 1 }
48
+ return optimal unless opts[:num_procs]
49
+
50
+ requested = opts[:num_procs].to_i
51
+ used = requested.clamp(1, Etc.nprocessors)
52
+ if used != requested
53
+ print_warning("--num-procs #{requested} is outside 1..#{Etc.nprocessors}; using #{used}.")
54
+ end
55
+ used
56
56
  end
57
57
 
58
58
  # Process articles using turbo mode (split-first architecture from v1.x)
@@ -244,6 +244,9 @@ class WpApp
244
244
  next if redirect_page?(text)
245
245
 
246
246
  article = Article.new(text, title, strip_tmarker)
247
+ ids = Wp2txt.page_ids(page_xml)
248
+ article.page_id = ids[:page_id]
249
+ article.revision_id = ids[:revision_id]
247
250
  result = format_article(article, config)
248
251
  next unless result
249
252
 
@@ -254,7 +257,9 @@ class WpApp
254
257
  out.puts(result)
255
258
  end
256
259
  article_count += 1
257
- rescue StandardError
260
+ rescue StandardError => e
261
+ # Keep going, but never lose an article silently
262
+ warn "wp2txt: skipped #{title.inspect}: #{e.class}: #{e.message}"
258
263
  next
259
264
  end
260
265
  end
@@ -355,8 +360,8 @@ class WpApp
355
360
  end
356
361
  puts
357
362
 
358
- stream.each_page do |title, text|
359
- pages << [title, text]
363
+ stream.each_page(with_ids: true) do |title, text, ids|
364
+ pages << [title, text, ids]
360
365
  page_count += 1
361
366
 
362
367
  # Process batch when full
@@ -449,8 +454,13 @@ class WpApp
449
454
  )
450
455
  else
451
456
  # Fall back to Parallel gem (process-based parallelism)
452
- Parallel.map(pages, in_processes: num_processes) do |title, text|
457
+ # Anything the writer still buffers would be inherited by every
458
+ # forked worker and written again as each one exits.
459
+ writer.flush
460
+ Parallel.map(pages, in_processes: num_processes) do |title, text, ids|
453
461
  article = Article.new(text, title, strip_tmarker)
462
+ article.page_id = ids&.dig(:page_id)
463
+ article.revision_id = ids&.dig(:revision_id)
454
464
  format_article(article, config)
455
465
  end
456
466
  end
@@ -26,7 +26,7 @@ module Wp2txt
26
26
  # an article contains elements, each of which is [TYPE, string]
27
27
  class Article
28
28
  include Wp2txt
29
- attr_accessor :elements, :title, :categories
29
+ attr_accessor :elements, :title, :categories, :page_id, :revision_id
30
30
 
31
31
  def initialize(text, title = "", strip_tmarker = false)
32
32
  @title = title.strip
@@ -6,6 +6,17 @@ module Wp2txt
6
6
  (value || "0").to_i
7
7
  end
8
8
 
9
+ # Page and revision IDs from one <page> element of a dump. The page's own
10
+ # <id> precedes <revision>; the revision's <id> is its first child (the
11
+ # contributor's <id> comes later). Only the header before <text> is read.
12
+ def self.page_ids(page_xml)
13
+ head = page_xml[0, page_xml.index("<text") || page_xml.size]
14
+ {
15
+ page_id: head[%r{<page>.*?<id>(\d+)</id>}m, 1]&.to_i,
16
+ revision_id: head[%r{<revision>\s*<id>(\d+)</id>}, 1]&.to_i
17
+ }
18
+ end
19
+
9
20
  # =========================================================================
10
21
  # Custom Exception Classes
11
22
  # =========================================================================
data/lib/wp2txt/corpus.rb CHANGED
@@ -401,6 +401,8 @@ module Wp2txt
401
401
  titles.each_slice(EXTRACT_BATCH_SIZE) do |batch|
402
402
  raise Cancelled if cancel_check&.call
403
403
 
404
+ # Forked workers inherit this file's buffer and would write it again on exit
405
+ f.flush
404
406
  pages = reader.extract_articles_parallel(batch, num_processes: num_processes)
405
407
  batch.each do |t|
406
408
  page = pages[t]
@@ -111,6 +111,8 @@ module Wp2txt
111
111
 
112
112
  if page
113
113
  article = Article.new(page[:text], page[:title], !config[:marker])
114
+ article.page_id = page[:id]
115
+ article.revision_id = page[:revision_id]
114
116
  result = format_article(article, config)
115
117
  writer.write(result)
116
118
  extracted_count += 1
@@ -480,6 +482,8 @@ module Wp2txt
480
482
 
481
483
  pages.each do |page|
482
484
  article = Article.new(page[:text], page[:title], !config[:marker])
485
+ article.page_id = page[:id]
486
+ article.revision_id = page[:revision_id]
483
487
  result = format_article(article, config)
484
488
  writer.write(result)
485
489
  extracted_count += 1
@@ -493,6 +497,8 @@ module Wp2txt
493
497
 
494
498
  if page
495
499
  article = Article.new(page[:text], page[:title], !config[:marker])
500
+ article.page_id = page[:id]
501
+ article.revision_id = page[:revision_id]
496
502
  result = format_article(article, config)
497
503
  writer.write(result)
498
504
  extracted_count += 1
@@ -14,6 +14,24 @@ module Wp2txt
14
14
 
15
15
  # Format article based on configuration and output format
16
16
  def format_article(article, config)
17
+ with_page_ids(format_article_body(article, config), article)
18
+ end
19
+
20
+ # JSON records carry the dump's page and revision IDs right after the title,
21
+ # when the reader supplied them, so a record can be traced back to its source.
22
+ def with_page_ids(result, article)
23
+ return result unless result.is_a?(Hash) && (article.page_id || article.revision_id)
24
+
25
+ result.each_with_object({}) do |(key, value), out|
26
+ out[key] = value
27
+ next unless key == "title"
28
+
29
+ out["page_id"] = article.page_id
30
+ out["revision_id"] = article.revision_id
31
+ end
32
+ end
33
+
34
+ def format_article_body(article, config)
17
35
  # Store original title for magic word expansion in content
18
36
  original_title = article.title.dup
19
37
  article.title = format_wiki(article.title, config)
@@ -95,6 +95,11 @@ module Wp2txt
95
95
  @early_terminated == true
96
96
  end
97
97
 
98
+ # Where the stream holding the last found article ends. Scanning stops as
99
+ # soon as every target is found, so without this the reader cannot tell how
100
+ # far that stream extends and would read the dump to its end.
101
+ attr_reader :stream_end_offset
102
+
98
103
  def find_by_title(title)
99
104
  @entries_by_title[title]
100
105
  end
@@ -179,6 +184,14 @@ module Wp2txt
179
184
  end
180
185
  end
181
186
 
187
+ def next_stream_offset_after(io, offset)
188
+ io.each_line do |line|
189
+ next_offset = line.split(":", 2).first.to_i
190
+ return next_offset if next_offset > offset
191
+ end
192
+ nil
193
+ end
194
+
182
195
  def parse_index_stream(io)
183
196
  count = 0
184
197
  io.each_line do |line|
@@ -205,6 +218,7 @@ module Wp2txt
205
218
  @found_targets << title if @target_titles.include?(title)
206
219
  if @found_targets.size == @target_titles.size
207
220
  @early_terminated = true
221
+ @stream_end_offset = next_stream_offset_after(io, offset)
208
222
  print "\r Found all #{@target_titles.size} target articles" if @show_progress
209
223
  puts if @show_progress
210
224
  break
@@ -223,6 +237,10 @@ module Wp2txt
223
237
 
224
238
  # Reads articles from multistream bz2 files
225
239
  class MultistreamReader
240
+ # A single multistream bz2 stream holds about 100 pages; anything this large
241
+ # past the last known offset is several streams, not one.
242
+ MAX_TAIL_STREAM_BYTES = 64 * 1024 * 1024
243
+
226
244
  attr_reader :multistream_path, :index
227
245
 
228
246
  # Initialize reader with multistream file and index
@@ -357,7 +375,14 @@ module Wp2txt
357
375
  if next_offset
358
376
  compressed_data = f.read(next_offset - offset)
359
377
  else
360
- # Last stream - read to end
378
+ # Last stream of the dump: the rest of the file is that one stream.
379
+ # A large remainder means the end was never recorded, and reading it
380
+ # in one call fails outright on some platforms (EINVAL on macOS).
381
+ remaining = File.size(@multistream_path) - offset
382
+ if remaining > MAX_TAIL_STREAM_BYTES
383
+ raise Wp2txt::Error, "cannot locate the end of the stream at offset #{offset} " \
384
+ "(#{remaining} bytes to end of file); the index may be incomplete"
385
+ end
361
386
  compressed_data = f.read
362
387
  end
363
388
 
@@ -372,7 +397,9 @@ module Wp2txt
372
397
  # here would silently read gigabytes to EOF, so fail fast instead
373
398
  raise "Stream offset #{current_offset} not found in index (#{offsets.size} streams known)" unless idx
374
399
 
375
- offsets[idx + 1]
400
+ return offsets[idx + 1] if idx + 1 < offsets.size
401
+
402
+ @index.respond_to?(:stream_end_offset) ? @index.stream_end_offset : nil
376
403
  end
377
404
 
378
405
  def decompress_bz2(data)
@@ -394,6 +421,7 @@ module Wp2txt
394
421
  return {
395
422
  title: page_title,
396
423
  id: page_node.at_xpath("id")&.text&.to_i,
424
+ revision_id: page_node.at_xpath("revision/id")&.text&.to_i,
397
425
  text: page_node.at_xpath(".//text")&.text || ""
398
426
  }
399
427
  end
@@ -409,6 +437,7 @@ module Wp2txt
409
437
  page = {
410
438
  title: page_node.at_xpath("title")&.text,
411
439
  id: page_node.at_xpath("id")&.text&.to_i,
440
+ revision_id: page_node.at_xpath("revision/id")&.text&.to_i,
412
441
  text: page_node.at_xpath(".//text")&.text || ""
413
442
  }
414
443
  yield page if page[:title]
@@ -100,6 +100,14 @@ module Wp2txt
100
100
  raise Wp2txt::FileIOError, "Write failed: #{e.message}"
101
101
  end
102
102
 
103
+ # Push buffered output to the file. Call before forking: a child process
104
+ # inherits the buffer and writes it out again when it exits.
105
+ def flush
106
+ @mutex.synchronize do
107
+ @current_file.flush if @current_file && !@current_file.closed?
108
+ end
109
+ end
110
+
103
111
  # Close current file and finalize
104
112
  def close
105
113
  @mutex.synchronize do
@@ -64,20 +64,28 @@ module Wp2txt
64
64
 
65
65
  # Iterate over each page in the input
66
66
  # Yields [title, text] for each page
67
- def each_page
68
- return enum_for(:each_page) unless block_given?
67
+ # Yields title and text of each article. With with_ids: true, also yields
68
+ # { page_id:, revision_id: } taken from the dump.
69
+ def each_page(with_ids: false, &block)
70
+ return enum_for(:each_page, with_ids: with_ids) unless block
71
+
72
+ emit = if with_ids
73
+ ->(title, text, ids) { block.call(title, text, ids) }
74
+ else
75
+ ->(title, text, _ids) { block.call(title, text) }
76
+ end
69
77
 
70
78
  if File.directory?(@input_path)
71
79
  # Process XML files in directory
72
80
  Dir.glob(File.join(@input_path, "*.xml")).sort.each do |xml_file|
73
- process_xml_file(xml_file) { |title, text| yield title, text }
81
+ process_xml_file(xml_file) { |title, text, ids| emit.call(title, text, ids) }
74
82
  end
75
83
  elsif @input_path.end_with?(".bz2")
76
84
  # Process bz2 compressed file with streaming
77
- process_bz2_stream { |title, text| yield title, text }
85
+ process_bz2_stream { |title, text, ids| emit.call(title, text, ids) }
78
86
  elsif @input_path.end_with?(".xml")
79
87
  # Process single XML file
80
- process_xml_file(@input_path) { |title, text| yield title, text }
88
+ process_xml_file(@input_path) { |title, text, ids| emit.call(title, text, ids) }
81
89
  else
82
90
  raise ArgumentError, "Unsupported input format: #{@input_path}"
83
91
  end
@@ -162,20 +170,32 @@ module Wp2txt
162
170
  def fill_buffer
163
171
  chunk = @file_pointer.read(@buffer_size)
164
172
  unless chunk
165
- @buffer << @pending_bytes.to_s.dup.force_encoding(Encoding::UTF_8).scrub("")
166
- @pending_bytes = +"".b
173
+ unless @pending_bytes.to_s.empty?
174
+ raise Wp2txt::EncodingError,
175
+ "input ends in the middle of a UTF-8 character (byte #{@bytes_read}): #{@input_path}"
176
+ end
167
177
  return false
168
178
  end
169
179
 
170
- @bytes_read += chunk.bytesize
171
- bytes = @pending_bytes.to_s.b + chunk.b
172
- # Retain a trailing UTF-8 sequence until the next read. Only complete
173
- # chunks are scrubbed, so valid characters split by read are preserved.
180
+ carried = @pending_bytes.to_s.b
181
+ start = @bytes_read - carried.bytesize # stream position of the first carried byte
182
+ bytes = carried + chunk.b
183
+ # A read can end partway through a multi-byte character; carry those
184
+ # bytes into the next read rather than judging half a character.
174
185
  tail = bytes[/[\xC2-\xF4][\x80-\xBF]{0,2}\z/n]
175
186
  width = tail && (tail.getbyte(0) < 0xE0 ? 2 : tail.getbyte(0) < 0xF0 ? 3 : 4)
176
187
  @pending_bytes = tail && tail.bytesize < width ? tail : +"".b
177
- bytes = bytes.byteslice(0, bytes.bytesize - @pending_bytes.bytesize)
178
- @buffer << bytes.force_encoding(Encoding::UTF_8).scrub("")
188
+ bytes = bytes.byteslice(0, bytes.bytesize - @pending_bytes.bytesize).force_encoding(Encoding::UTF_8)
189
+ # Anything still invalid is corrupt input. Dropping it would change the
190
+ # text without a trace, so stop and say where it is.
191
+ unless bytes.valid_encoding?
192
+ bad = bytes.each_char.find_index { |c| !c.valid_encoding? }.to_i
193
+ offset = start + bytes[0, bad].bytesize
194
+ raise Wp2txt::EncodingError,
195
+ "invalid UTF-8 near decompressed byte #{offset}: #{@input_path}"
196
+ end
197
+ @bytes_read += chunk.bytesize
198
+ @buffer << bytes
179
199
 
180
200
  # Adaptive buffer adjustment: if memory is low, reduce buffer size
181
201
  if @adaptive_buffer && MemoryMonitor.memory_low?
@@ -252,7 +272,7 @@ module Wp2txt
252
272
  end
253
273
 
254
274
  @pages_processed += 1
255
- [title, text]
275
+ [title, text, Wp2txt.page_ids(page_xml)]
256
276
  rescue Nokogiri::XML::SyntaxError
257
277
  # Skip malformed XML
258
278
  nil
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Wp2txt
4
- VERSION = "2.3.3"
4
+ VERSION = "2.3.4"
5
5
  end
data/lib/wp2txt.rb CHANGED
@@ -269,12 +269,14 @@ module Wp2txt
269
269
  end
270
270
  page << line if inside_page
271
271
  end
272
- if page.empty?
273
- false
274
- else
275
- page.force_encoding("utf-8")
272
+ return false if page.empty?
273
+
274
+ page.force_encoding("utf-8")
275
+ # Corrupt input must not be cleaned into different text without a trace
276
+ unless page.valid_encoding?
277
+ title = page[%r{<title>([^<]*)</title>}, 1]&.scrub("?")
278
+ raise Wp2txt::EncodingError, "invalid UTF-8 in page #{title.inspect} of #{@input_file}"
276
279
  end
277
- rescue ::Encoding::InvalidByteSequenceError, ::Encoding::UndefinedConversionError
278
280
  page
279
281
  end
280
282
 
@@ -0,0 +1,146 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "spec_helper"
4
+ require "open3"
5
+ require "json"
6
+ require "tmpdir"
7
+ require "parallel"
8
+ require_relative "support/multistream_fixture"
9
+ require "wp2txt/multistream"
10
+ require "wp2txt/output_writer"
11
+ require "wp2txt/stream_processor"
12
+
13
+ RSpec.describe "output integrity" do
14
+ include MultistreamFixture
15
+
16
+ around do |example|
17
+ Dir.mktmpdir("wp2txt-integrity-") do |dir|
18
+ @dir = dir
19
+ example.run
20
+ end
21
+ end
22
+
23
+ let(:cli) { File.expand_path("../bin/wp2txt", __dir__) }
24
+ let(:lib) { File.expand_path("../lib", __dir__) }
25
+
26
+ def run_cli(*args)
27
+ Open3.capture3(RbConfig.ruby, "-I", lib, cli, *args)
28
+ end
29
+
30
+ def json_records(dir)
31
+ Dir[File.join(dir, "*")].flat_map { |f| File.readlines(f) }.map { |line| JSON.parse(line) }
32
+ end
33
+
34
+ def write_pages(path, count)
35
+ body = (1..count).map { |i| page_xml(id: i, ns: 0, title: "記事#{i}", text: "本文#{i}。日本語の文。\n") }.join
36
+ File.write(path, "<mediawiki>\n#{body}</mediawiki>\n")
37
+ end
38
+
39
+ describe "forked workers" do
40
+ it "do not write the parent's buffered output again when they exit" do
41
+ writer = Wp2txt::OutputWriter.new(output_dir: @dir, base_name: "out", format: :text, file_size_mb: 10)
42
+ writer.write("one line still sitting in the buffer")
43
+ writer.flush
44
+ Parallel.map([1, 2, 3], in_processes: 3) { |x| x }
45
+ files = writer.close
46
+ expect(File.readlines(files.first).size).to eq(1)
47
+ end
48
+
49
+ it "leave every article exactly once in streaming output across batch boundaries" do
50
+ input = File.join(@dir, "pages.xml")
51
+ out = File.join(@dir, "out")
52
+ Dir.mkdir(out)
53
+ write_pages(input, 450) # batches of 200 with -n 2
54
+ _stdout, stderr, status = run_cli("-i", input, "-o", out, "--format", "json", "-n", "2")
55
+ expect(status.success?).to be(true), stderr
56
+
57
+ titles = json_records(out).map { |r| r["title"] }
58
+ expect(titles.size).to eq(450)
59
+ expect(titles.tally.select { |_, n| n > 1 }).to be_empty
60
+ end
61
+ end
62
+
63
+ describe "--num-procs" do
64
+ it "uses the requested process count" do
65
+ input = File.join(@dir, "pages.xml")
66
+ out = File.join(@dir, "out")
67
+ Dir.mkdir(out)
68
+ write_pages(input, 3)
69
+ stdout, stderr, status = run_cli("-i", input, "-o", out, "--format", "json", "-n", "1")
70
+ expect(status.success?).to be(true), stderr
71
+ expect(stdout.gsub(/\e\[[\d;]*m/, "")).to match(/CPU cores:\s*1\b/)
72
+ end
73
+ end
74
+
75
+ describe "page and revision IDs" do
76
+ it "appear in JSON records from the streaming path, including summaries" do
77
+ input = File.join(@dir, "pages.xml")
78
+ write_pages(input, 3)
79
+ [[], ["--summary-only"]].each do |extra|
80
+ out = File.join(@dir, "out#{extra.size}")
81
+ Dir.mkdir(out)
82
+ _stdout, stderr, status = run_cli("-i", input, "-o", out, "--format", "json", *extra)
83
+ expect(status.success?).to be(true), stderr
84
+ record = json_records(out).find { |r| r["title"] == "記事2" }
85
+ expect(record).to include("page_id" => 2, "revision_id" => 200)
86
+ expect(record.keys.first(3)).to eq(%w[title page_id revision_id])
87
+ end
88
+ end
89
+
90
+ it "appear in JSON records from the default path for bz2 dumps" do
91
+ dump, = create_fixture(@dir)
92
+ out = File.join(@dir, "out")
93
+ Dir.mkdir(out)
94
+ _stdout, stderr, status = run_cli("-i", dump, "-o", out, "--format", "json")
95
+ expect(status.success?).to be(true), stderr
96
+ record = json_records(out).find { |r| r["title"] == "Film B" }
97
+ expect(record).to include("page_id" => 2, "revision_id" => 200)
98
+ end
99
+
100
+ it "are read from the page header, not from the contributor" do
101
+ xml = "<page><title>X</title><ns>0</ns><id>12</id><revision><id>34</id>" \
102
+ "<contributor><id>9</id></contributor><text>t</text></revision></page>"
103
+ expect(Wp2txt.page_ids(xml)).to eq(page_id: 12, revision_id: 34)
104
+ end
105
+
106
+ it "are offered by StreamProcessor only on request, keeping each_page's shape" do
107
+ input = File.join(@dir, "pages.xml")
108
+ write_pages(input, 1)
109
+ processor = Wp2txt::StreamProcessor.new(input, adaptive_buffer: false)
110
+ expect(processor.each_page.to_a).to eq([["記事1", "本文1。日本語の文。\n"]])
111
+ with_ids = Wp2txt::StreamProcessor.new(input, adaptive_buffer: false).each_page(with_ids: true).to_a
112
+ expect(with_ids.first.last).to eq(page_id: 1, revision_id: 100)
113
+ end
114
+ end
115
+
116
+ describe "targeted extraction after an early index stop" do
117
+ it "knows where the last found article's stream ends" do
118
+ _dump, index_path = create_fixture(@dir)
119
+ index = Wp2txt::MultistreamIndex.new(index_path, use_cache: false, target_titles: ["Film A"],
120
+ show_progress: false)
121
+ expect(index.early_terminated?).to be(true)
122
+ expect(index.stream_offsets).to eq([0])
123
+ second_stream = File.readlines(index_path).map { |line| line.split(":").first.to_i }.uniq[1]
124
+ expect(index.stream_end_offset).to eq(second_stream)
125
+ end
126
+
127
+ it "reads only that stream" do
128
+ dump, index_path = create_fixture(@dir)
129
+ index = Wp2txt::MultistreamIndex.new(index_path, use_cache: false, target_titles: ["Film A"],
130
+ show_progress: false)
131
+ reader = Wp2txt::MultistreamReader.new(dump, index)
132
+ stub_const("Wp2txt::MultistreamReader::MAX_TAIL_STREAM_BYTES", 0)
133
+ expect(reader.extract_article("Film A")).to include(title: "Film A", id: 1, revision_id: 100)
134
+ end
135
+
136
+ it "refuses to read a large remainder in one call when the end is unknown" do
137
+ dump, index_path = create_fixture(@dir)
138
+ index = Wp2txt::MultistreamIndex.new(index_path, use_cache: false, target_titles: ["Film A"],
139
+ show_progress: false)
140
+ index.instance_variable_set(:@stream_end_offset, nil)
141
+ reader = Wp2txt::MultistreamReader.new(dump, index)
142
+ stub_const("Wp2txt::MultistreamReader::MAX_TAIL_STREAM_BYTES", 0)
143
+ expect { reader.extract_article("Film A") }.to raise_error(Wp2txt::Error, /cannot locate the end/)
144
+ end
145
+ end
146
+ end
@@ -40,10 +40,25 @@ RSpec.describe "P1 correctness contracts" do
40
40
  end
41
41
  end
42
42
 
43
- it "scrubs invalid bytes and incomplete EOF without losing adjacent valid text" do
44
- processor = processor_for("é\xFFあ😀\xE3\x81".b, 1)
45
- nil while processor.send(:fill_buffer)
46
- expect(processor.instance_variable_get(:@buffer)).to eq("éあ😀")
43
+ it "stops on an invalid byte and reports where it is" do
44
+ processor = processor_for("éあ\xFF😀".b, 1)
45
+ expect { nil while processor.send(:fill_buffer) }
46
+ .to raise_error(Wp2txt::EncodingError, /byte 5\b/)
47
+ end
48
+
49
+ it "stops when the input ends partway through a character" do
50
+ processor = processor_for("éあ\xE3\x81".b, 1)
51
+ expect { nil while processor.send(:fill_buffer) }
52
+ .to raise_error(Wp2txt::EncodingError, /middle of a UTF-8 character/)
53
+ end
54
+
55
+ it "passes valid text through unchanged at every buffer size" do
56
+ text = "é\nあ😀終\n" * 7
57
+ [1, 2, 3, 5, 64].each do |size|
58
+ processor = processor_for(text.b, size)
59
+ nil while processor.send(:fill_buffer)
60
+ expect(processor.instance_variable_get(:@buffer)).to eq(text)
61
+ end
47
62
  end
48
63
 
49
64
  it "handles empty input" do
@@ -105,7 +120,7 @@ RSpec.describe "P1 correctness contracts" do
105
120
  it "retains the historical ns=0 default for missing ns in both parsers" do
106
121
  xml = page_xml(id: 1, ns: 0, title: "作品: 東京", text: "本文").sub(/<ns>.*?<\/ns>/, "")
107
122
  processor = Wp2txt::StreamProcessor.new("unused.xml", adaptive_buffer: false)
108
- expect(processor.send(:parse_page_xml, xml)).to eq(["作品: 東京", "本文"])
123
+ expect(processor.send(:parse_page_xml, xml)).to eq(["作品: 東京", "本文", { page_id: 1, revision_id: 100 }])
109
124
  rows = { pages: [], categories: [], sections: [], hierarchy: [] }
110
125
  Wp2txt::MetadataIndexBuilder.scan_page(xml, rows)
111
126
  expect(rows[:pages].first[2]).to eq(0)
@@ -36,6 +36,18 @@ RSpec.describe "titles extraction and SQL file output" do
36
36
  File.readlines(path).map { |l| JSON.parse(l) }
37
37
  end
38
38
 
39
+ describe "extract_corpus across several batches with worker processes" do
40
+ it "writes each record once (workers must not re-flush the output buffer)" do
41
+ stub_const("Wp2txt::Corpus::EXTRACT_BATCH_SIZE", 1)
42
+ out = File.join(@dir, "batched.jsonl")
43
+ titles = ["Film A", "Film B", "Person X"]
44
+ result = @corpus.extract_corpus(output_path: out, content: "full", titles: titles, num_processes: 2)
45
+ records = read_jsonl(out)
46
+ expect(records.map { |r| r["title"] }.tally).to eq(titles.to_h { |t| [t, 1] })
47
+ expect(records.size).to eq(result[:records_written])
48
+ end
49
+ end
50
+
39
51
  describe "extract_corpus titles:" do
40
52
  it "extracts an explicit set with normalization, dedup, and input order" do
41
53
  out = File.join(@dir, "t.jsonl")
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: wp2txt
3
3
  version: !ruby/object:Gem::Version
4
- version: 2.3.3
4
+ version: 2.3.4
5
5
  platform: ruby
6
6
  authors:
7
7
  - Yoichiro Hasebe
@@ -284,6 +284,7 @@ files:
284
284
  - spec/metadata_index_spec.rb
285
285
  - spec/multi_dump_attach_spec.rb
286
286
  - spec/multistream_spec.rb
287
+ - spec/output_integrity_spec.rb
287
288
  - spec/output_writer_spec.rb
288
289
  - spec/p1_correctness_spec.rb
289
290
  - spec/parser_functions_spec.rb
@@ -351,6 +352,7 @@ test_files:
351
352
  - spec/metadata_index_spec.rb
352
353
  - spec/multi_dump_attach_spec.rb
353
354
  - spec/multistream_spec.rb
355
+ - spec/output_integrity_spec.rb
354
356
  - spec/output_writer_spec.rb
355
357
  - spec/p1_correctness_spec.rb
356
358
  - spec/parser_functions_spec.rb