wp2txt 2.3.3 → 2.3.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +9 -0
- data/README.md +5 -2
- data/README_ja.md +5 -2
- data/bin/wp2txt +22 -12
- data/lib/wp2txt/article.rb +1 -1
- data/lib/wp2txt/constants.rb +11 -0
- data/lib/wp2txt/corpus.rb +2 -0
- data/lib/wp2txt/extractor.rb +6 -0
- data/lib/wp2txt/formatter.rb +18 -0
- data/lib/wp2txt/multistream.rb +31 -2
- data/lib/wp2txt/output_writer.rb +8 -0
- data/lib/wp2txt/stream_processor.rb +34 -14
- data/lib/wp2txt/version.rb +1 -1
- data/lib/wp2txt.rb +7 -5
- data/spec/output_integrity_spec.rb +146 -0
- data/spec/p1_correctness_spec.rb +20 -5
- data/spec/titles_output_path_spec.rb +12 -0
- metadata +3 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 3972afe52897df60befbb96bed211e4fd05e33cce380ed46ced223637aed6954
|
|
4
|
+
data.tar.gz: 6674325567d9b21cfb8819110827a16327aaee2e7e0e509115953e3b97229b74
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: b7e36ac5cc6ab08699f3765768d9629b8cf193d1c134fcaeb9925228739b960bbc71c9ac5056c0bda5add9d1110cf993220b0a5a508e29caa9529eac84830278
|
|
7
|
+
data.tar.gz: abba849021598f7003a6fce9798517b2f497486853ce65522f687ea767eea1a1b91b87046c715f90d7908bff228cc441beec6cf4907c4885adeddd9967335165
|
data/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
8
8
|
Entries were rewritten in August 2026 to describe what changed for people using
|
|
9
9
|
wp2txt, rather than how it was implemented. The changes themselves are unaltered.
|
|
10
10
|
|
|
11
|
+
## [2.3.4] - 2026-09-30
|
|
12
|
+
|
|
13
|
+
- **Records are no longer duplicated in `--no-turbo` output or in large `extract_corpus` runs**: the last record written before each batch could appear once more for every worker process — up to eight identical lines in a row. The default mode for `.bz2` dumps was not affected. If you extracted more than 200 articles with `extract_corpus`, or used `--no-turbo`, check earlier output for repeated lines; the repeats are byte-for-byte identical, so removing exact duplicate lines restores it
|
|
14
|
+
- **`--num-procs` is now honoured**: any value smaller than the number of CPU cores was silently replaced with the automatic choice. The value you give is used, and one outside 1 to the core count is adjusted with a warning
|
|
15
|
+
- **`--articles` no longer fails on large dumps**: extracting specific articles could stop with `Invalid argument` (seen on macOS) when the last requested article sat early in a multi-gigabyte dump
|
|
16
|
+
- **JSON records include `page_id` and `revision_id`**: the dump's own IDs for the page and its revision, right after `title`, in every mode — so each record can be traced to the exact version it came from. From Ruby, `StreamProcessor#each_page(with_ids: true)` yields them as a third value; plain `each_page` is unchanged
|
|
17
|
+
- **Corrupt input now stops processing instead of being cleaned silently**: invalid UTF-8 raises `Wp2txt::EncodingError` naming where it was found. Official dumps are valid UTF-8, so this only triggers on damaged files
|
|
18
|
+
- **Articles that fail to render are reported**: the default mode skipped them without a word; each one is now named on stderr with the reason
|
|
19
|
+
|
|
11
20
|
## [2.3.3] - 2026-09-08
|
|
12
21
|
|
|
13
22
|
- **Short searches no longer report a false zero**: in Japanese, Chinese, and Korean indexes a phrase of one or two characters could match nothing and come back as `0 matches`, which reads exactly like a term that is genuinely absent from the dump. Such a search now fails with an explicit error instead. If you recorded a zero result for a short term with an earlier version, re-check it
|
data/README.md
CHANGED
|
@@ -172,13 +172,16 @@ CATEGORIES: Category1, Category2, Category3
|
|
|
172
172
|
Each line contains one JSON object:
|
|
173
173
|
|
|
174
174
|
```json
|
|
175
|
-
{"title": "Article Title", "categories": ["Cat1", "Cat2"], "text": "...", "redirect": null}
|
|
175
|
+
{"title": "Article Title", "page_id": 12345, "revision_id": 67890, "categories": ["Cat1", "Cat2"], "text": "...", "redirect": null}
|
|
176
176
|
```
|
|
177
177
|
|
|
178
|
+
`page_id` and `revision_id` are the dump's own `<id>` values for the page and its
|
|
179
|
+
revision, so a record can be traced back to the exact version of the article it came from.
|
|
180
|
+
|
|
178
181
|
For redirect articles:
|
|
179
182
|
|
|
180
183
|
```json
|
|
181
|
-
{"title": "NYC", "categories": [], "text": "", "redirect": "New York City"}
|
|
184
|
+
{"title": "NYC", "page_id": 23456, "revision_id": 78901, "categories": [], "text": "", "redirect": "New York City"}
|
|
182
185
|
```
|
|
183
186
|
|
|
184
187
|
## Cache Management
|
data/README_ja.md
CHANGED
|
@@ -156,13 +156,16 @@ CATEGORIES: カテゴリ1, カテゴリ2, カテゴリ3
|
|
|
156
156
|
各行に1つのJSONオブジェクト:
|
|
157
157
|
|
|
158
158
|
```json
|
|
159
|
-
{"title": "記事タイトル", "categories": ["カテゴリ1", "カテゴリ2"], "text": "...", "redirect": null}
|
|
159
|
+
{"title": "記事タイトル", "page_id": 12345, "revision_id": 67890, "categories": ["カテゴリ1", "カテゴリ2"], "text": "...", "redirect": null}
|
|
160
160
|
```
|
|
161
161
|
|
|
162
|
+
`page_id` と `revision_id` はダンプに記録されたページと版の `<id>` です。各レコードが
|
|
163
|
+
どの記事のどの版から取り出されたかを後から辿れます。
|
|
164
|
+
|
|
162
165
|
リダイレクト記事の場合:
|
|
163
166
|
|
|
164
167
|
```json
|
|
165
|
-
{"title": "NYC", "categories": [], "text": "", "redirect": "New York City"}
|
|
168
|
+
{"title": "NYC", "page_id": 23456, "revision_id": 78901, "categories": [], "text": "", "redirect": "New York City"}
|
|
166
169
|
```
|
|
167
170
|
|
|
168
171
|
## キャッシュ管理
|
data/bin/wp2txt
CHANGED
|
@@ -45,14 +45,14 @@ class WpApp
|
|
|
45
45
|
def calculate_num_processes(opts)
|
|
46
46
|
optimal = Wp2txt::MemoryMonitor.optimal_processes
|
|
47
47
|
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
48
|
+
return optimal unless opts[:num_procs]
|
|
49
|
+
|
|
50
|
+
requested = opts[:num_procs].to_i
|
|
51
|
+
used = requested.clamp(1, Etc.nprocessors)
|
|
52
|
+
if used != requested
|
|
53
|
+
print_warning("--num-procs #{requested} is outside 1..#{Etc.nprocessors}; using #{used}.")
|
|
54
|
+
end
|
|
55
|
+
used
|
|
56
56
|
end
|
|
57
57
|
|
|
58
58
|
# Process articles using turbo mode (split-first architecture from v1.x)
|
|
@@ -244,6 +244,9 @@ class WpApp
|
|
|
244
244
|
next if redirect_page?(text)
|
|
245
245
|
|
|
246
246
|
article = Article.new(text, title, strip_tmarker)
|
|
247
|
+
ids = Wp2txt.page_ids(page_xml)
|
|
248
|
+
article.page_id = ids[:page_id]
|
|
249
|
+
article.revision_id = ids[:revision_id]
|
|
247
250
|
result = format_article(article, config)
|
|
248
251
|
next unless result
|
|
249
252
|
|
|
@@ -254,7 +257,9 @@ class WpApp
|
|
|
254
257
|
out.puts(result)
|
|
255
258
|
end
|
|
256
259
|
article_count += 1
|
|
257
|
-
rescue StandardError
|
|
260
|
+
rescue StandardError => e
|
|
261
|
+
# Keep going, but never lose an article silently
|
|
262
|
+
warn "wp2txt: skipped #{title.inspect}: #{e.class}: #{e.message}"
|
|
258
263
|
next
|
|
259
264
|
end
|
|
260
265
|
end
|
|
@@ -355,8 +360,8 @@ class WpApp
|
|
|
355
360
|
end
|
|
356
361
|
puts
|
|
357
362
|
|
|
358
|
-
stream.each_page do |title, text|
|
|
359
|
-
pages << [title, text]
|
|
363
|
+
stream.each_page(with_ids: true) do |title, text, ids|
|
|
364
|
+
pages << [title, text, ids]
|
|
360
365
|
page_count += 1
|
|
361
366
|
|
|
362
367
|
# Process batch when full
|
|
@@ -449,8 +454,13 @@ class WpApp
|
|
|
449
454
|
)
|
|
450
455
|
else
|
|
451
456
|
# Fall back to Parallel gem (process-based parallelism)
|
|
452
|
-
|
|
457
|
+
# Anything the writer still buffers would be inherited by every
|
|
458
|
+
# forked worker and written again as each one exits.
|
|
459
|
+
writer.flush
|
|
460
|
+
Parallel.map(pages, in_processes: num_processes) do |title, text, ids|
|
|
453
461
|
article = Article.new(text, title, strip_tmarker)
|
|
462
|
+
article.page_id = ids&.dig(:page_id)
|
|
463
|
+
article.revision_id = ids&.dig(:revision_id)
|
|
454
464
|
format_article(article, config)
|
|
455
465
|
end
|
|
456
466
|
end
|
data/lib/wp2txt/article.rb
CHANGED
|
@@ -26,7 +26,7 @@ module Wp2txt
|
|
|
26
26
|
# an article contains elements, each of which is [TYPE, string]
|
|
27
27
|
class Article
|
|
28
28
|
include Wp2txt
|
|
29
|
-
attr_accessor :elements, :title, :categories
|
|
29
|
+
attr_accessor :elements, :title, :categories, :page_id, :revision_id
|
|
30
30
|
|
|
31
31
|
def initialize(text, title = "", strip_tmarker = false)
|
|
32
32
|
@title = title.strip
|
data/lib/wp2txt/constants.rb
CHANGED
|
@@ -6,6 +6,17 @@ module Wp2txt
|
|
|
6
6
|
(value || "0").to_i
|
|
7
7
|
end
|
|
8
8
|
|
|
9
|
+
# Page and revision IDs from one <page> element of a dump. The page's own
|
|
10
|
+
# <id> precedes <revision>; the revision's <id> is its first child (the
|
|
11
|
+
# contributor's <id> comes later). Only the header before <text> is read.
|
|
12
|
+
def self.page_ids(page_xml)
|
|
13
|
+
head = page_xml[0, page_xml.index("<text") || page_xml.size]
|
|
14
|
+
{
|
|
15
|
+
page_id: head[%r{<page>.*?<id>(\d+)</id>}m, 1]&.to_i,
|
|
16
|
+
revision_id: head[%r{<revision>\s*<id>(\d+)</id>}, 1]&.to_i
|
|
17
|
+
}
|
|
18
|
+
end
|
|
19
|
+
|
|
9
20
|
# =========================================================================
|
|
10
21
|
# Custom Exception Classes
|
|
11
22
|
# =========================================================================
|
data/lib/wp2txt/corpus.rb
CHANGED
|
@@ -401,6 +401,8 @@ module Wp2txt
|
|
|
401
401
|
titles.each_slice(EXTRACT_BATCH_SIZE) do |batch|
|
|
402
402
|
raise Cancelled if cancel_check&.call
|
|
403
403
|
|
|
404
|
+
# Forked workers inherit this file's buffer and would write it again on exit
|
|
405
|
+
f.flush
|
|
404
406
|
pages = reader.extract_articles_parallel(batch, num_processes: num_processes)
|
|
405
407
|
batch.each do |t|
|
|
406
408
|
page = pages[t]
|
data/lib/wp2txt/extractor.rb
CHANGED
|
@@ -111,6 +111,8 @@ module Wp2txt
|
|
|
111
111
|
|
|
112
112
|
if page
|
|
113
113
|
article = Article.new(page[:text], page[:title], !config[:marker])
|
|
114
|
+
article.page_id = page[:id]
|
|
115
|
+
article.revision_id = page[:revision_id]
|
|
114
116
|
result = format_article(article, config)
|
|
115
117
|
writer.write(result)
|
|
116
118
|
extracted_count += 1
|
|
@@ -480,6 +482,8 @@ module Wp2txt
|
|
|
480
482
|
|
|
481
483
|
pages.each do |page|
|
|
482
484
|
article = Article.new(page[:text], page[:title], !config[:marker])
|
|
485
|
+
article.page_id = page[:id]
|
|
486
|
+
article.revision_id = page[:revision_id]
|
|
483
487
|
result = format_article(article, config)
|
|
484
488
|
writer.write(result)
|
|
485
489
|
extracted_count += 1
|
|
@@ -493,6 +497,8 @@ module Wp2txt
|
|
|
493
497
|
|
|
494
498
|
if page
|
|
495
499
|
article = Article.new(page[:text], page[:title], !config[:marker])
|
|
500
|
+
article.page_id = page[:id]
|
|
501
|
+
article.revision_id = page[:revision_id]
|
|
496
502
|
result = format_article(article, config)
|
|
497
503
|
writer.write(result)
|
|
498
504
|
extracted_count += 1
|
data/lib/wp2txt/formatter.rb
CHANGED
|
@@ -14,6 +14,24 @@ module Wp2txt
|
|
|
14
14
|
|
|
15
15
|
# Format article based on configuration and output format
|
|
16
16
|
def format_article(article, config)
|
|
17
|
+
with_page_ids(format_article_body(article, config), article)
|
|
18
|
+
end
|
|
19
|
+
|
|
20
|
+
# JSON records carry the dump's page and revision IDs right after the title,
|
|
21
|
+
# when the reader supplied them, so a record can be traced back to its source.
|
|
22
|
+
def with_page_ids(result, article)
|
|
23
|
+
return result unless result.is_a?(Hash) && (article.page_id || article.revision_id)
|
|
24
|
+
|
|
25
|
+
result.each_with_object({}) do |(key, value), out|
|
|
26
|
+
out[key] = value
|
|
27
|
+
next unless key == "title"
|
|
28
|
+
|
|
29
|
+
out["page_id"] = article.page_id
|
|
30
|
+
out["revision_id"] = article.revision_id
|
|
31
|
+
end
|
|
32
|
+
end
|
|
33
|
+
|
|
34
|
+
def format_article_body(article, config)
|
|
17
35
|
# Store original title for magic word expansion in content
|
|
18
36
|
original_title = article.title.dup
|
|
19
37
|
article.title = format_wiki(article.title, config)
|
data/lib/wp2txt/multistream.rb
CHANGED
|
@@ -95,6 +95,11 @@ module Wp2txt
|
|
|
95
95
|
@early_terminated == true
|
|
96
96
|
end
|
|
97
97
|
|
|
98
|
+
# Where the stream holding the last found article ends. Scanning stops as
|
|
99
|
+
# soon as every target is found, so without this the reader cannot tell how
|
|
100
|
+
# far that stream extends and would read the dump to its end.
|
|
101
|
+
attr_reader :stream_end_offset
|
|
102
|
+
|
|
98
103
|
def find_by_title(title)
|
|
99
104
|
@entries_by_title[title]
|
|
100
105
|
end
|
|
@@ -179,6 +184,14 @@ module Wp2txt
|
|
|
179
184
|
end
|
|
180
185
|
end
|
|
181
186
|
|
|
187
|
+
def next_stream_offset_after(io, offset)
|
|
188
|
+
io.each_line do |line|
|
|
189
|
+
next_offset = line.split(":", 2).first.to_i
|
|
190
|
+
return next_offset if next_offset > offset
|
|
191
|
+
end
|
|
192
|
+
nil
|
|
193
|
+
end
|
|
194
|
+
|
|
182
195
|
def parse_index_stream(io)
|
|
183
196
|
count = 0
|
|
184
197
|
io.each_line do |line|
|
|
@@ -205,6 +218,7 @@ module Wp2txt
|
|
|
205
218
|
@found_targets << title if @target_titles.include?(title)
|
|
206
219
|
if @found_targets.size == @target_titles.size
|
|
207
220
|
@early_terminated = true
|
|
221
|
+
@stream_end_offset = next_stream_offset_after(io, offset)
|
|
208
222
|
print "\r Found all #{@target_titles.size} target articles" if @show_progress
|
|
209
223
|
puts if @show_progress
|
|
210
224
|
break
|
|
@@ -223,6 +237,10 @@ module Wp2txt
|
|
|
223
237
|
|
|
224
238
|
# Reads articles from multistream bz2 files
|
|
225
239
|
class MultistreamReader
|
|
240
|
+
# A single multistream bz2 stream holds about 100 pages; anything this large
|
|
241
|
+
# past the last known offset is several streams, not one.
|
|
242
|
+
MAX_TAIL_STREAM_BYTES = 64 * 1024 * 1024
|
|
243
|
+
|
|
226
244
|
attr_reader :multistream_path, :index
|
|
227
245
|
|
|
228
246
|
# Initialize reader with multistream file and index
|
|
@@ -357,7 +375,14 @@ module Wp2txt
|
|
|
357
375
|
if next_offset
|
|
358
376
|
compressed_data = f.read(next_offset - offset)
|
|
359
377
|
else
|
|
360
|
-
# Last stream
|
|
378
|
+
# Last stream of the dump: the rest of the file is that one stream.
|
|
379
|
+
# A large remainder means the end was never recorded, and reading it
|
|
380
|
+
# in one call fails outright on some platforms (EINVAL on macOS).
|
|
381
|
+
remaining = File.size(@multistream_path) - offset
|
|
382
|
+
if remaining > MAX_TAIL_STREAM_BYTES
|
|
383
|
+
raise Wp2txt::Error, "cannot locate the end of the stream at offset #{offset} " \
|
|
384
|
+
"(#{remaining} bytes to end of file); the index may be incomplete"
|
|
385
|
+
end
|
|
361
386
|
compressed_data = f.read
|
|
362
387
|
end
|
|
363
388
|
|
|
@@ -372,7 +397,9 @@ module Wp2txt
|
|
|
372
397
|
# here would silently read gigabytes to EOF, so fail fast instead
|
|
373
398
|
raise "Stream offset #{current_offset} not found in index (#{offsets.size} streams known)" unless idx
|
|
374
399
|
|
|
375
|
-
offsets[idx + 1]
|
|
400
|
+
return offsets[idx + 1] if idx + 1 < offsets.size
|
|
401
|
+
|
|
402
|
+
@index.respond_to?(:stream_end_offset) ? @index.stream_end_offset : nil
|
|
376
403
|
end
|
|
377
404
|
|
|
378
405
|
def decompress_bz2(data)
|
|
@@ -394,6 +421,7 @@ module Wp2txt
|
|
|
394
421
|
return {
|
|
395
422
|
title: page_title,
|
|
396
423
|
id: page_node.at_xpath("id")&.text&.to_i,
|
|
424
|
+
revision_id: page_node.at_xpath("revision/id")&.text&.to_i,
|
|
397
425
|
text: page_node.at_xpath(".//text")&.text || ""
|
|
398
426
|
}
|
|
399
427
|
end
|
|
@@ -409,6 +437,7 @@ module Wp2txt
|
|
|
409
437
|
page = {
|
|
410
438
|
title: page_node.at_xpath("title")&.text,
|
|
411
439
|
id: page_node.at_xpath("id")&.text&.to_i,
|
|
440
|
+
revision_id: page_node.at_xpath("revision/id")&.text&.to_i,
|
|
412
441
|
text: page_node.at_xpath(".//text")&.text || ""
|
|
413
442
|
}
|
|
414
443
|
yield page if page[:title]
|
data/lib/wp2txt/output_writer.rb
CHANGED
|
@@ -100,6 +100,14 @@ module Wp2txt
|
|
|
100
100
|
raise Wp2txt::FileIOError, "Write failed: #{e.message}"
|
|
101
101
|
end
|
|
102
102
|
|
|
103
|
+
# Push buffered output to the file. Call before forking: a child process
|
|
104
|
+
# inherits the buffer and writes it out again when it exits.
|
|
105
|
+
def flush
|
|
106
|
+
@mutex.synchronize do
|
|
107
|
+
@current_file.flush if @current_file && !@current_file.closed?
|
|
108
|
+
end
|
|
109
|
+
end
|
|
110
|
+
|
|
103
111
|
# Close current file and finalize
|
|
104
112
|
def close
|
|
105
113
|
@mutex.synchronize do
|
|
@@ -64,20 +64,28 @@ module Wp2txt
|
|
|
64
64
|
|
|
65
65
|
# Iterate over each page in the input
|
|
66
66
|
# Yields [title, text] for each page
|
|
67
|
-
|
|
68
|
-
|
|
67
|
+
# Yields title and text of each article. With with_ids: true, also yields
|
|
68
|
+
# { page_id:, revision_id: } taken from the dump.
|
|
69
|
+
def each_page(with_ids: false, &block)
|
|
70
|
+
return enum_for(:each_page, with_ids: with_ids) unless block
|
|
71
|
+
|
|
72
|
+
emit = if with_ids
|
|
73
|
+
->(title, text, ids) { block.call(title, text, ids) }
|
|
74
|
+
else
|
|
75
|
+
->(title, text, _ids) { block.call(title, text) }
|
|
76
|
+
end
|
|
69
77
|
|
|
70
78
|
if File.directory?(@input_path)
|
|
71
79
|
# Process XML files in directory
|
|
72
80
|
Dir.glob(File.join(@input_path, "*.xml")).sort.each do |xml_file|
|
|
73
|
-
process_xml_file(xml_file) { |title, text|
|
|
81
|
+
process_xml_file(xml_file) { |title, text, ids| emit.call(title, text, ids) }
|
|
74
82
|
end
|
|
75
83
|
elsif @input_path.end_with?(".bz2")
|
|
76
84
|
# Process bz2 compressed file with streaming
|
|
77
|
-
process_bz2_stream { |title, text|
|
|
85
|
+
process_bz2_stream { |title, text, ids| emit.call(title, text, ids) }
|
|
78
86
|
elsif @input_path.end_with?(".xml")
|
|
79
87
|
# Process single XML file
|
|
80
|
-
process_xml_file(@input_path) { |title, text|
|
|
88
|
+
process_xml_file(@input_path) { |title, text, ids| emit.call(title, text, ids) }
|
|
81
89
|
else
|
|
82
90
|
raise ArgumentError, "Unsupported input format: #{@input_path}"
|
|
83
91
|
end
|
|
@@ -162,20 +170,32 @@ module Wp2txt
|
|
|
162
170
|
def fill_buffer
|
|
163
171
|
chunk = @file_pointer.read(@buffer_size)
|
|
164
172
|
unless chunk
|
|
165
|
-
|
|
166
|
-
|
|
173
|
+
unless @pending_bytes.to_s.empty?
|
|
174
|
+
raise Wp2txt::EncodingError,
|
|
175
|
+
"input ends in the middle of a UTF-8 character (byte #{@bytes_read}): #{@input_path}"
|
|
176
|
+
end
|
|
167
177
|
return false
|
|
168
178
|
end
|
|
169
179
|
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
#
|
|
180
|
+
carried = @pending_bytes.to_s.b
|
|
181
|
+
start = @bytes_read - carried.bytesize # stream position of the first carried byte
|
|
182
|
+
bytes = carried + chunk.b
|
|
183
|
+
# A read can end partway through a multi-byte character; carry those
|
|
184
|
+
# bytes into the next read rather than judging half a character.
|
|
174
185
|
tail = bytes[/[\xC2-\xF4][\x80-\xBF]{0,2}\z/n]
|
|
175
186
|
width = tail && (tail.getbyte(0) < 0xE0 ? 2 : tail.getbyte(0) < 0xF0 ? 3 : 4)
|
|
176
187
|
@pending_bytes = tail && tail.bytesize < width ? tail : +"".b
|
|
177
|
-
bytes = bytes.byteslice(0, bytes.bytesize - @pending_bytes.bytesize)
|
|
178
|
-
|
|
188
|
+
bytes = bytes.byteslice(0, bytes.bytesize - @pending_bytes.bytesize).force_encoding(Encoding::UTF_8)
|
|
189
|
+
# Anything still invalid is corrupt input. Dropping it would change the
|
|
190
|
+
# text without a trace, so stop and say where it is.
|
|
191
|
+
unless bytes.valid_encoding?
|
|
192
|
+
bad = bytes.each_char.find_index { |c| !c.valid_encoding? }.to_i
|
|
193
|
+
offset = start + bytes[0, bad].bytesize
|
|
194
|
+
raise Wp2txt::EncodingError,
|
|
195
|
+
"invalid UTF-8 near decompressed byte #{offset}: #{@input_path}"
|
|
196
|
+
end
|
|
197
|
+
@bytes_read += chunk.bytesize
|
|
198
|
+
@buffer << bytes
|
|
179
199
|
|
|
180
200
|
# Adaptive buffer adjustment: if memory is low, reduce buffer size
|
|
181
201
|
if @adaptive_buffer && MemoryMonitor.memory_low?
|
|
@@ -252,7 +272,7 @@ module Wp2txt
|
|
|
252
272
|
end
|
|
253
273
|
|
|
254
274
|
@pages_processed += 1
|
|
255
|
-
[title, text]
|
|
275
|
+
[title, text, Wp2txt.page_ids(page_xml)]
|
|
256
276
|
rescue Nokogiri::XML::SyntaxError
|
|
257
277
|
# Skip malformed XML
|
|
258
278
|
nil
|
data/lib/wp2txt/version.rb
CHANGED
data/lib/wp2txt.rb
CHANGED
|
@@ -269,12 +269,14 @@ module Wp2txt
|
|
|
269
269
|
end
|
|
270
270
|
page << line if inside_page
|
|
271
271
|
end
|
|
272
|
-
if page.empty?
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
272
|
+
return false if page.empty?
|
|
273
|
+
|
|
274
|
+
page.force_encoding("utf-8")
|
|
275
|
+
# Corrupt input must not be cleaned into different text without a trace
|
|
276
|
+
unless page.valid_encoding?
|
|
277
|
+
title = page[%r{<title>([^<]*)</title>}, 1]&.scrub("?")
|
|
278
|
+
raise Wp2txt::EncodingError, "invalid UTF-8 in page #{title.inspect} of #{@input_file}"
|
|
276
279
|
end
|
|
277
|
-
rescue ::Encoding::InvalidByteSequenceError, ::Encoding::UndefinedConversionError
|
|
278
280
|
page
|
|
279
281
|
end
|
|
280
282
|
|
|
@@ -0,0 +1,146 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "spec_helper"
|
|
4
|
+
require "open3"
|
|
5
|
+
require "json"
|
|
6
|
+
require "tmpdir"
|
|
7
|
+
require "parallel"
|
|
8
|
+
require_relative "support/multistream_fixture"
|
|
9
|
+
require "wp2txt/multistream"
|
|
10
|
+
require "wp2txt/output_writer"
|
|
11
|
+
require "wp2txt/stream_processor"
|
|
12
|
+
|
|
13
|
+
RSpec.describe "output integrity" do
|
|
14
|
+
include MultistreamFixture
|
|
15
|
+
|
|
16
|
+
around do |example|
|
|
17
|
+
Dir.mktmpdir("wp2txt-integrity-") do |dir|
|
|
18
|
+
@dir = dir
|
|
19
|
+
example.run
|
|
20
|
+
end
|
|
21
|
+
end
|
|
22
|
+
|
|
23
|
+
let(:cli) { File.expand_path("../bin/wp2txt", __dir__) }
|
|
24
|
+
let(:lib) { File.expand_path("../lib", __dir__) }
|
|
25
|
+
|
|
26
|
+
def run_cli(*args)
|
|
27
|
+
Open3.capture3(RbConfig.ruby, "-I", lib, cli, *args)
|
|
28
|
+
end
|
|
29
|
+
|
|
30
|
+
def json_records(dir)
|
|
31
|
+
Dir[File.join(dir, "*")].flat_map { |f| File.readlines(f) }.map { |line| JSON.parse(line) }
|
|
32
|
+
end
|
|
33
|
+
|
|
34
|
+
def write_pages(path, count)
|
|
35
|
+
body = (1..count).map { |i| page_xml(id: i, ns: 0, title: "記事#{i}", text: "本文#{i}。日本語の文。\n") }.join
|
|
36
|
+
File.write(path, "<mediawiki>\n#{body}</mediawiki>\n")
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
describe "forked workers" do
|
|
40
|
+
it "do not write the parent's buffered output again when they exit" do
|
|
41
|
+
writer = Wp2txt::OutputWriter.new(output_dir: @dir, base_name: "out", format: :text, file_size_mb: 10)
|
|
42
|
+
writer.write("one line still sitting in the buffer")
|
|
43
|
+
writer.flush
|
|
44
|
+
Parallel.map([1, 2, 3], in_processes: 3) { |x| x }
|
|
45
|
+
files = writer.close
|
|
46
|
+
expect(File.readlines(files.first).size).to eq(1)
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
it "leave every article exactly once in streaming output across batch boundaries" do
|
|
50
|
+
input = File.join(@dir, "pages.xml")
|
|
51
|
+
out = File.join(@dir, "out")
|
|
52
|
+
Dir.mkdir(out)
|
|
53
|
+
write_pages(input, 450) # batches of 200 with -n 2
|
|
54
|
+
_stdout, stderr, status = run_cli("-i", input, "-o", out, "--format", "json", "-n", "2")
|
|
55
|
+
expect(status.success?).to be(true), stderr
|
|
56
|
+
|
|
57
|
+
titles = json_records(out).map { |r| r["title"] }
|
|
58
|
+
expect(titles.size).to eq(450)
|
|
59
|
+
expect(titles.tally.select { |_, n| n > 1 }).to be_empty
|
|
60
|
+
end
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
describe "--num-procs" do
|
|
64
|
+
it "uses the requested process count" do
|
|
65
|
+
input = File.join(@dir, "pages.xml")
|
|
66
|
+
out = File.join(@dir, "out")
|
|
67
|
+
Dir.mkdir(out)
|
|
68
|
+
write_pages(input, 3)
|
|
69
|
+
stdout, stderr, status = run_cli("-i", input, "-o", out, "--format", "json", "-n", "1")
|
|
70
|
+
expect(status.success?).to be(true), stderr
|
|
71
|
+
expect(stdout.gsub(/\e\[[\d;]*m/, "")).to match(/CPU cores:\s*1\b/)
|
|
72
|
+
end
|
|
73
|
+
end
|
|
74
|
+
|
|
75
|
+
describe "page and revision IDs" do
|
|
76
|
+
it "appear in JSON records from the streaming path, including summaries" do
|
|
77
|
+
input = File.join(@dir, "pages.xml")
|
|
78
|
+
write_pages(input, 3)
|
|
79
|
+
[[], ["--summary-only"]].each do |extra|
|
|
80
|
+
out = File.join(@dir, "out#{extra.size}")
|
|
81
|
+
Dir.mkdir(out)
|
|
82
|
+
_stdout, stderr, status = run_cli("-i", input, "-o", out, "--format", "json", *extra)
|
|
83
|
+
expect(status.success?).to be(true), stderr
|
|
84
|
+
record = json_records(out).find { |r| r["title"] == "記事2" }
|
|
85
|
+
expect(record).to include("page_id" => 2, "revision_id" => 200)
|
|
86
|
+
expect(record.keys.first(3)).to eq(%w[title page_id revision_id])
|
|
87
|
+
end
|
|
88
|
+
end
|
|
89
|
+
|
|
90
|
+
it "appear in JSON records from the default path for bz2 dumps" do
|
|
91
|
+
dump, = create_fixture(@dir)
|
|
92
|
+
out = File.join(@dir, "out")
|
|
93
|
+
Dir.mkdir(out)
|
|
94
|
+
_stdout, stderr, status = run_cli("-i", dump, "-o", out, "--format", "json")
|
|
95
|
+
expect(status.success?).to be(true), stderr
|
|
96
|
+
record = json_records(out).find { |r| r["title"] == "Film B" }
|
|
97
|
+
expect(record).to include("page_id" => 2, "revision_id" => 200)
|
|
98
|
+
end
|
|
99
|
+
|
|
100
|
+
it "are read from the page header, not from the contributor" do
|
|
101
|
+
xml = "<page><title>X</title><ns>0</ns><id>12</id><revision><id>34</id>" \
|
|
102
|
+
"<contributor><id>9</id></contributor><text>t</text></revision></page>"
|
|
103
|
+
expect(Wp2txt.page_ids(xml)).to eq(page_id: 12, revision_id: 34)
|
|
104
|
+
end
|
|
105
|
+
|
|
106
|
+
it "are offered by StreamProcessor only on request, keeping each_page's shape" do
|
|
107
|
+
input = File.join(@dir, "pages.xml")
|
|
108
|
+
write_pages(input, 1)
|
|
109
|
+
processor = Wp2txt::StreamProcessor.new(input, adaptive_buffer: false)
|
|
110
|
+
expect(processor.each_page.to_a).to eq([["記事1", "本文1。日本語の文。\n"]])
|
|
111
|
+
with_ids = Wp2txt::StreamProcessor.new(input, adaptive_buffer: false).each_page(with_ids: true).to_a
|
|
112
|
+
expect(with_ids.first.last).to eq(page_id: 1, revision_id: 100)
|
|
113
|
+
end
|
|
114
|
+
end
|
|
115
|
+
|
|
116
|
+
describe "targeted extraction after an early index stop" do
|
|
117
|
+
it "knows where the last found article's stream ends" do
|
|
118
|
+
_dump, index_path = create_fixture(@dir)
|
|
119
|
+
index = Wp2txt::MultistreamIndex.new(index_path, use_cache: false, target_titles: ["Film A"],
|
|
120
|
+
show_progress: false)
|
|
121
|
+
expect(index.early_terminated?).to be(true)
|
|
122
|
+
expect(index.stream_offsets).to eq([0])
|
|
123
|
+
second_stream = File.readlines(index_path).map { |line| line.split(":").first.to_i }.uniq[1]
|
|
124
|
+
expect(index.stream_end_offset).to eq(second_stream)
|
|
125
|
+
end
|
|
126
|
+
|
|
127
|
+
it "reads only that stream" do
|
|
128
|
+
dump, index_path = create_fixture(@dir)
|
|
129
|
+
index = Wp2txt::MultistreamIndex.new(index_path, use_cache: false, target_titles: ["Film A"],
|
|
130
|
+
show_progress: false)
|
|
131
|
+
reader = Wp2txt::MultistreamReader.new(dump, index)
|
|
132
|
+
stub_const("Wp2txt::MultistreamReader::MAX_TAIL_STREAM_BYTES", 0)
|
|
133
|
+
expect(reader.extract_article("Film A")).to include(title: "Film A", id: 1, revision_id: 100)
|
|
134
|
+
end
|
|
135
|
+
|
|
136
|
+
it "refuses to read a large remainder in one call when the end is unknown" do
|
|
137
|
+
dump, index_path = create_fixture(@dir)
|
|
138
|
+
index = Wp2txt::MultistreamIndex.new(index_path, use_cache: false, target_titles: ["Film A"],
|
|
139
|
+
show_progress: false)
|
|
140
|
+
index.instance_variable_set(:@stream_end_offset, nil)
|
|
141
|
+
reader = Wp2txt::MultistreamReader.new(dump, index)
|
|
142
|
+
stub_const("Wp2txt::MultistreamReader::MAX_TAIL_STREAM_BYTES", 0)
|
|
143
|
+
expect { reader.extract_article("Film A") }.to raise_error(Wp2txt::Error, /cannot locate the end/)
|
|
144
|
+
end
|
|
145
|
+
end
|
|
146
|
+
end
|
data/spec/p1_correctness_spec.rb
CHANGED
|
@@ -40,10 +40,25 @@ RSpec.describe "P1 correctness contracts" do
|
|
|
40
40
|
end
|
|
41
41
|
end
|
|
42
42
|
|
|
43
|
-
it "
|
|
44
|
-
processor = processor_for("
|
|
45
|
-
nil while processor.send(:fill_buffer)
|
|
46
|
-
|
|
43
|
+
it "stops on an invalid byte and reports where it is" do
|
|
44
|
+
processor = processor_for("éあ\xFF😀".b, 1)
|
|
45
|
+
expect { nil while processor.send(:fill_buffer) }
|
|
46
|
+
.to raise_error(Wp2txt::EncodingError, /byte 5\b/)
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
it "stops when the input ends partway through a character" do
|
|
50
|
+
processor = processor_for("éあ\xE3\x81".b, 1)
|
|
51
|
+
expect { nil while processor.send(:fill_buffer) }
|
|
52
|
+
.to raise_error(Wp2txt::EncodingError, /middle of a UTF-8 character/)
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
it "passes valid text through unchanged at every buffer size" do
|
|
56
|
+
text = "é\nあ😀終\n" * 7
|
|
57
|
+
[1, 2, 3, 5, 64].each do |size|
|
|
58
|
+
processor = processor_for(text.b, size)
|
|
59
|
+
nil while processor.send(:fill_buffer)
|
|
60
|
+
expect(processor.instance_variable_get(:@buffer)).to eq(text)
|
|
61
|
+
end
|
|
47
62
|
end
|
|
48
63
|
|
|
49
64
|
it "handles empty input" do
|
|
@@ -105,7 +120,7 @@ RSpec.describe "P1 correctness contracts" do
|
|
|
105
120
|
it "retains the historical ns=0 default for missing ns in both parsers" do
|
|
106
121
|
xml = page_xml(id: 1, ns: 0, title: "作品: 東京", text: "本文").sub(/<ns>.*?<\/ns>/, "")
|
|
107
122
|
processor = Wp2txt::StreamProcessor.new("unused.xml", adaptive_buffer: false)
|
|
108
|
-
expect(processor.send(:parse_page_xml, xml)).to eq(["作品: 東京", "本文"])
|
|
123
|
+
expect(processor.send(:parse_page_xml, xml)).to eq(["作品: 東京", "本文", { page_id: 1, revision_id: 100 }])
|
|
109
124
|
rows = { pages: [], categories: [], sections: [], hierarchy: [] }
|
|
110
125
|
Wp2txt::MetadataIndexBuilder.scan_page(xml, rows)
|
|
111
126
|
expect(rows[:pages].first[2]).to eq(0)
|
|
@@ -36,6 +36,18 @@ RSpec.describe "titles extraction and SQL file output" do
|
|
|
36
36
|
File.readlines(path).map { |l| JSON.parse(l) }
|
|
37
37
|
end
|
|
38
38
|
|
|
39
|
+
describe "extract_corpus across several batches with worker processes" do
|
|
40
|
+
it "writes each record once (workers must not re-flush the output buffer)" do
|
|
41
|
+
stub_const("Wp2txt::Corpus::EXTRACT_BATCH_SIZE", 1)
|
|
42
|
+
out = File.join(@dir, "batched.jsonl")
|
|
43
|
+
titles = ["Film A", "Film B", "Person X"]
|
|
44
|
+
result = @corpus.extract_corpus(output_path: out, content: "full", titles: titles, num_processes: 2)
|
|
45
|
+
records = read_jsonl(out)
|
|
46
|
+
expect(records.map { |r| r["title"] }.tally).to eq(titles.to_h { |t| [t, 1] })
|
|
47
|
+
expect(records.size).to eq(result[:records_written])
|
|
48
|
+
end
|
|
49
|
+
end
|
|
50
|
+
|
|
39
51
|
describe "extract_corpus titles:" do
|
|
40
52
|
it "extracts an explicit set with normalization, dedup, and input order" do
|
|
41
53
|
out = File.join(@dir, "t.jsonl")
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: wp2txt
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 2.3.
|
|
4
|
+
version: 2.3.4
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Yoichiro Hasebe
|
|
@@ -284,6 +284,7 @@ files:
|
|
|
284
284
|
- spec/metadata_index_spec.rb
|
|
285
285
|
- spec/multi_dump_attach_spec.rb
|
|
286
286
|
- spec/multistream_spec.rb
|
|
287
|
+
- spec/output_integrity_spec.rb
|
|
287
288
|
- spec/output_writer_spec.rb
|
|
288
289
|
- spec/p1_correctness_spec.rb
|
|
289
290
|
- spec/parser_functions_spec.rb
|
|
@@ -351,6 +352,7 @@ test_files:
|
|
|
351
352
|
- spec/metadata_index_spec.rb
|
|
352
353
|
- spec/multi_dump_attach_spec.rb
|
|
353
354
|
- spec/multistream_spec.rb
|
|
355
|
+
- spec/output_integrity_spec.rb
|
|
354
356
|
- spec/output_writer_spec.rb
|
|
355
357
|
- spec/p1_correctness_spec.rb
|
|
356
358
|
- spec/parser_functions_spec.rb
|