fetch_util 0.8.0 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +153 -0
- data/README.md +10 -0
- data/lib/fetch_util/assets/extract.js +1 -1
- data/lib/fetch_util/fetcher.rb +1 -0
- data/lib/fetch_util/version.rb +1 -1
- metadata +2 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: cd00c416e7059bb87a1550c79781737e65e206956302e171b81178acb28209bc
|
|
4
|
+
data.tar.gz: f4df7dafba7184a2dc622e1e92e5f7617626491a2e87b7491bd94a2efe043e7a
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 551d18d6974814b9edc158f631db550831bfcb226653663bbeac18813076360fbe4b9105783ac96427b2a7bf1a1d885da109a7701bdbc889b45097cfd0b7f7a6
|
|
7
|
+
data.tar.gz: 42c9da118fc910f1e2a138c6de3978fed436f1eb5b5b6f69e06f088446b745bdc4eac394b1efe0fa21e5bb95c5393b1d85c3d6122a85c05eb4a9424645d4ec94
|
data/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,159 @@
|
|
|
2
2
|
|
|
3
3
|
## Unreleased
|
|
4
4
|
|
|
5
|
+
## v0.9.0 - 2026-10-04
|
|
6
|
+
|
|
7
|
+
### Changed
|
|
8
|
+
|
|
9
|
+
- Replace 25 runtime host-profile handlers with shared extraction, leaving 70 registrations. This includes the Index.hu, Sabah, Interia and Abril handlers alongside the families listed below; WordPress host-override removals are counted separately.
|
|
10
|
+
- Require source-owned linked summaries and editorial context for inferred list digests, bind article context to its visible source owner, and preserve independent multi-topic warnings.
|
|
11
|
+
- Preserve distinct source result records that share one destination URL.
|
|
12
|
+
- Keep every table cell from a proven source-owned article body when a product or repeated-record list would omit it.
|
|
13
|
+
- Preserve complete source-owned About and page-introduction paragraphs in repeated-record list results without borrowing neighboring prose.
|
|
14
|
+
- Retire the Sky News Arabia profile while preserving owned Arabic reporting, complete prose excerpts, literal attribution, media, and truthful news-age warnings through shared extraction.
|
|
15
|
+
- Retire the CSDN profile in favor of shared article and declared-description extraction, preserving public prose and visible access notices without fabricated fallback text.
|
|
16
|
+
- Retire the Walla article profile in favor of shared extraction that preserves source headings, visible attribution, complete prose, and labelled topic destinations.
|
|
17
|
+
- Select a semantic article body only when its unique headline, matching source description, and owned header establish one source record; preserve complete article material and unclassified sibling prose.
|
|
18
|
+
- Retire the YLE live-article profile in favor of source-owned liveblog extraction, retaining complete updates and coherent parent attribution when the requested entry is not independently identified.
|
|
19
|
+
- Retire the Jang profile while preserving complete reporting, owned media, explicitly addressed updates, and full liveblog context when update ownership is unproved.
|
|
20
|
+
- Retire the GOV.UK guidance profile while preserving source-owned Contents, plain article layouts, publication precision, complete sections, and public update history through shared extraction.
|
|
21
|
+
- Retire the Federal Register profile in favor of source-owned notice extraction, preserving formal sections, agency and date attribution, complete summaries, and linked document resources.
|
|
22
|
+
- Retire the Marca article profile while preserving complete public reporting and owned media, and retain truthful warnings for explicit premium previews.
|
|
23
|
+
- Retire the legal-reference article profile while preserving source-owned encyclopedia headings, complete prose, lists, and linked legal resources through shared extraction.
|
|
24
|
+
- Retire the Xinhua weekly profile in favor of shared list ownership, preserving every visible issue and its linked cover image.
|
|
25
|
+
- Retire the SegmentFault article profile while preserving complete CJK articles, code, images, and source-owned reader comments through shared extraction.
|
|
26
|
+
- Retire the 20minutos live-article profile while preserving every visible update, source order, and complete numeric lead sentences through shared extraction.
|
|
27
|
+
- Retire the Dnevnik article profile while preserving visible public teasers, source language, author attribution, and subscription warnings through shared extraction.
|
|
28
|
+
- Retire the Al Masry Al Youm article profile, preserving complete digest stories and applying normal news-age warnings instead of a documentation exemption.
|
|
29
|
+
- Retire the Tempo article profile while preserving source-owned reporting, public previews, captions, and access warnings through shared extraction.
|
|
30
|
+
- Retire the arXiv abstract profile in favor of shared scholarly extraction that preserves source-owned authors, dates, version history, and download links.
|
|
31
|
+
- Retire the IEEE Xplore abstract profile and obsolete publisher registration module in favor of complete shared scholarly records.
|
|
32
|
+
- Retire the Elsevier article profile and its unused configuration pipeline in favor of shared scholarly frontmatter and body ownership.
|
|
33
|
+
- Retire the PLOS article profile while preserving complete source-owned sections through shared scholarly extraction.
|
|
34
|
+
- Retire the ACS abstract profile in favor of shared scholarly ownership and complete abstract excerpts.
|
|
35
|
+
- Retire the Maroela-specific WordPress override in favor of shared title, prose-excerpt, discussion-status, and post-owned comment handling.
|
|
36
|
+
- Retire the Ghost article handler in favor of shared extraction that preserves owned headers, complete leads, editorial headings, linked-card attribution, and embedded media.
|
|
37
|
+
- Retire the NewsIT-specific WordPress override in favor of shared extraction that preserves owned leads, media, quotations, links, and reader discussions.
|
|
38
|
+
|
|
39
|
+
### Fixed
|
|
40
|
+
|
|
41
|
+
- Select uniquely source-owned posts for permalink URLs.
|
|
42
|
+
- Select complete authored opening paragraphs only from the mapped article owner when a broader fallback root preserves the article body.
|
|
43
|
+
- Keep a complete source-mapped service article when a product-tab list would drop its owned sections, prose, features and media; require unique source ownership and compare encoded image URLs by DOM identity.
|
|
44
|
+
- Keep complete visible product-service reviews only when the selected service and review collection have matching, visible product-review links; decline foreign or ambiguous review owners.
|
|
45
|
+
- Preserve a precise source-owned publication timestamp only when its calendar day agrees with the visible date, without guessing timezone conversions.
|
|
46
|
+
- Keep contributor-header dates only when they belong to the uniquely selected article body and do not conflict with other owned dates.
|
|
47
|
+
- Preserve uniquely owned focal media articles, their complete summaries, publication times and related records at the original source position.
|
|
48
|
+
- Scope Reader summaries to the proved article body and header even with a publisher-suffixed document title or an independent neighboring section.
|
|
49
|
+
- Keep complete source-owned Reader article summaries independent of byline classification while excluding summary copies in neighboring or nested articles.
|
|
50
|
+
- Retain original same-page headline links only when the complete selected article body matches the source owner.
|
|
51
|
+
- Bind standalone static-generator title brands only to a source-owned article heading, preserving genuine title extensions.
|
|
52
|
+
- Preserve publication dates only when the uniquely owned source article, header, and publication records agree.
|
|
53
|
+
- Preserve complete source-owned article body leads, including short Reader paragraphs with unique article and body proof.
|
|
54
|
+
- Preserve complete query-bound material results and decline list takeover when any visible owner material is unaccounted for.
|
|
55
|
+
- Preserve every source-owned search-history occurrence in order, including repeated destinations, and decline incomplete histories.
|
|
56
|
+
- Preserve sparse focal media articles with their complete independently owned related collections and media resources.
|
|
57
|
+
- Restore source-owned primary-service and calculator sections when interactive headings displace the article.
|
|
58
|
+
- Avoid repeated source-article ownership scans while pruning related rails from an already-proven clone.
|
|
59
|
+
- Preserve source-owned query directories when a separate focal article is not a fully proven NewsArticle.
|
|
60
|
+
- Remove terminal recommendation furniture only when complete source-to-clone ownership is proven, preserving authored citations, lists, captions, and article prose.
|
|
61
|
+
- Preserve publication labels only when independently confirmed as body-owned dates; retain surrounding complete article prose and reject conflicting or foreign date values.
|
|
62
|
+
- Return the uniquely matched visible source H1 for article titles, requiring linked source proof before borrowing a publisher-owned heading.
|
|
63
|
+
- Keep complete source-owned excerpt paragraphs beneath publication actions; preserve short authored introductions and standfirsts while declining hidden or unrelated prose.
|
|
64
|
+
- Recognize a complete, source-owned primary headline collection as a credible generic homepage list when it has no editorial section wrapper; retain mismatch warnings for incidental navigation and error pages.
|
|
65
|
+
- Retain linked image, label, title and date fields for every proven repeated article card in generic lists without flattening source context.
|
|
66
|
+
- Keep a typed runtime detail record ahead of a list candidate only when its URL-bound record evidence would otherwise be lost.
|
|
67
|
+
- Preserve the complete summary of a proved focal article while declining related-card, quoted, or contradictory summaries.
|
|
68
|
+
- Use a uniquely attributed external article H1 only when its complete selected body belongs to that local source record.
|
|
69
|
+
- Preserve source-proved service context, calculator controls, visible FAQ triggers and owned media in order, while declining hidden or unrelated material.
|
|
70
|
+
- Keep proved result counts and continuation controls in source order with their linked records; treat unproved numeric-leading paragraphs as authored context.
|
|
71
|
+
- Render linked cover images from source-owned repeated records alongside their ordered issue metadata when structured-card Markdown drops the image destination.
|
|
72
|
+
- Keep the actual unlinked source H2 as the Markdown heading for its search-history collection, even when record metadata repeats its label.
|
|
73
|
+
- Bind article and question-answer publication dates to their local source owner, preserving article headlines beside related cards and declining ambiguous or unrelated answer records.
|
|
74
|
+
- Recognize source-owned institutional case indexes with complete case-card fields while preserving long legal titles and focal articles.
|
|
75
|
+
- Render complete source-owned structured card headings and fields for named container and block-anchor cards.
|
|
76
|
+
- Remove a related-story rail only when every selected article paragraph and destination is proven source-owned.
|
|
77
|
+
- Remove structurally empty article ad placeholders only inside a proven source-owned article body.
|
|
78
|
+
- Remove source-owned Google publisher action prompts without rewriting neighboring Reader text or citations.
|
|
79
|
+
- Preserve a source-owned Reader headline and use a complete, visibly labelled header summary only when its entire article body is source-proven.
|
|
80
|
+
- Remove empty comment-form headings only within a proven article owner, preserving authored headings and published replies.
|
|
81
|
+
- Map repeated-list records and their introduction into one source-owned cleaned main, excluding unrelated page context.
|
|
82
|
+
- Keep verified focal articles and typed detail pages ahead of unrelated repeated records while preserving genuine list introductions and distinct source-record occurrences.
|
|
83
|
+
- Carry verified public-preview ownership into canonical article rendering and excerpts, preserving surrounding public material while excluding locked and foreign content.
|
|
84
|
+
- Select complete, source-owned prose after an explicit photo credit while retaining the credit, media, metadata, and legacy dated-caption behavior.
|
|
85
|
+
- Preserve substantive visible editorial paragraphs inside ad-labelled containers when promotional text and media do not establish an advertisement.
|
|
86
|
+
- Preserve visible, source-owned card headers and their destinations without promoting hidden title or record evidence, or headers inside explicit navigation.
|
|
87
|
+
- Preserve a source-owned headline's actual self-link in article Markdown while rejecting unrelated, hidden, or invented heading destinations.
|
|
88
|
+
- Resolve article titles only from a unique visible source owner whose complete selected prose matches in order, preserving ambiguity and excluding hidden or unrelated title witnesses.
|
|
89
|
+
- Restore uniquely owned Reader media headers without duplicating images, separating captions, or borrowing unrelated article content.
|
|
90
|
+
- Apply contact-link permissions after final content classification, retaining safe article and notice contacts while keeping list, documentation, opaque HTML, and media destinations HTTP-only.
|
|
91
|
+
- Exclude explicitly identified advertisement owners from repeated-list evidence while preserving ordinary editorial cards.
|
|
92
|
+
- Score complete list owners before applying repeated-record coverage checks, retaining minority sections, controlled continuations, and source-owned metadata.
|
|
93
|
+
- Preserve the optional-metadata contract when resolving source-owned fallback headers, without changing explicitly supplied attribution.
|
|
94
|
+
- Replace an excerpt proved to be the article's byline with a complete owned summary or first prose paragraph, preserving short leads and rejecting foreign or ambiguous text.
|
|
95
|
+
- Restore removed, source-owned Reader credits in their original order without changing native author selection or dropping already retained coauthors.
|
|
96
|
+
- Remove inert custom share and related-content labels while preserving material custom components, source prose, headings, resources, and shadow content.
|
|
97
|
+
- Recognize named site-link navigation so it does not displace the page's primary list records.
|
|
98
|
+
- Remove unsupported link destinations from generic list HTML without dropping visible labels or changing safe inline article contacts.
|
|
99
|
+
- Preserve independently linked live-update homepage lists while retaining complete inline liveblog parents and same-document update permalinks.
|
|
100
|
+
- Require distinct linked sections before classifying short news subsections as a newsletter, preserving topic tags and measuring prose independently of destination length.
|
|
101
|
+
- Restore Reader bylines only from headers bound to its actual source paragraphs, preserving conflicting author metadata and rejecting foreign or ambiguous paragraph edges.
|
|
102
|
+
- Preserve declared page descriptions separately from article excerpts while keeping extractor-owned descriptions, including explicitly empty ones, authoritative.
|
|
103
|
+
- Preserve literal bylines from the selected article's owned header without replacing conflicting author attribution, and normalize localized publication digits on the final chosen timestamp.
|
|
104
|
+
- Keep liveblog titles consistent with the proved parent or update scope, including parent headings corroborated by source descriptions when headline metadata names a child update.
|
|
105
|
+
- Preserve source-owned liveblog parents and explicitly addressed updates, including material links, per-update attribution, and reader discussions without borrowing child metadata for the parent.
|
|
106
|
+
- Preserve complete canonical article bodies, independent header contributions, owned attribution, media, and reader discussions while rejecting ambiguous ownership and unrepresented source material.
|
|
107
|
+
- Prefer visible publication-date values over adjacent section labels while retaining localized, yearless, relative, and existing fallback date formats.
|
|
108
|
+
- Require WordPress-specific producer evidence before applying shared WordPress extraction, rather than treating ordinary article body classes as a CMS signature.
|
|
109
|
+
- Preserve declared guidance publication timestamps when they agree with the source-owned publication day, including GOV.UK producer metadata on any host.
|
|
110
|
+
- Recover article author and publication metadata from source-owned headers without attributing lists or mixed-source Reader content to that article.
|
|
111
|
+
- Preserve source-owned Contents navigation, complete guidance sections, publication metadata, and ordered update history without promoting unrelated components over an independently selected article.
|
|
112
|
+
- Extract formal notices through source-owned headings, agency and publication metadata, preserving complete sections, linked contacts, document resources, and reader discussions.
|
|
113
|
+
- Preserve source headline hierarchy when Reader splits explicit paragraph breaks, requiring complete, contiguous, source-identified paragraph groups.
|
|
114
|
+
- Restore retained article headlines through exact source-node provenance and ordered body ownership while preserving original anchors and excluding hidden header text.
|
|
115
|
+
- Preserve source-owned repeated collections, short record titles, and linked covers through shared list extraction while keeping distinct destinations and narrative ownership boundaries.
|
|
116
|
+
- Recognize unheaded scholarly records only when a unique journal citation, document title, abstract, and visible scientific sections share one owner.
|
|
117
|
+
- Keep independently owned directories ahead of canonical introduction metadata on collection routes, without promoting hidden, sidebar, or related record groups.
|
|
118
|
+
- Count prose across writing systems and distinguish linked product offers from digest stories when classifying article formats.
|
|
119
|
+
- Preserve homepage landing copy, service links, and partner media when a visible external heading owns one unambiguous prose article.
|
|
120
|
+
- Restore visible author-profile links for source-owned fallback articles using the same paragraph and identity proof as Reader output.
|
|
121
|
+
- Preserve complete source-owned focal article paragraphs in excerpts instead of clipping words or merging the next paragraph.
|
|
122
|
+
- Prefer explicit discussion threads over broad comment-layout wrappers so adjacent recommendations are not appended as reader comments.
|
|
123
|
+
- Preserve source-owned article discussions beside verified focal bodies, including short replies, unlinked authors, comment media, footer attribution, and thread destinations.
|
|
124
|
+
- Keep source-owned article bodies and visible public previews ahead of unrelated lists, while preserving surrounding editorial material and rejecting ambiguous, hidden, and social-post content.
|
|
125
|
+
- Preserve distinct source-owned scholarly highlights alongside the primary abstract without duplicating responsive summaries or unrelated recommendations.
|
|
126
|
+
- Extract scholarly records from shared abstract, article-body, citation, author, and version-history ownership while preserving complete abstracts and source-ordered links.
|
|
127
|
+
- Distinguish Arabic writer labels from author names, retaining named authors and rejecting empty attribution labels.
|
|
128
|
+
- Choose WordPress body prose ahead of caption fields and standalone photo-credit lines for excerpts while retaining all captions, images, and links in the article.
|
|
129
|
+
- Remove complete inert closed-discussion notices while preserving quoted statements, editorial text, reader comments, links, and media.
|
|
130
|
+
- Normalize WordPress fallback titles against an explicitly declared publisher suffix while preserving actual headlines and unmatched title text.
|
|
131
|
+
- Preserve closed WordPress discussions and replies inside their owning post even when no comment form supplies the post ID, while rejecting conflicting ownership declarations.
|
|
132
|
+
- Preserve Reader article-header ownership across transformed linked cards when their source text, destination, and position agree.
|
|
133
|
+
- Keep linked resource-card attribution in article content without mistaking it for the article's author.
|
|
134
|
+
- Preserve prose-supported editorial headings whose generated anchor IDs contain newsletter, advertising, or other cleanup keywords.
|
|
135
|
+
- Use complete, explicitly marked source-owned Reader header leads as excerpts while preserving existing excerpts for ambiguous header text.
|
|
136
|
+
- Restore visible article-header context when its unique headline and ordered body paragraphs establish Reader ownership, preserving taxonomy, reading time, and distinct header leads.
|
|
137
|
+
- Restrict duration-player label cleanup to standalone timing lines so editorial prose and image descriptions keep the word “duration”.
|
|
138
|
+
- Retain linked images and charts inside Markdown table cells and separate block-level cell labels without losing destinations.
|
|
139
|
+
- Preserve owned inline header dates when the same article repeats its exact headline in the body, while rejecting competing article headings.
|
|
140
|
+
- Preserve complete structured discussion tables in list extraction, including short topic labels, author links, and per-row metadata.
|
|
141
|
+
- Skip standalone image-strip labels when choosing WordPress prose excerpts while retaining their text and images in the full article.
|
|
142
|
+
- Recover explicitly labeled inline publication and update dates from owned article headers without treating the header or body as a date.
|
|
143
|
+
- Preserve visible external WordPress comments and nested replies when the thread's declared post ID matches the selected article, including author links and timestamps.
|
|
144
|
+
- Retain uniquely owned WordPress header headlines in article HTML as well as Markdown and title metadata.
|
|
145
|
+
- Derive shared WordPress excerpts from complete owned header leads or body prose instead of image Markdown and captions.
|
|
146
|
+
- Reject whole article and body containers as visible publication-date fields while retaining localized dates and explicit timestamps.
|
|
147
|
+
- Preserve visible, source-labeled editorial iframe destinations as links when article cleanup removes the embedded frames.
|
|
148
|
+
- Retain source-owned WordPress featured images and captions before the article body without importing hidden or unrelated media.
|
|
149
|
+
- Preserve separate WordPress article-header leads when unique headline and body ownership establish their relationship, excluding hidden, sidebar, and ambiguous summaries.
|
|
150
|
+
- Exclude complete inert Italian automatic-audio notices from article text and excerpts while preserving quoted notices, additional reporting, links, and recordings.
|
|
151
|
+
|
|
152
|
+
### Known Limitations
|
|
153
|
+
|
|
154
|
+
- The 2026-10-04 frozen-source comparison of 300 requested website hosts against the published v0.8.0 found 101 reviewed improvements, 180 ties, 15 mixed outcomes with confirmed regressions, zero release-better outcomes, and four inconclusive cases. Of the improvements, 46 are description-only metadata additions. Strict no-regression acceptance did not pass; gains do not offset the confirmed losses.
|
|
155
|
+
- Remaining regressions include focal-story and tutorial-body selection, missing Markdown paragraphs, FAQ and service records, media links and bibliographic details, publication dates, author attribution, and article/list/newsletter/liveblog classification.
|
|
156
|
+
- Both versions reached the benchmark's 120-second native-extraction deadline on WeWorkRemotely and Rezultati. GogoRoyal and Aktuality remain inconclusive because substantive shadow content was not fully represented in the frozen replay.
|
|
157
|
+
|
|
5
158
|
## v0.8.0 - 2026-09-24
|
|
6
159
|
|
|
7
160
|
### Changed
|
data/README.md
CHANGED
|
@@ -20,6 +20,12 @@ The easiest way to explain `fetch_util` is in three steps:
|
|
|
20
20
|
|
|
21
21
|
This is bounded native browser execution, not a custom challenge solver. If the challenge does not resolve within the bound, the result remains an explicit interstitial with warnings.
|
|
22
22
|
|
|
23
|
+
For an article page with a unique source headline and semantic article record, shared extraction verifies their source relationship before selecting the body. It retains the complete article and its owned media, downloads, and access notices; a separate linked-record group is excluded only when its heading/link structure is complete and leaves no other visible material behind. Ambiguous sibling notes and references remain available.
|
|
24
|
+
|
|
25
|
+
When that source-owned article also has a uniquely attributed header, extraction preserves its visible author names and publication date. Unambiguous localized numerals in publication timestamps are normalized without changing the source text shown in the article.
|
|
26
|
+
|
|
27
|
+
A source-visible byline fills a missing or publisher-only author field, or extends the same already identified person. It does not replace a conflicting person name declared for the page. Multiple marked author names are combined only when they share one visibly owned byline group.
|
|
28
|
+
|
|
23
29
|
## Installation
|
|
24
30
|
|
|
25
31
|
Add the gem to your Gemfile:
|
|
@@ -106,6 +112,8 @@ fetch_util fetch https://example.com/selected --format json
|
|
|
106
112
|
Useful result fields:
|
|
107
113
|
|
|
108
114
|
- `title`
|
|
115
|
+
- `excerpt`
|
|
116
|
+
- `description` — the declared page description for article results, separate from the selected excerpt; extractor-provided descriptions remain authoritative
|
|
109
117
|
- `markdown`
|
|
110
118
|
- `final_url`
|
|
111
119
|
- `canonical_url`
|
|
@@ -113,6 +121,8 @@ Useful result fields:
|
|
|
113
121
|
- `suspect`
|
|
114
122
|
- `warnings`
|
|
115
123
|
|
|
124
|
+
For article results, `description` exposes declared page metadata independently of the selected article excerpt. Extractor-owned descriptions remain authoritative for every content type; non-article results do not inherit page descriptions.
|
|
125
|
+
|
|
116
126
|
## Common Options
|
|
117
127
|
|
|
118
128
|
- `timeout:` browser timeout in seconds; it is also the bounded observation budget for a delivered Anubis challenge
|