fetch_util 0.6.2 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 48a3acb88a4254f45ac57c889570f0f70b2fcae20e42957164d7b23114f18226
4
- data.tar.gz: 21525e77346c57d53ce78070654462817990bb6bbdb6a229e6e9ad283940cc0d
3
+ metadata.gz: 579a3f35c401910fa8fb6713370113b8158b170438716be04b224542fd19b46d
4
+ data.tar.gz: 80c3444cfe7cf2ad43ebaa47a46ec0ffb28fb12915c188672beba142e30bfd6a
5
5
  SHA512:
6
- metadata.gz: 9b511e0a3a8603b3d577595fe98870b2b2563123efb99c2afd8a2b807eda326f0d62a110d6d3262285fcac3744b0ba1647046da20c1f669ca6405d05ac5a5e21
7
- data.tar.gz: 2840c6daad1e744558d6c9c7ee82ce0c873113437f7ed28ff511316322bdc0ce871de63bb0cff836607083841d930dee7af94f50292b8e7509435917162fda1b
6
+ metadata.gz: 54999188b4a18ae81c17733689c7f050e85b4d98dd30b5f65933d17745b8e399add16972d10186edd3bc3d5bce8438a06a72373ce61739f87d9293800c787b42
7
+ data.tar.gz: e1c52a6939a13a9e1c70d310c0dfb4ce206c7f4161979d4ca5c8fa35b0715105d985a43cd30aa1e0874ab66713bd58617e992761836c81374db063048dc43f39
data/CHANGELOG.md CHANGED
@@ -2,6 +2,166 @@
2
2
 
3
3
  ## Unreleased
4
4
 
5
+ ## v0.7.0 - 2026-09-20
6
+
7
+ ### Changed
8
+
9
+ - Remove only short exact-token `most-read` recommendation furniture from cleaned article clones while preserving substantive or semantic owners, similarly named content, intentional page overviews, and the source DOM.
10
+ - Preserve source-owned reader excerpts for structured detail articles whose body uses non-semantic wrappers, without promoting category/date chrome or ambiguous page prose.
11
+ - Preserve exact-root homepage lead coverage when generalized clone mapping cannot prove correspondence, without dropping original lead destinations.
12
+ - Preserve substantive standalone prose and control labels owned by homepage lead collections while excluding unavailable interface controls and card-local duplicates.
13
+ - Preserve source-owned homepage context when a body-wide lead collection maps uniquely into a narrower cleaned list root.
14
+ - Preserve coherent developer-product homepage narratives when repeated substantive feature sections own a complete set of product destinations.
15
+ - Preserve the inactive suffix of a complete materialized Slick record carousel while excluding cloned presentation copies and ambiguous hidden inventories.
16
+ - Preserve image, category, title, and date relationships in repeated article cards whose outer link wraps every field.
17
+ - Keep a substantial structured detail article when late list detection sees only distributed portal rails, while preserving independent directories and list routes.
18
+ - Preserve safe article-owned supplemental asset and video destinations that reader mode omits.
19
+ - Prefer an explicitly named author field over sibling contact controls when collecting visible article bylines.
20
+ - Keep source-visible article summaries intact instead of expanding short reader excerpts into truncated flattened text.
21
+
22
+ - Preserve a visible localized publication date when it was embedded in a rejected date-only byline instead of falling back to stale structured metadata.
23
+ - Prefer article-specific author metadata over generic site-level author declarations.
24
+ - Ignore relative author-profile paths in metadata so a visible human byline can win.
25
+ - Omit supplemental list-card text when the exact same value is already rendered as local record context.
26
+
27
+ - Do not promote a byline owned by one list record into page-level metadata when the page does not independently declare it.
28
+ - Omit repeated action labels from list-card supplemental text when the same-label destination is already represented by a peer record.
29
+ - Fail fast on ambiguous trailing-action cards before running ownership scans across large list collections.
30
+ - Preserve compact card metadata rows in source order instead of extracting their timestamps ahead of adjacent local labels.
31
+ - Suppress a supplemental tracking alias only when its canonical record is already represented under the same visible label, while retaining distinct-label references to shared destinations.
32
+ - Keep a proved trailing-action card heading adjacent to its description without adding presentation punctuation absent from the source.
33
+ - Exclude an exact page hostname from article bylines while preserving human and organization author names.
34
+ - Remove an accidental host-only diagnostic trap, keep generic carousel supplements out of profile-owned homepage lists, and recover proved carousel records whose titles are separate from short action links.
35
+ - Preserve complete media-backed record inventories from explicitly controlled ARIA carousels while leaving uncontrolled hidden slides excluded.
36
+ - Preserve source-owned local heading destinations when the same URLs also identify other list records.
37
+ - Preserve a uniquely owned heading before its description when a structured card ends with a trailing action link.
38
+ - Preserve every substantive card in a proved reader-mode article carousel when normal visibility extraction already retained an exact source prefix.
39
+ - Preserve every unique safe record from explicit continuous news tickers on not-found pages without relaxing normal visibility clipping.
40
+
41
+ - Keep nested homepage authors attached to their story while separating them from the story's declared title.
42
+ - Render list-card bylines and exact time fields once when those fields also appear in raw card detail.
43
+ - Keep localized publication dates separate from person or site-identity bylines.
44
+ - Preserve title-owned media resources while stripping generic portal record actions and duplicate title fields.
45
+ - Separate each generic portal record's visibility-pruned supplemental card from its metadata owner so discarded action controls stay removed without losing prose, media, captions, authors, or dates.
46
+ - Preserve visible semantic record collections as lists when some record destinations are intentionally inert, without exposing unsafe links or weakening ordinary record-heading ownership.
47
+ - Preserve structured metadata from each generic portal lead's proved local record instead of flattening it into residual text.
48
+ - Preserve exact visible page-owned prose from a broader list root when generic portal lead extraction narrows to a nested main owner.
49
+ - Retire the dedicated Danas article profile after shared WordPress extraction preserves the complete article, media, cited references, category tags, metadata, and clean source DOM.
50
+ - Preserve separation between directly adjacent visible tag, topic, and chip labels so compact link groups do not collapse into one Markdown token.
51
+ - Remove short publisher follow/subscribe notes only when exact reusable owner and sentence-level action evidence prove CTA furniture.
52
+ - Remove heading-only generic reply prompts from empty comment forms while preserving corrections, published comments, media, and source DOM.
53
+ - Build fallback article excerpts from the first substantial cleaned body paragraph instead of title, byline, related, or control chrome.
54
+ - Retire the Blic profile after shared article, author, metadata, homepage-list, recommendation, audio, and empty-ad handling preserve or improve its visible output.
55
+ - Preserve collected site metadata in shared fallback article extraction instead of replacing it with the request hostname.
56
+ - Remove only structurally empty article ad placeholders while preserving referenced, styled, shadow-owned, nested, or otherwise material content and the source DOM.
57
+ - Record scripted and declarative shadow roots so generic composed-DOM extraction preserves visible closed-component content without mutating source DOM.
58
+ - Invoke the browser extraction bundle through a private evaluated API so page-owned compatibility globals cannot intercept extraction options.
59
+ - Remove only bounded imperative prompts from exact article audio-player components while preserving episode titles, transcripts, media, controls, and source DOM.
60
+ - Apply proved terminal article-recommendation cleanup to shared fallback extraction when reader mode is disabled.
61
+ - Preserve a visible focal article's exact localized author-profile destination when reader mode keeps the author name but omits its locally owned link.
62
+ - Remove exact terminal `news-box` link collections during final article rendering only when a multilingual recommendation heading and substantial focal prose prove related-news ownership, while preserving source-owned prose, resources, structured content, nested articles, and nonterminal sections.
63
+ - Keep multi-section or otherwise proved portal root lists distinct from newsletter or digest formats.
64
+ - Handle Hindustan Times articles through shared reader extraction while preserving their visible byline and publication time.
65
+ - Preserve a verified visible focal-article byline when generic reader extraction omits its unlinked source node.
66
+ - Remove explicit article comment forms and zero-count continuation controls while preserving material comments, replies, and nonempty discussion links.
67
+ - Handle Trend articles through shared WordPress or generic article extraction without host-specific slug-warning suppression.
68
+ - Handle Protothema articles through shared reader extraction while preserving the complete visible body, lead media, metadata, warnings, and source DOM.
69
+ - Remove short terminal exact article-flow-note widgets only after a substantial article body, while preserving linked, attributed, repeated, nonterminal, or structurally substantive editorial notes.
70
+ - Handle Zeit articles through shared reader extraction while preserving the visible body, complete author identity, author destination, headings, warnings, and source DOM.
71
+
72
+ - Preserve a verified focal article's safe author-profile destination when reader mode expands an opaque author handle to its full structured name.
73
+
74
+ - Remove short browser-capability fallback messages from proved article audio-player components while preserving labels, transcripts, media, controls, and explanatory prose.
75
+
76
+ - Expand very short Readability excerpts from substantial parsed article bodies instead of returning isolated disclaimers as summaries.
77
+ - Remove bounded external-content consent placeholders through exact component markers while preserving substantive reporting and real embedded media.
78
+ - Prefer one title-owned nested article over a broader localized story rail only when the wrapper's remaining material is fully proved to be repeated linked records, using parent-chain indexes rather than pairwise fallback-candidate scans.
79
+ - Remove exact `ads` class-token furniture from generic article clones while preserving substantive editorial prose and semantic article owners.
80
+ - Handle Kaler Kantho articles through shared article extraction while preserving their body, classification, warnings, and source DOM.
81
+ - Handle Blick articles and live tickers through shared article extraction while preserving summaries, updates, authors, warnings, and source DOM.
82
+ - Handle Kurir articles through shared reader extraction while preserving the visible headline, lead media, standfirst, body, and metadata.
83
+ - Treat explicitly named section-link groups as secondary navigation without hiding similarly named story containers.
84
+ ### Fixed
85
+
86
+ - Keep article-scale homepage prose out of compact list supplements while retaining independently owned lead controls and short context.
87
+ - Preserve explicit source-owned article summaries before expanding reader excerpts from generic body prose.
88
+ - Remove publisher follow and subscribe furniture before adding reader-excerpt provenance attributes, without trusting page-authored marker lookalikes.
89
+
90
+ - Prefer a substantially more informative metadata author name to an opaque short reader-mode handle while preserving meaningful visible bylines and mononyms.
91
+ - Remove short, structurally identified article audio-control bars without discarding media, transcripts, episode metadata, or podcast/catalog content.
92
+ - Recover one unambiguous visible article lead and its immediately attached image when Readability selects the later body from the same article owner.
93
+ - Recognize localized author-profile destinations as focal article bylines without borrowing authors from related sidebars.
94
+ - Handle Onet's homepage through shared list extraction while preserving visible sections, stories, aliases, authors, and ordering.
95
+ - Keep a materially different visible story alias linked when only tracking parameters distinguish it from an already selected destination.
96
+ - Preserve a locally authored visible headline when a sectioned homepage also links the same story under a different title.
97
+ - Keep a story author owned by the selected nearest semantic list record even when its inner link is chosen as the record card.
98
+ - Remove structurally proved dropdown actions from list section headings without deleting material change-related titles or selected values.
99
+ - Keep linked collection headings out of broad record owners when multiple independent semantic records prove the collection boundary.
100
+ - Exclude list records owned by explicit advertisement tokens without treating ordinary substring lookalikes as ads.
101
+ - Preserve an unambiguous local author when duplicate card layouts represent the same story.
102
+ - Keep authors on homepage story cards out of page-level bylines while retaining a focal homepage article's author.
103
+ - Do not infer liveblog format from repeated timestamps on a root homepage list without explicit liveblog evidence.
104
+ - Preserve Markdown block boundaries when removing empty image and link placeholders.
105
+ - Preserve linked and unlinked collection headings through complete flat-list coverage, while keeping peer-proven one-link semantic cards as records instead of promoting their titles to section labels.
106
+ - Keep explicit article bylines attached to their local stories even when author-link classes resemble cards, while preserving author-directory records and excluding comment authors.
107
+ - Preserve unlinked collection headings when unrelated record metadata repeats the same label, using local heading ownership rather than global text equality.
108
+ - Handle Glassdoor-style community landing pages through shared extraction, preserving the visible summary without leaking collapsed directory panels.
109
+ - Exclude content inside zero-sized explicitly clipping panels while preserving visible overflow, scrollable collections and controlled-panel visibility contracts.
110
+ - Keep nested article authors and other record fields out of an enclosing collection link's metadata while retaining the fields on their actual stories.
111
+ - Use shared article extraction for Index.hr, retaining the full article's headings, metadata, images and citations while removing only proved empty interface widgets.
112
+ - Remove empty article-owned recommendation and comment loaders and short support-action widgets while retaining substantive discussions, explanations and citations.
113
+ - Prefer a selected article's own heading when the metadata title differs only by a verified publication-name suffix.
114
+ - Handle HighWire-style abstract and bodymatter articles through the shared scholarly engines, preserving their body, title and citation resources without a duplicate publisher registration.
115
+ - Use shared CMS extraction for Avesta articles so an unrelated news-card heading cannot replace the requested article title.
116
+ - Handle Hatena-style blog entries through the shared article engine, preserving owned content and metadata without a dedicated hostname profile.
117
+ - Preserve article headings, links and block structure when enriching sports, property and product content with structured details.
118
+ - Preserve visible children of translated overflow-visible containers instead of clipping an entire carousel by its track's border box.
119
+ - Use word segmentation for non-space-delimited headlines so meaningful Chinese, Thai and related-script collections are not mistaken for navigation.
120
+ - Keep visible topic and category destinations as local list context when they are not already represented by a record or description.
121
+ - Preserve named title fields in record-owned headers even when publishers use divs or spans instead of heading tags.
122
+ - Exclude fully clipped horizontal carousel content while preserving partially visible cards, scrollable collections and complete code examples.
123
+ - Keep mixed image-backed and plain sibling headlines local to their own records rather than repeating the surrounding collection as detail.
124
+ - Keep alignment-only layout wrappers distinct from story cards, preserving independent headlines and local record ownership.
125
+ - Keep article/list arbitration independent of added list presentation context, preserving substantive articles beside smaller directories.
126
+ - Recognize editorial columns split across sibling one-story asides without admitting hidden or unowned panels.
127
+ - Preserve explicit card headings when their lazy media remains hidden.
128
+ - Preserve section-owned header labels and links while removing their navigation controls.
129
+ - Preserve named overlay destinations that match a visible card headline without confusing nested follow-up stories with the primary record.
130
+ - Preserve primary story headings inside card-owned headers while retaining site navigation cleanup.
131
+ - Preserve short linked category headings and their destinations when they own an extracted collection.
132
+ - Preserve proved homepage editorial columns through list cleanup, keeping visible headline collections and their local context despite sidebar presentation.
133
+ - Preserve locally owned collection headings through header cleanup and render selected list context in source order with consistent record-reference ownership.
134
+ - Keep document-level layout state distinct from navigation when extracting cloned pages, preserving records beneath header or sidebar presentation classes.
135
+ - Match abbreviated consent-vendor namespaces as class or id tokens, preserving ordinary sticky layouts and CSS-variable references during cookie cleanup.
136
+ - Extract visible open-component content and rendered slots through shared DOM cloning, preserving nested stories, article text and references without exposing hidden or unassigned content.
137
+ - Use shared extraction for WP homepages, preserving untitled lead grids and complete section context instead of the narrower host profile.
138
+ - Preserve sparse section headings and unlinked notices when complete flat records supplement a partial sectioned list.
139
+ - Match actual map-library containers during cleanup, preserving unrelated leaflet cards and roadmap sections.
140
+ - Preserve code-local blank lines and literal markup through Markdown normalization, including longer and nested fences.
141
+ - Preserve fenced samples inside inline-code wrappers instead of flattening their lines.
142
+ - Preserve qualified software download and documentation groups despite header, sidebar, or call-to-action presentation classes.
143
+ - Keep text links valid when decorative icons are wrapped in block-level figures.
144
+ - Preserve accessible names and destinations on icon-only resource links while retaining ordinary control cleanup.
145
+ - Preserve named download artifacts and supporting HTTP links in ordinary Markdown tables.
146
+ - Preserve prose and list instructions containing formatted inline code, including legacy typewriter examples, during JavaScript and debug-noise cleanup.
147
+ - Preserve compact image-led list records when converting standalone heading-based resource cards.
148
+ - Keep pure documentation-link indexes distinct from software project overviews by requiring locally owned descriptive prose.
149
+ - Preserve locally owned software release versions and release-note links embedded in project footers.
150
+ - Keep heading-based resource cards linked while preserving their heading levels, paragraphs, and images.
151
+ - Recognize software overview prose in lists and inline layout blocks, preserving owned feature sections across heading levels.
152
+ - Require badge-specific evidence before removing short linked images, preserving named editor and project resources.
153
+ - Preserve resource destinations when linked cards contain block-formatted labels.
154
+ - Preserve page-owned preformatted examples during generic homepage and index-list arbitration, keeping instructional prose and code together while retaining real linked-record feeds.
155
+ - Recover short commands and omitted instructional context from a visibility-cleaned fallback when it preserves the selected article's ordered text, resources, and exact code, retaining its metadata and reader-mode provenance.
156
+ - Preserve syntax-highlighted identifiers and literal code text during shared UI/docs cleanup, including `print`, `copy`, and `help`, while still removing actual copy and print controls.
157
+ - Keep sectioned directories in list mode when their complete page-owned descriptions are rendered after classification; restrict the new late-portal protection to code coverage.
158
+ - Preserve complete materialized code-editor examples, block-based line breaks, and highlighted code before article cleanup, without collapsing multiple examples into the first code block.
159
+ - Extract evidenced software-project homepages as complete documentation overviews, preserving feature prose, examples, releases, and resource lists together instead of reducing them to unrelated links.
160
+
161
+ ### Known Limitations
162
+
163
+ - A fresh immutable-DOM comparison of 300 websites against published v0.6.2 found 121 improvements, 148 ties, no regressions, 28 mixed outcomes, and three inconclusive cases after repair. Mixed and inconclusive outcomes are not wins, so this does not establish a universal strict improvement.
164
+
5
165
  ## v0.6.2 - 2026-09-10
6
166
 
7
167
  ### Changed
data/README.md CHANGED
@@ -184,7 +184,8 @@ pp FetchUtil.regulatory(
184
184
 
185
185
  ## Behavior
186
186
 
187
- - Extracts articles, list/index pages, and search pages into compact markdown.
187
+ - Extracts articles, list/index pages, search pages, glossaries, and structured detail content into compact markdown.
188
+ - Handles dynamic pages, SPAs, and open or closed component content in a real browser.
188
189
  - Uses page classification to select extraction logic appropriate to the rendered page type.
189
190
  - Detects consent prompts, login-required pages, and challenge/interstitial screens and reports them with concise summaries and warning tags. A delivered JavaScript/WebAssembly proof-of-work may complete in the normal browser session within `timeout`; unresolved challenges remain explicit interstitials.
190
191
  - Cleans up docs/reference pages aggressively enough for agent consumption.