pydfdoi 0.1.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,24 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ *$py.class
4
+
5
+ .pytest_cache/
6
+ .coverage
7
+ htmlcov/
8
+
9
+ .venv/
10
+ venv/
11
+ env/
12
+
13
+ build/
14
+ dist/
15
+ *.egg-info/
16
+
17
+ .mypy_cache/
18
+ .ruff_cache/
19
+ .pyright/
20
+
21
+ .DS_Store
22
+ Thumbs.db
23
+
24
+ *.pdf
@@ -0,0 +1,48 @@
1
+ # Changelog
2
+
3
+ All notable changes to `pydfdoi` are documented in this file.
4
+
5
+ The project follows semantic versioning before `1.0.0` with the usual
6
+ pre-1.0 caveat: public APIs may still change while the package is stabilizing.
7
+
8
+ ## [0.1.1] - 2026-10-01
9
+
10
+ ### Changed
11
+
12
+ - Renamed the public package, import package, CLI command, and GitHub repository
13
+ target from `pdfdoi` to `pydfdoi` after PyPI rejected `pdfdoi` as too similar
14
+ to an existing project name.
15
+ - Kept `pdfdoi` as a compatibility console command.
16
+
17
+ ## [0.1.0] - 2026-10-01
18
+
19
+ ### Added
20
+
21
+ - Initial package structure using the `src/pydfdoi` layout.
22
+ - Command-line entry point: `pydfdoi`.
23
+ - DOI extraction from PDF metadata and the first pages of PDF text.
24
+ - Command-line DOI scanning defaults to 20 pages to better support ebooks.
25
+ - PDF metadata writing for `/doi`, `/doiURL`, `/Title`, viewer title behavior,
26
+ and first-page opening.
27
+ - Python 3.14 is the target development and runtime environment.
28
+ - `pikepdf` is the primary PDF writing backend for metadata and page-label
29
+ updates, with `pypdf` retained for text extraction and verification.
30
+ - `capture_and_write_doi_metadata()` for metadata-only PDF cleanup without
31
+ renaming.
32
+ - Crossref lookup support for DOI metadata.
33
+ - Citation-based PDF renaming with filenames such as
34
+ `Author et al. Year Title.pdf`.
35
+ - Journal article PDF page-label alignment from Crossref page metadata.
36
+ - Journal/book/general DOI metadata capture functions:
37
+ `capture_journal_doi()`, `capture_book_doi()`, and `capture_doi_metadata()`.
38
+ - Curated publisher helper table with common DOI prefixes, abbreviations, and
39
+ aliases.
40
+ - John Benjamins DOI prefix support.
41
+ - Windows context-menu and SendTo helper scripts.
42
+ - Test suite covering DOI normalization, citation filenames, publisher helpers,
43
+ DOI metadata capture, and PDF metadata writing.
44
+
45
+ ### Notes
46
+
47
+ - Author: Yihtsy <yihtsy@outlook.com>
48
+ - This is the first public-ready development release.
pydfdoi-0.1.1/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Yihtsy
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
pydfdoi-0.1.1/PKG-INFO ADDED
@@ -0,0 +1,392 @@
1
+ Metadata-Version: 2.5
2
+ Name: pydfdoi
3
+ Version: 0.1.1
4
+ Summary: Extract DOI values from PDFs, write DOI metadata, and rename papers from Crossref citation data.
5
+ Project-URL: Homepage, https://github.com/Yihtsy/pydfdoi
6
+ Project-URL: Repository, https://github.com/Yihtsy/pydfdoi
7
+ Project-URL: Issues, https://github.com/Yihtsy/pydfdoi/issues
8
+ Project-URL: Changelog, https://github.com/Yihtsy/pydfdoi/blob/main/CHANGELOG.md
9
+ Author-email: Yihtsy <yihtsy@outlook.com>
10
+ License: MIT License
11
+
12
+ Copyright (c) 2026 Yihtsy
13
+
14
+ Permission is hereby granted, free of charge, to any person obtaining a copy
15
+ of this software and associated documentation files (the "Software"), to deal
16
+ in the Software without restriction, including without limitation the rights
17
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
18
+ copies of the Software, and to permit persons to whom the Software is
19
+ furnished to do so, subject to the following conditions:
20
+
21
+ The above copyright notice and this permission notice shall be included in all
22
+ copies or substantial portions of the Software.
23
+
24
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
25
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
26
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
27
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
28
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
29
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
30
+ SOFTWARE.
31
+ License-File: LICENSE
32
+ Keywords: crossref,doi,metadata,pdf,rename
33
+ Classifier: Development Status :: 3 - Alpha
34
+ Classifier: Environment :: Console
35
+ Classifier: Intended Audience :: Science/Research
36
+ Classifier: Programming Language :: Python :: 3
37
+ Classifier: Programming Language :: Python :: 3 :: Only
38
+ Classifier: Programming Language :: Python :: 3.14
39
+ Classifier: Topic :: Scientific/Engineering :: Information Analysis
40
+ Classifier: Topic :: Utilities
41
+ Requires-Python: >=3.14
42
+ Requires-Dist: pikepdf>=10
43
+ Requires-Dist: pypdf>=5
44
+ Requires-Dist: requests>=2.31
45
+ Provides-Extra: dev
46
+ Requires-Dist: build>=1.2; extra == 'dev'
47
+ Requires-Dist: pytest>=8; extra == 'dev'
48
+ Requires-Dist: twine>=5; extra == 'dev'
49
+ Description-Content-Type: text/markdown
50
+
51
+ # pydfdoi
52
+
53
+ `pydfdoi` is a small Python package for DOI-aware PDF cleanup. It extracts DOI
54
+ values from PDF metadata or PDF text, writes DOI-related PDF metadata, and can
55
+ rename academic PDFs from Crossref citation data.
56
+
57
+ Author: Yihtsy <yihtsy@outlook.com>
58
+
59
+ ## Features
60
+
61
+ - Extract DOI values from existing PDF metadata and the first pages of text.
62
+ - Write `/doi` and `/doiURL` into PDF metadata.
63
+ - Set `/Title` to the file name for metadata-only cleanup.
64
+ - Configure PDFs to open on the first page.
65
+ - Ask PDF viewers to show the file name in the title bar.
66
+ - Query Crossref metadata for DOI records.
67
+ - Rename journal article PDFs as `Author et al. Year Title.pdf`.
68
+ - Align journal article PDF page labels with Crossref page metadata.
69
+ - Capture journal DOI metadata, book DOI metadata, or general DOI metadata.
70
+ - Look up common publisher abbreviations and DOI prefixes.
71
+ - Install optional Windows context-menu and SendTo shortcuts.
72
+
73
+ ## Package Name
74
+
75
+ The project uses one consistent package name:
76
+
77
+ - PyPI distribution: `pydfdoi`
78
+ - Python import: `pydfdoi`
79
+ - Console command: `pydfdoi`
80
+ - Compatibility command: `pdfdoi`
81
+
82
+ This avoids the common mismatch where a package is installed with a hyphenated
83
+ name but imported with an underscore.
84
+
85
+ ## Installation
86
+
87
+ After the package is published to PyPI:
88
+
89
+ ```powershell
90
+ python -m pip install pydfdoi
91
+ ```
92
+
93
+ For local development from a cloned repository:
94
+
95
+ ```powershell
96
+ C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m venv .venv
97
+ .\.venv\Scripts\Activate.ps1
98
+ python -m pip install -U pip
99
+ python -m pip install -e ".[dev]"
100
+ ```
101
+
102
+ Editable installation means changes under `src\pydfdoi` are immediately used by
103
+ the command line and tests. Do not develop directly inside Python's
104
+ `site-packages`; keep the source in a normal project folder and use editable
105
+ installation.
106
+
107
+ ## Command Line
108
+
109
+ Update DOI metadata only:
110
+
111
+ ```powershell
112
+ pydfdoi --mode metadata-only "D:\path\paper.pdf"
113
+ ```
114
+
115
+ Rename one or more PDFs from Crossref citation metadata:
116
+
117
+ ```powershell
118
+ pydfdoi --mode rename "D:\path\paper1.pdf" "D:\path\paper2.pdf"
119
+ ```
120
+
121
+ Scan more pages for DOI text:
122
+
123
+ ```powershell
124
+ pydfdoi --mode metadata-only --max-pages 5 "D:\path\paper.pdf"
125
+ ```
126
+
127
+ The command-line default is 20 pages, which is friendlier to ebooks whose DOI
128
+ often appears on the title or copyright pages rather than the first page.
129
+
130
+ Preview actions without writing changes:
131
+
132
+ ```powershell
133
+ pydfdoi --mode rename --dry-run "D:\path\paper.pdf"
134
+ ```
135
+
136
+ Print JSON output:
137
+
138
+ ```powershell
139
+ pydfdoi --mode metadata-only --json "D:\path\paper.pdf"
140
+ ```
141
+
142
+ Align PDF page labels with Crossref journal article page metadata:
143
+
144
+ ```powershell
145
+ pydfdoi --mode page-labels "D:\path\article.pdf"
146
+ ```
147
+
148
+ If the PDF has extra non-article pages, they are labeled with `skip-` by
149
+ default. Extra pages are assumed to be at the front unless you choose another
150
+ placement:
151
+
152
+ ```powershell
153
+ pydfdoi --mode page-labels --extra-placement front "D:\path\article.pdf"
154
+ pydfdoi --mode page-labels --extra-placement back "D:\path\article.pdf"
155
+ pydfdoi --mode page-labels --extra-placement split "D:\path\article.pdf"
156
+ ```
157
+
158
+ You can customize the special label prefix:
159
+
160
+ ```powershell
161
+ pydfdoi --mode page-labels --skipped-label-prefix "nonarticle-" "D:\path\article.pdf"
162
+ ```
163
+
164
+ If no file is passed, the command opens a file picker.
165
+
166
+ ## Python API
167
+
168
+ ### Metadata-only cleanup
169
+
170
+ Use `capture_and_write_doi_metadata()` when you only want to capture DOI
171
+ information and update the PDF metadata.
172
+
173
+ ```python
174
+ from pydfdoi import capture_and_write_doi_metadata
175
+
176
+ doi = capture_and_write_doi_metadata("paper.pdf")
177
+ ```
178
+
179
+ This function:
180
+
181
+ - extracts the DOI from the PDF;
182
+ - writes `/doi`;
183
+ - writes `/doiURL`;
184
+ - sets `/Title` to the file name without `.pdf`;
185
+ - configures the PDF to open on the first page;
186
+ - does not query Crossref;
187
+ - does not rename the file.
188
+
189
+ ### Rename by citation
190
+
191
+ ```python
192
+ from pathlib import Path
193
+
194
+ from pydfdoi import extract_doi, rename_by_citation
195
+
196
+ path = Path("paper.pdf")
197
+ doi = extract_doi(path)
198
+ if doi:
199
+ renamed_path = rename_by_citation(path, doi)
200
+ ```
201
+
202
+ ### Journal, book, and general DOI capture
203
+
204
+ ```python
205
+ from pydfdoi import capture_book_doi, capture_doi_metadata, capture_journal_doi
206
+
207
+ journal = capture_journal_doi("article.pdf")
208
+ book = capture_book_doi("book.pdf")
209
+ any_work = capture_doi_metadata("unknown.pdf")
210
+ ```
211
+
212
+ - `capture_journal_doi()` accepts Crossref journal-like work types.
213
+ - `capture_book_doi()` accepts Crossref book-like work types.
214
+ - `capture_doi_metadata()` is the general DOI metadata entry point.
215
+
216
+ `capture_book_doi()` scans more pages by default than the journal helper because
217
+ books often put DOI information on title, copyright, or series pages rather than
218
+ the first page. Its default is 20 pages:
219
+
220
+ ```python
221
+ book = capture_book_doi("book.pdf")
222
+ book = capture_book_doi("book.pdf", max_pages=30)
223
+ book = capture_book_doi("book.pdf", strict_type=True)
224
+ ```
225
+
226
+ When `strict_type=False`, which is the default, `capture_book_doi()` can still
227
+ accept a DOI if Crossref returns an unknown work type but the page-count
228
+ heuristic says the PDF is book-like. If Crossref clearly identifies the work as
229
+ journal-like, it is still rejected by the book helper.
230
+
231
+ The general entry point can also infer a coarse work category:
232
+
233
+ ```python
234
+ result = capture_doi_metadata("unknown.pdf", article_page_limit=80)
235
+
236
+ print(result.inferred_kind) # "journal", "book", or None
237
+ print(result.classification_source) # "page-count", "crossref", or None
238
+ print(result.page_count)
239
+ ```
240
+
241
+ By default, `capture_doi_metadata()` first makes a simple page-count guess:
242
+ PDFs with 80 pages or fewer are treated as journal-like, and longer PDFs are
243
+ treated as book-like. If Crossref returns a known work type, the Crossref type
244
+ overrides the page-count guess.
245
+
246
+ ### Publisher helpers
247
+
248
+ ```python
249
+ from pydfdoi import PUBLISHERS, publisher_for_doi, publisher_for_name
250
+
251
+ publisher = publisher_for_doi("10.1038/s41586-020-2649-2")
252
+ publisher = publisher_for_name("John Wiley & Sons")
253
+ ```
254
+
255
+ The publisher table is a curated starter list of common publishers, DOI
256
+ prefixes, abbreviations, and aliases. It is intentionally not exhaustive:
257
+ DOI prefixes are assigned to registrants and may remain in use after publisher
258
+ mergers, platform changes, or imprint transfers.
259
+
260
+ ### Journal article page labels
261
+
262
+ Use `align_journal_article_page_labels()` to align PDF page labels with Crossref
263
+ article page metadata.
264
+
265
+ ```python
266
+ from pydfdoi import align_journal_article_page_labels
267
+
268
+ plan = align_journal_article_page_labels(
269
+ "article.pdf",
270
+ extra_placement="front",
271
+ skipped_label_prefix="skip-",
272
+ )
273
+ ```
274
+
275
+ For a Crossref page range such as `10-12`, a four-page PDF is labeled as:
276
+
277
+ ```text
278
+ skip-1, 10, 11, 12
279
+ ```
280
+
281
+ The function can also use Crossref `article-number` when a conventional page
282
+ range is absent. For article-number-only records, labels look like `e12345-1`,
283
+ `e12345-2`, and so on.
284
+
285
+ Because Crossref usually does not say whether extra PDF pages are at the front
286
+ or the back, `extra_placement` is explicit:
287
+
288
+ - `front`: extra pages are before the article content.
289
+ - `back`: extra pages are after the article content.
290
+ - `split`: extra pages are split between front and back.
291
+
292
+ ## Windows Context Menu
293
+
294
+ The repository includes compatibility scripts for Windows Explorer workflows.
295
+
296
+ Run PowerShell in the project folder:
297
+
298
+ ```powershell
299
+ .\install_context_menu.ps1
300
+ ```
301
+
302
+ This adds two PDF right-click entries:
303
+
304
+ - `PDF DOI: rename by citation`
305
+ - `PDF DOI: metadata only`
306
+
307
+ It also installs SendTo shortcuts, which are usually the most reliable route
308
+ for processing multiple selected PDFs in Windows Explorer.
309
+
310
+ To uninstall:
311
+
312
+ ```powershell
313
+ .\install_context_menu.ps1 -Uninstall
314
+ ```
315
+
316
+ `pdf_doi_tool.py` is kept as a compatibility wrapper for these scripts and for
317
+ older local workflows.
318
+
319
+ ## Development
320
+
321
+ Project layout:
322
+
323
+ ```text
324
+ src/
325
+ pydfdoi/
326
+ __init__.py
327
+ citation.py
328
+ cli.py
329
+ crossref.py
330
+ doi.py
331
+ metadata.py
332
+ page_labels.py
333
+ pdf.py
334
+ processing.py
335
+ publishers.py
336
+ tests/
337
+ ```
338
+
339
+ Run tests:
340
+
341
+ ```powershell
342
+ C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m pytest
343
+ ```
344
+
345
+ Build local distributions:
346
+
347
+ ```powershell
348
+ C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m build
349
+ ```
350
+
351
+ Check built distributions before publishing:
352
+
353
+ ```powershell
354
+ C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m twine check dist/*
355
+ ```
356
+
357
+ Publisher-specific observations and heuristic notes live in
358
+ [docs/development-notes.md](docs/development-notes.md).
359
+
360
+ ## Release Notes
361
+
362
+ Current version: `0.1.1`
363
+
364
+ See [CHANGELOG.md](CHANGELOG.md) for the full version history.
365
+
366
+ ## Publishing
367
+
368
+ Before publishing a new release:
369
+
370
+ 1. Update `version` in `pyproject.toml`.
371
+ 2. Update `__version__` in `src/pydfdoi/__init__.py`.
372
+ 3. Add a new entry to `CHANGELOG.md`.
373
+ 4. Run `python -m pytest`.
374
+ 5. Run `python -m build`.
375
+ 6. Run `python -m twine check dist/*`.
376
+ 7. Tag the release in Git.
377
+ 8. Upload to PyPI with `python -m twine upload dist/*`.
378
+
379
+ For the first PyPI upload, create an account-wide API token on PyPI and run the
380
+ local helper:
381
+
382
+ ```powershell
383
+ .\scripts\publish_pypi.ps1 -SkipBuild
384
+ ```
385
+
386
+ The helper prompts for the token without echoing it, sets Twine credentials only
387
+ for that process, uploads the built `0.1.0` distributions, and clears the
388
+ temporary environment variables afterward.
389
+
390
+ ## License
391
+
392
+ `pydfdoi` is released under the MIT License. See [LICENSE](LICENSE).