pdfua 0.1.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- pdfua-0.1.1/.gitignore +35 -0
- pdfua-0.1.1/CHANGELOG.md +52 -0
- pdfua-0.1.1/CONTRIBUTING.md +85 -0
- pdfua-0.1.1/LICENSE +21 -0
- pdfua-0.1.1/PKG-INFO +353 -0
- pdfua-0.1.1/README.md +321 -0
- pdfua-0.1.1/pyproject.toml +78 -0
- pdfua-0.1.1/src/pdfua/__init__.py +46 -0
- pdfua-0.1.1/src/pdfua/__main__.py +10 -0
- pdfua-0.1.1/src/pdfua/catalog.py +185 -0
- pdfua-0.1.1/src/pdfua/cli.py +278 -0
- pdfua-0.1.1/src/pdfua/data/__init__.py +1 -0
- pdfua-0.1.1/src/pdfua/data/pdfua1_rules.json +537 -0
- pdfua-0.1.1/src/pdfua/document.py +511 -0
- pdfua-0.1.1/src/pdfua/errors.py +19 -0
- pdfua-0.1.1/src/pdfua/model.py +207 -0
- pdfua-0.1.1/src/pdfua/py.typed +0 -0
- pdfua-0.1.1/src/pdfua/reporters.py +230 -0
- pdfua-0.1.1/src/pdfua/rules/__init__.py +39 -0
- pdfua-0.1.1/src/pdfua/rules/base.py +119 -0
- pdfua-0.1.1/src/pdfua/rules/language.py +140 -0
- pdfua-0.1.1/src/pdfua/rules/structure.py +221 -0
- pdfua-0.1.1/src/pdfua/rules/tagging.py +142 -0
- pdfua-0.1.1/src/pdfua/rules/wcag.py +125 -0
- pdfua-0.1.1/src/pdfua/validator.py +99 -0
- pdfua-0.1.1/tests/__init__.py +0 -0
- pdfua-0.1.1/tests/conftest.py +64 -0
- pdfua-0.1.1/tests/fixtures/README.md +61 -0
- pdfua-0.1.1/tests/fixtures/__init__.py +0 -0
- pdfua-0.1.1/tests/fixtures/build.py +310 -0
- pdfua-0.1.1/tests/test_cli.py +205 -0
- pdfua-0.1.1/tests/test_coverage.py +133 -0
- pdfua-0.1.1/tests/test_document.py +216 -0
- pdfua-0.1.1/tests/test_model.py +170 -0
- pdfua-0.1.1/tests/test_real_documents.py +133 -0
- pdfua-0.1.1/tests/test_reporters.py +147 -0
- pdfua-0.1.1/tests/test_rules_contract.py +118 -0
pdfua-0.1.1/.gitignore
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Python
|
|
2
|
+
__pycache__/
|
|
3
|
+
*.py[cod]
|
|
4
|
+
*.egg-info/
|
|
5
|
+
.eggs/
|
|
6
|
+
build/
|
|
7
|
+
dist/
|
|
8
|
+
.venv/
|
|
9
|
+
venv/
|
|
10
|
+
.pytest_cache/
|
|
11
|
+
.mypy_cache/
|
|
12
|
+
.ruff_cache/
|
|
13
|
+
.coverage
|
|
14
|
+
coverage.xml
|
|
15
|
+
htmlcov/
|
|
16
|
+
|
|
17
|
+
# Secrets — never commit these
|
|
18
|
+
.env
|
|
19
|
+
.env.*
|
|
20
|
+
*.key
|
|
21
|
+
*.pem
|
|
22
|
+
*.p12
|
|
23
|
+
*.pfx
|
|
24
|
+
secrets/
|
|
25
|
+
|
|
26
|
+
# Downloaded test corpora (regenerate with scripts/fetch_corpus.py)
|
|
27
|
+
corpus/
|
|
28
|
+
tests/fixtures/downloaded/
|
|
29
|
+
*.pdf
|
|
30
|
+
!tests/fixtures/**/*.pdf
|
|
31
|
+
|
|
32
|
+
# Editors / OS
|
|
33
|
+
.DS_Store
|
|
34
|
+
.idea/
|
|
35
|
+
.vscode/
|
pdfua-0.1.1/CHANGELOG.md
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented in this file.
|
|
4
|
+
|
|
5
|
+
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
6
|
+
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
|
+
|
|
8
|
+
## [Unreleased]
|
|
9
|
+
|
|
10
|
+
## [0.1.1] - 2026-10-04
|
|
11
|
+
|
|
12
|
+
### Changed
|
|
13
|
+
|
|
14
|
+
- Packaging and release metadata only; no code or behaviour changes. Added
|
|
15
|
+
complete PyPI metadata (PEP 639 `license` expression, author, project URLs
|
|
16
|
+
and classifiers) and a Trusted Publishing (OIDC) release workflow that
|
|
17
|
+
publishes on `v*` tags.
|
|
18
|
+
|
|
19
|
+
## [0.1.0] - 2026-09-29
|
|
20
|
+
|
|
21
|
+
Initial release.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- `pdfua check <file>` with `--format text|json|sarif`, `--rules`,
|
|
26
|
+
`--min-severity`, `--certain-only`, `--fail-fast`, `--quiet`.
|
|
27
|
+
- `pdfua rules` listing implemented rules and the clauses with no coverage.
|
|
28
|
+
- 12 checks covering 16 of the 106 machine-checkable PDF/UA-1 rules (15%):
|
|
29
|
+
- `UA-01-005` structure tree present
|
|
30
|
+
- `UA-01-002` `/MarkInfo` `/Marked` true
|
|
31
|
+
- `UA-01-003` page content tagged or marked as artifact
|
|
32
|
+
- `UA-06-001` document language
|
|
33
|
+
- `UA-07-001` document title
|
|
34
|
+
- `UA-07-002` `DisplayDocTitle`
|
|
35
|
+
- `UA-13-004` figure alternative text
|
|
36
|
+
- `UA-14-001` non-empty headings
|
|
37
|
+
- `UA-15-003` determinable table structure
|
|
38
|
+
- `UA-18-001` annotation descriptions
|
|
39
|
+
- `UA-18-005` link descriptions
|
|
40
|
+
- `UA-28-004` PDF/UA identification in XMP
|
|
41
|
+
- SARIF 2.1.0 output for GitHub code scanning and Azure DevOps.
|
|
42
|
+
- Severity-ordered exit codes (0 clean, 1 info, 2 warning, 3 error).
|
|
43
|
+
- Full type annotations, `py.typed` marker, strict `mypy` configuration.
|
|
44
|
+
|
|
45
|
+
### Notes
|
|
46
|
+
|
|
47
|
+
- This is 12 checks covering 16 of the 106 machine-checkable rules in the
|
|
48
|
+
PDF/UA-1 profile. The other 90 are not checked. See the README section
|
|
49
|
+
"What this does NOT check".
|
|
50
|
+
- Development was validated against veraPDF 1.30.2 on a 16-file corpus (15
|
|
51
|
+
distinct documents) of real `govinfo.gov`, `irs.gov` and `ada.gov` PDFs. See
|
|
52
|
+
`docs/FALSIFICATION.md` for the evidence.
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
# Contributing to pdfua
|
|
2
|
+
|
|
3
|
+
Thanks for looking. This project has one governing value: **do not overclaim.**
|
|
4
|
+
|
|
5
|
+
A validator that says "PASS" is making a claim about everything it did not check.
|
|
6
|
+
Every design decision here follows from that.
|
|
7
|
+
|
|
8
|
+
## The one rule
|
|
9
|
+
|
|
10
|
+
**If you add a check, you must declare exactly which PDF/UA-1 rule identifiers it
|
|
11
|
+
decides, and you must add a fixture that proves it fires.**
|
|
12
|
+
|
|
13
|
+
A rule declares its coverage with the `covers` attribute:
|
|
14
|
+
|
|
15
|
+
```python
|
|
16
|
+
class MyRule(Rule):
|
|
17
|
+
id = "UA-99-001"
|
|
18
|
+
title = "Something decidable"
|
|
19
|
+
severity = Severity.ERROR
|
|
20
|
+
pdfua_clause = "7.9"
|
|
21
|
+
covers = ("7.9-1",) # clause-test, as enumerated in the PDF/UA-1 profile
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
`tests/test_coverage.py` fails if a declared identifier is not a real PDF/UA-1
|
|
25
|
+
rule, and `tests/test_rules_contract.py` fails if a registered rule has no
|
|
26
|
+
positive fixture. Both exist to stop the coverage number in the README from
|
|
27
|
+
drifting away from the truth.
|
|
28
|
+
|
|
29
|
+
## Adding a rule
|
|
30
|
+
|
|
31
|
+
1. Implement it in the appropriate `src/pdfua/rules/*.py` module.
|
|
32
|
+
2. Set `covers` to the identifiers it decides. If it does not correspond to a
|
|
33
|
+
single profile rule, leave `covers` empty — it then contributes nothing to the
|
|
34
|
+
coverage figure, which is the honest outcome.
|
|
35
|
+
3. Add a builder to `tests/fixtures/build.py` that changes **exactly one thing**
|
|
36
|
+
from `valid_document()`.
|
|
37
|
+
4. Add the rule to `RULE_CASES` in `tests/test_rules_contract.py`.
|
|
38
|
+
5. Run the suite.
|
|
39
|
+
|
|
40
|
+
## The fixture design rule
|
|
41
|
+
|
|
42
|
+
`valid_document()` satisfies every rule. Each `*_violating` builder changes one
|
|
43
|
+
thing. If your builder changes two, a test for either rule can pass for the wrong
|
|
44
|
+
reason, which is worse than no test.
|
|
45
|
+
|
|
46
|
+
## What will be rejected
|
|
47
|
+
|
|
48
|
+
- **A check that guesses.** If a rule needs human judgement, it must set
|
|
49
|
+
`confidence=Confidence.HEURISTIC` and say in its message that a human must
|
|
50
|
+
confirm. See `TableHeaderRule` for the pattern.
|
|
51
|
+
- **A check that is stricter than the standard without saying so.** If you flag
|
|
52
|
+
something the specification permits, the finding must state that explicitly.
|
|
53
|
+
See the whitespace-`/Alt` branch in `FigureAltTextRule`.
|
|
54
|
+
- **A rule that reports per-instance when one finding suffices.** Reporting an
|
|
55
|
+
untagged document's defect once per page buries the signal.
|
|
56
|
+
- **A silent skip.** A rule that cannot evaluate a document must raise
|
|
57
|
+
`PdfuaError` with a reason, so it appears in `report.skipped`. Never return
|
|
58
|
+
nothing and let it look like a pass.
|
|
59
|
+
|
|
60
|
+
## Development
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
uv venv
|
|
64
|
+
uv pip install -e ".[dev]"
|
|
65
|
+
.venv/bin/pytest
|
|
66
|
+
.venv/bin/mypy src/pdfua
|
|
67
|
+
.venv/bin/ruff check .
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
The suite must stay green, coverage at or above 80%, and `mypy --strict` must
|
|
71
|
+
pass with no `Any`. The package ships `py.typed`; annotations are part of the
|
|
72
|
+
public interface.
|
|
73
|
+
|
|
74
|
+
## Commits
|
|
75
|
+
|
|
76
|
+
Conventional Commits (`feat:`, `fix:`, `docs:`, `test:`, `refactor:`, `chore:`).
|
|
77
|
+
Small and focused. **Never use `--no-verify`** — if a hook fails, fix the cause.
|
|
78
|
+
|
|
79
|
+
## Reporting a false positive
|
|
80
|
+
|
|
81
|
+
This matters more than a missed detection. If the tool reports a problem in a
|
|
82
|
+
document that is actually fine, please open an issue with the document (or a
|
|
83
|
+
minimal reproduction) and the rule ID. A false positive is how accessibility
|
|
84
|
+
tooling loses its users' trust, and it is treated as a bug of the highest
|
|
85
|
+
severity.
|
pdfua-0.1.1/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 beduldul
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
pdfua-0.1.1/PKG-INFO
ADDED
|
@@ -0,0 +1,353 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: pdfua
|
|
3
|
+
Version: 0.1.1
|
|
4
|
+
Summary: Check the machine-checkable subset of PDF/UA-1 and WCAG 2.1 for PDFs in Python, built on pikepdf with no JVM
|
|
5
|
+
Project-URL: Homepage, https://github.com/beduldul/pdfua
|
|
6
|
+
Project-URL: Repository, https://github.com/beduldul/pdfua
|
|
7
|
+
Project-URL: Issues, https://github.com/beduldul/pdfua/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/beduldul/pdfua/blob/main/CHANGELOG.md
|
|
9
|
+
Author-email: Abdul Afif Al Kaysan <alkaysan07@gmail.com>
|
|
10
|
+
License-Expression: MIT
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Keywords: a11y,accessibility,ada,pdf,pdf-ua,pdfua,sarif,section508,validator,wcag
|
|
13
|
+
Classifier: Development Status :: 3 - Alpha
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Operating System :: OS Independent
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
21
|
+
Classifier: Topic :: Software Development :: Quality Assurance
|
|
22
|
+
Classifier: Topic :: Text Processing :: Markup
|
|
23
|
+
Classifier: Typing :: Typed
|
|
24
|
+
Requires-Python: >=3.10
|
|
25
|
+
Requires-Dist: pikepdf>=8.0
|
|
26
|
+
Provides-Extra: dev
|
|
27
|
+
Requires-Dist: mypy>=1.11; extra == 'dev'
|
|
28
|
+
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
|
|
29
|
+
Requires-Dist: pytest>=8.0; extra == 'dev'
|
|
30
|
+
Requires-Dist: ruff>=0.6; extra == 'dev'
|
|
31
|
+
Description-Content-Type: text/markdown
|
|
32
|
+
|
|
33
|
+
[](https://github.com/beduldul/pdfua/actions/workflows/ci.yml)
|
|
34
|
+
# pdfua
|
|
35
|
+
|
|
36
|
+
**Check the machine-checkable subset of PDF/UA-1 and WCAG 2.1 for PDFs — a small
|
|
37
|
+
Python package built on `pikepdf` (QPDF), with no JVM.**
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
pip install "git+https://github.com/beduldul/pdfua.git"
|
|
41
|
+
pdfua check report.pdf --format sarif
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**This implements 12 checks covering 16 of the 106 machine-checkable rules in the
|
|
45
|
+
PDF/UA-1 profile — 15%.** Read
|
|
46
|
+
[What this does NOT check](#what-this-does-not-check) before you rely on it. It is
|
|
47
|
+
a fast triage gate, not a conformance certifier. For a conformance claim, run
|
|
48
|
+
[veraPDF](https://verapdf.org).
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## The problem
|
|
53
|
+
|
|
54
|
+
There is a dated legal forcing function and no tool that fits a CI pipeline.
|
|
55
|
+
|
|
56
|
+
The U.S. Department of Justice's Title II rule requires state and local
|
|
57
|
+
governments to meet **WCAG 2.1 Level AA** for their web content and mobile apps.
|
|
58
|
+
PDFs are explicitly named as covered content. The compliance dates, from
|
|
59
|
+
`ada.gov`:
|
|
60
|
+
|
|
61
|
+
| State and local government size | Compliance date |
|
|
62
|
+
|---|---|
|
|
63
|
+
| 50,000 or more persons | **April 26, 2027** |
|
|
64
|
+
| 0 to 49,999 persons | **April 26, 2028** |
|
|
65
|
+
| Special district governments | **April 26, 2028** |
|
|
66
|
+
|
|
67
|
+
> "UPDATE: On April 20, 2026, the Department published an Interim Final Rule (IFR)
|
|
68
|
+
> extending the compliance date for State and local government entities with a
|
|
69
|
+
> total population of 50,000 or more to April 26, 2027. The compliance date for
|
|
70
|
+
> public entities with a total population of less than 50,000, or any special
|
|
71
|
+
> district government, is extended to April 26, 2028."
|
|
72
|
+
>
|
|
73
|
+
> — <https://www.ada.gov/title-ii-web-rule/>
|
|
74
|
+
|
|
75
|
+
> "The documents are word processing, presentation, PDF, or spreadsheet files"
|
|
76
|
+
> — <https://www.ada.gov/resources/2024-03-08-web-rule/>
|
|
77
|
+
|
|
78
|
+
So there is a hard deadline, a large population of PDFs that must comply, and
|
|
79
|
+
consequently a need to *check* PDFs automatically.
|
|
80
|
+
|
|
81
|
+
### Why the existing options do not fit
|
|
82
|
+
|
|
83
|
+
| Option | Why it does not work in CI |
|
|
84
|
+
|---|---|
|
|
85
|
+
| **veraPDF** — the authoritative validator | **Java.** Requires a JVM on every build agent. GPL-3.0 (dual MPLv2+). |
|
|
86
|
+
| **opendataloader-pdf** | The PyPI package is *"a Python wrapper for the opendataloader-pdf **Java CLI**"*. Its free tier is auto-tagging (producing a Tagged PDF); **PDF/UA export is an enterprise add-on**. It is a *remediator*, not an *auditor*. |
|
|
87
|
+
| **pdf-a11y** | A remediation pipeline whose `validator.py` is a `subprocess` wrapper around the veraPDF binary. It needs the JVM too. |
|
|
88
|
+
| **wcag_pdf_pytest** | Regex-scans raw PDF bytes for `/StructTreeRoot` and returns hardcoded passes for most criteria ("Heuristic pass: … not applicable"). It reports PASS on criteria it never examined. |
|
|
89
|
+
| Commercial services | Per-document pricing; not a build gate. |
|
|
90
|
+
|
|
91
|
+
`pdfua` occupies the empty slot: **a `pip install` from GitHub, no JVM, structured
|
|
92
|
+
output, exit codes.** It checks the failure classes that actually dominate real government
|
|
93
|
+
PDFs, and it says plainly what it did not check.
|
|
94
|
+
|
|
95
|
+
### What the evidence showed
|
|
96
|
+
|
|
97
|
+
This library was built only after a falsification pass against veraPDF 1.30.2 on
|
|
98
|
+
16 real government PDF files from `govinfo.gov`, `irs.gov` and `ada.gov` (15
|
|
99
|
+
distinct documents — see below). The full
|
|
100
|
+
evidence is in [`docs/FALSIFICATION.md`](docs/FALSIFICATION.md). The two findings
|
|
101
|
+
that shaped the design:
|
|
102
|
+
|
|
103
|
+
- **veraPDF reports 0 of 16 compliant.** All 16 fail PDF/UA-1.
|
|
104
|
+
- **99.55% of the 643,471 failed checks are two rules** — `7.1-3` ("content shall
|
|
105
|
+
be marked as Artifact or tagged as real content") and `7.2-34` ("natural
|
|
106
|
+
language for text in page content shall be determined"). Both are reachable in
|
|
107
|
+
Python without a JVM, and both are implemented here. A prototype's untagged-text counter
|
|
108
|
+
tracked veraPDF's `7.1-3` at 92–99.9% agreement (e.g. 86,118 vs 86,131 on one
|
|
109
|
+
1,039-page document).
|
|
110
|
+
|
|
111
|
+
Those totals cover **16 validated files, but only 15 distinct documents**: two
|
|
112
|
+
corpus files (`p08.pdf` and `7ef8534308.pdf`) are byte-identical (sha256
|
|
113
|
+
`9f5806…a853`), so the 643,471 figure double-counts 172,257 checks. Across the 15
|
|
114
|
+
distinct documents the de-duplicated total is **471,214**, of which the same two
|
|
115
|
+
rules account for 468,314 — **99.38%**, down from 99.55%. The concentration claim
|
|
116
|
+
holds either way, which is what the design argument rests on.
|
|
117
|
+
|
|
118
|
+
That is the honest case for this tool: the failures that dominate real documents
|
|
119
|
+
are cheap to detect and currently require a JVM to detect.
|
|
120
|
+
|
|
121
|
+
---
|
|
122
|
+
|
|
123
|
+
## Install
|
|
124
|
+
|
|
125
|
+
**Not yet on PyPI — install from GitHub for now.**
|
|
126
|
+
|
|
127
|
+
```bash
|
|
128
|
+
pip install "git+https://github.com/beduldul/pdfua.git"
|
|
129
|
+
# or with uv:
|
|
130
|
+
uv pip install "git+https://github.com/beduldul/pdfua.git"
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
Requires Python 3.10+ and `pikepdf`. No JVM, no Java, no external binary.
|
|
134
|
+
|
|
135
|
+
### Why `pikepdf`
|
|
136
|
+
|
|
137
|
+
A PDF/UA check has to walk the COS object graph: `/StructTreeRoot`, `/K` child
|
|
138
|
+
chains, indirect references, content streams. `pikepdf` binds QPDF, ships binary
|
|
139
|
+
wheels for Linux/macOS/Windows (so `pip install` needs no compiler), and exposes
|
|
140
|
+
references rather than flattening them. `pypdf` was the alternative; it hides
|
|
141
|
+
indirect references, which PDF/UA rules care about. **You are not expected to
|
|
142
|
+
write a PDF parser, and this package does not contain one.**
|
|
143
|
+
|
|
144
|
+
---
|
|
145
|
+
|
|
146
|
+
## Usage
|
|
147
|
+
|
|
148
|
+
### CLI
|
|
149
|
+
|
|
150
|
+
```bash
|
|
151
|
+
pdfua check document.pdf
|
|
152
|
+
pdfua check *.pdf --format sarif > results.sarif
|
|
153
|
+
pdfua check document.pdf --rules UA-01,UA-13-004
|
|
154
|
+
pdfua check document.pdf --min-severity warning --certain-only
|
|
155
|
+
pdfua rules # what is implemented, and what is not
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
### Exit codes
|
|
159
|
+
|
|
160
|
+
| Code | Meaning |
|
|
161
|
+
|---|---|
|
|
162
|
+
| `0` | No findings |
|
|
163
|
+
| `1` | Findings at INFO severity |
|
|
164
|
+
| `2` | Findings at WARNING severity |
|
|
165
|
+
| `3` | Findings at ERROR severity |
|
|
166
|
+
| `4` | The file could not be read as a PDF |
|
|
167
|
+
|
|
168
|
+
With `--min-severity` set, any surviving finding exits `1`. Without it, the exit
|
|
169
|
+
code is ordered by severity, so the obvious thing works:
|
|
170
|
+
|
|
171
|
+
```bash
|
|
172
|
+
pdfua check report.pdf || echo "PDF/UA problems found"
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
### Library
|
|
176
|
+
|
|
177
|
+
```python
|
|
178
|
+
from pdfua import validate, Severity
|
|
179
|
+
|
|
180
|
+
report = validate("document.pdf")
|
|
181
|
+
|
|
182
|
+
print(report.exit_code) # 0 clean, 1 info, 2 warning, 3 error
|
|
183
|
+
print(report.worst_severity()) # Severity.ERROR or None
|
|
184
|
+
|
|
185
|
+
for finding in report.findings:
|
|
186
|
+
print(finding.rule_id, finding.severity.name, finding.location.describe())
|
|
187
|
+
print(" ", finding.message)
|
|
188
|
+
print(" fix:", finding.remediation)
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
Filtering, and the coverage figure:
|
|
192
|
+
|
|
193
|
+
```python
|
|
194
|
+
from pdfua import coverage, describe_rules
|
|
195
|
+
from pdfua.model import filter_findings, Confidence
|
|
196
|
+
|
|
197
|
+
cov = coverage()
|
|
198
|
+
print(cov.summary()) # "16 of 106 PDF/UA-1 rules (15%) across 12 implemented checks"
|
|
199
|
+
|
|
200
|
+
certain = filter_findings(report.findings, min_confidence=Confidence.CERTAIN)
|
|
201
|
+
```
|
|
202
|
+
|
|
203
|
+
---
|
|
204
|
+
|
|
205
|
+
## What this checks
|
|
206
|
+
|
|
207
|
+
12 checks, covering 16 PDF/UA-1 rule identifiers. Each is decidable from the PDF
|
|
208
|
+
alone, with no interpretation. `pdfua rules` prints this table from the registry,
|
|
209
|
+
so the numbers cannot drift from the code.
|
|
210
|
+
|
|
211
|
+
| Rule | PDF/UA-1 clause | WCAG | Severity | What it decides |
|
|
212
|
+
|---|---|---|---|---|
|
|
213
|
+
| `UA-01-005` | §7.1 | 1.3.1 | error | `/StructTreeRoot` is present |
|
|
214
|
+
| `UA-01-002` | §6.2 | 1.3.1 | error | `/MarkInfo /Marked` is true |
|
|
215
|
+
| `UA-01-003` | §7.1 | 1.3.1 | error | Page content is tagged or marked as an artifact |
|
|
216
|
+
| `UA-06-001` | §7.2 | 3.1.1 | error | Catalog `/Lang` is present and well-formed |
|
|
217
|
+
| `UA-07-001` | §7.1 | 2.4.2 | warning | A document `/Title` exists |
|
|
218
|
+
| `UA-07-002` | §7.1 | 2.4.2 | warning | `DisplayDocTitle` is true |
|
|
219
|
+
| `UA-13-004` | §7.3 | 1.1.1 | error | Every `/Figure` has non-empty `/Alt` or `/ActualText` |
|
|
220
|
+
| `UA-14-001` | §7.4 | 1.3.1, 2.4.6 | warning | No empty heading elements |
|
|
221
|
+
| `UA-15-003` | §7.5 | 1.3.1 | error | Table structure is determinable |
|
|
222
|
+
| `UA-18-001` | §7.18.1 | 1.1.1, 4.1.2 | warning | Annotations have descriptions |
|
|
223
|
+
| `UA-18-005` | §7.18.5 | 2.4.4 | warning | Link annotations have descriptions |
|
|
224
|
+
| `UA-28-004` | §5 | — | warning | PDF/UA identification in XMP |
|
|
225
|
+
|
|
226
|
+
---
|
|
227
|
+
|
|
228
|
+
## What this does NOT check
|
|
229
|
+
|
|
230
|
+
**A validator that overstates its coverage is worse than no validator.** This
|
|
231
|
+
section is the most important one in this file.
|
|
232
|
+
|
|
233
|
+
### It implements 12 checks, covering 16 of 106 rules
|
|
234
|
+
|
|
235
|
+
The public veraPDF PDF/UA-1 profile enumerates **106 machine-checkable rules** for
|
|
236
|
+
ISO 14289-1. This tool implements **12 checks**, which decide **16** of those
|
|
237
|
+
rules. `pdfua rules` prints both counts from the registry, so neither can drift.
|
|
238
|
+
|
|
239
|
+
**The other 90 rules are not checked.** A clean run means "these 16 rules found
|
|
240
|
+
nothing", never "this document conforms".
|
|
241
|
+
|
|
242
|
+
Coverage is counted at *rule* granularity, not per section, and that distinction
|
|
243
|
+
is the point. Clause 7.2 alone contains **41** rules; implementing one `/Lang`
|
|
244
|
+
check does not make the other forty checked. A tool that reported "7.2: 41/41"
|
|
245
|
+
because of a single check would be producing exactly the false comfort this
|
|
246
|
+
project exists to avoid.
|
|
247
|
+
|
|
248
|
+
Notably **not** checked:
|
|
249
|
+
|
|
250
|
+
- **Font embedding, `CIDSet`, glyph widths, font descriptors** (§7.21). These are
|
|
251
|
+
the rules veraPDF uses to catch professionally-remediated files — `ada.gov`'s
|
|
252
|
+
own web-rule PDF is `105/106` clean and fails only `7.21.4.2-2` (CIDSet). This
|
|
253
|
+
tool will report that document as clean, because it does not look at fonts.
|
|
254
|
+
- **Reading order correctness** (§7.2). The tool confirms content *is* marked; it
|
|
255
|
+
cannot judge whether the order is *meaningful*.
|
|
256
|
+
- **Colour spaces, output intents, transparency, ICC profiles.**
|
|
257
|
+
- **WTPDF 1.0** and **PDF/UA-2** (ISO 14289-2 / ISO 32005).
|
|
258
|
+
- **PDF/A** of any flavour.
|
|
259
|
+
- **Encrypted documents.** They raise, rather than guess.
|
|
260
|
+
|
|
261
|
+
### It cannot judge quality, only presence
|
|
262
|
+
|
|
263
|
+
This is the fundamental line, and the tool never crosses it:
|
|
264
|
+
|
|
265
|
+
| Machine-checkable (this tool) | Requires a human |
|
|
266
|
+
|---|---|
|
|
267
|
+
| A `/Figure` has an `/Alt` entry | Whether the alt text is *accurate* |
|
|
268
|
+
| The `/Alt` entry is non-empty | Whether it is *appropriate in context* |
|
|
269
|
+
| A table has `/TH` cells | Whether the header association is *correct* |
|
|
270
|
+
| Content is marked | Whether the reading order is *sensible* |
|
|
271
|
+
| A title exists | Whether the title is *descriptive* |
|
|
272
|
+
|
|
273
|
+
A figure whose alt text is `" "` (a single space) is structurally present and
|
|
274
|
+
semantically worthless. The tool flags the empty case (`UA-13-004`) but cannot
|
|
275
|
+
tell you that `"image"` or `"figure 3"` is a bad description. **Automated checks
|
|
276
|
+
reduce manual review; they do not replace it.**
|
|
277
|
+
|
|
278
|
+
### Findings marked `heuristic` are questions, not verdicts
|
|
279
|
+
|
|
280
|
+
Three finding sites carry `confidence: heuristic`, and in each case it is the
|
|
281
|
+
*site*, not the whole rule, that is a judgement call:
|
|
282
|
+
|
|
283
|
+
- **`UA-07-001`** (`DocumentTitleRule`) — every finding it emits is heuristic
|
|
284
|
+
(`src/pdfua/rules/language.py:60` sets the class-level confidence).
|
|
285
|
+
- **`UA-13-004`** (`FigureAltTextRule`) — only the *empty or whitespace-only
|
|
286
|
+
`/Alt`* branch, which is stricter than the letter of §7.3 and that veraPDF
|
|
287
|
+
passes (`src/pdfua/rules/structure.py:73`). A figure with *no* `/Alt` at all
|
|
288
|
+
is a certain finding.
|
|
289
|
+
- **`UA-15-003`** (`TableHeaderRule`) — only the *association* branch, where
|
|
290
|
+
header cells exist but none carry `/Scope` or `/Headers`
|
|
291
|
+
(`src/pdfua/rules/structure.py:144`). The rule itself has class-level severity
|
|
292
|
+
`ERROR`, and its other branches are certain.
|
|
293
|
+
|
|
294
|
+
They narrow the question for a human. Use `--certain-only` to see only findings
|
|
295
|
+
the tool is certain about.
|
|
296
|
+
|
|
297
|
+
### A mistake this project made, recorded on purpose
|
|
298
|
+
|
|
299
|
+
An early version reported `"TH cells lack Scope"` on documents veraPDF passes.
|
|
300
|
+
That was **wrong about the standard**: ISO 14289-1 §7.5 requires `Scope` only *if
|
|
301
|
+
the structure is not determinable via `Headers` and `IDs`*. It is a fallback, not
|
|
302
|
+
a universal requirement. The IRS W-9 has `/TH` cells with neither, and is valid.
|
|
303
|
+
|
|
304
|
+
The rule was removed. It is documented in `docs/FALSIFICATION.md` and in the
|
|
305
|
+
`structure.py` module docstring because this is exactly the failure mode that
|
|
306
|
+
makes accessibility tooling distrusted, and the credibility of this tool depends
|
|
307
|
+
on not repeating it.
|
|
308
|
+
|
|
309
|
+
---
|
|
310
|
+
|
|
311
|
+
## For a conformance claim, use veraPDF
|
|
312
|
+
|
|
313
|
+
`pdfua` and veraPDF are complements, not competitors.
|
|
314
|
+
|
|
315
|
+
```bash
|
|
316
|
+
# Every commit: fast, no JVM
|
|
317
|
+
pdfua check report.pdf --format sarif || true
|
|
318
|
+
|
|
319
|
+
# Before you claim conformance: the authority
|
|
320
|
+
verapdf --flavour ua1 --format json report.pdf
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
`pdfua` gives you a JVM-free gate that produces SARIF for code scanning. Runtime
|
|
324
|
+
scales with document size and structure, so treat any single figure with care:
|
|
325
|
+
measured on this project's corpus (median of five runs) it is **~2–3 ms for a
|
|
326
|
+
184 KB document** and **~1.4 s for the 3.9 MB `ada.gov` web-rule PDF** — fast on
|
|
327
|
+
typical documents, seconds on very large ones, and always far below the cost of
|
|
328
|
+
starting a JVM. veraPDF remains the reference implementation. **If your CI has a JVM and
|
|
329
|
+
you need a conformance claim, use veraPDF and do not install this.** The tool
|
|
330
|
+
exists for the pipelines where a JVM is the reason nothing is checked at all.
|
|
331
|
+
|
|
332
|
+
---
|
|
333
|
+
|
|
334
|
+
## Development
|
|
335
|
+
|
|
336
|
+
```bash
|
|
337
|
+
uv venv && uv pip install -e ".[dev]"
|
|
338
|
+
.venv/bin/pytest
|
|
339
|
+
.venv/bin/mypy src/pdfua
|
|
340
|
+
.venv/bin/ruff check .
|
|
341
|
+
```
|
|
342
|
+
|
|
343
|
+
Test fixtures that violate exactly one rule each are **generated in code** by
|
|
344
|
+
`tests/fixtures/build.py` using `pikepdf`; no binary PDFs are committed. See
|
|
345
|
+
`tests/fixtures/README.md`.
|
|
346
|
+
|
|
347
|
+
## License
|
|
348
|
+
|
|
349
|
+
MIT. See [LICENSE](LICENSE).
|
|
350
|
+
|
|
351
|
+
The bundled `pdfua1_rules.json` contains **rule identifiers only** (ISO clause
|
|
352
|
+
number, test number, validation-object name) — no descriptive text is reproduced
|
|
353
|
+
from any validation profile, because that text belongs to its authors.
|