pdfua 0.1.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. pdfua-0.1.1/.gitignore +35 -0
  2. pdfua-0.1.1/CHANGELOG.md +52 -0
  3. pdfua-0.1.1/CONTRIBUTING.md +85 -0
  4. pdfua-0.1.1/LICENSE +21 -0
  5. pdfua-0.1.1/PKG-INFO +353 -0
  6. pdfua-0.1.1/README.md +321 -0
  7. pdfua-0.1.1/pyproject.toml +78 -0
  8. pdfua-0.1.1/src/pdfua/__init__.py +46 -0
  9. pdfua-0.1.1/src/pdfua/__main__.py +10 -0
  10. pdfua-0.1.1/src/pdfua/catalog.py +185 -0
  11. pdfua-0.1.1/src/pdfua/cli.py +278 -0
  12. pdfua-0.1.1/src/pdfua/data/__init__.py +1 -0
  13. pdfua-0.1.1/src/pdfua/data/pdfua1_rules.json +537 -0
  14. pdfua-0.1.1/src/pdfua/document.py +511 -0
  15. pdfua-0.1.1/src/pdfua/errors.py +19 -0
  16. pdfua-0.1.1/src/pdfua/model.py +207 -0
  17. pdfua-0.1.1/src/pdfua/py.typed +0 -0
  18. pdfua-0.1.1/src/pdfua/reporters.py +230 -0
  19. pdfua-0.1.1/src/pdfua/rules/__init__.py +39 -0
  20. pdfua-0.1.1/src/pdfua/rules/base.py +119 -0
  21. pdfua-0.1.1/src/pdfua/rules/language.py +140 -0
  22. pdfua-0.1.1/src/pdfua/rules/structure.py +221 -0
  23. pdfua-0.1.1/src/pdfua/rules/tagging.py +142 -0
  24. pdfua-0.1.1/src/pdfua/rules/wcag.py +125 -0
  25. pdfua-0.1.1/src/pdfua/validator.py +99 -0
  26. pdfua-0.1.1/tests/__init__.py +0 -0
  27. pdfua-0.1.1/tests/conftest.py +64 -0
  28. pdfua-0.1.1/tests/fixtures/README.md +61 -0
  29. pdfua-0.1.1/tests/fixtures/__init__.py +0 -0
  30. pdfua-0.1.1/tests/fixtures/build.py +310 -0
  31. pdfua-0.1.1/tests/test_cli.py +205 -0
  32. pdfua-0.1.1/tests/test_coverage.py +133 -0
  33. pdfua-0.1.1/tests/test_document.py +216 -0
  34. pdfua-0.1.1/tests/test_model.py +170 -0
  35. pdfua-0.1.1/tests/test_real_documents.py +133 -0
  36. pdfua-0.1.1/tests/test_reporters.py +147 -0
  37. pdfua-0.1.1/tests/test_rules_contract.py +118 -0
pdfua-0.1.1/.gitignore ADDED
@@ -0,0 +1,35 @@
1
+ # Python
2
+ __pycache__/
3
+ *.py[cod]
4
+ *.egg-info/
5
+ .eggs/
6
+ build/
7
+ dist/
8
+ .venv/
9
+ venv/
10
+ .pytest_cache/
11
+ .mypy_cache/
12
+ .ruff_cache/
13
+ .coverage
14
+ coverage.xml
15
+ htmlcov/
16
+
17
+ # Secrets — never commit these
18
+ .env
19
+ .env.*
20
+ *.key
21
+ *.pem
22
+ *.p12
23
+ *.pfx
24
+ secrets/
25
+
26
+ # Downloaded test corpora (regenerate with scripts/fetch_corpus.py)
27
+ corpus/
28
+ tests/fixtures/downloaded/
29
+ *.pdf
30
+ !tests/fixtures/**/*.pdf
31
+
32
+ # Editors / OS
33
+ .DS_Store
34
+ .idea/
35
+ .vscode/
@@ -0,0 +1,52 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented in this file.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
+ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
+
8
+ ## [Unreleased]
9
+
10
+ ## [0.1.1] - 2026-10-04
11
+
12
+ ### Changed
13
+
14
+ - Packaging and release metadata only; no code or behaviour changes. Added
15
+ complete PyPI metadata (PEP 639 `license` expression, author, project URLs
16
+ and classifiers) and a Trusted Publishing (OIDC) release workflow that
17
+ publishes on `v*` tags.
18
+
19
+ ## [0.1.0] - 2026-09-29
20
+
21
+ Initial release.
22
+
23
+ ### Added
24
+
25
+ - `pdfua check <file>` with `--format text|json|sarif`, `--rules`,
26
+ `--min-severity`, `--certain-only`, `--fail-fast`, `--quiet`.
27
+ - `pdfua rules` listing implemented rules and the clauses with no coverage.
28
+ - 12 checks covering 16 of the 106 machine-checkable PDF/UA-1 rules (15%):
29
+ - `UA-01-005` structure tree present
30
+ - `UA-01-002` `/MarkInfo` `/Marked` true
31
+ - `UA-01-003` page content tagged or marked as artifact
32
+ - `UA-06-001` document language
33
+ - `UA-07-001` document title
34
+ - `UA-07-002` `DisplayDocTitle`
35
+ - `UA-13-004` figure alternative text
36
+ - `UA-14-001` non-empty headings
37
+ - `UA-15-003` determinable table structure
38
+ - `UA-18-001` annotation descriptions
39
+ - `UA-18-005` link descriptions
40
+ - `UA-28-004` PDF/UA identification in XMP
41
+ - SARIF 2.1.0 output for GitHub code scanning and Azure DevOps.
42
+ - Severity-ordered exit codes (0 clean, 1 info, 2 warning, 3 error).
43
+ - Full type annotations, `py.typed` marker, strict `mypy` configuration.
44
+
45
+ ### Notes
46
+
47
+ - This is 12 checks covering 16 of the 106 machine-checkable rules in the
48
+ PDF/UA-1 profile. The other 90 are not checked. See the README section
49
+ "What this does NOT check".
50
+ - Development was validated against veraPDF 1.30.2 on a 16-file corpus (15
51
+ distinct documents) of real `govinfo.gov`, `irs.gov` and `ada.gov` PDFs. See
52
+ `docs/FALSIFICATION.md` for the evidence.
@@ -0,0 +1,85 @@
1
+ # Contributing to pdfua
2
+
3
+ Thanks for looking. This project has one governing value: **do not overclaim.**
4
+
5
+ A validator that says "PASS" is making a claim about everything it did not check.
6
+ Every design decision here follows from that.
7
+
8
+ ## The one rule
9
+
10
+ **If you add a check, you must declare exactly which PDF/UA-1 rule identifiers it
11
+ decides, and you must add a fixture that proves it fires.**
12
+
13
+ A rule declares its coverage with the `covers` attribute:
14
+
15
+ ```python
16
+ class MyRule(Rule):
17
+ id = "UA-99-001"
18
+ title = "Something decidable"
19
+ severity = Severity.ERROR
20
+ pdfua_clause = "7.9"
21
+ covers = ("7.9-1",) # clause-test, as enumerated in the PDF/UA-1 profile
22
+ ```
23
+
24
+ `tests/test_coverage.py` fails if a declared identifier is not a real PDF/UA-1
25
+ rule, and `tests/test_rules_contract.py` fails if a registered rule has no
26
+ positive fixture. Both exist to stop the coverage number in the README from
27
+ drifting away from the truth.
28
+
29
+ ## Adding a rule
30
+
31
+ 1. Implement it in the appropriate `src/pdfua/rules/*.py` module.
32
+ 2. Set `covers` to the identifiers it decides. If it does not correspond to a
33
+ single profile rule, leave `covers` empty — it then contributes nothing to the
34
+ coverage figure, which is the honest outcome.
35
+ 3. Add a builder to `tests/fixtures/build.py` that changes **exactly one thing**
36
+ from `valid_document()`.
37
+ 4. Add the rule to `RULE_CASES` in `tests/test_rules_contract.py`.
38
+ 5. Run the suite.
39
+
40
+ ## The fixture design rule
41
+
42
+ `valid_document()` satisfies every rule. Each `*_violating` builder changes one
43
+ thing. If your builder changes two, a test for either rule can pass for the wrong
44
+ reason, which is worse than no test.
45
+
46
+ ## What will be rejected
47
+
48
+ - **A check that guesses.** If a rule needs human judgement, it must set
49
+ `confidence=Confidence.HEURISTIC` and say in its message that a human must
50
+ confirm. See `TableHeaderRule` for the pattern.
51
+ - **A check that is stricter than the standard without saying so.** If you flag
52
+ something the specification permits, the finding must state that explicitly.
53
+ See the whitespace-`/Alt` branch in `FigureAltTextRule`.
54
+ - **A rule that reports per-instance when one finding suffices.** Reporting an
55
+ untagged document's defect once per page buries the signal.
56
+ - **A silent skip.** A rule that cannot evaluate a document must raise
57
+ `PdfuaError` with a reason, so it appears in `report.skipped`. Never return
58
+ nothing and let it look like a pass.
59
+
60
+ ## Development
61
+
62
+ ```bash
63
+ uv venv
64
+ uv pip install -e ".[dev]"
65
+ .venv/bin/pytest
66
+ .venv/bin/mypy src/pdfua
67
+ .venv/bin/ruff check .
68
+ ```
69
+
70
+ The suite must stay green, coverage at or above 80%, and `mypy --strict` must
71
+ pass with no `Any`. The package ships `py.typed`; annotations are part of the
72
+ public interface.
73
+
74
+ ## Commits
75
+
76
+ Conventional Commits (`feat:`, `fix:`, `docs:`, `test:`, `refactor:`, `chore:`).
77
+ Small and focused. **Never use `--no-verify`** — if a hook fails, fix the cause.
78
+
79
+ ## Reporting a false positive
80
+
81
+ This matters more than a missed detection. If the tool reports a problem in a
82
+ document that is actually fine, please open an issue with the document (or a
83
+ minimal reproduction) and the rule ID. A false positive is how accessibility
84
+ tooling loses its users' trust, and it is treated as a bug of the highest
85
+ severity.
pdfua-0.1.1/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 beduldul
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
pdfua-0.1.1/PKG-INFO ADDED
@@ -0,0 +1,353 @@
1
+ Metadata-Version: 2.5
2
+ Name: pdfua
3
+ Version: 0.1.1
4
+ Summary: Check the machine-checkable subset of PDF/UA-1 and WCAG 2.1 for PDFs in Python, built on pikepdf with no JVM
5
+ Project-URL: Homepage, https://github.com/beduldul/pdfua
6
+ Project-URL: Repository, https://github.com/beduldul/pdfua
7
+ Project-URL: Issues, https://github.com/beduldul/pdfua/issues
8
+ Project-URL: Changelog, https://github.com/beduldul/pdfua/blob/main/CHANGELOG.md
9
+ Author-email: Abdul Afif Al Kaysan <alkaysan07@gmail.com>
10
+ License-Expression: MIT
11
+ License-File: LICENSE
12
+ Keywords: a11y,accessibility,ada,pdf,pdf-ua,pdfua,sarif,section508,validator,wcag
13
+ Classifier: Development Status :: 3 - Alpha
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Operating System :: OS Independent
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.10
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Topic :: Software Development :: Quality Assurance
22
+ Classifier: Topic :: Text Processing :: Markup
23
+ Classifier: Typing :: Typed
24
+ Requires-Python: >=3.10
25
+ Requires-Dist: pikepdf>=8.0
26
+ Provides-Extra: dev
27
+ Requires-Dist: mypy>=1.11; extra == 'dev'
28
+ Requires-Dist: pytest-cov>=5.0; extra == 'dev'
29
+ Requires-Dist: pytest>=8.0; extra == 'dev'
30
+ Requires-Dist: ruff>=0.6; extra == 'dev'
31
+ Description-Content-Type: text/markdown
32
+
33
+ [![CI](https://github.com/beduldul/pdfua/actions/workflows/ci.yml/badge.svg)](https://github.com/beduldul/pdfua/actions/workflows/ci.yml)
34
+ # pdfua
35
+
36
+ **Check the machine-checkable subset of PDF/UA-1 and WCAG 2.1 for PDFs — a small
37
+ Python package built on `pikepdf` (QPDF), with no JVM.**
38
+
39
+ ```bash
40
+ pip install "git+https://github.com/beduldul/pdfua.git"
41
+ pdfua check report.pdf --format sarif
42
+ ```
43
+
44
+ **This implements 12 checks covering 16 of the 106 machine-checkable rules in the
45
+ PDF/UA-1 profile — 15%.** Read
46
+ [What this does NOT check](#what-this-does-not-check) before you rely on it. It is
47
+ a fast triage gate, not a conformance certifier. For a conformance claim, run
48
+ [veraPDF](https://verapdf.org).
49
+
50
+ ---
51
+
52
+ ## The problem
53
+
54
+ There is a dated legal forcing function and no tool that fits a CI pipeline.
55
+
56
+ The U.S. Department of Justice's Title II rule requires state and local
57
+ governments to meet **WCAG 2.1 Level AA** for their web content and mobile apps.
58
+ PDFs are explicitly named as covered content. The compliance dates, from
59
+ `ada.gov`:
60
+
61
+ | State and local government size | Compliance date |
62
+ |---|---|
63
+ | 50,000 or more persons | **April 26, 2027** |
64
+ | 0 to 49,999 persons | **April 26, 2028** |
65
+ | Special district governments | **April 26, 2028** |
66
+
67
+ > "UPDATE: On April 20, 2026, the Department published an Interim Final Rule (IFR)
68
+ > extending the compliance date for State and local government entities with a
69
+ > total population of 50,000 or more to April 26, 2027. The compliance date for
70
+ > public entities with a total population of less than 50,000, or any special
71
+ > district government, is extended to April 26, 2028."
72
+ >
73
+ > — <https://www.ada.gov/title-ii-web-rule/>
74
+
75
+ > "The documents are word processing, presentation, PDF, or spreadsheet files"
76
+ > — <https://www.ada.gov/resources/2024-03-08-web-rule/>
77
+
78
+ So there is a hard deadline, a large population of PDFs that must comply, and
79
+ consequently a need to *check* PDFs automatically.
80
+
81
+ ### Why the existing options do not fit
82
+
83
+ | Option | Why it does not work in CI |
84
+ |---|---|
85
+ | **veraPDF** — the authoritative validator | **Java.** Requires a JVM on every build agent. GPL-3.0 (dual MPLv2+). |
86
+ | **opendataloader-pdf** | The PyPI package is *"a Python wrapper for the opendataloader-pdf **Java CLI**"*. Its free tier is auto-tagging (producing a Tagged PDF); **PDF/UA export is an enterprise add-on**. It is a *remediator*, not an *auditor*. |
87
+ | **pdf-a11y** | A remediation pipeline whose `validator.py` is a `subprocess` wrapper around the veraPDF binary. It needs the JVM too. |
88
+ | **wcag_pdf_pytest** | Regex-scans raw PDF bytes for `/StructTreeRoot` and returns hardcoded passes for most criteria ("Heuristic pass: … not applicable"). It reports PASS on criteria it never examined. |
89
+ | Commercial services | Per-document pricing; not a build gate. |
90
+
91
+ `pdfua` occupies the empty slot: **a `pip install` from GitHub, no JVM, structured
92
+ output, exit codes.** It checks the failure classes that actually dominate real government
93
+ PDFs, and it says plainly what it did not check.
94
+
95
+ ### What the evidence showed
96
+
97
+ This library was built only after a falsification pass against veraPDF 1.30.2 on
98
+ 16 real government PDF files from `govinfo.gov`, `irs.gov` and `ada.gov` (15
99
+ distinct documents — see below). The full
100
+ evidence is in [`docs/FALSIFICATION.md`](docs/FALSIFICATION.md). The two findings
101
+ that shaped the design:
102
+
103
+ - **veraPDF reports 0 of 16 compliant.** All 16 fail PDF/UA-1.
104
+ - **99.55% of the 643,471 failed checks are two rules** — `7.1-3` ("content shall
105
+ be marked as Artifact or tagged as real content") and `7.2-34` ("natural
106
+ language for text in page content shall be determined"). Both are reachable in
107
+ Python without a JVM, and both are implemented here. A prototype's untagged-text counter
108
+ tracked veraPDF's `7.1-3` at 92–99.9% agreement (e.g. 86,118 vs 86,131 on one
109
+ 1,039-page document).
110
+
111
+ Those totals cover **16 validated files, but only 15 distinct documents**: two
112
+ corpus files (`p08.pdf` and `7ef8534308.pdf`) are byte-identical (sha256
113
+ `9f5806…a853`), so the 643,471 figure double-counts 172,257 checks. Across the 15
114
+ distinct documents the de-duplicated total is **471,214**, of which the same two
115
+ rules account for 468,314 — **99.38%**, down from 99.55%. The concentration claim
116
+ holds either way, which is what the design argument rests on.
117
+
118
+ That is the honest case for this tool: the failures that dominate real documents
119
+ are cheap to detect and currently require a JVM to detect.
120
+
121
+ ---
122
+
123
+ ## Install
124
+
125
+ **Not yet on PyPI — install from GitHub for now.**
126
+
127
+ ```bash
128
+ pip install "git+https://github.com/beduldul/pdfua.git"
129
+ # or with uv:
130
+ uv pip install "git+https://github.com/beduldul/pdfua.git"
131
+ ```
132
+
133
+ Requires Python 3.10+ and `pikepdf`. No JVM, no Java, no external binary.
134
+
135
+ ### Why `pikepdf`
136
+
137
+ A PDF/UA check has to walk the COS object graph: `/StructTreeRoot`, `/K` child
138
+ chains, indirect references, content streams. `pikepdf` binds QPDF, ships binary
139
+ wheels for Linux/macOS/Windows (so `pip install` needs no compiler), and exposes
140
+ references rather than flattening them. `pypdf` was the alternative; it hides
141
+ indirect references, which PDF/UA rules care about. **You are not expected to
142
+ write a PDF parser, and this package does not contain one.**
143
+
144
+ ---
145
+
146
+ ## Usage
147
+
148
+ ### CLI
149
+
150
+ ```bash
151
+ pdfua check document.pdf
152
+ pdfua check *.pdf --format sarif > results.sarif
153
+ pdfua check document.pdf --rules UA-01,UA-13-004
154
+ pdfua check document.pdf --min-severity warning --certain-only
155
+ pdfua rules # what is implemented, and what is not
156
+ ```
157
+
158
+ ### Exit codes
159
+
160
+ | Code | Meaning |
161
+ |---|---|
162
+ | `0` | No findings |
163
+ | `1` | Findings at INFO severity |
164
+ | `2` | Findings at WARNING severity |
165
+ | `3` | Findings at ERROR severity |
166
+ | `4` | The file could not be read as a PDF |
167
+
168
+ With `--min-severity` set, any surviving finding exits `1`. Without it, the exit
169
+ code is ordered by severity, so the obvious thing works:
170
+
171
+ ```bash
172
+ pdfua check report.pdf || echo "PDF/UA problems found"
173
+ ```
174
+
175
+ ### Library
176
+
177
+ ```python
178
+ from pdfua import validate, Severity
179
+
180
+ report = validate("document.pdf")
181
+
182
+ print(report.exit_code) # 0 clean, 1 info, 2 warning, 3 error
183
+ print(report.worst_severity()) # Severity.ERROR or None
184
+
185
+ for finding in report.findings:
186
+ print(finding.rule_id, finding.severity.name, finding.location.describe())
187
+ print(" ", finding.message)
188
+ print(" fix:", finding.remediation)
189
+ ```
190
+
191
+ Filtering, and the coverage figure:
192
+
193
+ ```python
194
+ from pdfua import coverage, describe_rules
195
+ from pdfua.model import filter_findings, Confidence
196
+
197
+ cov = coverage()
198
+ print(cov.summary()) # "16 of 106 PDF/UA-1 rules (15%) across 12 implemented checks"
199
+
200
+ certain = filter_findings(report.findings, min_confidence=Confidence.CERTAIN)
201
+ ```
202
+
203
+ ---
204
+
205
+ ## What this checks
206
+
207
+ 12 checks, covering 16 PDF/UA-1 rule identifiers. Each is decidable from the PDF
208
+ alone, with no interpretation. `pdfua rules` prints this table from the registry,
209
+ so the numbers cannot drift from the code.
210
+
211
+ | Rule | PDF/UA-1 clause | WCAG | Severity | What it decides |
212
+ |---|---|---|---|---|
213
+ | `UA-01-005` | §7.1 | 1.3.1 | error | `/StructTreeRoot` is present |
214
+ | `UA-01-002` | §6.2 | 1.3.1 | error | `/MarkInfo /Marked` is true |
215
+ | `UA-01-003` | §7.1 | 1.3.1 | error | Page content is tagged or marked as an artifact |
216
+ | `UA-06-001` | §7.2 | 3.1.1 | error | Catalog `/Lang` is present and well-formed |
217
+ | `UA-07-001` | §7.1 | 2.4.2 | warning | A document `/Title` exists |
218
+ | `UA-07-002` | §7.1 | 2.4.2 | warning | `DisplayDocTitle` is true |
219
+ | `UA-13-004` | §7.3 | 1.1.1 | error | Every `/Figure` has non-empty `/Alt` or `/ActualText` |
220
+ | `UA-14-001` | §7.4 | 1.3.1, 2.4.6 | warning | No empty heading elements |
221
+ | `UA-15-003` | §7.5 | 1.3.1 | error | Table structure is determinable |
222
+ | `UA-18-001` | §7.18.1 | 1.1.1, 4.1.2 | warning | Annotations have descriptions |
223
+ | `UA-18-005` | §7.18.5 | 2.4.4 | warning | Link annotations have descriptions |
224
+ | `UA-28-004` | §5 | — | warning | PDF/UA identification in XMP |
225
+
226
+ ---
227
+
228
+ ## What this does NOT check
229
+
230
+ **A validator that overstates its coverage is worse than no validator.** This
231
+ section is the most important one in this file.
232
+
233
+ ### It implements 12 checks, covering 16 of 106 rules
234
+
235
+ The public veraPDF PDF/UA-1 profile enumerates **106 machine-checkable rules** for
236
+ ISO 14289-1. This tool implements **12 checks**, which decide **16** of those
237
+ rules. `pdfua rules` prints both counts from the registry, so neither can drift.
238
+
239
+ **The other 90 rules are not checked.** A clean run means "these 16 rules found
240
+ nothing", never "this document conforms".
241
+
242
+ Coverage is counted at *rule* granularity, not per section, and that distinction
243
+ is the point. Clause 7.2 alone contains **41** rules; implementing one `/Lang`
244
+ check does not make the other forty checked. A tool that reported "7.2: 41/41"
245
+ because of a single check would be producing exactly the false comfort this
246
+ project exists to avoid.
247
+
248
+ Notably **not** checked:
249
+
250
+ - **Font embedding, `CIDSet`, glyph widths, font descriptors** (§7.21). These are
251
+ the rules veraPDF uses to catch professionally-remediated files — `ada.gov`'s
252
+ own web-rule PDF is `105/106` clean and fails only `7.21.4.2-2` (CIDSet). This
253
+ tool will report that document as clean, because it does not look at fonts.
254
+ - **Reading order correctness** (§7.2). The tool confirms content *is* marked; it
255
+ cannot judge whether the order is *meaningful*.
256
+ - **Colour spaces, output intents, transparency, ICC profiles.**
257
+ - **WTPDF 1.0** and **PDF/UA-2** (ISO 14289-2 / ISO 32005).
258
+ - **PDF/A** of any flavour.
259
+ - **Encrypted documents.** They raise, rather than guess.
260
+
261
+ ### It cannot judge quality, only presence
262
+
263
+ This is the fundamental line, and the tool never crosses it:
264
+
265
+ | Machine-checkable (this tool) | Requires a human |
266
+ |---|---|
267
+ | A `/Figure` has an `/Alt` entry | Whether the alt text is *accurate* |
268
+ | The `/Alt` entry is non-empty | Whether it is *appropriate in context* |
269
+ | A table has `/TH` cells | Whether the header association is *correct* |
270
+ | Content is marked | Whether the reading order is *sensible* |
271
+ | A title exists | Whether the title is *descriptive* |
272
+
273
+ A figure whose alt text is `" "` (a single space) is structurally present and
274
+ semantically worthless. The tool flags the empty case (`UA-13-004`) but cannot
275
+ tell you that `"image"` or `"figure 3"` is a bad description. **Automated checks
276
+ reduce manual review; they do not replace it.**
277
+
278
+ ### Findings marked `heuristic` are questions, not verdicts
279
+
280
+ Three finding sites carry `confidence: heuristic`, and in each case it is the
281
+ *site*, not the whole rule, that is a judgement call:
282
+
283
+ - **`UA-07-001`** (`DocumentTitleRule`) — every finding it emits is heuristic
284
+ (`src/pdfua/rules/language.py:60` sets the class-level confidence).
285
+ - **`UA-13-004`** (`FigureAltTextRule`) — only the *empty or whitespace-only
286
+ `/Alt`* branch, which is stricter than the letter of §7.3 and that veraPDF
287
+ passes (`src/pdfua/rules/structure.py:73`). A figure with *no* `/Alt` at all
288
+ is a certain finding.
289
+ - **`UA-15-003`** (`TableHeaderRule`) — only the *association* branch, where
290
+ header cells exist but none carry `/Scope` or `/Headers`
291
+ (`src/pdfua/rules/structure.py:144`). The rule itself has class-level severity
292
+ `ERROR`, and its other branches are certain.
293
+
294
+ They narrow the question for a human. Use `--certain-only` to see only findings
295
+ the tool is certain about.
296
+
297
+ ### A mistake this project made, recorded on purpose
298
+
299
+ An early version reported `"TH cells lack Scope"` on documents veraPDF passes.
300
+ That was **wrong about the standard**: ISO 14289-1 §7.5 requires `Scope` only *if
301
+ the structure is not determinable via `Headers` and `IDs`*. It is a fallback, not
302
+ a universal requirement. The IRS W-9 has `/TH` cells with neither, and is valid.
303
+
304
+ The rule was removed. It is documented in `docs/FALSIFICATION.md` and in the
305
+ `structure.py` module docstring because this is exactly the failure mode that
306
+ makes accessibility tooling distrusted, and the credibility of this tool depends
307
+ on not repeating it.
308
+
309
+ ---
310
+
311
+ ## For a conformance claim, use veraPDF
312
+
313
+ `pdfua` and veraPDF are complements, not competitors.
314
+
315
+ ```bash
316
+ # Every commit: fast, no JVM
317
+ pdfua check report.pdf --format sarif || true
318
+
319
+ # Before you claim conformance: the authority
320
+ verapdf --flavour ua1 --format json report.pdf
321
+ ```
322
+
323
+ `pdfua` gives you a JVM-free gate that produces SARIF for code scanning. Runtime
324
+ scales with document size and structure, so treat any single figure with care:
325
+ measured on this project's corpus (median of five runs) it is **~2–3 ms for a
326
+ 184 KB document** and **~1.4 s for the 3.9 MB `ada.gov` web-rule PDF** — fast on
327
+ typical documents, seconds on very large ones, and always far below the cost of
328
+ starting a JVM. veraPDF remains the reference implementation. **If your CI has a JVM and
329
+ you need a conformance claim, use veraPDF and do not install this.** The tool
330
+ exists for the pipelines where a JVM is the reason nothing is checked at all.
331
+
332
+ ---
333
+
334
+ ## Development
335
+
336
+ ```bash
337
+ uv venv && uv pip install -e ".[dev]"
338
+ .venv/bin/pytest
339
+ .venv/bin/mypy src/pdfua
340
+ .venv/bin/ruff check .
341
+ ```
342
+
343
+ Test fixtures that violate exactly one rule each are **generated in code** by
344
+ `tests/fixtures/build.py` using `pikepdf`; no binary PDFs are committed. See
345
+ `tests/fixtures/README.md`.
346
+
347
+ ## License
348
+
349
+ MIT. See [LICENSE](LICENSE).
350
+
351
+ The bundled `pdfua1_rules.json` contains **rule identifiers only** (ISO clause
352
+ number, test number, validation-object name) — no descriptive text is reproduced
353
+ from any validation profile, because that text belongs to its authors.