on-the-list 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. on_the_list-0.1.0/LICENSE +21 -0
  2. on_the_list-0.1.0/PKG-INFO +417 -0
  3. on_the_list-0.1.0/README.md +394 -0
  4. on_the_list-0.1.0/on_the_list/__init__.py +3 -0
  5. on_the_list-0.1.0/on_the_list/__main__.py +4 -0
  6. on_the_list-0.1.0/on_the_list/analyse.py +162 -0
  7. on_the_list-0.1.0/on_the_list/annexes.py +331 -0
  8. on_the_list-0.1.0/on_the_list/checks.py +521 -0
  9. on_the_list-0.1.0/on_the_list/cli.py +233 -0
  10. on_the_list-0.1.0/on_the_list/data/annex_II.csv +2488 -0
  11. on_the_list-0.1.0/on_the_list/data/annex_III.csv +1766 -0
  12. on_the_list-0.1.0/on_the_list/data/annex_IV.csv +175 -0
  13. on_the_list-0.1.0/on_the_list/data/annex_V.csv +138 -0
  14. on_the_list-0.1.0/on_the_list/data/annex_VI.csv +111 -0
  15. on_the_list-0.1.0/on_the_list/fetch.py +84 -0
  16. on_the_list-0.1.0/on_the_list/ingredients.py +450 -0
  17. on_the_list-0.1.0/on_the_list/models.py +213 -0
  18. on_the_list-0.1.0/on_the_list/normalise.py +94 -0
  19. on_the_list-0.1.0/on_the_list/paths.py +41 -0
  20. on_the_list-0.1.0/on_the_list/register.py +152 -0
  21. on_the_list-0.1.0/on_the_list/registry_manifest.py +98 -0
  22. on_the_list-0.1.0/on_the_list/report.py +448 -0
  23. on_the_list-0.1.0/on_the_list.egg-info/PKG-INFO +417 -0
  24. on_the_list-0.1.0/on_the_list.egg-info/SOURCES.txt +33 -0
  25. on_the_list-0.1.0/on_the_list.egg-info/dependency_links.txt +1 -0
  26. on_the_list-0.1.0/on_the_list.egg-info/entry_points.txt +2 -0
  27. on_the_list-0.1.0/on_the_list.egg-info/top_level.txt +1 -0
  28. on_the_list-0.1.0/pyproject.toml +46 -0
  29. on_the_list-0.1.0/setup.cfg +4 -0
  30. on_the_list-0.1.0/tests/test_annexes.py +245 -0
  31. on_the_list-0.1.0/tests/test_checks.py +359 -0
  32. on_the_list-0.1.0/tests/test_cli.py +260 -0
  33. on_the_list-0.1.0/tests/test_ingredients.py +413 -0
  34. on_the_list-0.1.0/tests/test_offline.py +393 -0
  35. on_the_list-0.1.0/tests/test_register.py +175 -0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Waiga Arya
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,417 @@
1
+ Metadata-Version: 2.4
2
+ Name: on-the-list
3
+ Version: 0.1.0
4
+ Summary: Reads a cosmetic ingredient list and reports what the EU's official annexes say about the ingredients on it.
5
+ Author: Waiga Arya
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/Waiga/on-the-list
8
+ Project-URL: Issues, https://github.com/Waiga/on-the-list/issues
9
+ Keywords: cosmetics,inci,ingredients,cosing,labelling,cli
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Environment :: Console
12
+ Classifier: Intended Audience :: Manufacturing
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.11
15
+ Classifier: Programming Language :: Python :: 3.12
16
+ Classifier: Programming Language :: Python :: 3.13
17
+ Classifier: Programming Language :: Python :: 3.14
18
+ Classifier: Topic :: Text Processing :: Linguistic
19
+ Requires-Python: >=3.11
20
+ Description-Content-Type: text/markdown
21
+ License-File: LICENSE
22
+ Dynamic: license-file
23
+
24
+ # on-the-list
25
+
26
+ Reads a cosmetic ingredient list and reports what the EU's official annexes say
27
+ about the ingredients on it.
28
+
29
+ Runs entirely on your machine. No account, no API key, no upload, no
30
+ dependencies beyond Python itself. The annexes are shipped with the package;
31
+ downloading a fresh copy is a separate command you have to type.
32
+
33
+ What came back when it was pointed at 16,635 real published labels, including
34
+ the checks it gets wrong and how often, is written up in
35
+ [An ingredient list cannot tell you most of what you want to know](https://medium.com/@aryawaiga0/an-ingredient-list-cannot-tell-you-most-of-what-you-want-to-know-f3807f357837).
36
+
37
+ ```
38
+ $ on-the-list examples/shade-stick.txt
39
+
40
+ on-the-list 0.1.0 — examples/shade-stick.txt
41
+ CosIng annexes II-VI, Commission last update 28/08/2026, 1913 distinct names,
42
+ read from the copy shipped with this package
43
+
44
+ 10 ingredients read. 2 prohibited, 1 colourant order, 1 repeated entry.
45
+
46
+ NAMES THAT MATCH AN ANNEX II ENTRY
47
+ Annex II is the list of substances prohibited in cosmetic products. A match
48
+ is reported here as a match. Whether it means anything about this product
49
+ is for someone with the formulation in front of them.
50
+
51
+ Butylphenyl Methylpropional matches Annex II, entry 1666, in the
52
+ Commission's list of substances prohibited in cosmetic products.
53
+ printed as: Butylphenyl Methylpropional
54
+ register name: BUTYLPHENYL METHYLPROPIONAL (read from the 'identified'
55
+ column)
56
+ annex entry: 2-(4-tert-butylbenzyl) propionaldehyde
57
+ ```
58
+
59
+ ## The thing to understand before anything else
60
+
61
+ **An ingredient that appears in none of the annexes is not an error, and this
62
+ tool never reports it as one.**
63
+
64
+ Annexes II to VI are lists of *restricted* substances: prohibited, restricted,
65
+ permitted colourants, permitted preservatives, permitted UV filters. They are
66
+ not an inventory of valid INCI names. Aqua and Glycerin appear in none of them,
67
+ because nothing restricts Aqua or Glycerin.
68
+
69
+ So *"that is not a real INCI name"* is not a check this tool can make, and it
70
+ does not attempt one. Absence from the annexes means **these annexes say
71
+ nothing about it**, and the report says exactly that, in those words. There is
72
+ no `unknown` and no `unrecognised` anywhere in the output.
73
+
74
+ This is the most likely way to misread the tool, which is why it is the first
75
+ thing on the page. The reasoning is set out in full in
76
+ [`docs/superpowers/specs/2026-09-09-on-the-list-design.md`](docs/superpowers/specs/2026-09-09-on-the-list-design.md).
77
+
78
+ ## What it does not do
79
+
80
+ It never says a product is compliant, safe, legal, permitted or clean. It
81
+ compares names printed on a label with names printed in a published annex, and
82
+ reports where they coincide. Whether that means anything about a particular
83
+ product depends on its concentration, its product type, its route of exposure
84
+ and its formulation, none of which an ingredient list states.
85
+
86
+ That is not modesty. It is the only claim the evidence supports.
87
+
88
+ ## Install
89
+
90
+ ```bash
91
+ pip install .
92
+ ```
93
+
94
+ Python 3.11 or newer. Nothing else.
95
+
96
+ ## Use
97
+
98
+ ```bash
99
+ on-the-list label.txt # whole pack text; the list is found in it
100
+ cat product-page.txt | on-the-list -
101
+ on-the-list --ingredients inci.txt # you already have just the list
102
+ on-the-list --ingredients inci.txt --pack-text pack.txt # runs the warning check too
103
+ on-the-list label.txt --format markdown # to paste into an email
104
+ on-the-list label.txt --format json # for a pipeline
105
+ on-the-list label.txt --skip colourant-order
106
+ on-the-list --list-checks
107
+ on-the-list register # which register is in use, and its hashes
108
+ on-the-list update-register # the only command that uses the network
109
+ ```
110
+
111
+ Exit codes: `0` nothing found, `1` at least one finding, `2` could not run. A
112
+ label whose ingredient list did not parse exits `2`, never `0`, and so does a
113
+ run with every check switched off: a green CI job over a label that nothing was
114
+ compared to is the one thing this tool exists not to do.
115
+
116
+ ## The four checks
117
+
118
+ Each one can be switched off on its own with `--skip`, and the report always
119
+ prints which ran and which did not, with a reason. A check that did not run must
120
+ never read as a check that found nothing.
121
+
122
+ | | |
123
+ |---|---|
124
+ | **prohibited** | An ingredient's name matches an entry in **Annex II**. Reported as a match with the entry number, split into entries that say something unconditional about the substance and entries whose own wording sets a condition. |
125
+ | **colourant-order** | A `CI NNNNN` colourant is printed before a non-colourant. Article 19(1)(g) allows colourants in any order *after* the other ingredients, so this is a positional observation with no judgement in it. |
126
+ | **repeated-entry** | The same name appears twice in the declared list. Needs no register at all. |
127
+ | **warning-wording** | An annex attaches a `Contains …` statement to an ingredient on the list, and that statement is not in the pack text you supplied. Reported as a **gap between two documents**, never as a violation: the wording may be printed somewhere the supplied text does not cover, in another language, or the entry's condition may not apply. It needs pack text to search — a whole label file counts as that, `--ingredients` on its own does not — and it says which when it does not run. |
128
+
129
+ Two things that are not checks and change no exit code:
130
+
131
+ **What the annexes name.** Every entry in Annexes II to VI that names an
132
+ ingredient on the list, with what that annex is — prohibited, restricted, a
133
+ permitted colourant, preservative or UV filter — and the product types the entry
134
+ is limited to. This is the tool's own title, and for a while it was the one
135
+ thing the report did not print: a label containing Phenoxyethanol produced no
136
+ finding, and the reader was told one of three ingredients was named somewhere
137
+ and never told where.
138
+
139
+ **Coverage.** How many ingredients no annex entry names, labelled *not
140
+ restricted by these annexes*.
141
+
142
+ ### Four narrowings that cost coverage on purpose
143
+
144
+ **The position check uses only bare `CI NNNNN` forms.** Named Annex IV
145
+ substances are not used, because most do more than one job: Titanium Dioxide is
146
+ an Annex IV colourant, an Annex VI UV filter and an opacifier, and where it
147
+ sits in a list proves nothing about which of those it is doing. A CI number is
148
+ unambiguous.
149
+
150
+ **Only `Contains …` statements are searched for in pack text.** The annexes'
151
+ wording column mixes label text with conditions that have nothing to do with a
152
+ label — *"Purity criteria as set out in Commission Directive 95/45/EC (E 129)"*,
153
+ *"Only nanomaterials having the following characteristics are allowed"*.
154
+ Treating the whole column as label text would report a missing warning on every
155
+ sunscreen containing zinc oxide. 38 of the 627 entries in Annexes III to VI
156
+ yield a statement this tool will look for; the rest are shown to the reader and
157
+ not searched for, and the report says so.
158
+
159
+ **A match on part of a printed name is not reported against Annex II.**
160
+ `CI 77288 / CHROMIUM` is a colourant printed beside its element name, and the
161
+ fragment `CHROMIUM` matches the Annex II entry for chromium metal. Over 16,635
162
+ real labels, fragment matches produced no correct Annex II finding at all, so
163
+ they are shown under *considered and not counted* with the reason instead. A
164
+ name with a bracketed aside removed — `Titanium Dioxide (nano)` → `Titanium
165
+ Dioxide` — is still the name of the ingredient, and does match.
166
+
167
+ **Nothing is built on a concentration.** A label does not state one. Every
168
+ maximum-concentration column in the annexes is unused.
169
+
170
+ ## Reproducibility
171
+
172
+ The register shipped in this package is the European Commission's own CosIng
173
+ export, byte for byte. Every report names the version it used, and you can
174
+ check it:
175
+
176
+ ```bash
177
+ $ on-the-list register
178
+ CosIng annexes II-VI, Commission last update 28/08/2026, 1913 distinct names,
179
+ read from the copy shipped with this package
180
+
181
+ CosIng, © European Union, 1995-2026. Annexes II to VI of Regulation (EC)
182
+ No 1223/2009, retrieved from the European Commission CosIng API. Reused under
183
+ Commission Decision 2011/833/EU (CC BY 4.0). Unmodified.
184
+
185
+ Annex II — list of substances prohibited in cosmetic products
186
+ 1758 entries, Commission last update 28/08/2026
187
+ sha256 b7105a05bf724bb10cebaac3a3813b6146c3153ae1d07097d35bb1f73cd7283b
188
+ https://api.tech.ec.europa.eu/cosing20/1.0/api/annexes/II/export-csv
189
+ ```
190
+
191
+ ```bash
192
+ curl -s https://api.tech.ec.europa.eu/cosing20/1.0/api/annexes/II/export-csv | shasum -a 256
193
+ ```
194
+
195
+ If that hash no longer matches, the Commission has republished the annex, and a
196
+ result produced against the old one is a result about a document that no longer
197
+ exists. `on-the-list update-register` downloads the current files; the report
198
+ then names the new hashes instead.
199
+
200
+ Full manifest, including the corpus used for the measurements below:
201
+ [`docs/corpus-manifest.md`](docs/corpus-manifest.md).
202
+
203
+ ## Measured against real labels
204
+
205
+ Unit tests pass on the inputs their author imagined, which proves very little.
206
+ This was run over **16,635 real published cosmetic labels** from an Open Beauty
207
+ Facts export — real packs, real messiness, none of it written by this project.
208
+ The selection rule is every record in the export whose `ingredients_text` field
209
+ is at least 50 characters. Nothing else is filtered or sampled.
210
+
211
+ `tools/measure_corpus.py` is the script that produced every number here.
212
+
213
+ | | |
214
+ |---|---|
215
+ | Records in the export | 64,237 |
216
+ | Labels selected | 16,635 |
217
+ | Labels whose list parsed | 16,535 |
218
+ | Crashes | 0 |
219
+ | Ingredients parsed | 298,423 |
220
+ | **Annex II matches** | **2,408** on 2,037 labels |
221
+ | — where the annex entry is unconditional | 1,066 on 942 labels |
222
+ | — where the annex entry sets a condition | 1,342 on 1,167 labels |
223
+ | Colourant position | 975 on 645 labels |
224
+ | Repeated entries | 712 on 291 labels |
225
+ | Warning wording not found | 1,170 on 1,027 labels |
226
+ | Considered and not counted | 554 |
227
+ | Labels with pack prose inside the ingredient panel | 4,288 (25.8%) |
228
+ | Labels whose ingredient field carries a section heading | 1,631 (9.8%) |
229
+
230
+ Every one of those comes out of `tools/measure_corpus.py`, including the
231
+ per-entry and per-statement breakdowns quoted below. `--samples` draws the
232
+ audit samples with the seed and size the manifest names.
233
+
234
+ ### How wrong each check is
235
+
236
+ Hand-audited against the published annex text or the label's own list. These
237
+ are the numbers, not a summary of them.
238
+
239
+ **prohibited — 9.3% wrong on the unconditional group.** All 19 Annex II entries
240
+ behind the 1,066 unconditional findings were read against the annex text.
241
+ Eighteen are correct: Butylphenyl Methylpropional (651 findings), Zinc
242
+ Pyrithione (105), Hydroxyisohexyl 3-Cyclohexene Carboxaldehyde (100),
243
+ Pentasodium Pentetate (65), Ergocalciferol and Cholecalciferol (11), borates
244
+ (7), 4-Methylbenzylidene Camphor (5), Formaldehyde (4), and a tail of one to
245
+ three each. One is wrong, and it is the tool's largest single known false
246
+ positive: **Annex II entry 1388 is `Octamethylcyclotetrasiloxane; D4`, and the
247
+ Commission's own identified-ingredients column gives its INCI name as
248
+ `CYCLOMETHICONE`**. Cyclomethicone names a *mixture* of cyclic siloxanes, so a
249
+ label printing it has not said it contains D4. That produced 99 findings. There
250
+ is no mechanical signal separating this from a correct match, so it is reported
251
+ here rather than patched around.
252
+
253
+ The 1,342 conditional findings are counted neither right nor wrong, because an
254
+ ingredient list cannot settle them. They are dominated by entries whose
255
+ condition no label states: petrolatum *"except if the full refining history is
256
+ known"*, the furocoumarin entry that names ordinary citrus oils *"except for
257
+ normal content in natural essences used"*, and the colourants prohibited only
258
+ *"when used as a substance in hair dye products"*. Each is printed with its own
259
+ wording so a reader can see the condition and decide.
260
+
261
+ **colourant-order — 4 of 30 wrong.** Thirty findings drawn by
262
+ `tools/measure_corpus.py --samples out.json` (seed 11, one row per label), each
263
+ read against the label's own list. Two failures are pack prose that the parser
264
+ kept — Spanish marketing copy after the colourants, and an OCR'd panel where
265
+ the list has no commas at all. One is a make-up palette declaring several
266
+ shades in one field. One is an OCR'd label whose text before `INGREDIENTS` is
267
+ unreadable.
268
+
269
+ The underlying weakness is deeper than that rate suggests, and it is stated in
270
+ [`docs/limitations.md`](docs/limitations.md): the Regulation *permits*
271
+ colourants after the other ingredients but does not *require* it, so a colourant
272
+ in weight order is not necessarily out of place. The check reports where a name
273
+ is printed and nothing more.
274
+
275
+ **repeated-entry — 16 of 30 wrong.** Thirty findings from the same draw, read
276
+ the same way. This is the weakest check in the tool and the number is not a
277
+ typo.
278
+
279
+ Eight of the sixteen are **multi-component packs**: a hair colour kit declaring
280
+ the crème, the developer and the conditioner in one field; a 2-in-1 shampoo. An
281
+ ingredient appearing once in each component is not a repeated ingredient. The
282
+ other eight are fields that are not one product's ingredient list at all — an
283
+ alphabetical glossary of every ingredient a brand uses, a food supplement, a
284
+ certification mark parsed as an ingredient, and OCR damage that split one name
285
+ into two.
286
+
287
+ The tool detects the multi-component case when the pack prints a heading it can
288
+ see — `Gel N°1:`, `MASQUE:`, a second `Ingredients:` — and then says so and does
289
+ not run either order-dependent check. It flagged 1,631 of the 16,635 labels that
290
+ way. It cannot detect it when the pack prints no heading, which is most of the
291
+ time, and no heuristic tried here improved that materially without losing real
292
+ findings.
293
+
294
+ **On a single product's own ingredient list — which is what the tool is for —
295
+ none of the sixteen failure modes applies.** All twelve sound findings in the
296
+ same sample were single-product lists. The corpus number is still the honest one
297
+ to publish, and it is 53%.
298
+
299
+ **warning-wording — not measured.** Open Beauty Facts carries no field holding
300
+ the text printed on a pack, so every finding in the corpus run is a gap by
301
+ construction and the rate means nothing. What the run does establish is that the
302
+ check fires on the right population: of the 1,170 findings, 692 are
303
+ `Contains sodium fluoride` on fluoride toothpastes, 132
304
+ `Contains sodium monofluorophosphate`, 127 `Contains hydrogen peroxide` on
305
+ developers, 93 `Contains ammonia` on hair colour and 48
306
+ `Contains Benzophenone-3` on sunscreens. Its accuracy is untested, and this
307
+ README will say so until a corpus with real pack text exists.
308
+
309
+ ### What the measurement found that the tests did not
310
+
311
+ Every one of these was a real defect, found only by meeting 16,635 real labels
312
+ or by reading the annexes against their own text:
313
+
314
+ - **Feeding `csv.reader` a list of lines destroys newlines inside quoted
315
+ fields.** The annexes put ten-line warning blocks in one cell.
316
+ `Contains selenium disulphide\nAvoid contact with eyes` silently became one
317
+ string. The extractor found 32 statements instead of 38, and 19 of the ones it
318
+ did find ran on into the sentence after them. No error, all tests green. The
319
+ reader uses `io.StringIO`.
320
+ - **The identified-ingredients column is comma-separated and its values contain
321
+ commas.** A plain `split(",")` gives a different answer from the parser on 52
322
+ rows, and on 38 of them it leaves a one-character fragment: a substance called
323
+ `N`, from `N,N-DIETHYL-m-AMINOPHENOL`.
324
+ - **Annex II entry 1725 is `Styrene/Acrylates copolymer (nano)` and lists the
325
+ ordinary INCI name beside it.** 421 labels printing the ordinary polymer
326
+ matched a prohibited entry that is not about it. They are now reported as
327
+ considered and not counted, with the reason.
328
+ - **HTML entities were resolved after splitting, not before.** The `;` inside
329
+ `<` is a separator: one Russian pack produced three ingredients called
330
+ `&lt`, an American one produced three called `FD`, and each set was reported
331
+ as a repeated ingredient.
332
+ - **`[+/- CI 77491, CI 77492]` closes its bracket after the colourants, not
333
+ after the marker.** A pattern that expected `[+/-]` as a unit matched none of
334
+ it, so a whole shade-range block was read as declared ingredients.
335
+ - **Toothpastes print `Contient du fluorure de sodium (1450 ppm de fluor)`
336
+ immediately after the colourants with no break.** Read as an ingredient, it
337
+ was the commonest reason a colourant appeared to be printed in front of a
338
+ non-colourant.
339
+ - **`Cl 77492` with a lowercase L is endemic in scanned panels**, and appears
340
+ beside a correctly read `CI` in the same list.
341
+ - **`Styrene / Acrylates Copolymer` is one name, not two.** Treating a spaced
342
+ slash as a synonym separator left the fragment `Styrene`, which matches the
343
+ styrene monomer in Annex II, on 35 real labels. Requiring a space on both
344
+ sides of the slash left 21 of them; a polymer names its monomers with slashes,
345
+ and the rule now says so, which leaves none.
346
+ - **Hydroquinone is prohibited by Annex II "with the exception of entry 14 in
347
+ Annex III"**, where it is allowed in professional nail products. Reporting
348
+ only the prohibition is a half-truth, so a match that also appears in another
349
+ annex now says which.
350
+ - **Only 314 of Annex II's 1,758 rows carry an INCI name at all.** The rest are
351
+ chemical or CAS identifiers with no cosmetic-glossary equivalent. The
352
+ prohibited check can only see 18% of the prohibited list.
353
+ - **A file handed to `--ingredients` usually starts with the word
354
+ "Ingredients:".** It was not stripped on that path, so the first ingredient
355
+ became `Ingredients: Formaldehyde`, matched nothing, and the run exited 0 —
356
+ while the same text through the whole-label path reported the match.
357
+ - **`Chromium (CI 77288)` was reported as prohibited and `CI 77288 / CHROMIUM`
358
+ was not.** Same substance, two print orders, opposite answers. What survives a
359
+ bracket being removed is the name only if it is at least half the words.
360
+ - **`Glycerin +/- 0.5%` opened a shade-range block.** Every declared ingredient
361
+ after a printed tolerance dropped out of both order-dependent checks, and the
362
+ report said only that the list had a "may contain" block.
363
+ - **The position check was quadratic.** 20,000 colour index numbers took 64
364
+ seconds; the hostile-input tests never called the checks, only the parser.
365
+ - **A colourant found only in Annex II was described as "listed in Annex IV".**
366
+ Four are: CI 12150, CI 20170, CI 27290 and CI 45425, each prohibited in hair
367
+ dye. The finding now names the entries the register actually holds.
368
+ - **A malformed `--register` directory raised a traceback and exited 1** — the
369
+ code for "findings were reported" — and the message about the Commission
370
+ having changed the export format was unreachable.
371
+ - **`--skip` on all four checks exited 0.** Nothing was compared, and the report
372
+ said "nothing found".
373
+
374
+ The corpus itself is not redistributed here. Open Beauty Facts is ODbL 1.0,
375
+ which is share-alike and incompatible with this repository's MIT licence. Only
376
+ the measurements are published, with the export's hash so you can obtain the
377
+ same file.
378
+
379
+ ## What it misses
380
+
381
+ Stated plainly, because a checking tool that hides its blind spots is worse than
382
+ none. The long version is in [`docs/limitations.md`](docs/limitations.md).
383
+
384
+ - **It sees 18% of Annex II.** Most rows have no INCI name, so most prohibited
385
+ substances cannot be matched against a label at all.
386
+ - **It cannot see concentration.** Most annex restrictions are limits, and a
387
+ limit is not something a name can breach.
388
+ - **It cannot see product type.** A restriction on hair dye says nothing about a
389
+ face cream, and an ingredient list does not say which the product is.
390
+ - **It matches exact names only, after folding.** No fuzzy matching, no
391
+ substring search, no edit distance. A misspelling, a supplier's trade name, or
392
+ a name the Commission writes differently will not match, and the tool will not
393
+ tell you it missed one. It does look a printed name up under a bracketed or
394
+ slash-separated part of itself — that is how `CI 77891` is found inside
395
+ `Titanium Dioxide (CI 77891)` — but **only the Annex II check refuses a match
396
+ found that way**; the other checks and the coverage count accept it.
397
+ - **It reads five annexes.** Annex I (the safety report) and Annex VII are not
398
+ read, and neither is any national requirement, retailer standard, or rule
399
+ outside the EU.
400
+ - **A pack panel inside the ingredient field derails everything.** A quarter of
401
+ the corpus had prose in the panel. The tool says so when it can tell.
402
+ - **A "may contain" block is excluded from the position check** and is printed
403
+ across a whole shade range, so nothing positional can be said about it.
404
+ - **Crowd-sourced label data can be wrong.** In the measurement above, a finding
405
+ may be about a mistyped record rather than about a pack.
406
+
407
+ ## Licence
408
+
409
+ MIT, for the code. See `LICENSE`.
410
+
411
+ The annex data in `on_the_list/data/` is European Commission material: CosIng,
412
+ © European Union, 1995-2026, Annexes II to VI of Regulation (EC) No 1223/2009,
413
+ retrieved from the CosIng API and redistributed unmodified. Reuse is governed by
414
+ Commission Decision 2011/833/EU; the Commission's legal notice states that
415
+ content it owns is licensed under CC BY 4.0.
416
+
417
+ Nothing in this repository is legal or regulatory advice.