on-the-list 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- on_the_list-0.1.0/LICENSE +21 -0
- on_the_list-0.1.0/PKG-INFO +417 -0
- on_the_list-0.1.0/README.md +394 -0
- on_the_list-0.1.0/on_the_list/__init__.py +3 -0
- on_the_list-0.1.0/on_the_list/__main__.py +4 -0
- on_the_list-0.1.0/on_the_list/analyse.py +162 -0
- on_the_list-0.1.0/on_the_list/annexes.py +331 -0
- on_the_list-0.1.0/on_the_list/checks.py +521 -0
- on_the_list-0.1.0/on_the_list/cli.py +233 -0
- on_the_list-0.1.0/on_the_list/data/annex_II.csv +2488 -0
- on_the_list-0.1.0/on_the_list/data/annex_III.csv +1766 -0
- on_the_list-0.1.0/on_the_list/data/annex_IV.csv +175 -0
- on_the_list-0.1.0/on_the_list/data/annex_V.csv +138 -0
- on_the_list-0.1.0/on_the_list/data/annex_VI.csv +111 -0
- on_the_list-0.1.0/on_the_list/fetch.py +84 -0
- on_the_list-0.1.0/on_the_list/ingredients.py +450 -0
- on_the_list-0.1.0/on_the_list/models.py +213 -0
- on_the_list-0.1.0/on_the_list/normalise.py +94 -0
- on_the_list-0.1.0/on_the_list/paths.py +41 -0
- on_the_list-0.1.0/on_the_list/register.py +152 -0
- on_the_list-0.1.0/on_the_list/registry_manifest.py +98 -0
- on_the_list-0.1.0/on_the_list/report.py +448 -0
- on_the_list-0.1.0/on_the_list.egg-info/PKG-INFO +417 -0
- on_the_list-0.1.0/on_the_list.egg-info/SOURCES.txt +33 -0
- on_the_list-0.1.0/on_the_list.egg-info/dependency_links.txt +1 -0
- on_the_list-0.1.0/on_the_list.egg-info/entry_points.txt +2 -0
- on_the_list-0.1.0/on_the_list.egg-info/top_level.txt +1 -0
- on_the_list-0.1.0/pyproject.toml +46 -0
- on_the_list-0.1.0/setup.cfg +4 -0
- on_the_list-0.1.0/tests/test_annexes.py +245 -0
- on_the_list-0.1.0/tests/test_checks.py +359 -0
- on_the_list-0.1.0/tests/test_cli.py +260 -0
- on_the_list-0.1.0/tests/test_ingredients.py +413 -0
- on_the_list-0.1.0/tests/test_offline.py +393 -0
- on_the_list-0.1.0/tests/test_register.py +175 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Waiga Arya
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,417 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: on-the-list
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Reads a cosmetic ingredient list and reports what the EU's official annexes say about the ingredients on it.
|
|
5
|
+
Author: Waiga Arya
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/Waiga/on-the-list
|
|
8
|
+
Project-URL: Issues, https://github.com/Waiga/on-the-list/issues
|
|
9
|
+
Keywords: cosmetics,inci,ingredients,cosing,labelling,cli
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Environment :: Console
|
|
12
|
+
Classifier: Intended Audience :: Manufacturing
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
18
|
+
Classifier: Topic :: Text Processing :: Linguistic
|
|
19
|
+
Requires-Python: >=3.11
|
|
20
|
+
Description-Content-Type: text/markdown
|
|
21
|
+
License-File: LICENSE
|
|
22
|
+
Dynamic: license-file
|
|
23
|
+
|
|
24
|
+
# on-the-list
|
|
25
|
+
|
|
26
|
+
Reads a cosmetic ingredient list and reports what the EU's official annexes say
|
|
27
|
+
about the ingredients on it.
|
|
28
|
+
|
|
29
|
+
Runs entirely on your machine. No account, no API key, no upload, no
|
|
30
|
+
dependencies beyond Python itself. The annexes are shipped with the package;
|
|
31
|
+
downloading a fresh copy is a separate command you have to type.
|
|
32
|
+
|
|
33
|
+
What came back when it was pointed at 16,635 real published labels, including
|
|
34
|
+
the checks it gets wrong and how often, is written up in
|
|
35
|
+
[An ingredient list cannot tell you most of what you want to know](https://medium.com/@aryawaiga0/an-ingredient-list-cannot-tell-you-most-of-what-you-want-to-know-f3807f357837).
|
|
36
|
+
|
|
37
|
+
```
|
|
38
|
+
$ on-the-list examples/shade-stick.txt
|
|
39
|
+
|
|
40
|
+
on-the-list 0.1.0 — examples/shade-stick.txt
|
|
41
|
+
CosIng annexes II-VI, Commission last update 28/08/2026, 1913 distinct names,
|
|
42
|
+
read from the copy shipped with this package
|
|
43
|
+
|
|
44
|
+
10 ingredients read. 2 prohibited, 1 colourant order, 1 repeated entry.
|
|
45
|
+
|
|
46
|
+
NAMES THAT MATCH AN ANNEX II ENTRY
|
|
47
|
+
Annex II is the list of substances prohibited in cosmetic products. A match
|
|
48
|
+
is reported here as a match. Whether it means anything about this product
|
|
49
|
+
is for someone with the formulation in front of them.
|
|
50
|
+
|
|
51
|
+
Butylphenyl Methylpropional matches Annex II, entry 1666, in the
|
|
52
|
+
Commission's list of substances prohibited in cosmetic products.
|
|
53
|
+
printed as: Butylphenyl Methylpropional
|
|
54
|
+
register name: BUTYLPHENYL METHYLPROPIONAL (read from the 'identified'
|
|
55
|
+
column)
|
|
56
|
+
annex entry: 2-(4-tert-butylbenzyl) propionaldehyde
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## The thing to understand before anything else
|
|
60
|
+
|
|
61
|
+
**An ingredient that appears in none of the annexes is not an error, and this
|
|
62
|
+
tool never reports it as one.**
|
|
63
|
+
|
|
64
|
+
Annexes II to VI are lists of *restricted* substances: prohibited, restricted,
|
|
65
|
+
permitted colourants, permitted preservatives, permitted UV filters. They are
|
|
66
|
+
not an inventory of valid INCI names. Aqua and Glycerin appear in none of them,
|
|
67
|
+
because nothing restricts Aqua or Glycerin.
|
|
68
|
+
|
|
69
|
+
So *"that is not a real INCI name"* is not a check this tool can make, and it
|
|
70
|
+
does not attempt one. Absence from the annexes means **these annexes say
|
|
71
|
+
nothing about it**, and the report says exactly that, in those words. There is
|
|
72
|
+
no `unknown` and no `unrecognised` anywhere in the output.
|
|
73
|
+
|
|
74
|
+
This is the most likely way to misread the tool, which is why it is the first
|
|
75
|
+
thing on the page. The reasoning is set out in full in
|
|
76
|
+
[`docs/superpowers/specs/2026-09-09-on-the-list-design.md`](docs/superpowers/specs/2026-09-09-on-the-list-design.md).
|
|
77
|
+
|
|
78
|
+
## What it does not do
|
|
79
|
+
|
|
80
|
+
It never says a product is compliant, safe, legal, permitted or clean. It
|
|
81
|
+
compares names printed on a label with names printed in a published annex, and
|
|
82
|
+
reports where they coincide. Whether that means anything about a particular
|
|
83
|
+
product depends on its concentration, its product type, its route of exposure
|
|
84
|
+
and its formulation, none of which an ingredient list states.
|
|
85
|
+
|
|
86
|
+
That is not modesty. It is the only claim the evidence supports.
|
|
87
|
+
|
|
88
|
+
## Install
|
|
89
|
+
|
|
90
|
+
```bash
|
|
91
|
+
pip install .
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
Python 3.11 or newer. Nothing else.
|
|
95
|
+
|
|
96
|
+
## Use
|
|
97
|
+
|
|
98
|
+
```bash
|
|
99
|
+
on-the-list label.txt # whole pack text; the list is found in it
|
|
100
|
+
cat product-page.txt | on-the-list -
|
|
101
|
+
on-the-list --ingredients inci.txt # you already have just the list
|
|
102
|
+
on-the-list --ingredients inci.txt --pack-text pack.txt # runs the warning check too
|
|
103
|
+
on-the-list label.txt --format markdown # to paste into an email
|
|
104
|
+
on-the-list label.txt --format json # for a pipeline
|
|
105
|
+
on-the-list label.txt --skip colourant-order
|
|
106
|
+
on-the-list --list-checks
|
|
107
|
+
on-the-list register # which register is in use, and its hashes
|
|
108
|
+
on-the-list update-register # the only command that uses the network
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Exit codes: `0` nothing found, `1` at least one finding, `2` could not run. A
|
|
112
|
+
label whose ingredient list did not parse exits `2`, never `0`, and so does a
|
|
113
|
+
run with every check switched off: a green CI job over a label that nothing was
|
|
114
|
+
compared to is the one thing this tool exists not to do.
|
|
115
|
+
|
|
116
|
+
## The four checks
|
|
117
|
+
|
|
118
|
+
Each one can be switched off on its own with `--skip`, and the report always
|
|
119
|
+
prints which ran and which did not, with a reason. A check that did not run must
|
|
120
|
+
never read as a check that found nothing.
|
|
121
|
+
|
|
122
|
+
| | |
|
|
123
|
+
|---|---|
|
|
124
|
+
| **prohibited** | An ingredient's name matches an entry in **Annex II**. Reported as a match with the entry number, split into entries that say something unconditional about the substance and entries whose own wording sets a condition. |
|
|
125
|
+
| **colourant-order** | A `CI NNNNN` colourant is printed before a non-colourant. Article 19(1)(g) allows colourants in any order *after* the other ingredients, so this is a positional observation with no judgement in it. |
|
|
126
|
+
| **repeated-entry** | The same name appears twice in the declared list. Needs no register at all. |
|
|
127
|
+
| **warning-wording** | An annex attaches a `Contains …` statement to an ingredient on the list, and that statement is not in the pack text you supplied. Reported as a **gap between two documents**, never as a violation: the wording may be printed somewhere the supplied text does not cover, in another language, or the entry's condition may not apply. It needs pack text to search — a whole label file counts as that, `--ingredients` on its own does not — and it says which when it does not run. |
|
|
128
|
+
|
|
129
|
+
Two things that are not checks and change no exit code:
|
|
130
|
+
|
|
131
|
+
**What the annexes name.** Every entry in Annexes II to VI that names an
|
|
132
|
+
ingredient on the list, with what that annex is — prohibited, restricted, a
|
|
133
|
+
permitted colourant, preservative or UV filter — and the product types the entry
|
|
134
|
+
is limited to. This is the tool's own title, and for a while it was the one
|
|
135
|
+
thing the report did not print: a label containing Phenoxyethanol produced no
|
|
136
|
+
finding, and the reader was told one of three ingredients was named somewhere
|
|
137
|
+
and never told where.
|
|
138
|
+
|
|
139
|
+
**Coverage.** How many ingredients no annex entry names, labelled *not
|
|
140
|
+
restricted by these annexes*.
|
|
141
|
+
|
|
142
|
+
### Four narrowings that cost coverage on purpose
|
|
143
|
+
|
|
144
|
+
**The position check uses only bare `CI NNNNN` forms.** Named Annex IV
|
|
145
|
+
substances are not used, because most do more than one job: Titanium Dioxide is
|
|
146
|
+
an Annex IV colourant, an Annex VI UV filter and an opacifier, and where it
|
|
147
|
+
sits in a list proves nothing about which of those it is doing. A CI number is
|
|
148
|
+
unambiguous.
|
|
149
|
+
|
|
150
|
+
**Only `Contains …` statements are searched for in pack text.** The annexes'
|
|
151
|
+
wording column mixes label text with conditions that have nothing to do with a
|
|
152
|
+
label — *"Purity criteria as set out in Commission Directive 95/45/EC (E 129)"*,
|
|
153
|
+
*"Only nanomaterials having the following characteristics are allowed"*.
|
|
154
|
+
Treating the whole column as label text would report a missing warning on every
|
|
155
|
+
sunscreen containing zinc oxide. 38 of the 627 entries in Annexes III to VI
|
|
156
|
+
yield a statement this tool will look for; the rest are shown to the reader and
|
|
157
|
+
not searched for, and the report says so.
|
|
158
|
+
|
|
159
|
+
**A match on part of a printed name is not reported against Annex II.**
|
|
160
|
+
`CI 77288 / CHROMIUM` is a colourant printed beside its element name, and the
|
|
161
|
+
fragment `CHROMIUM` matches the Annex II entry for chromium metal. Over 16,635
|
|
162
|
+
real labels, fragment matches produced no correct Annex II finding at all, so
|
|
163
|
+
they are shown under *considered and not counted* with the reason instead. A
|
|
164
|
+
name with a bracketed aside removed — `Titanium Dioxide (nano)` → `Titanium
|
|
165
|
+
Dioxide` — is still the name of the ingredient, and does match.
|
|
166
|
+
|
|
167
|
+
**Nothing is built on a concentration.** A label does not state one. Every
|
|
168
|
+
maximum-concentration column in the annexes is unused.
|
|
169
|
+
|
|
170
|
+
## Reproducibility
|
|
171
|
+
|
|
172
|
+
The register shipped in this package is the European Commission's own CosIng
|
|
173
|
+
export, byte for byte. Every report names the version it used, and you can
|
|
174
|
+
check it:
|
|
175
|
+
|
|
176
|
+
```bash
|
|
177
|
+
$ on-the-list register
|
|
178
|
+
CosIng annexes II-VI, Commission last update 28/08/2026, 1913 distinct names,
|
|
179
|
+
read from the copy shipped with this package
|
|
180
|
+
|
|
181
|
+
CosIng, © European Union, 1995-2026. Annexes II to VI of Regulation (EC)
|
|
182
|
+
No 1223/2009, retrieved from the European Commission CosIng API. Reused under
|
|
183
|
+
Commission Decision 2011/833/EU (CC BY 4.0). Unmodified.
|
|
184
|
+
|
|
185
|
+
Annex II — list of substances prohibited in cosmetic products
|
|
186
|
+
1758 entries, Commission last update 28/08/2026
|
|
187
|
+
sha256 b7105a05bf724bb10cebaac3a3813b6146c3153ae1d07097d35bb1f73cd7283b
|
|
188
|
+
https://api.tech.ec.europa.eu/cosing20/1.0/api/annexes/II/export-csv
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
```bash
|
|
192
|
+
curl -s https://api.tech.ec.europa.eu/cosing20/1.0/api/annexes/II/export-csv | shasum -a 256
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
If that hash no longer matches, the Commission has republished the annex, and a
|
|
196
|
+
result produced against the old one is a result about a document that no longer
|
|
197
|
+
exists. `on-the-list update-register` downloads the current files; the report
|
|
198
|
+
then names the new hashes instead.
|
|
199
|
+
|
|
200
|
+
Full manifest, including the corpus used for the measurements below:
|
|
201
|
+
[`docs/corpus-manifest.md`](docs/corpus-manifest.md).
|
|
202
|
+
|
|
203
|
+
## Measured against real labels
|
|
204
|
+
|
|
205
|
+
Unit tests pass on the inputs their author imagined, which proves very little.
|
|
206
|
+
This was run over **16,635 real published cosmetic labels** from an Open Beauty
|
|
207
|
+
Facts export — real packs, real messiness, none of it written by this project.
|
|
208
|
+
The selection rule is every record in the export whose `ingredients_text` field
|
|
209
|
+
is at least 50 characters. Nothing else is filtered or sampled.
|
|
210
|
+
|
|
211
|
+
`tools/measure_corpus.py` is the script that produced every number here.
|
|
212
|
+
|
|
213
|
+
| | |
|
|
214
|
+
|---|---|
|
|
215
|
+
| Records in the export | 64,237 |
|
|
216
|
+
| Labels selected | 16,635 |
|
|
217
|
+
| Labels whose list parsed | 16,535 |
|
|
218
|
+
| Crashes | 0 |
|
|
219
|
+
| Ingredients parsed | 298,423 |
|
|
220
|
+
| **Annex II matches** | **2,408** on 2,037 labels |
|
|
221
|
+
| — where the annex entry is unconditional | 1,066 on 942 labels |
|
|
222
|
+
| — where the annex entry sets a condition | 1,342 on 1,167 labels |
|
|
223
|
+
| Colourant position | 975 on 645 labels |
|
|
224
|
+
| Repeated entries | 712 on 291 labels |
|
|
225
|
+
| Warning wording not found | 1,170 on 1,027 labels |
|
|
226
|
+
| Considered and not counted | 554 |
|
|
227
|
+
| Labels with pack prose inside the ingredient panel | 4,288 (25.8%) |
|
|
228
|
+
| Labels whose ingredient field carries a section heading | 1,631 (9.8%) |
|
|
229
|
+
|
|
230
|
+
Every one of those comes out of `tools/measure_corpus.py`, including the
|
|
231
|
+
per-entry and per-statement breakdowns quoted below. `--samples` draws the
|
|
232
|
+
audit samples with the seed and size the manifest names.
|
|
233
|
+
|
|
234
|
+
### How wrong each check is
|
|
235
|
+
|
|
236
|
+
Hand-audited against the published annex text or the label's own list. These
|
|
237
|
+
are the numbers, not a summary of them.
|
|
238
|
+
|
|
239
|
+
**prohibited — 9.3% wrong on the unconditional group.** All 19 Annex II entries
|
|
240
|
+
behind the 1,066 unconditional findings were read against the annex text.
|
|
241
|
+
Eighteen are correct: Butylphenyl Methylpropional (651 findings), Zinc
|
|
242
|
+
Pyrithione (105), Hydroxyisohexyl 3-Cyclohexene Carboxaldehyde (100),
|
|
243
|
+
Pentasodium Pentetate (65), Ergocalciferol and Cholecalciferol (11), borates
|
|
244
|
+
(7), 4-Methylbenzylidene Camphor (5), Formaldehyde (4), and a tail of one to
|
|
245
|
+
three each. One is wrong, and it is the tool's largest single known false
|
|
246
|
+
positive: **Annex II entry 1388 is `Octamethylcyclotetrasiloxane; D4`, and the
|
|
247
|
+
Commission's own identified-ingredients column gives its INCI name as
|
|
248
|
+
`CYCLOMETHICONE`**. Cyclomethicone names a *mixture* of cyclic siloxanes, so a
|
|
249
|
+
label printing it has not said it contains D4. That produced 99 findings. There
|
|
250
|
+
is no mechanical signal separating this from a correct match, so it is reported
|
|
251
|
+
here rather than patched around.
|
|
252
|
+
|
|
253
|
+
The 1,342 conditional findings are counted neither right nor wrong, because an
|
|
254
|
+
ingredient list cannot settle them. They are dominated by entries whose
|
|
255
|
+
condition no label states: petrolatum *"except if the full refining history is
|
|
256
|
+
known"*, the furocoumarin entry that names ordinary citrus oils *"except for
|
|
257
|
+
normal content in natural essences used"*, and the colourants prohibited only
|
|
258
|
+
*"when used as a substance in hair dye products"*. Each is printed with its own
|
|
259
|
+
wording so a reader can see the condition and decide.
|
|
260
|
+
|
|
261
|
+
**colourant-order — 4 of 30 wrong.** Thirty findings drawn by
|
|
262
|
+
`tools/measure_corpus.py --samples out.json` (seed 11, one row per label), each
|
|
263
|
+
read against the label's own list. Two failures are pack prose that the parser
|
|
264
|
+
kept — Spanish marketing copy after the colourants, and an OCR'd panel where
|
|
265
|
+
the list has no commas at all. One is a make-up palette declaring several
|
|
266
|
+
shades in one field. One is an OCR'd label whose text before `INGREDIENTS` is
|
|
267
|
+
unreadable.
|
|
268
|
+
|
|
269
|
+
The underlying weakness is deeper than that rate suggests, and it is stated in
|
|
270
|
+
[`docs/limitations.md`](docs/limitations.md): the Regulation *permits*
|
|
271
|
+
colourants after the other ingredients but does not *require* it, so a colourant
|
|
272
|
+
in weight order is not necessarily out of place. The check reports where a name
|
|
273
|
+
is printed and nothing more.
|
|
274
|
+
|
|
275
|
+
**repeated-entry — 16 of 30 wrong.** Thirty findings from the same draw, read
|
|
276
|
+
the same way. This is the weakest check in the tool and the number is not a
|
|
277
|
+
typo.
|
|
278
|
+
|
|
279
|
+
Eight of the sixteen are **multi-component packs**: a hair colour kit declaring
|
|
280
|
+
the crème, the developer and the conditioner in one field; a 2-in-1 shampoo. An
|
|
281
|
+
ingredient appearing once in each component is not a repeated ingredient. The
|
|
282
|
+
other eight are fields that are not one product's ingredient list at all — an
|
|
283
|
+
alphabetical glossary of every ingredient a brand uses, a food supplement, a
|
|
284
|
+
certification mark parsed as an ingredient, and OCR damage that split one name
|
|
285
|
+
into two.
|
|
286
|
+
|
|
287
|
+
The tool detects the multi-component case when the pack prints a heading it can
|
|
288
|
+
see — `Gel N°1:`, `MASQUE:`, a second `Ingredients:` — and then says so and does
|
|
289
|
+
not run either order-dependent check. It flagged 1,631 of the 16,635 labels that
|
|
290
|
+
way. It cannot detect it when the pack prints no heading, which is most of the
|
|
291
|
+
time, and no heuristic tried here improved that materially without losing real
|
|
292
|
+
findings.
|
|
293
|
+
|
|
294
|
+
**On a single product's own ingredient list — which is what the tool is for —
|
|
295
|
+
none of the sixteen failure modes applies.** All twelve sound findings in the
|
|
296
|
+
same sample were single-product lists. The corpus number is still the honest one
|
|
297
|
+
to publish, and it is 53%.
|
|
298
|
+
|
|
299
|
+
**warning-wording — not measured.** Open Beauty Facts carries no field holding
|
|
300
|
+
the text printed on a pack, so every finding in the corpus run is a gap by
|
|
301
|
+
construction and the rate means nothing. What the run does establish is that the
|
|
302
|
+
check fires on the right population: of the 1,170 findings, 692 are
|
|
303
|
+
`Contains sodium fluoride` on fluoride toothpastes, 132
|
|
304
|
+
`Contains sodium monofluorophosphate`, 127 `Contains hydrogen peroxide` on
|
|
305
|
+
developers, 93 `Contains ammonia` on hair colour and 48
|
|
306
|
+
`Contains Benzophenone-3` on sunscreens. Its accuracy is untested, and this
|
|
307
|
+
README will say so until a corpus with real pack text exists.
|
|
308
|
+
|
|
309
|
+
### What the measurement found that the tests did not
|
|
310
|
+
|
|
311
|
+
Every one of these was a real defect, found only by meeting 16,635 real labels
|
|
312
|
+
or by reading the annexes against their own text:
|
|
313
|
+
|
|
314
|
+
- **Feeding `csv.reader` a list of lines destroys newlines inside quoted
|
|
315
|
+
fields.** The annexes put ten-line warning blocks in one cell.
|
|
316
|
+
`Contains selenium disulphide\nAvoid contact with eyes` silently became one
|
|
317
|
+
string. The extractor found 32 statements instead of 38, and 19 of the ones it
|
|
318
|
+
did find ran on into the sentence after them. No error, all tests green. The
|
|
319
|
+
reader uses `io.StringIO`.
|
|
320
|
+
- **The identified-ingredients column is comma-separated and its values contain
|
|
321
|
+
commas.** A plain `split(",")` gives a different answer from the parser on 52
|
|
322
|
+
rows, and on 38 of them it leaves a one-character fragment: a substance called
|
|
323
|
+
`N`, from `N,N-DIETHYL-m-AMINOPHENOL`.
|
|
324
|
+
- **Annex II entry 1725 is `Styrene/Acrylates copolymer (nano)` and lists the
|
|
325
|
+
ordinary INCI name beside it.** 421 labels printing the ordinary polymer
|
|
326
|
+
matched a prohibited entry that is not about it. They are now reported as
|
|
327
|
+
considered and not counted, with the reason.
|
|
328
|
+
- **HTML entities were resolved after splitting, not before.** The `;` inside
|
|
329
|
+
`<` is a separator: one Russian pack produced three ingredients called
|
|
330
|
+
`<`, an American one produced three called `FD`, and each set was reported
|
|
331
|
+
as a repeated ingredient.
|
|
332
|
+
- **`[+/- CI 77491, CI 77492]` closes its bracket after the colourants, not
|
|
333
|
+
after the marker.** A pattern that expected `[+/-]` as a unit matched none of
|
|
334
|
+
it, so a whole shade-range block was read as declared ingredients.
|
|
335
|
+
- **Toothpastes print `Contient du fluorure de sodium (1450 ppm de fluor)`
|
|
336
|
+
immediately after the colourants with no break.** Read as an ingredient, it
|
|
337
|
+
was the commonest reason a colourant appeared to be printed in front of a
|
|
338
|
+
non-colourant.
|
|
339
|
+
- **`Cl 77492` with a lowercase L is endemic in scanned panels**, and appears
|
|
340
|
+
beside a correctly read `CI` in the same list.
|
|
341
|
+
- **`Styrene / Acrylates Copolymer` is one name, not two.** Treating a spaced
|
|
342
|
+
slash as a synonym separator left the fragment `Styrene`, which matches the
|
|
343
|
+
styrene monomer in Annex II, on 35 real labels. Requiring a space on both
|
|
344
|
+
sides of the slash left 21 of them; a polymer names its monomers with slashes,
|
|
345
|
+
and the rule now says so, which leaves none.
|
|
346
|
+
- **Hydroquinone is prohibited by Annex II "with the exception of entry 14 in
|
|
347
|
+
Annex III"**, where it is allowed in professional nail products. Reporting
|
|
348
|
+
only the prohibition is a half-truth, so a match that also appears in another
|
|
349
|
+
annex now says which.
|
|
350
|
+
- **Only 314 of Annex II's 1,758 rows carry an INCI name at all.** The rest are
|
|
351
|
+
chemical or CAS identifiers with no cosmetic-glossary equivalent. The
|
|
352
|
+
prohibited check can only see 18% of the prohibited list.
|
|
353
|
+
- **A file handed to `--ingredients` usually starts with the word
|
|
354
|
+
"Ingredients:".** It was not stripped on that path, so the first ingredient
|
|
355
|
+
became `Ingredients: Formaldehyde`, matched nothing, and the run exited 0 —
|
|
356
|
+
while the same text through the whole-label path reported the match.
|
|
357
|
+
- **`Chromium (CI 77288)` was reported as prohibited and `CI 77288 / CHROMIUM`
|
|
358
|
+
was not.** Same substance, two print orders, opposite answers. What survives a
|
|
359
|
+
bracket being removed is the name only if it is at least half the words.
|
|
360
|
+
- **`Glycerin +/- 0.5%` opened a shade-range block.** Every declared ingredient
|
|
361
|
+
after a printed tolerance dropped out of both order-dependent checks, and the
|
|
362
|
+
report said only that the list had a "may contain" block.
|
|
363
|
+
- **The position check was quadratic.** 20,000 colour index numbers took 64
|
|
364
|
+
seconds; the hostile-input tests never called the checks, only the parser.
|
|
365
|
+
- **A colourant found only in Annex II was described as "listed in Annex IV".**
|
|
366
|
+
Four are: CI 12150, CI 20170, CI 27290 and CI 45425, each prohibited in hair
|
|
367
|
+
dye. The finding now names the entries the register actually holds.
|
|
368
|
+
- **A malformed `--register` directory raised a traceback and exited 1** — the
|
|
369
|
+
code for "findings were reported" — and the message about the Commission
|
|
370
|
+
having changed the export format was unreachable.
|
|
371
|
+
- **`--skip` on all four checks exited 0.** Nothing was compared, and the report
|
|
372
|
+
said "nothing found".
|
|
373
|
+
|
|
374
|
+
The corpus itself is not redistributed here. Open Beauty Facts is ODbL 1.0,
|
|
375
|
+
which is share-alike and incompatible with this repository's MIT licence. Only
|
|
376
|
+
the measurements are published, with the export's hash so you can obtain the
|
|
377
|
+
same file.
|
|
378
|
+
|
|
379
|
+
## What it misses
|
|
380
|
+
|
|
381
|
+
Stated plainly, because a checking tool that hides its blind spots is worse than
|
|
382
|
+
none. The long version is in [`docs/limitations.md`](docs/limitations.md).
|
|
383
|
+
|
|
384
|
+
- **It sees 18% of Annex II.** Most rows have no INCI name, so most prohibited
|
|
385
|
+
substances cannot be matched against a label at all.
|
|
386
|
+
- **It cannot see concentration.** Most annex restrictions are limits, and a
|
|
387
|
+
limit is not something a name can breach.
|
|
388
|
+
- **It cannot see product type.** A restriction on hair dye says nothing about a
|
|
389
|
+
face cream, and an ingredient list does not say which the product is.
|
|
390
|
+
- **It matches exact names only, after folding.** No fuzzy matching, no
|
|
391
|
+
substring search, no edit distance. A misspelling, a supplier's trade name, or
|
|
392
|
+
a name the Commission writes differently will not match, and the tool will not
|
|
393
|
+
tell you it missed one. It does look a printed name up under a bracketed or
|
|
394
|
+
slash-separated part of itself — that is how `CI 77891` is found inside
|
|
395
|
+
`Titanium Dioxide (CI 77891)` — but **only the Annex II check refuses a match
|
|
396
|
+
found that way**; the other checks and the coverage count accept it.
|
|
397
|
+
- **It reads five annexes.** Annex I (the safety report) and Annex VII are not
|
|
398
|
+
read, and neither is any national requirement, retailer standard, or rule
|
|
399
|
+
outside the EU.
|
|
400
|
+
- **A pack panel inside the ingredient field derails everything.** A quarter of
|
|
401
|
+
the corpus had prose in the panel. The tool says so when it can tell.
|
|
402
|
+
- **A "may contain" block is excluded from the position check** and is printed
|
|
403
|
+
across a whole shade range, so nothing positional can be said about it.
|
|
404
|
+
- **Crowd-sourced label data can be wrong.** In the measurement above, a finding
|
|
405
|
+
may be about a mistyped record rather than about a pack.
|
|
406
|
+
|
|
407
|
+
## Licence
|
|
408
|
+
|
|
409
|
+
MIT, for the code. See `LICENSE`.
|
|
410
|
+
|
|
411
|
+
The annex data in `on_the_list/data/` is European Commission material: CosIng,
|
|
412
|
+
© European Union, 1995-2026, Annexes II to VI of Regulation (EC) No 1223/2009,
|
|
413
|
+
retrieved from the CosIng API and redistributed unmodified. Reuse is governed by
|
|
414
|
+
Commission Decision 2011/833/EU; the Commission's legal notice states that
|
|
415
|
+
content it owns is licensed under CC BY 4.0.
|
|
416
|
+
|
|
417
|
+
Nothing in this repository is legal or regulatory advice.
|