swedish-pii 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +476 -0
- package/dist/index.cjs +178495 -0
- package/dist/index.d.cts +163 -0
- package/dist/index.d.ts +163 -0
- package/dist/index.js +178457 -0
- package/package.json +71 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 okasi
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,476 @@
|
|
|
1
|
+
# 🇸🇪 Swedish PII Detection
|
|
2
|
+
|
|
3
|
+
Detects and masks Personal Identifiable Information (PII) in Swedish text.
|
|
4
|
+
|
|
5
|
+
**▶ Try it live: <https://okasi.github.io/swedish-pii/>** — everything runs
|
|
6
|
+
in your browser; text never leaves the page.
|
|
7
|
+
|
|
8
|
+
Given free text, the engine finds names, identity numbers, addresses, financial
|
|
9
|
+
data, contact details and sensitive attributes, and replaces each occurrence
|
|
10
|
+
with a stable placeholder:
|
|
11
|
+
|
|
12
|
+
```text
|
|
13
|
+
Anna Andersson bor på Storgatan 12, 114 55 Stockholm. Pnr 811218-9876.
|
|
14
|
+
→
|
|
15
|
+
<PER_FIRST_1> <PER_LAST_1> bor på <SE_STREET_ADDRESS_1>, <SE_POSTAL_CODE_1> Stockholm. Pnr <SE_PERSONAL_IDENTITY_NUMBER_MALE_1>.
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
## 🚀 Quick Start
|
|
21
|
+
|
|
22
|
+
```zsh
|
|
23
|
+
npm install
|
|
24
|
+
npm run dev # demo UI at http://localhost:3000
|
|
25
|
+
npm test # run the test suite
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
### 📦 Library Usage
|
|
29
|
+
|
|
30
|
+
The core lives in [`src/lib`](src/lib/index.ts), has no framework
|
|
31
|
+
dependencies, and ships as the `swedish-pii` npm package (ESM + CJS +
|
|
32
|
+
TypeScript types, datasets bundled in — zero runtime dependencies):
|
|
33
|
+
|
|
34
|
+
```zsh
|
|
35
|
+
npm install swedish-pii
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
```ts
|
|
39
|
+
import { maskPII, detectPII } from "swedish-pii"; // or "@/lib" inside this repo
|
|
40
|
+
|
|
41
|
+
const { maskedText, maskedData, entities } = maskPII(text);
|
|
42
|
+
// entities carry exact { start, end } offsets and a confidence `score`
|
|
43
|
+
|
|
44
|
+
// Presidio-style confidence: checksum-validated matches score ~0.95,
|
|
45
|
+
// gazetteer hits ~0.85–0.9, plain shapes ~0.6, failed checksums ~0.45,
|
|
46
|
+
// context-starved shapes ~0.25. Tune the cutoff per use case:
|
|
47
|
+
maskPII(text, { scoreThreshold: 0.2 }); // recall-first (e.g. raw CSV columns)
|
|
48
|
+
maskPII(text, { scoreThreshold: 0.9 }); // precision-first
|
|
49
|
+
// scoreThreshold must be a finite number from 0 through 1
|
|
50
|
+
|
|
51
|
+
// Strict mode drops anything that fails its checksum (Luhn for cards,
|
|
52
|
+
// personnummer and org numbers; mod-97 for IBANs) or calendar check:
|
|
53
|
+
const strict = maskPII(text, { strict: true });
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
### 🌐 HTTP API
|
|
57
|
+
|
|
58
|
+
```zsh
|
|
59
|
+
curl -X POST localhost:3000/api/mask \
|
|
60
|
+
-H 'Content-Type: application/json' \
|
|
61
|
+
-d '{"text": "Anna bor i Åre kommun", "strict": false, "scoreThreshold": 0.4}'
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
Returns `{ maskedText, maskedData, entities }` (each entity carries a
|
|
65
|
+
confidence `score`).
|
|
66
|
+
|
|
67
|
+
### ⚙️ How It Works
|
|
68
|
+
|
|
69
|
+
Every detector runs against the **original** text and yields entity spans
|
|
70
|
+
with a confidence score (Presidio-style: validators always run and adjust
|
|
71
|
+
the score, context cues boost or dim it, and `scoreThreshold` decides what
|
|
72
|
+
survives). Overlapping spans are resolved by detector priority (most
|
|
73
|
+
specific first — e.g. a personnummer wins over the generic bank-account
|
|
74
|
+
pattern), then the masked text is built in a single pass. Detectors never
|
|
75
|
+
see each other's placeholders, so masked output can't be corrupted by
|
|
76
|
+
later patterns. Patterns — including the large list-based alternations —
|
|
77
|
+
are compiled once and reused, so throughput stays in the MB/s range
|
|
78
|
+
(~0.1 ms per short text, ~15 ms for a 100 kB document).
|
|
79
|
+
|
|
80
|
+
Names are matched against the SCB name lists: "First Last" pairs are
|
|
81
|
+
fuzzy-matched with Jaro–Winkler (length-bucketed and memoized), single
|
|
82
|
+
capitalized words by exact lookup. Compound given names absent from the
|
|
83
|
+
list ("Ecenur" = Ece + Nur) are recognized by decomposing them into two
|
|
84
|
+
registered same-gender names, using tail-suffixes learned from the name
|
|
85
|
+
data itself — guarded by stopword, morpheme and place-name exclusions so
|
|
86
|
+
prose words ("Engine", "Semester") never decompose.
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## 📦 Current Version
|
|
91
|
+
|
|
92
|
+
### 🔍 What Can Be Detected?
|
|
93
|
+
|
|
94
|
+
<details>
|
|
95
|
+
<summary>💳 Financial</summary>
|
|
96
|
+
|
|
97
|
+
- American Express credit card numbers
|
|
98
|
+
- Mastercard credit card numbers
|
|
99
|
+
- Visa credit card numbers
|
|
100
|
+
- Swedish IBAN codes
|
|
101
|
+
- Swedish BIC codes
|
|
102
|
+
- Swedish bank account numbers
|
|
103
|
+
- Swedish Bankgiro & Plusgiro numbers (Luhn-validated)
|
|
104
|
+
- Swedish VAT numbers
|
|
105
|
+
- Cryptocurrency wallet addresses (BTC, ETH)
|
|
106
|
+
|
|
107
|
+
</details>
|
|
108
|
+
|
|
109
|
+
<details>
|
|
110
|
+
<summary>🆔 Identification Numbers</summary>
|
|
111
|
+
|
|
112
|
+
- Swedish personal identity numbers (male)
|
|
113
|
+
- Swedish personal identity numbers (female)
|
|
114
|
+
- Swedish coordination numbers (male)
|
|
115
|
+
- Swedish coordination numbers (female)
|
|
116
|
+
- Swedish passport / national ID card numbers
|
|
117
|
+
|
|
118
|
+
</details>
|
|
119
|
+
|
|
120
|
+
<details>
|
|
121
|
+
<summary>📞 Contact</summary>
|
|
122
|
+
|
|
123
|
+
- Email addresses
|
|
124
|
+
- Phone numbers
|
|
125
|
+
- Social media information
|
|
126
|
+
|
|
127
|
+
</details>
|
|
128
|
+
|
|
129
|
+
<details>
|
|
130
|
+
<summary>📍 Location</summary>
|
|
131
|
+
|
|
132
|
+
- Swedish street addresses
|
|
133
|
+
- Swedish postal codes
|
|
134
|
+
- Swedish municipalities
|
|
135
|
+
- Swedish counties
|
|
136
|
+
- Swedish cities
|
|
137
|
+
- Property designations (fastighetsbeteckningar)
|
|
138
|
+
- GPS coordinates
|
|
139
|
+
|
|
140
|
+
</details>
|
|
141
|
+
|
|
142
|
+
<details>
|
|
143
|
+
<summary>🏢 Work / Education</summary>
|
|
144
|
+
|
|
145
|
+
- Swedish work organizations
|
|
146
|
+
- Swedish education organizations
|
|
147
|
+
- Swedish education programs
|
|
148
|
+
- Swedish work professions
|
|
149
|
+
- Swedish organization numbers
|
|
150
|
+
|
|
151
|
+
</details>
|
|
152
|
+
|
|
153
|
+
<details>
|
|
154
|
+
<summary>🔒 Sensitive Attributes</summary>
|
|
155
|
+
|
|
156
|
+
- Marital status information
|
|
157
|
+
- Genetic sex information
|
|
158
|
+
- Disability information
|
|
159
|
+
- Religion information
|
|
160
|
+
- Sexual orientation information
|
|
161
|
+
- Demographic information
|
|
162
|
+
- Political ideologies information
|
|
163
|
+
- Labor union membership (GDPR Art. 9)
|
|
164
|
+
|
|
165
|
+
</details>
|
|
166
|
+
|
|
167
|
+
<details>
|
|
168
|
+
<summary>🧩 Misc</summary>
|
|
169
|
+
|
|
170
|
+
- Swedish license plate information
|
|
171
|
+
- IP addresses (IPv4 & IPv6)
|
|
172
|
+
- MAC addresses
|
|
173
|
+
- Date information
|
|
174
|
+
- Time information
|
|
175
|
+
- Court & authority case numbers (dnr, mål nr)
|
|
176
|
+
- Age information
|
|
177
|
+
|
|
178
|
+
</details>
|
|
179
|
+
|
|
180
|
+
<details>
|
|
181
|
+
<summary>👤 Names</summary>
|
|
182
|
+
|
|
183
|
+
- Top `20,524` male first names with at least 10 bearers in Sweden 1999-2020
|
|
184
|
+
- Top `23,347` female first names with at least 10 bearers in Sweden 1999-2020
|
|
185
|
+
- Top `107,762` last names with at least 10 bearers in Sweden 1999-2020
|
|
186
|
+
|
|
187
|
+
</details>
|
|
188
|
+
|
|
189
|
+
For the exact label emitted per category (e.g. `PER_FIRST`, `SE_BANK_NUMBER`,
|
|
190
|
+
`MARITAL_STATUS`), see [Detector Labels](#-detector-labels) below.
|
|
191
|
+
|
|
192
|
+
## 🛣️ Roadmap
|
|
193
|
+
|
|
194
|
+
- Patterns for:
|
|
195
|
+
- ~~Passport numbers~~ ✅
|
|
196
|
+
- Residence Permit Number?
|
|
197
|
+
- ~~Bank Account Number (Bankgiro/Plusgiro)~~ ✅
|
|
198
|
+
- Set lookups for:
|
|
199
|
+
- Localities 🏘️
|
|
200
|
+
- ~~Cities 🏙️~~ ✅
|
|
201
|
+
- ~~Labor Unions~~ ✅ (`SE_LABOR_UNION`, context-aware for ambiguous names like "Vision")
|
|
202
|
+
- Reduce false positives
|
|
203
|
+
- Improve performance & simplify
|
|
204
|
+
- ~~Comprehensive tests~~ ✅ (`npm test`, 300+ tests across detectors, validators, matching, and the engine)
|
|
205
|
+
- ~~Make it to a npm package (library)~~ ✅ (`npm run build:lib`, publish with `npm publish`)
|
|
206
|
+
- Make a documentation page (frontend)
|
|
207
|
+
|
|
208
|
+
---
|
|
209
|
+
|
|
210
|
+
## 📚 Dataset Sources
|
|
211
|
+
|
|
212
|
+
- **Names:**
|
|
213
|
+
- SCB 2020 First names:
|
|
214
|
+
<https://www.statistikdatabasen.scb.se/pxweb/en/ssd/START__BE__BE0001__BE0001G/BE0001FNamn10/>
|
|
215
|
+
- SCB 2020 Last names:
|
|
216
|
+
<https://www.statistikdatabasen.scb.se/pxweb/en/ssd/START__BE__BE0001__BE0001G/BE0001ENamn10/>
|
|
217
|
+
- Skatteverket newborn names (you can write "*" to select all):
|
|
218
|
+
<https://www6.skatteverket.se/sense/app/c13f8ffe-f90d-4c38-b426-646ee1226b75/sheet/50b8c57d-23f9-4f74-bd6f-7a18d8096226/state/analysis>
|
|
219
|
+
- Skatteverket popular surnames:
|
|
220
|
+
<https://skatteverket.se/privat/folkbokforing/namn/bytaefternamn/sokblanddevanligasteefternamnen.4.515a6be615c637b9aa48e09.html>
|
|
221
|
+
|
|
222
|
+
- **Education Programs:**
|
|
223
|
+
- <https://github.com/swedishdata/education-work-social/blob/master/scb-sun-2000.csv>
|
|
224
|
+
|
|
225
|
+
- **Professions:**
|
|
226
|
+
- <https://github.com/swedishdata/education-work-social/blob/master/arbetsformedlingen-job-titles.csv>
|
|
227
|
+
|
|
228
|
+
- **Marital Status:**
|
|
229
|
+
- <https://github.com/swedishdata/education-work-social/blob/master/scb-family-marital-status-and-consensual-union-terms.csv>
|
|
230
|
+
|
|
231
|
+
- **Sexual Orientation:**
|
|
232
|
+
- <https://glaad.org/reference/terms>
|
|
233
|
+
|
|
234
|
+
- **Addresses:**
|
|
235
|
+
- Street & Postal (partial):
|
|
236
|
+
<https://github.com/beshrkayali/sverige_postnummer>
|
|
237
|
+
|
|
238
|
+
- **Most up-to-date addresses:**
|
|
239
|
+
- <https://www.postnummerservice.se/>
|
|
240
|
+
- <https://www.postnord.se/en/our-tools/search-postcode-and-address/>
|
|
241
|
+
- <https://download.geofabrik.de/europe/sweden.html>
|
|
242
|
+
|
|
243
|
+
---
|
|
244
|
+
|
|
245
|
+
## 🗺️ Extracting Data from OpenStreetMap via Osmium
|
|
246
|
+
|
|
247
|
+
All of the commands below are also wrapped in
|
|
248
|
+
[`scripts/extract-osm-data.sh`](scripts/extract-osm-data.sh), which downloads
|
|
249
|
+
and caches `sweden-latest.osm.pbf` automatically, validates each output
|
|
250
|
+
before overwriting the previous file, and cleans up all intermediates:
|
|
251
|
+
|
|
252
|
+
```zsh
|
|
253
|
+
scripts/extract-osm-data.sh # everything
|
|
254
|
+
scripts/extract-osm-data.sh counties # or one dataset:
|
|
255
|
+
# counties | cities | municipalities
|
|
256
|
+
# | areas | streets
|
|
257
|
+
```
|
|
258
|
+
|
|
259
|
+
The `streets` target also derives `data/streets.json` (unique street
|
|
260
|
+
names), which powers the exact-lookup part of the `SE_STREET_ADDRESS`
|
|
261
|
+
detector.
|
|
262
|
+
|
|
263
|
+
The raw commands it runs are documented here for reference:
|
|
264
|
+
|
|
265
|
+
### 🛠️ Install Tools via Homebrew
|
|
266
|
+
|
|
267
|
+
```zsh
|
|
268
|
+
brew install osmium-tool
|
|
269
|
+
brew install jq
|
|
270
|
+
```
|
|
271
|
+
|
|
272
|
+
### 🏞️ Extract Counties
|
|
273
|
+
|
|
274
|
+
```zsh
|
|
275
|
+
osmium tags-filter sweden-latest.osm.pbf "nwr/admin_level=4" -o counties.osm.pbf && \
|
|
276
|
+
osmium export counties.osm.pbf -o counties.geojson && \
|
|
277
|
+
jq -r '
|
|
278
|
+
[
|
|
279
|
+
.features[]
|
|
280
|
+
| select(
|
|
281
|
+
.properties.admin_level == "4"
|
|
282
|
+
and .properties.name != null
|
|
283
|
+
and (.properties.name | test("län$"))
|
|
284
|
+
)
|
|
285
|
+
| .properties.name
|
|
286
|
+
]
|
|
287
|
+
| sort
|
|
288
|
+
| unique
|
|
289
|
+
' counties.geojson > counties.json && \
|
|
290
|
+
rm counties.osm.pbf counties.geojson
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
### 🏙️ Extract Cities
|
|
294
|
+
|
|
295
|
+
```zsh
|
|
296
|
+
osmium tags-filter sweden-latest.osm.pbf "nwr/place=city" -o cities.osm.pbf && \
|
|
297
|
+
osmium export cities.osm.pbf -o cities.geojson && \
|
|
298
|
+
jq -r '
|
|
299
|
+
[
|
|
300
|
+
.features[]
|
|
301
|
+
| select(
|
|
302
|
+
.properties.name != null
|
|
303
|
+
and (.properties.name | test("^[0-9]+$") | not)
|
|
304
|
+
and (.properties.name | test("vägen$|kommun$") | not)
|
|
305
|
+
)
|
|
306
|
+
| .properties.name
|
|
307
|
+
]
|
|
308
|
+
| sort
|
|
309
|
+
| unique
|
|
310
|
+
' cities.geojson > cities.json && \
|
|
311
|
+
rm cities.osm.pbf cities.geojson
|
|
312
|
+
```
|
|
313
|
+
|
|
314
|
+
### 🏛️ Extract Municipalities
|
|
315
|
+
|
|
316
|
+
```zsh
|
|
317
|
+
osmium tags-filter sweden-latest.osm.pbf nwr/boundary=administrative -o municipalities.osm.pbf && \
|
|
318
|
+
osmium export municipalities.osm.pbf -o municipalities.geojson && \
|
|
319
|
+
jq -r '
|
|
320
|
+
.features
|
|
321
|
+
| map(select(.properties.admin_level == "7" and .properties.name != null and .properties.name != "Svartån"))
|
|
322
|
+
| map(.properties.name)
|
|
323
|
+
| map(if endswith(" kommun") or . == "Göteborgs Stad" then . else . + " kommun" end)
|
|
324
|
+
| sort
|
|
325
|
+
| unique
|
|
326
|
+
' municipalities.geojson > municipalities.json && \
|
|
327
|
+
rm municipalities.osm.pbf municipalities.geojson
|
|
328
|
+
```
|
|
329
|
+
|
|
330
|
+
### 🏘️ Extract Suburbs & Neighborhoods
|
|
331
|
+
|
|
332
|
+
```zsh
|
|
333
|
+
osmium tags-filter sweden-latest.osm.pbf "nwr/place=suburb" -o suburbs.osm.pbf && \
|
|
334
|
+
osmium tags-filter sweden-latest.osm.pbf "nwr/place=neighborhood" -o neighborhoods.osm.pbf && \
|
|
335
|
+
osmium tags-filter sweden-latest.osm.pbf "nwr/place=town" -o towns.osm.pbf && \
|
|
336
|
+
osmium merge suburbs.osm.pbf neighborhoods.osm.pbf towns.osm.pbf -o areas.osm.pbf && \
|
|
337
|
+
osmium export areas.osm.pbf -o areas.geojson && \
|
|
338
|
+
jq -r '
|
|
339
|
+
[
|
|
340
|
+
.features[]
|
|
341
|
+
| select(
|
|
342
|
+
.properties.name != null
|
|
343
|
+
and (.properties.name | test("^[0-9]+$") | not)
|
|
344
|
+
)
|
|
345
|
+
| .properties.name
|
|
346
|
+
]
|
|
347
|
+
| sort
|
|
348
|
+
| unique
|
|
349
|
+
' areas.geojson > areas.json && \
|
|
350
|
+
rm suburbs.osm.pbf neighborhoods.osm.pbf towns.osm.pbf areas.osm.pbf areas.geojson
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
### 🚦 Extract Streets & Postalcodes
|
|
354
|
+
|
|
355
|
+
```zsh
|
|
356
|
+
osmium tags-filter sweden-latest.osm.pbf 'addr:street=*' 'addr:postcode=*' -o addresses.osm.pbf && \
|
|
357
|
+
osmium export addresses.osm.pbf -o addresses.geojson && \
|
|
358
|
+
jq -r '
|
|
359
|
+
.features
|
|
360
|
+
| map(
|
|
361
|
+
select(
|
|
362
|
+
(.properties["addr:street"] != null) and
|
|
363
|
+
(.properties["addr:postcode"] != null) and
|
|
364
|
+
(.properties["addr:street"] | test("^[0-9]+$") | not)
|
|
365
|
+
)
|
|
366
|
+
)
|
|
367
|
+
| map({
|
|
368
|
+
street: .properties["addr:street"],
|
|
369
|
+
postcode: (
|
|
370
|
+
.properties["addr:postcode"]
|
|
371
|
+
| gsub(" "; "")
|
|
372
|
+
| if test("^[0-9]{5}$") then .[:3] + " " + .[3:] else . end
|
|
373
|
+
),
|
|
374
|
+
housenumber: (
|
|
375
|
+
(.properties["addr:housenumber"] // "")
|
|
376
|
+
| split(",")
|
|
377
|
+
| map(gsub("^ +| +$"; ""))
|
|
378
|
+
)
|
|
379
|
+
})
|
|
380
|
+
| group_by(.street)
|
|
381
|
+
| map(
|
|
382
|
+
{
|
|
383
|
+
street: .[0].street,
|
|
384
|
+
postalcodes: (map(.postcode) | unique),
|
|
385
|
+
housenumbers: (map(.housenumber) | add | unique)
|
|
386
|
+
}
|
|
387
|
+
)
|
|
388
|
+
| sort_by(.street)
|
|
389
|
+
' addresses.geojson > streets_postalcodes.json && \
|
|
390
|
+
rm addresses.osm.pbf addresses.geojson
|
|
391
|
+
```
|
|
392
|
+
|
|
393
|
+
---
|
|
394
|
+
|
|
395
|
+
## 🧪 RegEx Ideas
|
|
396
|
+
|
|
397
|
+
- All allowed characters in names:
|
|
398
|
+
<https://skatteverket.se/privat/folkbokforing/namn.4.18e1b10334ebe8bc80004083.html#Accordionrubrik>
|
|
399
|
+
|
|
400
|
+
---
|
|
401
|
+
|
|
402
|
+
## 🏷️ Detector Labels
|
|
403
|
+
|
|
404
|
+
| Category | Labels |
|
|
405
|
+
| --- | --- |
|
|
406
|
+
| 👤 Names | `PER_FIRST`, `PER_LAST` |
|
|
407
|
+
| 🆔 Identity numbers | `SE_PERSONAL_IDENTITY_NUMBER_MALE/FEMALE`, `SE_COORDINATION_NUMBER_MALE/FEMALE`, `SE_PASSPORT_NUMBER` |
|
|
408
|
+
| 💳 Financial | `AMEX_CREDIT_CARD`, `MASTERCARD_CREDIT_CARD`, `VISA_CREDIT_CARD`, `IBAN_CODE`, `BIC_CODE`, `SE_BANK_NUMBER`, `SE_BANKGIRO`, `SE_PLUSGIRO`, `SE_VAT_NUMBER`, `CRYPTO_WALLET` |
|
|
409
|
+
| 📞 Contact | `EMAIL_ADDRESS`, `PHONE_NUMBER`, `SOCIAL_MEDIA` |
|
|
410
|
+
| 📍 Location | `SE_STREET_ADDRESS`, `SE_POSTAL_CODE`, `SE_MUNICIPALITY`, `SE_COUNTY`, `SE_CITY`, `SE_PROPERTY_DESIGNATION`, `COORDINATE` |
|
|
411
|
+
| 🏢 Work / education | `SE_WORK_ORGANIZATION`, `SE_EDUCATION_ORGANIZATION`, `SE_EDUCATION_PROGRAM`, `SE_WORK_PROFESSION`, `SE_ORGANIZATION_NUMBER` |
|
|
412
|
+
| 🔒 Sensitive attributes | `MARITAL_STATUS`, `GENETIC_SEX`, `DISABILITY`, `RELIGION`, `SEXUAL_ORIENTATION`, `DEMOGRAPHIC`, `POLITICAL_IDEOLOGIES`, `SE_LABOR_UNION` |
|
|
413
|
+
| 🧩 Misc | `SE_LICENSE_PLATE`, `IP_ADDRESS` (v4+v6), `MAC_ADDRESS`, `DATE`, `TIME`, `SE_CASE_NUMBER`, `AGE` |
|
|
414
|
+
|
|
415
|
+
---
|
|
416
|
+
|
|
417
|
+
## 🗂️ Project Layout
|
|
418
|
+
|
|
419
|
+
```
|
|
420
|
+
src/lib/ framework-free detection & masking core
|
|
421
|
+
engine.ts span-based masking engine + detector priority order
|
|
422
|
+
detectors/ one module per category
|
|
423
|
+
validation/ Luhn checksums, calendar-date validation
|
|
424
|
+
matching/ Jaro–Winkler similarity
|
|
425
|
+
src/app/ Next.js demo UI + POST /api/mask route
|
|
426
|
+
data/ name/profession/program lists (SCB, Arbetsförmedlingen)
|
|
427
|
+
data/raw/ location datasets extracted from OpenStreetMap
|
|
428
|
+
scripts/ data extraction pipeline
|
|
429
|
+
tests/ vitest suite
|
|
430
|
+
```
|
|
431
|
+
|
|
432
|
+
## 🧑💻 Development
|
|
433
|
+
|
|
434
|
+
```zsh
|
|
435
|
+
npm run dev # dev server
|
|
436
|
+
npm test # tests (vitest)
|
|
437
|
+
npm run test:watch # tests in watch mode
|
|
438
|
+
npm run typecheck # tsc --noEmit
|
|
439
|
+
npm run lint # eslint
|
|
440
|
+
npm run build # production build
|
|
441
|
+
```
|
|
442
|
+
|
|
443
|
+
CI runs lint, typecheck, tests and the build on every push and PR
|
|
444
|
+
([.github/workflows/ci.yml](.github/workflows/ci.yml)).
|
|
445
|
+
|
|
446
|
+
## ⚠️ Known Limitations
|
|
447
|
+
|
|
448
|
+
- Regex + list lookup, not NLP: single capitalized words that happen to be
|
|
449
|
+
registered names are masked even out of name context. Several precision
|
|
450
|
+
mechanisms keep this in check:
|
|
451
|
+
- a stoplist of ~90 function words ("Vi", "Han", "Men" and "The" are
|
|
452
|
+
all registered SCB names);
|
|
453
|
+
- **context gating** — low-structure patterns only fire near a cue:
|
|
454
|
+
bank numbers need "konto"/"bank"/"account" nearby, BICs need
|
|
455
|
+
"BIC"/"SWIFT", postal codes need an address cue or a following
|
|
456
|
+
capitalized place name ("114 55 Stockholm");
|
|
457
|
+
- BIC/IBAN candidates must carry a real ISO 3166 country code;
|
|
458
|
+
- English demonyms match case-sensitively ("Polish" yes, "polish the
|
|
459
|
+
furniture" no) while Swedish ones stay case-insensitive ("svensk");
|
|
460
|
+
- "gift"/"single" skip English determiner and compound contexts
|
|
461
|
+
("a gift", "single sign-on") but keep the marital sense ("är gift").
|
|
462
|
+
|
|
463
|
+
Homographs that are genuinely common names ("Stig", "Bo") stay maskable
|
|
464
|
+
by design. The remaining accepted tradeoffs are pinned in
|
|
465
|
+
[tests/false-positives.test.ts](tests/false-positives.test.ts).
|
|
466
|
+
- Default mode favors recall (synthetic/example numbers are masked even with
|
|
467
|
+
invalid checksums); use `strict: true` to favor precision.
|
|
468
|
+
- Street addresses combine an exact lookup against the OSM street list
|
|
469
|
+
(~22k names in `data/streets.json`, including suffix-less streets like
|
|
470
|
+
"Aftonsången") with a heuristic fallback (capitalized words ending in
|
|
471
|
+
a street suffix) for streets missing from OSM. Single-word street
|
|
472
|
+
names under 5 characters are ignored — they collide with prose.
|
|
473
|
+
|
|
474
|
+
## 📄 License
|
|
475
|
+
|
|
476
|
+
[MIT](LICENSE)
|