localewarden 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,5 +1,7 @@
1
1
  # localewarden
2
2
 
3
+ [![npm](https://img.shields.io/npm/v/localewarden)](https://www.npmjs.com/package/localewarden) [![CI](https://github.com/martinb207/localewarden/actions/workflows/ci.yml/badge.svg)](https://github.com/martinb207/localewarden/actions/workflows/ci.yml) [![license](https://img.shields.io/npm/l/localewarden)](LICENSE)
4
+
3
5
  **Incremental AI translation for JSON locale files.** It translates only what changed, never overwrites a translation a person fixed, and checks every result before it is written.
4
6
 
5
7
  ```bash
@@ -9,7 +11,7 @@ npx localewarden # translate new and changed strings
9
11
  npx localewarden check # quality check, no API calls (use it in CI)
10
12
  ```
11
13
 
12
- Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files. Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
14
+ Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files, with Flutter (`.arb` files) and with fastlane's App Store / Play Store metadata (`.txt` files). Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
13
15
 
14
16
  ## Why
15
17
 
@@ -20,14 +22,22 @@ Translating locale files with a language model is easy once. Keeping 20 language
20
22
  - **Models are inconsistent across batches.** French screens mix "tu" and "vous", Spanish copies English Title Case ("Configure Su Cuenta"), and Polish or Russian address every user as a man.
21
23
  - **Broken output ships silently.** A translated placeholder (`{heures}` instead of `{hours}`) shows raw braces in your app. A dropped `</strong>` breaks the layout. Stray Cyrillic letters end up in a Danish sentence.
22
24
 
23
- localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output.
25
+ localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output. The checks are tuned against that app's real texts (UI, website, long-form learning content and store listings, about 165 MB) to report problems without flooding you with false alarms. On those texts they still find things that slipped through earlier pipelines: sections cut off after the English grew, stray letters from other alphabets, a trial notice left in English.
26
+
27
+ ## How it compares
28
+
29
+ - **Translation platforms** (Crowdin, Lokalise, Phrase, Weblate) are hosted services with editors, translator workflows and review for teams. localewarden is a small CLI that runs in your repository and CI, with no account and no server. If you have professional translators, a platform fits better. If a model translates and people only fix the odd string, this is the lighter setup.
30
+ - **"Translate my JSON with GPT" scripts** usually send every string on every run and overwrite whatever is there. localewarden keeps state, so it only sends what changed, keeps human fixes, and checks the output.
31
+ - **Editor extensions** (such as i18n Ally) help you write and look up keys while coding. localewarden is about filling and maintaining 10 to 40 languages afterwards. The two work well together.
24
32
 
25
33
  ## What it does
26
34
 
27
35
  - **Translates only what changed.** It remembers a hash of each source string per language. New strings are translated. Changed strings are *revised*: the model gets the existing translation and changes only what the source change requires. Removed strings are deleted from every language.
28
36
  - **Protects hand edits.** If someone edited a translation, localewarden detects it, keeps it, and lists it for review. If the source of a hand-edited string changes later, the string is flagged instead of overwritten.
29
- - **Checks every result before writing it.** Broken placeholders, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems are retried and reported.
37
+ - **Checks every result before writing it.** Broken placeholders, injected HTML or scripts, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems (too long, content missing, words left in English) are retried and reported.
30
38
  - **Consistent style per language.** It enforces formal or informal address per language (`du`/`Sie`, `tu`/`vous`, `ты`/`вы` and 16 more), uses sentence case where the language does, avoids gendered forms for "you", and applies local typography (French spacing, `92 %` in German, CJK quotation marks).
39
+ - **Plural forms per language.** For i18next-style keys (`item_one`, `item_other`) it adds the forms a language needs but English lacks, such as Polish `_few` and `_many` or Arabic `_zero`, `_two`, `_few` and `_many` (CLDR plural rules).
40
+ - **Data files and store listings.** Fields like `id`, `type` or `image` are copied instead of translated (`ignoreKeys`), and so are URLs, email addresses and file paths. Length limits per key (`maxLength`) are passed to the model and checked: App Store names, SEO titles, buttons.
31
41
  - **Glossary and protected names.** You choose fixed renderings ("Privacy Policy" -> "Politique de confidentialité") and names that must never be translated. The check accepts grammatical case endings.
32
42
  - **Quality check for CI.** `localewarden check` runs all checks without any API calls and exits non-zero on errors.
33
43
  - **Targeted repair.** `--fix-flagged` asks the model to fix only what the check flagged. The fix is accepted only if the problem is gone and little else changed.
@@ -76,7 +86,7 @@ es Mantén vivas tus plantas sin tener que pensar en ello
76
86
  ja 何も考えなくても、植物を元気に保てます
77
87
  ```
78
88
 
79
- German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic).
89
+ German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic). There are also examples for [Flutter ARB files](examples/flutter) and [App Store / Play Store texts with fastlane](examples/fastlane).
80
90
 
81
91
  ## Quick start
82
92
 
@@ -131,24 +141,29 @@ The same checks run in two places. Right after each model answer, a failed hard
131
141
  | Check | Finds | Severity |
132
142
  | --- | --- | --- |
133
143
  | `placeholder` | `{name}`, `{{count}}`, `%s`, `%1$d`, `%{x}`, `${x}`, `<0></0>` renamed, translated, added or dropped. ICU `plural`/`select` arguments are compared, while plural categories may differ per language. | error |
134
- | `script` | Letters from an alphabet the language does not use ("刺激" in German), or a word that mixes Latin with Cyrillic/Greek lookalikes ("Вarda") | error |
144
+ | `unsafe` | HTML tags, attributes, event handlers or `javascript:`/`data:` URLs that the source does not have. Translations are often rendered as raw HTML, so this would be a script injection. | error |
145
+ | `script` | Letters from an alphabet the language does not use ("刺激" in German), a word that mixes Latin with Cyrillic/Greek lookalikes ("Вarda"), or Simplified characters in Traditional Chinese (`zh-TW`) and the reverse | error |
135
146
  | `markup` | Changed link targets, different number of tags, unclosed or misnested tags, dropped list items | warning (broken tags and changed links: never written) |
136
147
  | `years` | A year from the source missing or changed (citations, dates) | warning |
137
148
  | `formality` | The other form of address than configured, both forms in one string, or masculine-only forms for "you" | warning |
149
+ | `length` | Longer than the `maxLength` configured for the key | warning |
138
150
  | `titlecase` | English Title Case copied into a language that uses sentence case | warning |
139
151
  | `ampersand` | "&" in languages that write the word | warning |
140
152
  | `glossary` | A glossary rendering missing (case endings allowed), or a `doNotTranslate` name translated | warning |
141
153
  | `untranslated` | Identical to the source (prose of 3+ words; "OK" and names are fine) | warning |
142
- | `partial` | Source-language words left inside the translation, an untranslated bold lead-in, or a hedge that became certainty ("tend to" stated as fact) | warning |
154
+ | `partial` | Source-language words left inside the translation, an untranslated bold lead-in, a hedge that became certainty ("tend to" stated as fact), a translation much shorter than its source (content cut off, or the source grew after it was translated), or sibling options that got the same translation although the source differs ("Rarely" and "Occasionally" both "Selten") | warning |
143
155
 
144
156
  ```bash
145
157
  npx localewarden check # counts per language and check
146
158
  npx localewarden check -v # with examples
147
159
  npx localewarden check --strict # exit 1 on warnings too
148
160
  npx localewarden check --json # for scripts
161
+ npx localewarden check --fix # repair placeholders with one possible fix, no API calls
149
162
  ```
150
163
 
151
- Approved hand edits are skipped, except for placeholder and script errors, which break the app either way.
164
+ `--fix` repairs a translated placeholder when the source has exactly one and the translation renamed it (`{stunden}` back to `{hours}`). The file is edited in place, so its formatting stays as it is. Anything less certain is left for `--fix-flagged` or a person.
165
+
166
+ Approved strings are skipped, except for errors (placeholder, unsafe, script), which break the app either way. You can approve any string, not only hand edits: `npx localewarden review --approve de:home.title` tells the check that a person looked at it (for example a pun on a brand name that is correct without the name), and runs leave it alone.
152
167
 
153
168
  ### In CI
154
169
 
@@ -166,6 +181,39 @@ jobs:
166
181
  - run: npx localewarden check
167
182
  ```
168
183
 
184
+ ### Translating automatically
185
+
186
+ When the source language changes on `main`, translate and open a pull request for review:
187
+
188
+ ```yaml
189
+ # .github/workflows/translate.yml
190
+ name: translate
191
+ on:
192
+ push:
193
+ branches: [main]
194
+ paths: ['locales/en.json'] # your source files
195
+ permissions:
196
+ contents: write
197
+ pull-requests: write
198
+ jobs:
199
+ translate:
200
+ runs-on: ubuntu-latest
201
+ steps:
202
+ - uses: actions/checkout@v4
203
+ - uses: actions/setup-node@v4
204
+ with: { node-version: 22 }
205
+ - run: npx localewarden --max-tokens 200000
206
+ env:
207
+ OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
208
+ - uses: peter-evans/create-pull-request@v7
209
+ with:
210
+ branch: localewarden/translations
211
+ title: Update translations
212
+ commit-message: Update translations
213
+ ```
214
+
215
+ The pull request contains the locale files and `.localewarden/`, so a reviewer sees exactly which strings changed.
216
+
169
217
  ### Fixing what the check finds
170
218
 
171
219
  ```bash
@@ -203,6 +251,9 @@ npx localewarden review --release de:home.title # hand it back: next run revis
203
251
  | `glossary` | `{}` | `{"fr": {"Terms of Service": "Conditions d'utilisation"}}` |
204
252
  | `termNotes` | `{}` | Meanings of ambiguous terms, sent only with strings that contain them: `{"snooze": "postpone a reminder"}` |
205
253
  | `instructions` | `{}` | Extra instructions per language, `"*"` for all: `{"es": "Use neutral Latin American Spanish."}` |
254
+ | `ignoreKeys` | `[]` | Keys that are not text, copied from the source: `["id", "type", "**.sources.*"]`. `*` matches within a key segment, `**` across segments; a pattern without a dot matches the last segment anywhere. URLs, emails, file paths and numbers are always copied. |
255
+ | `exclude` | `[]` | Source files to skip: `["locales/{lang}/nav.json"]` |
256
+ | `maxLength` | `{}` | Character limits per key pattern: `{"**.meta.title": 60, "name": 30}`. The model is told the limit; longer results are retried once and reported by the `length` check. |
206
257
  | `placeholders` | built-in | Regular expressions (strings) that match your placeholders. Replaces the built-in list. |
207
258
  | `model` | `"gpt-5.4-mini"` | Any chat model your endpoint offers |
208
259
  | `baseUrl` | `"https://api.openai.com/v1"` | Any OpenAI-compatible endpoint |
@@ -214,6 +265,37 @@ npx localewarden review --release de:home.title # hand it back: next run revis
214
265
  | `batchSize` | `20` | Strings per request (smaller for scripts that need many tokens) |
215
266
  | `stateDir` | `".localewarden"` | Where state and the review list live |
216
267
 
268
+ ### App Store and Play Store listings (fastlane)
269
+
270
+ ```json
271
+ {
272
+ "sourceLanguage": "en-US",
273
+ "targetLanguages": ["de-DE", "fr-FR", "ja"],
274
+ "files": "fastlane/metadata/{lang}/*.txt",
275
+ "exclude": ["fastlane/metadata/{lang}/*_url.txt"],
276
+ "maxLength": { "name": 30, "subtitle": 30, "keywords": 100, "promotional_text": 170, "description": 4000 },
277
+ "termNotes": { "keywords": "a comma-separated keyword list for store search, not a sentence" }
278
+ }
279
+ ```
280
+
281
+ Each `.txt` file is one string, keyed by its file name, so the limits above apply to `name.txt`, `subtitle.txt` and so on.
282
+
283
+ ### Flutter (ARB)
284
+
285
+ ```json
286
+ { "files": "lib/l10n/app_{lang}.arb", "targetLanguages": ["de", "fr", "pt_BR"] }
287
+ ```
288
+
289
+ Metadata (`@@locale`, `@key` descriptions and placeholders) is copied, not translated, and `@@locale` is set to the target language. ICU plurals and selects keep their structure, and each language gets the plural categories it needs.
290
+
291
+ ### Data files
292
+
293
+ For content JSON with ids, types and links, list the non-text keys:
294
+
295
+ ```json
296
+ { "files": "content/**/*.{lang}.json", "ignoreKeys": ["id", "type", "category", "image", "**.sources.*"] }
297
+ ```
298
+
217
299
  ### Other providers
218
300
 
219
301
  ```json
@@ -231,7 +313,7 @@ Small local models make noticeably more mistakes. The checks catch the mechanica
231
313
  ```text
232
314
  localewarden [translate] --dry-run --lang de,fr --fix-flagged --retranslate-all
233
315
  --overwrite-manual --max-tokens <n> --verbose
234
- localewarden check --lang de,fr --verbose --limit <n> --strict --json
316
+ localewarden check --lang de,fr --verbose --limit <n> --strict --json --fix
235
317
  localewarden review --all --approve <sel>... --release <sel>...
236
318
  localewarden init
237
319
  Global: --config <path> --help --version
@@ -257,13 +339,21 @@ const findings = checkProject(config);
257
339
  - Unchanged strings cost nothing. In the example above, changing two English strings and updating four languages took 4 requests and about 4,000 tokens.
258
340
  - Strings, keys, your `context` and glossary are sent to the API you configure. Nothing else is sent anywhere. There is no telemetry.
259
341
 
342
+ ## Security
343
+
344
+ Translations are treated as untrusted: any markup the source does not have is blocked, files are only written inside the project, and API errors are redacted before printing. Details and how to report a problem: [SECURITY.md](SECURITY.md).
345
+
260
346
  ## Limitations
261
347
 
262
- - JSON only (nested objects, arrays, flat keys). YAML, PO, XLIFF and ARB are not supported yet.
348
+ - JSON (nested objects, arrays, flat keys), Flutter ARB and plain `.txt` files. YAML, PO and XLIFF are not supported yet.
263
349
  - The checks catch mechanical problems, not every wrong meaning. Have a native speaker look at important screens, then approve their edits with `review`.
264
350
  - Rules for form of address, gender and typography exist for the languages listed above. Other languages are translated with the general rules.
265
351
  - A run that is interrupted keeps everything written so far. Unwritten strings are picked up on the next run.
266
352
 
353
+ ## Contributing
354
+
355
+ Bug reports with a concrete example (source, language, output, expected) help most. See [CONTRIBUTING.md](CONTRIBUTING.md).
356
+
267
357
  ## License
268
358
 
269
359
  [MIT](LICENSE)
package/dist/checks.d.ts CHANGED
@@ -1,22 +1,27 @@
1
1
  import type { Config } from './config.js';
2
+ import { Scope } from './scope.js';
2
3
  /**
3
4
  * Deterministic quality checks. No API calls, so they can run in CI on every commit.
4
5
  *
5
6
  * placeholder placeholder set differs from the source error
7
+ * unsafe HTML tags, attributes, event handlers or javascript:/data: URLs
8
+ * that the source does not have (script injection) error
6
9
  * script letters from a script the language does not use, or a word
7
10
  * mixing Latin with Cyrillic/Greek lookalikes error
8
11
  * markup links, tags or list items differ from the source; broken tags
9
12
  * years a year from the source is missing or changed (citations, dates)
10
13
  * formality the other form of address than configured, or both mixed;
11
14
  * masculine-only forms for "you" when genderNeutral is on
15
+ * length longer than the maxLength configured for the key
12
16
  * titlecase English Title Case copied into a sentence-case language
13
17
  * ampersand "&" in a language that writes the word
14
18
  * glossary a glossary rendering or a doNotTranslate name is missing
15
19
  * untranslated identical to the source (prose of 3+ words)
16
20
  * partial source-language words left inside an otherwise translated string,
17
- * or a dropped hedge ("tend to" stated as certain)
21
+ * a dropped hedge ("tend to" stated as certain), or a translation much
22
+ * shorter than its source (content missing or cut off)
18
23
  */
19
- export type CheckName = 'placeholder' | 'script' | 'markup' | 'years' | 'formality' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
24
+ export type CheckName = 'placeholder' | 'unsafe' | 'script' | 'markup' | 'years' | 'formality' | 'length' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
20
25
  export declare const CHECKS: CheckName[];
21
26
  export declare const ERROR_CHECKS: Set<CheckName>;
22
27
  /** Checks a targeted repair (--fix-flagged) may try to fix. */
@@ -28,7 +33,14 @@ export interface Issue {
28
33
  export declare class Checker {
29
34
  readonly config: Config;
30
35
  readonly placeholderRe: RegExp;
36
+ readonly scope: Scope;
37
+ /** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
38
+ readonly keptWords: Set<string>;
31
39
  constructor(config: Config);
40
+ /** The text without doNotTranslate names, which stay the same in every language. */
41
+ withoutNames(text: string): string;
42
+ /** "62 characters, limit 60", or null. */
43
+ tooLong(key: string, text: string): string | null;
32
44
  get englishSource(): boolean;
33
45
  placeholdersMatch(key: string, source: string, text: string): boolean;
34
46
  placeholderNote(source: string, text: string): string;
@@ -48,8 +60,17 @@ export declare class Checker {
48
60
  }
49
61
  /** Non-Latin scripts each language is written in. Latin is always allowed (names, codes). */
50
62
  export declare const NATIVE_SCRIPTS: Record<string, string[]>;
63
+ /** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
64
+ export declare function wrongChineseScript(lang: string, text: string): string | null;
51
65
  /** Why `text` contains letters that cannot belong to `lang`, or null. */
52
66
  export declare function foreignScript(lang: string, text: string): string | null;
67
+ /**
68
+ * Markup the translation adds that the source does not have: a new tag type, a new or changed
69
+ * attribute, an event handler or a script URL. Translations are often rendered as raw HTML
70
+ * (dangerouslySetInnerHTML, v-html), so such an addition is a script-injection risk, whether
71
+ * it comes from a model mistake or from a manipulated source string or response.
72
+ */
73
+ export declare function unsafeAdditions(source: string, text: string): string | null;
53
74
  export declare const hrefSignature: (text: string) => string;
54
75
  /** A malformed, unclosed or misnested tag, or null. */
55
76
  export declare function brokenMarkup(text: string): string | null;
@@ -73,3 +94,8 @@ export declare function isUnchangedProse(source: string, text: string, placehold
73
94
  * allowed) copied into the translation, or null.
74
95
  */
75
96
  export declare function sourceRun(source: string, text: string): string | null;
97
+ /**
98
+ * A translation with a fraction of the source's length lost content: cut off by the model, or
99
+ * the source grew after it was translated. Markup and placeholders are not counted.
100
+ */
101
+ export declare function muchShorter(lang: string, source: string, text: string): string | null;
package/dist/checks.js CHANGED
@@ -1,25 +1,29 @@
1
1
  import { placeholderRegExp, placeholderSignature, placeholdersMatch } from './placeholders.js';
2
2
  import { GENDERED_FORMS, NO_AMPERSAND_LANGUAGES, SENTENCE_CASE_LANGUAGES, formalityRule, registerFor, withoutQuotedSpeech, } from './style.js';
3
- import { baseLanguage, escapeRegExp } from './util.js';
3
+ import { Scope } from './scope.js';
4
+ import { baseLanguage, escapeRegExp, isTraditionalChinese } from './util.js';
4
5
  export const CHECKS = [
5
6
  'placeholder',
7
+ 'unsafe',
6
8
  'script',
7
9
  'markup',
8
10
  'years',
9
11
  'formality',
12
+ 'length',
10
13
  'titlecase',
11
14
  'ampersand',
12
15
  'glossary',
13
16
  'untranslated',
14
17
  'partial',
15
18
  ];
16
- export const ERROR_CHECKS = new Set(['placeholder', 'script']);
19
+ export const ERROR_CHECKS = new Set(['placeholder', 'unsafe', 'script']);
17
20
  /** Checks a targeted repair (--fix-flagged) may try to fix. */
18
21
  export const FIXABLE_CHECKS = new Set([
19
22
  'script',
20
23
  'markup',
21
24
  'years',
22
25
  'formality',
26
+ 'length',
23
27
  'titlecase',
24
28
  'ampersand',
25
29
  'glossary',
@@ -28,9 +32,24 @@ export const FIXABLE_CHECKS = new Set([
28
32
  export class Checker {
29
33
  config;
30
34
  placeholderRe;
35
+ scope;
36
+ /** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
37
+ keptWords;
31
38
  constructor(config) {
32
39
  this.config = config;
33
40
  this.placeholderRe = placeholderRegExp(config.placeholders);
41
+ this.scope = new Scope(config);
42
+ this.keptWords = new Set([...Object.keys(config.termNotes), ...config.doNotTranslate].flatMap(term => term.toLowerCase().split(/[^\p{L}]+/u)).filter(Boolean));
43
+ }
44
+ /** The text without doNotTranslate names, which stay the same in every language. */
45
+ withoutNames(text) {
46
+ return this.config.doNotTranslate.reduce((value, name) => value.split(name).join(' '), text);
47
+ }
48
+ /** "62 characters, limit 60", or null. */
49
+ tooLong(key, text) {
50
+ const max = this.scope.maxLength(key);
51
+ const length = [...text].length;
52
+ return max !== undefined && length > max ? `${length} characters, limit ${max}` : null;
34
53
  }
35
54
  get englishSource() {
36
55
  return baseLanguage(this.config.sourceLanguage) === 'en';
@@ -48,6 +67,9 @@ export class Checker {
48
67
  const base = baseLanguage(lang);
49
68
  if (!this.placeholdersMatch(key, source, text))
50
69
  add('placeholder', this.placeholderNote(source, text));
70
+ const unsafe = unsafeAdditions(source, text);
71
+ if (unsafe)
72
+ add('unsafe', unsafe);
51
73
  const foreign = foreignScript(this.config.sourceLanguage, source) ? null : foreignScript(lang, text);
52
74
  if (foreign)
53
75
  add('script', foreign);
@@ -84,16 +106,22 @@ export class Checker {
84
106
  this.config.sentenceCase &&
85
107
  SENTENCE_CASE_LANGUAGES.has(base) &&
86
108
  isEnglishTitleCase(source)) {
87
- const capitals = midCapitals(text, source, this.config.doNotTranslate);
109
+ const capitals = midCapitals(text, source, this.config.doNotTranslate, lang);
88
110
  if (capitals.length >= (source.trim().split(/\s+/).length <= 3 ? 1 : 2)) {
89
111
  add('titlecase', `capitalised: ${capitals.join(' ')}`);
90
112
  }
91
113
  }
114
+ const long = this.tooLong(key, text);
115
+ if (long)
116
+ add('length', long);
92
117
  const ampersands = (value) => value.split(' & ').length - 1;
93
118
  if (NO_AMPERSAND_LANGUAGES.has(base) && ampersands(text) > ampersands(source))
94
119
  add('ampersand');
95
- if (isUnchangedProse(source, text, this.placeholderRe))
120
+ if (isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe))
96
121
  add('untranslated');
122
+ const short = muchShorter(lang, source, text);
123
+ if (short)
124
+ add('partial', short);
97
125
  if (this.englishSource && text !== source) {
98
126
  const copied = sourceRun(source, text);
99
127
  if (copied)
@@ -101,7 +129,7 @@ export class Checker {
101
129
  const lead = !copied ? sourceBoldLeadIn(source, text, this.config.doNotTranslate) : null;
102
130
  if (lead)
103
131
  add('partial', `bold lead-in still in the source language: "${lead}"`);
104
- const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe) : null;
132
+ const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe, this.keptWords) : null;
105
133
  if (mixed)
106
134
  add('partial', mixed);
107
135
  if (droppedHedge(lang, source, text)) {
@@ -123,15 +151,15 @@ export class Checker {
123
151
  const placeholders = this.placeholdersMatch(key, source, text) ? null : `placeholder mismatch: ${this.placeholderNote(source, text)}`;
124
152
  const foreign = foreignScript(this.config.sourceLanguage, source) ? null : foreignScript(lang, text);
125
153
  const links = hrefSignature(source) !== hrefSignature(text) ? `links changed: [${hrefSignature(source)}] -> [${hrefSignature(text)}]` : null;
126
- const echoed = isUnchangedProse(source, text, this.placeholderRe) ? 'returned the source text unchanged' : null;
154
+ const echoed = isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe) ? 'returned the source text unchanged' : null;
127
155
  const broken = brokenMarkup(source) ? null : brokenMarkup(text);
128
156
  const boldMarkers = (value) => (value.match(/\*\*/g) ?? []).length % 2;
129
157
  const brokenBold = boldMarkers(text) === 1 && boldMarkers(source) === 0 ? 'unbalanced ** markers' : null;
130
158
  const leaked = /^(here('s| is) the translation|translation:)/i.test(text.trim()) ? 'model commentary in the output' : null;
131
- const hard = placeholders ?? foreign ?? links ?? echoed ?? broken ?? brokenBold ?? droppedBullets(source, text) ?? leaked;
159
+ const hard = placeholders ?? unsafeAdditions(source, text) ?? foreign ?? links ?? echoed ?? broken ?? brokenBold ?? droppedBullets(source, text) ?? leaked;
132
160
  const tags = tagCount(source) !== tagCount(text) ? `${tagCount(source)} tags in the source, got ${tagCount(text)}` : null;
133
161
  const copied = this.englishSource && text !== source ? sourceRun(source, text) : null;
134
- return { hard, soft: hard ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
162
+ return { hard, soft: hard ?? this.tooLong(key, text) ?? muchShorter(lang, source, text) ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
135
163
  }
136
164
  }
137
165
  // ---------------------------------------------------------------------------------------
@@ -184,8 +212,25 @@ const LATIN_LANGUAGES = new Set([
184
212
  ]);
185
213
  // Latin glued to a lookalike alphabet inside one word ("Вarda": Cyrillic В + Latin arda).
186
214
  const HOMOGLYPH_WORD = /(?=\p{L}*\p{Script=Latin})(?=\p{L}*[\p{Script=Cyrillic}\p{Script=Greek}])\p{L}+/u;
215
+ // Common characters that exist in only one of the two Chinese scripts (pairs at the same index).
216
+ const SIMPLIFIED_ONLY = '们这说时会来对个为发过还让现实动门问间题体关点应开东头书长见认学页电话语读写买卖钱网设计习惯觉机帮爱车钟儿气无边进选择检样经验数据项结种类业务环节';
217
+ const TRADITIONAL_ONLY = '們這說時會來對個為發過還讓現實動門問間題體關點應開東頭書長見認學頁電話語讀寫買賣錢網設計習慣覺機幫愛車鐘兒氣無邊進選擇檢樣經驗數據項結種類業務環節';
218
+ /** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
219
+ export function wrongChineseScript(lang, text) {
220
+ if (baseLanguage(lang) !== 'zh')
221
+ return null;
222
+ const traditional = isTraditionalChinese(lang);
223
+ const wrong = traditional ? SIMPLIFIED_ONLY : TRADITIONAL_ONLY;
224
+ const found = [...new Set([...text].filter(ch => wrong.includes(ch)))];
225
+ return found.length >= 2
226
+ ? `${traditional ? 'Simplified' : 'Traditional'} Chinese characters in ${lang}: ${found.slice(0, 6).join('')}`
227
+ : null;
228
+ }
187
229
  /** Why `text` contains letters that cannot belong to `lang`, or null. */
188
230
  export function foreignScript(lang, text) {
231
+ const chinese = wrongChineseScript(lang, text);
232
+ if (chinese)
233
+ return chinese;
189
234
  const base = baseLanguage(lang);
190
235
  const native = NATIVE_SCRIPTS[base] ?? (LATIN_LANGUAGES.has(base) ? [] : null);
191
236
  if (native === null)
@@ -207,6 +252,63 @@ export function foreignScript(lang, text) {
207
252
  return mixed ? `mixed-alphabet word: ${mixed[0]}` : null;
208
253
  }
209
254
  // ---------------------------------------------------------------------------------------
255
+ // Unsafe additions
256
+ const unescapeEntities = (text) => text
257
+ .replace(/&lt;|&#0*60;|&#x0*3c;/gi, '<')
258
+ .replace(/&gt;|&#0*62;|&#x0*3e;/gi, '>')
259
+ .replace(/&quot;|&#0*34;|&#x0*22;/gi, '"')
260
+ .replace(/&#0*39;|&#x0*27;|&apos;/gi, "'")
261
+ .replace(/&colon;|&#0*58;|&#x0*3a;/gi, ':');
262
+ // HTML elements, so text in angle brackets ("<minutes>", "<your name>") is not taken for markup.
263
+ const HTML_ELEMENTS = new Set(('a abbr address area article aside audio b base bdi bdo blockquote body br button canvas caption cite code col colgroup data datalist dd del details dfn dialog div dl dt em embed fieldset figcaption figure footer form frame frameset h1 h2 h3 h4 h5 h6 head header hr html i iframe img input ins kbd label legend li link main map mark math meta meter nav noscript object ol optgroup option output p param picture pre progress q rp rt ruby s samp script section select slot small source span strong style sub summary sup svg table tbody td template textarea tfoot th thead time title tr track u ul var video wbr animate foreignobject use image set').split(' '));
264
+ /** HTML element names, in lower case. Numbered <0> tags (react-i18next) are placeholders. */
265
+ const tagNames = (html) => new Set([...html.matchAll(/<\/?([a-z][\w-]*)/gi)].map(m => m[1].toLowerCase()).filter(tag => HTML_ELEMENTS.has(tag)));
266
+ const TEXT_ATTRIBUTES = new Set(['title', 'alt', 'aria-label', 'aria-description', 'placeholder']);
267
+ /** Every attribute as "name=value" (value without quotes), in lower case. */
268
+ const attributes = (html) => {
269
+ const found = new Set();
270
+ for (const [, tag, inner] of html.matchAll(/<([a-z][\w-]*)\s([^>]*)>?/gi)) {
271
+ // "<dakika 20)" (sw: under 20 minutes) is text, not a tag.
272
+ if (!HTML_ELEMENTS.has(tag.toLowerCase()))
273
+ continue;
274
+ for (const [, name, v1, v2, v3] of inner.matchAll(/([^\s"'=<>\/]+)(?:\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s"'>]+)))?/g)) {
275
+ const attr = name.toLowerCase();
276
+ // Text attributes are translated along with the visible text; only their presence counts.
277
+ const value = TEXT_ATTRIBUTES.has(attr) ? '*' : (v1 ?? v2 ?? v3 ?? '').trim().toLowerCase();
278
+ found.add(`${attr}=${value}`);
279
+ }
280
+ }
281
+ return found;
282
+ };
283
+ // Script URLs in a link target or Markdown link. Plain "data:" in prose is a word (pt/it "date:").
284
+ const DANGEROUS_URL = /(?:(?:href|src|action|formaction|xlink:href|poster|background)\s*=\s*["']?|\]\()\s*(?:javascript|vbscript|data)\s*:/gi;
285
+ // Formatting a translator may add for emphasis or a title (no attributes): a markup warning, not a risk.
286
+ const HARMLESS_TAGS = new Set(['br', 'i', 'b', 'em', 'strong', 'u', 's', 'sub', 'sup', 'small', 'mark', 'q', 'cite']);
287
+ /**
288
+ * Markup the translation adds that the source does not have: a new tag type, a new or changed
289
+ * attribute, an event handler or a script URL. Translations are often rendered as raw HTML
290
+ * (dangerouslySetInnerHTML, v-html), so such an addition is a script-injection risk, whether
291
+ * it comes from a model mistake or from a manipulated source string or response.
292
+ */
293
+ export function unsafeAdditions(source, text) {
294
+ const [src, out] = [unescapeEntities(source), unescapeEntities(text)];
295
+ const srcTags = tagNames(src);
296
+ const newTags = [...tagNames(out)].filter(tag => !srcTags.has(tag) && !HARMLESS_TAGS.has(tag));
297
+ if (newTags.length > 0)
298
+ return `HTML tag not in the source: <${newTags.join('>, <')}>`;
299
+ const srcAttrs = attributes(src);
300
+ const newAttrs = [...attributes(out)].filter(attr => !srcAttrs.has(attr));
301
+ const handler = newAttrs.find(attr => /^on/.test(attr));
302
+ if (handler)
303
+ return `event handler not in the source: ${handler.split('=')[0]}`;
304
+ if (newAttrs.length > 0)
305
+ return `HTML attribute not in the source: ${newAttrs[0]}`;
306
+ const urls = (value) => (value.match(DANGEROUS_URL) ?? []).length;
307
+ if (urls(out) > urls(src))
308
+ return 'javascript:, vbscript: or data: URL not in the source';
309
+ return null;
310
+ }
311
+ // ---------------------------------------------------------------------------------------
210
312
  // Markup
211
313
  const unescapeHtml = (text) => text.replace(/&lt;/g, '<').replace(/&gt;/g, '>').replace(/&quot;/g, '"');
212
314
  export const hrefSignature = (text) => (unescapeHtml(text).match(/href\s*=\s*"[^"]*"/g) ?? []).sort().join(' ');
@@ -259,7 +361,8 @@ const asciiDigits = (text) => text.replace(/\p{Nd}/gu, ch => {
259
361
  export function yearDifference(source, text) {
260
362
  // Decades ("the 2020s") are written in words or with suffixes in many languages.
261
363
  const decades = new Set(asciiDigits(source).match(/(?<!\d)(?:19|20)\d0(?=['’]?s\b)/g) ?? []);
262
- const years = (value) => (asciiDigits(value).match(/(?<!\d)(?:19|20)\d\d(?!\d)/g) ?? []).filter(y => !decades.has(y)).sort().join(',');
364
+ // "1,900", "1.900" and "1 900" are numbers, not years: drop thousands separators first.
365
+ const years = (value) => (asciiDigits(value).replace(/(\d)[,.\u00a0\u202f ](?=\d{3}(?!\d))/g, '$1').match(/(?<!\d)(?:19|20)\d\d(?!\d)/g) ?? []).filter(y => !decades.has(y)).sort().join(',');
263
366
  const [a, b] = [years(source), years(text)];
264
367
  return a === b ? null : `source years [${a}], got [${b}]`;
265
368
  }
@@ -306,20 +409,25 @@ export function keptNameMissing(name, source, text) {
306
409
  // ---------------------------------------------------------------------------------------
307
410
  // Title case
308
411
  const COMMON_NAMES = new Set(['iOS', 'Android', 'Apple', 'Google', 'iPhone', 'iPad', 'Mac', 'Windows', 'Linux', 'AI', 'API', 'URL', 'PDF', 'FAQ', 'OK']);
309
- function midCapitals(text, source, names) {
412
+ // Polish capitalises "you" pronouns as a sign of respect ("W Twoim planie"): correct, not Title Case.
413
+ const POLISH_RESPECT = /^(?:Ty|Twój|Twoja|Twoje|Twojego|Twojej|Twoim|Twoją|Twoich|Twoimi|Ciebie|Cię|Tobie|Tobą|Wy|Wasz|Wasza|Wasze|Wam|Was|Wami)$/u;
414
+ function midCapitals(text, source, names, lang = '') {
310
415
  const allowed = new Set([...COMMON_NAMES, ...names.flatMap(name => name.split(/\s+/))]);
311
416
  const isName = (word) => allowed.has(word) ||
312
417
  new RegExp(`(?<!\\p{L})${escapeRegExp(word)}(?!\\p{L})`, 'u').test(source) ||
313
418
  [...allowed].some(name => name.length >= 4 && word.startsWith(name));
314
419
  return text
315
- .split(/[.!?:—–\n•|]+/)
420
+ // Commas and brackets start segments too: list items ("SMART: Specific, Measurable") and
421
+ // bracketed words ("Customize (Optional)") are capitalised in many languages.
422
+ .split(/[.!?:—–\n•|,;()]+/)
316
423
  .flatMap(segment => {
317
424
  const words = segment.trim().split(/\s+/);
318
425
  const first = words.findIndex(word => /\p{L}/u.test(word));
319
426
  return first === -1 ? [] : words.slice(first + 1);
320
427
  })
321
428
  .map(word => word.replace(/^[^\p{L}]+|[^\p{L}]+$/gu, ''))
322
- .filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word));
429
+ .filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word))
430
+ .filter(word => !(baseLanguage(lang) === 'pl' && POLISH_RESPECT.test(word)));
323
431
  }
324
432
  function isEnglishTitleCase(source) {
325
433
  const rest = source
@@ -377,26 +485,34 @@ function sourceBoldLeadIn(source, text, names) {
377
485
  for (const lead of bold(source)) {
378
486
  if (lead.length <= 15 || !/\p{Ll}{3}/u.test(lead) || names.some(name => lead.includes(name)))
379
487
  continue;
488
+ // Only capitalised words ("Apple App Store:", "Google Play"): a name, kept on purpose.
489
+ if (lead.split(/\s+/).every(word => !/\p{L}/u.test(word) || /^[^\p{L}]*\p{Lu}/u.test(word)))
490
+ continue;
380
491
  if (translated.has(lead))
381
492
  return lead;
382
493
  }
383
494
  return null;
384
495
  }
496
+ const ICU_HEADER = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/g;
497
+ const ICU_HEADER_TEST = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/;
385
498
  // Loanwords commonly written in Latin script inside non-Latin text.
386
499
  const LATIN_LOANWORDS = new Set(['email', 'online', 'offline', 'emoji', 'smartphone', 'podcast', 'podcasts', 'wifi', 'blog', 'login', 'like', 'likes']);
387
500
  /** Source-language words left inside a non-Latin-script translation ("Settings → Privacy"). */
388
- function englishInNativeScript(lang, source, text, placeholderRe) {
501
+ function englishInNativeScript(lang, source, text, placeholderRe, allowed) {
389
502
  const base = baseLanguage(lang);
390
503
  // Greek writes many anglicisms in Latin script; not checked.
391
504
  if (!NATIVE_SCRIPTS[base] || base === 'el' || base === 'sr' || text === source)
392
505
  return null;
393
506
  const strip = (value) => withoutUrls(value.replace(/[\w.+-]+@[\w.-]+/g, ' ').replace(/<[^>]+>/g, ' '))
507
+ // ICU syntax ("{count, plural, one {…} other {…}}") is code, not English text.
508
+ .replace(ICU_HEADER, ' ')
509
+ .replace(ICU_HEADER_TEST.test(value) ? /(?:^|[\s}])(?:zero|one|two|few|many|other|=\d+|[\w-]+)\s*(?=\{)/g : /$^/g, ' ')
394
510
  .replace(placeholderRe, ' ')
395
511
  .replace(/[((][^))]*[))]/g, ' ')
396
512
  .replace(/["“„«「『‘'][^"”“»」』’']*["”“»」』’']/g, ' ');
397
513
  const sourceWords = new Set(strip(source).match(/\b[a-z]{4,}\b/g) ?? []);
398
514
  const left = [
399
- ...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word))),
515
+ ...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word) && !allowed.has(word))),
400
516
  ];
401
517
  const menu = /\b[A-Z][a-z]+ → [A-Z][a-z]+/.exec(text);
402
518
  if (menu && source.includes(menu[0]))
@@ -425,9 +541,22 @@ const HEDGE_MARKERS = {
425
541
  cs: /tendenc|obvykle|často|zpravidla|většinou|bývá|sklon|snadno|častěji|mív/iu,
426
542
  sk: /tendenc|obvykle|často|zvyčajne|väčšinou|býva|sklon|ľahko|častejšie|zvyk|skôr/iu,
427
543
  el: /τείν|συχνά|συνήθως|τάση|συνήθ|εύκολα|συχνότερα/iu,
428
- fi: /taipu|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
544
+ fi: /taipu|tapana|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
429
545
  ca: /tendeix|tendència|sol|sovint|generalment|acostum|normalment|fàcilment|freqüent/iu,
430
546
  };
547
+ // Chinese, Japanese and Korean need far fewer characters than English; Thai has no spaces.
548
+ const COMPACT_SCRIPTS = new Set(['zh', 'ja', 'ko']);
549
+ /**
550
+ * A translation with a fraction of the source's length lost content: cut off by the model, or
551
+ * the source grew after it was translated. Markup and placeholders are not counted.
552
+ */
553
+ export function muchShorter(lang, source, text) {
554
+ const visible = (value) => [...value.replace(/<[^>]+>|\{\{?[^{}]*\}?\}/g, '').replace(/\s+/g, ' ').trim()].length;
555
+ const [a, b] = [visible(source), visible(text)];
556
+ // Real translations from English rarely drop below 60% (Chinese/Japanese/Korean: 25%).
557
+ const min = COMPACT_SCRIPTS.has(baseLanguage(lang)) ? 0.2 : 0.5;
558
+ return a >= 200 && b < a * min ? `translation much shorter than the source (${b} of ${a} characters): content missing?` : null;
559
+ }
431
560
  function droppedHedge(lang, source, text) {
432
561
  const markers = HEDGE_MARKERS[baseLanguage(lang)];
433
562
  return Boolean(markers) && /\btends? to\b/i.test(source) && !markers.test(text);