localewarden 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +99 -9
- package/dist/checks.d.ts +28 -2
- package/dist/checks.js +144 -15
- package/dist/cli.js +25 -4
- package/dist/config.d.ts +10 -0
- package/dist/config.js +21 -1
- package/dist/files.d.ts +29 -0
- package/dist/files.js +105 -1
- package/dist/llm.js +3 -1
- package/dist/placeholders.js +2 -1
- package/dist/project.d.ts +15 -0
- package/dist/project.js +86 -5
- package/dist/prompt.d.ts +4 -0
- package/dist/prompt.js +15 -3
- package/dist/review.js +25 -5
- package/dist/scope.d.ts +15 -0
- package/dist/scope.js +26 -0
- package/dist/state.d.ts +1 -1
- package/dist/translate.js +226 -156
- package/package.json +12 -1
package/README.md
CHANGED
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
# localewarden
|
|
2
2
|
|
|
3
|
+
[](https://www.npmjs.com/package/localewarden) [](https://github.com/martinb207/localewarden/actions/workflows/ci.yml) [](LICENSE)
|
|
4
|
+
|
|
3
5
|
**Incremental AI translation for JSON locale files.** It translates only what changed, never overwrites a translation a person fixed, and checks every result before it is written.
|
|
4
6
|
|
|
5
7
|
```bash
|
|
@@ -9,7 +11,7 @@ npx localewarden # translate new and changed strings
|
|
|
9
11
|
npx localewarden check # quality check, no API calls (use it in CI)
|
|
10
12
|
```
|
|
11
13
|
|
|
12
|
-
Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files. Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
|
|
14
|
+
Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files, with Flutter (`.arb` files) and with fastlane's App Store / Play Store metadata (`.txt` files). Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
|
|
13
15
|
|
|
14
16
|
## Why
|
|
15
17
|
|
|
@@ -20,14 +22,22 @@ Translating locale files with a language model is easy once. Keeping 20 language
|
|
|
20
22
|
- **Models are inconsistent across batches.** French screens mix "tu" and "vous", Spanish copies English Title Case ("Configure Su Cuenta"), and Polish or Russian address every user as a man.
|
|
21
23
|
- **Broken output ships silently.** A translated placeholder (`{heures}` instead of `{hours}`) shows raw braces in your app. A dropped `</strong>` breaks the layout. Stray Cyrillic letters end up in a Danish sentence.
|
|
22
24
|
|
|
23
|
-
localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output.
|
|
25
|
+
localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output. The checks are tuned against that app's real texts (UI, website, long-form learning content and store listings, about 165 MB) to report problems without flooding you with false alarms. On those texts they still find things that slipped through earlier pipelines: sections cut off after the English grew, stray letters from other alphabets, a trial notice left in English.
|
|
26
|
+
|
|
27
|
+
## How it compares
|
|
28
|
+
|
|
29
|
+
- **Translation platforms** (Crowdin, Lokalise, Phrase, Weblate) are hosted services with editors, translator workflows and review for teams. localewarden is a small CLI that runs in your repository and CI, with no account and no server. If you have professional translators, a platform fits better. If a model translates and people only fix the odd string, this is the lighter setup.
|
|
30
|
+
- **"Translate my JSON with GPT" scripts** usually send every string on every run and overwrite whatever is there. localewarden keeps state, so it only sends what changed, keeps human fixes, and checks the output.
|
|
31
|
+
- **Editor extensions** (such as i18n Ally) help you write and look up keys while coding. localewarden is about filling and maintaining 10 to 40 languages afterwards. The two work well together.
|
|
24
32
|
|
|
25
33
|
## What it does
|
|
26
34
|
|
|
27
35
|
- **Translates only what changed.** It remembers a hash of each source string per language. New strings are translated. Changed strings are *revised*: the model gets the existing translation and changes only what the source change requires. Removed strings are deleted from every language.
|
|
28
36
|
- **Protects hand edits.** If someone edited a translation, localewarden detects it, keeps it, and lists it for review. If the source of a hand-edited string changes later, the string is flagged instead of overwritten.
|
|
29
|
-
- **Checks every result before writing it.** Broken placeholders, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems are retried and reported.
|
|
37
|
+
- **Checks every result before writing it.** Broken placeholders, injected HTML or scripts, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems (too long, content missing, words left in English) are retried and reported.
|
|
30
38
|
- **Consistent style per language.** It enforces formal or informal address per language (`du`/`Sie`, `tu`/`vous`, `ты`/`вы` and 16 more), uses sentence case where the language does, avoids gendered forms for "you", and applies local typography (French spacing, `92 %` in German, CJK quotation marks).
|
|
39
|
+
- **Plural forms per language.** For i18next-style keys (`item_one`, `item_other`) it adds the forms a language needs but English lacks, such as Polish `_few` and `_many` or Arabic `_zero`, `_two`, `_few` and `_many` (CLDR plural rules).
|
|
40
|
+
- **Data files and store listings.** Fields like `id`, `type` or `image` are copied instead of translated (`ignoreKeys`), and so are URLs, email addresses and file paths. Length limits per key (`maxLength`) are passed to the model and checked: App Store names, SEO titles, buttons.
|
|
31
41
|
- **Glossary and protected names.** You choose fixed renderings ("Privacy Policy" -> "Politique de confidentialité") and names that must never be translated. The check accepts grammatical case endings.
|
|
32
42
|
- **Quality check for CI.** `localewarden check` runs all checks without any API calls and exits non-zero on errors.
|
|
33
43
|
- **Targeted repair.** `--fix-flagged` asks the model to fix only what the check flagged. The fix is accepted only if the problem is gone and little else changed.
|
|
@@ -76,7 +86,7 @@ es Mantén vivas tus plantas sin tener que pensar en ello
|
|
|
76
86
|
ja 何も考えなくても、植物を元気に保てます
|
|
77
87
|
```
|
|
78
88
|
|
|
79
|
-
German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic).
|
|
89
|
+
German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic). There are also examples for [Flutter ARB files](examples/flutter) and [App Store / Play Store texts with fastlane](examples/fastlane).
|
|
80
90
|
|
|
81
91
|
## Quick start
|
|
82
92
|
|
|
@@ -131,24 +141,29 @@ The same checks run in two places. Right after each model answer, a failed hard
|
|
|
131
141
|
| Check | Finds | Severity |
|
|
132
142
|
| --- | --- | --- |
|
|
133
143
|
| `placeholder` | `{name}`, `{{count}}`, `%s`, `%1$d`, `%{x}`, `${x}`, `<0></0>` renamed, translated, added or dropped. ICU `plural`/`select` arguments are compared, while plural categories may differ per language. | error |
|
|
134
|
-
| `
|
|
144
|
+
| `unsafe` | HTML tags, attributes, event handlers or `javascript:`/`data:` URLs that the source does not have. Translations are often rendered as raw HTML, so this would be a script injection. | error |
|
|
145
|
+
| `script` | Letters from an alphabet the language does not use ("刺激" in German), a word that mixes Latin with Cyrillic/Greek lookalikes ("Вarda"), or Simplified characters in Traditional Chinese (`zh-TW`) and the reverse | error |
|
|
135
146
|
| `markup` | Changed link targets, different number of tags, unclosed or misnested tags, dropped list items | warning (broken tags and changed links: never written) |
|
|
136
147
|
| `years` | A year from the source missing or changed (citations, dates) | warning |
|
|
137
148
|
| `formality` | The other form of address than configured, both forms in one string, or masculine-only forms for "you" | warning |
|
|
149
|
+
| `length` | Longer than the `maxLength` configured for the key | warning |
|
|
138
150
|
| `titlecase` | English Title Case copied into a language that uses sentence case | warning |
|
|
139
151
|
| `ampersand` | "&" in languages that write the word | warning |
|
|
140
152
|
| `glossary` | A glossary rendering missing (case endings allowed), or a `doNotTranslate` name translated | warning |
|
|
141
153
|
| `untranslated` | Identical to the source (prose of 3+ words; "OK" and names are fine) | warning |
|
|
142
|
-
| `partial` | Source-language words left inside the translation, an untranslated bold lead-in,
|
|
154
|
+
| `partial` | Source-language words left inside the translation, an untranslated bold lead-in, a hedge that became certainty ("tend to" stated as fact), a translation much shorter than its source (content cut off, or the source grew after it was translated), or sibling options that got the same translation although the source differs ("Rarely" and "Occasionally" both "Selten") | warning |
|
|
143
155
|
|
|
144
156
|
```bash
|
|
145
157
|
npx localewarden check # counts per language and check
|
|
146
158
|
npx localewarden check -v # with examples
|
|
147
159
|
npx localewarden check --strict # exit 1 on warnings too
|
|
148
160
|
npx localewarden check --json # for scripts
|
|
161
|
+
npx localewarden check --fix # repair placeholders with one possible fix, no API calls
|
|
149
162
|
```
|
|
150
163
|
|
|
151
|
-
|
|
164
|
+
`--fix` repairs a translated placeholder when the source has exactly one and the translation renamed it (`{stunden}` back to `{hours}`). The file is edited in place, so its formatting stays as it is. Anything less certain is left for `--fix-flagged` or a person.
|
|
165
|
+
|
|
166
|
+
Approved strings are skipped, except for errors (placeholder, unsafe, script), which break the app either way. You can approve any string, not only hand edits: `npx localewarden review --approve de:home.title` tells the check that a person looked at it (for example a pun on a brand name that is correct without the name), and runs leave it alone.
|
|
152
167
|
|
|
153
168
|
### In CI
|
|
154
169
|
|
|
@@ -166,6 +181,39 @@ jobs:
|
|
|
166
181
|
- run: npx localewarden check
|
|
167
182
|
```
|
|
168
183
|
|
|
184
|
+
### Translating automatically
|
|
185
|
+
|
|
186
|
+
When the source language changes on `main`, translate and open a pull request for review:
|
|
187
|
+
|
|
188
|
+
```yaml
|
|
189
|
+
# .github/workflows/translate.yml
|
|
190
|
+
name: translate
|
|
191
|
+
on:
|
|
192
|
+
push:
|
|
193
|
+
branches: [main]
|
|
194
|
+
paths: ['locales/en.json'] # your source files
|
|
195
|
+
permissions:
|
|
196
|
+
contents: write
|
|
197
|
+
pull-requests: write
|
|
198
|
+
jobs:
|
|
199
|
+
translate:
|
|
200
|
+
runs-on: ubuntu-latest
|
|
201
|
+
steps:
|
|
202
|
+
- uses: actions/checkout@v4
|
|
203
|
+
- uses: actions/setup-node@v4
|
|
204
|
+
with: { node-version: 22 }
|
|
205
|
+
- run: npx localewarden --max-tokens 200000
|
|
206
|
+
env:
|
|
207
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
208
|
+
- uses: peter-evans/create-pull-request@v7
|
|
209
|
+
with:
|
|
210
|
+
branch: localewarden/translations
|
|
211
|
+
title: Update translations
|
|
212
|
+
commit-message: Update translations
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
The pull request contains the locale files and `.localewarden/`, so a reviewer sees exactly which strings changed.
|
|
216
|
+
|
|
169
217
|
### Fixing what the check finds
|
|
170
218
|
|
|
171
219
|
```bash
|
|
@@ -203,6 +251,9 @@ npx localewarden review --release de:home.title # hand it back: next run revis
|
|
|
203
251
|
| `glossary` | `{}` | `{"fr": {"Terms of Service": "Conditions d'utilisation"}}` |
|
|
204
252
|
| `termNotes` | `{}` | Meanings of ambiguous terms, sent only with strings that contain them: `{"snooze": "postpone a reminder"}` |
|
|
205
253
|
| `instructions` | `{}` | Extra instructions per language, `"*"` for all: `{"es": "Use neutral Latin American Spanish."}` |
|
|
254
|
+
| `ignoreKeys` | `[]` | Keys that are not text, copied from the source: `["id", "type", "**.sources.*"]`. `*` matches within a key segment, `**` across segments; a pattern without a dot matches the last segment anywhere. URLs, emails, file paths and numbers are always copied. |
|
|
255
|
+
| `exclude` | `[]` | Source files to skip: `["locales/{lang}/nav.json"]` |
|
|
256
|
+
| `maxLength` | `{}` | Character limits per key pattern: `{"**.meta.title": 60, "name": 30}`. The model is told the limit; longer results are retried once and reported by the `length` check. |
|
|
206
257
|
| `placeholders` | built-in | Regular expressions (strings) that match your placeholders. Replaces the built-in list. |
|
|
207
258
|
| `model` | `"gpt-5.4-mini"` | Any chat model your endpoint offers |
|
|
208
259
|
| `baseUrl` | `"https://api.openai.com/v1"` | Any OpenAI-compatible endpoint |
|
|
@@ -214,6 +265,37 @@ npx localewarden review --release de:home.title # hand it back: next run revis
|
|
|
214
265
|
| `batchSize` | `20` | Strings per request (smaller for scripts that need many tokens) |
|
|
215
266
|
| `stateDir` | `".localewarden"` | Where state and the review list live |
|
|
216
267
|
|
|
268
|
+
### App Store and Play Store listings (fastlane)
|
|
269
|
+
|
|
270
|
+
```json
|
|
271
|
+
{
|
|
272
|
+
"sourceLanguage": "en-US",
|
|
273
|
+
"targetLanguages": ["de-DE", "fr-FR", "ja"],
|
|
274
|
+
"files": "fastlane/metadata/{lang}/*.txt",
|
|
275
|
+
"exclude": ["fastlane/metadata/{lang}/*_url.txt"],
|
|
276
|
+
"maxLength": { "name": 30, "subtitle": 30, "keywords": 100, "promotional_text": 170, "description": 4000 },
|
|
277
|
+
"termNotes": { "keywords": "a comma-separated keyword list for store search, not a sentence" }
|
|
278
|
+
}
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
Each `.txt` file is one string, keyed by its file name, so the limits above apply to `name.txt`, `subtitle.txt` and so on.
|
|
282
|
+
|
|
283
|
+
### Flutter (ARB)
|
|
284
|
+
|
|
285
|
+
```json
|
|
286
|
+
{ "files": "lib/l10n/app_{lang}.arb", "targetLanguages": ["de", "fr", "pt_BR"] }
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Metadata (`@@locale`, `@key` descriptions and placeholders) is copied, not translated, and `@@locale` is set to the target language. ICU plurals and selects keep their structure, and each language gets the plural categories it needs.
|
|
290
|
+
|
|
291
|
+
### Data files
|
|
292
|
+
|
|
293
|
+
For content JSON with ids, types and links, list the non-text keys:
|
|
294
|
+
|
|
295
|
+
```json
|
|
296
|
+
{ "files": "content/**/*.{lang}.json", "ignoreKeys": ["id", "type", "category", "image", "**.sources.*"] }
|
|
297
|
+
```
|
|
298
|
+
|
|
217
299
|
### Other providers
|
|
218
300
|
|
|
219
301
|
```json
|
|
@@ -231,7 +313,7 @@ Small local models make noticeably more mistakes. The checks catch the mechanica
|
|
|
231
313
|
```text
|
|
232
314
|
localewarden [translate] --dry-run --lang de,fr --fix-flagged --retranslate-all
|
|
233
315
|
--overwrite-manual --max-tokens <n> --verbose
|
|
234
|
-
localewarden check --lang de,fr --verbose --limit <n> --strict --json
|
|
316
|
+
localewarden check --lang de,fr --verbose --limit <n> --strict --json --fix
|
|
235
317
|
localewarden review --all --approve <sel>... --release <sel>...
|
|
236
318
|
localewarden init
|
|
237
319
|
Global: --config <path> --help --version
|
|
@@ -257,13 +339,21 @@ const findings = checkProject(config);
|
|
|
257
339
|
- Unchanged strings cost nothing. In the example above, changing two English strings and updating four languages took 4 requests and about 4,000 tokens.
|
|
258
340
|
- Strings, keys, your `context` and glossary are sent to the API you configure. Nothing else is sent anywhere. There is no telemetry.
|
|
259
341
|
|
|
342
|
+
## Security
|
|
343
|
+
|
|
344
|
+
Translations are treated as untrusted: any markup the source does not have is blocked, files are only written inside the project, and API errors are redacted before printing. Details and how to report a problem: [SECURITY.md](SECURITY.md).
|
|
345
|
+
|
|
260
346
|
## Limitations
|
|
261
347
|
|
|
262
|
-
- JSON
|
|
348
|
+
- JSON (nested objects, arrays, flat keys), Flutter ARB and plain `.txt` files. YAML, PO and XLIFF are not supported yet.
|
|
263
349
|
- The checks catch mechanical problems, not every wrong meaning. Have a native speaker look at important screens, then approve their edits with `review`.
|
|
264
350
|
- Rules for form of address, gender and typography exist for the languages listed above. Other languages are translated with the general rules.
|
|
265
351
|
- A run that is interrupted keeps everything written so far. Unwritten strings are picked up on the next run.
|
|
266
352
|
|
|
353
|
+
## Contributing
|
|
354
|
+
|
|
355
|
+
Bug reports with a concrete example (source, language, output, expected) help most. See [CONTRIBUTING.md](CONTRIBUTING.md).
|
|
356
|
+
|
|
267
357
|
## License
|
|
268
358
|
|
|
269
359
|
[MIT](LICENSE)
|
package/dist/checks.d.ts
CHANGED
|
@@ -1,22 +1,27 @@
|
|
|
1
1
|
import type { Config } from './config.js';
|
|
2
|
+
import { Scope } from './scope.js';
|
|
2
3
|
/**
|
|
3
4
|
* Deterministic quality checks. No API calls, so they can run in CI on every commit.
|
|
4
5
|
*
|
|
5
6
|
* placeholder placeholder set differs from the source error
|
|
7
|
+
* unsafe HTML tags, attributes, event handlers or javascript:/data: URLs
|
|
8
|
+
* that the source does not have (script injection) error
|
|
6
9
|
* script letters from a script the language does not use, or a word
|
|
7
10
|
* mixing Latin with Cyrillic/Greek lookalikes error
|
|
8
11
|
* markup links, tags or list items differ from the source; broken tags
|
|
9
12
|
* years a year from the source is missing or changed (citations, dates)
|
|
10
13
|
* formality the other form of address than configured, or both mixed;
|
|
11
14
|
* masculine-only forms for "you" when genderNeutral is on
|
|
15
|
+
* length longer than the maxLength configured for the key
|
|
12
16
|
* titlecase English Title Case copied into a sentence-case language
|
|
13
17
|
* ampersand "&" in a language that writes the word
|
|
14
18
|
* glossary a glossary rendering or a doNotTranslate name is missing
|
|
15
19
|
* untranslated identical to the source (prose of 3+ words)
|
|
16
20
|
* partial source-language words left inside an otherwise translated string,
|
|
17
|
-
*
|
|
21
|
+
* a dropped hedge ("tend to" stated as certain), or a translation much
|
|
22
|
+
* shorter than its source (content missing or cut off)
|
|
18
23
|
*/
|
|
19
|
-
export type CheckName = 'placeholder' | 'script' | 'markup' | 'years' | 'formality' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
|
|
24
|
+
export type CheckName = 'placeholder' | 'unsafe' | 'script' | 'markup' | 'years' | 'formality' | 'length' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
|
|
20
25
|
export declare const CHECKS: CheckName[];
|
|
21
26
|
export declare const ERROR_CHECKS: Set<CheckName>;
|
|
22
27
|
/** Checks a targeted repair (--fix-flagged) may try to fix. */
|
|
@@ -28,7 +33,14 @@ export interface Issue {
|
|
|
28
33
|
export declare class Checker {
|
|
29
34
|
readonly config: Config;
|
|
30
35
|
readonly placeholderRe: RegExp;
|
|
36
|
+
readonly scope: Scope;
|
|
37
|
+
/** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
|
|
38
|
+
readonly keptWords: Set<string>;
|
|
31
39
|
constructor(config: Config);
|
|
40
|
+
/** The text without doNotTranslate names, which stay the same in every language. */
|
|
41
|
+
withoutNames(text: string): string;
|
|
42
|
+
/** "62 characters, limit 60", or null. */
|
|
43
|
+
tooLong(key: string, text: string): string | null;
|
|
32
44
|
get englishSource(): boolean;
|
|
33
45
|
placeholdersMatch(key: string, source: string, text: string): boolean;
|
|
34
46
|
placeholderNote(source: string, text: string): string;
|
|
@@ -48,8 +60,17 @@ export declare class Checker {
|
|
|
48
60
|
}
|
|
49
61
|
/** Non-Latin scripts each language is written in. Latin is always allowed (names, codes). */
|
|
50
62
|
export declare const NATIVE_SCRIPTS: Record<string, string[]>;
|
|
63
|
+
/** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
|
|
64
|
+
export declare function wrongChineseScript(lang: string, text: string): string | null;
|
|
51
65
|
/** Why `text` contains letters that cannot belong to `lang`, or null. */
|
|
52
66
|
export declare function foreignScript(lang: string, text: string): string | null;
|
|
67
|
+
/**
|
|
68
|
+
* Markup the translation adds that the source does not have: a new tag type, a new or changed
|
|
69
|
+
* attribute, an event handler or a script URL. Translations are often rendered as raw HTML
|
|
70
|
+
* (dangerouslySetInnerHTML, v-html), so such an addition is a script-injection risk, whether
|
|
71
|
+
* it comes from a model mistake or from a manipulated source string or response.
|
|
72
|
+
*/
|
|
73
|
+
export declare function unsafeAdditions(source: string, text: string): string | null;
|
|
53
74
|
export declare const hrefSignature: (text: string) => string;
|
|
54
75
|
/** A malformed, unclosed or misnested tag, or null. */
|
|
55
76
|
export declare function brokenMarkup(text: string): string | null;
|
|
@@ -73,3 +94,8 @@ export declare function isUnchangedProse(source: string, text: string, placehold
|
|
|
73
94
|
* allowed) copied into the translation, or null.
|
|
74
95
|
*/
|
|
75
96
|
export declare function sourceRun(source: string, text: string): string | null;
|
|
97
|
+
/**
|
|
98
|
+
* A translation with a fraction of the source's length lost content: cut off by the model, or
|
|
99
|
+
* the source grew after it was translated. Markup and placeholders are not counted.
|
|
100
|
+
*/
|
|
101
|
+
export declare function muchShorter(lang: string, source: string, text: string): string | null;
|
package/dist/checks.js
CHANGED
|
@@ -1,25 +1,29 @@
|
|
|
1
1
|
import { placeholderRegExp, placeholderSignature, placeholdersMatch } from './placeholders.js';
|
|
2
2
|
import { GENDERED_FORMS, NO_AMPERSAND_LANGUAGES, SENTENCE_CASE_LANGUAGES, formalityRule, registerFor, withoutQuotedSpeech, } from './style.js';
|
|
3
|
-
import {
|
|
3
|
+
import { Scope } from './scope.js';
|
|
4
|
+
import { baseLanguage, escapeRegExp, isTraditionalChinese } from './util.js';
|
|
4
5
|
export const CHECKS = [
|
|
5
6
|
'placeholder',
|
|
7
|
+
'unsafe',
|
|
6
8
|
'script',
|
|
7
9
|
'markup',
|
|
8
10
|
'years',
|
|
9
11
|
'formality',
|
|
12
|
+
'length',
|
|
10
13
|
'titlecase',
|
|
11
14
|
'ampersand',
|
|
12
15
|
'glossary',
|
|
13
16
|
'untranslated',
|
|
14
17
|
'partial',
|
|
15
18
|
];
|
|
16
|
-
export const ERROR_CHECKS = new Set(['placeholder', 'script']);
|
|
19
|
+
export const ERROR_CHECKS = new Set(['placeholder', 'unsafe', 'script']);
|
|
17
20
|
/** Checks a targeted repair (--fix-flagged) may try to fix. */
|
|
18
21
|
export const FIXABLE_CHECKS = new Set([
|
|
19
22
|
'script',
|
|
20
23
|
'markup',
|
|
21
24
|
'years',
|
|
22
25
|
'formality',
|
|
26
|
+
'length',
|
|
23
27
|
'titlecase',
|
|
24
28
|
'ampersand',
|
|
25
29
|
'glossary',
|
|
@@ -28,9 +32,24 @@ export const FIXABLE_CHECKS = new Set([
|
|
|
28
32
|
export class Checker {
|
|
29
33
|
config;
|
|
30
34
|
placeholderRe;
|
|
35
|
+
scope;
|
|
36
|
+
/** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
|
|
37
|
+
keptWords;
|
|
31
38
|
constructor(config) {
|
|
32
39
|
this.config = config;
|
|
33
40
|
this.placeholderRe = placeholderRegExp(config.placeholders);
|
|
41
|
+
this.scope = new Scope(config);
|
|
42
|
+
this.keptWords = new Set([...Object.keys(config.termNotes), ...config.doNotTranslate].flatMap(term => term.toLowerCase().split(/[^\p{L}]+/u)).filter(Boolean));
|
|
43
|
+
}
|
|
44
|
+
/** The text without doNotTranslate names, which stay the same in every language. */
|
|
45
|
+
withoutNames(text) {
|
|
46
|
+
return this.config.doNotTranslate.reduce((value, name) => value.split(name).join(' '), text);
|
|
47
|
+
}
|
|
48
|
+
/** "62 characters, limit 60", or null. */
|
|
49
|
+
tooLong(key, text) {
|
|
50
|
+
const max = this.scope.maxLength(key);
|
|
51
|
+
const length = [...text].length;
|
|
52
|
+
return max !== undefined && length > max ? `${length} characters, limit ${max}` : null;
|
|
34
53
|
}
|
|
35
54
|
get englishSource() {
|
|
36
55
|
return baseLanguage(this.config.sourceLanguage) === 'en';
|
|
@@ -48,6 +67,9 @@ export class Checker {
|
|
|
48
67
|
const base = baseLanguage(lang);
|
|
49
68
|
if (!this.placeholdersMatch(key, source, text))
|
|
50
69
|
add('placeholder', this.placeholderNote(source, text));
|
|
70
|
+
const unsafe = unsafeAdditions(source, text);
|
|
71
|
+
if (unsafe)
|
|
72
|
+
add('unsafe', unsafe);
|
|
51
73
|
const foreign = foreignScript(this.config.sourceLanguage, source) ? null : foreignScript(lang, text);
|
|
52
74
|
if (foreign)
|
|
53
75
|
add('script', foreign);
|
|
@@ -84,16 +106,22 @@ export class Checker {
|
|
|
84
106
|
this.config.sentenceCase &&
|
|
85
107
|
SENTENCE_CASE_LANGUAGES.has(base) &&
|
|
86
108
|
isEnglishTitleCase(source)) {
|
|
87
|
-
const capitals = midCapitals(text, source, this.config.doNotTranslate);
|
|
109
|
+
const capitals = midCapitals(text, source, this.config.doNotTranslate, lang);
|
|
88
110
|
if (capitals.length >= (source.trim().split(/\s+/).length <= 3 ? 1 : 2)) {
|
|
89
111
|
add('titlecase', `capitalised: ${capitals.join(' ')}`);
|
|
90
112
|
}
|
|
91
113
|
}
|
|
114
|
+
const long = this.tooLong(key, text);
|
|
115
|
+
if (long)
|
|
116
|
+
add('length', long);
|
|
92
117
|
const ampersands = (value) => value.split(' & ').length - 1;
|
|
93
118
|
if (NO_AMPERSAND_LANGUAGES.has(base) && ampersands(text) > ampersands(source))
|
|
94
119
|
add('ampersand');
|
|
95
|
-
if (isUnchangedProse(source, text, this.placeholderRe))
|
|
120
|
+
if (isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe))
|
|
96
121
|
add('untranslated');
|
|
122
|
+
const short = muchShorter(lang, source, text);
|
|
123
|
+
if (short)
|
|
124
|
+
add('partial', short);
|
|
97
125
|
if (this.englishSource && text !== source) {
|
|
98
126
|
const copied = sourceRun(source, text);
|
|
99
127
|
if (copied)
|
|
@@ -101,7 +129,7 @@ export class Checker {
|
|
|
101
129
|
const lead = !copied ? sourceBoldLeadIn(source, text, this.config.doNotTranslate) : null;
|
|
102
130
|
if (lead)
|
|
103
131
|
add('partial', `bold lead-in still in the source language: "${lead}"`);
|
|
104
|
-
const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe) : null;
|
|
132
|
+
const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe, this.keptWords) : null;
|
|
105
133
|
if (mixed)
|
|
106
134
|
add('partial', mixed);
|
|
107
135
|
if (droppedHedge(lang, source, text)) {
|
|
@@ -123,15 +151,15 @@ export class Checker {
|
|
|
123
151
|
const placeholders = this.placeholdersMatch(key, source, text) ? null : `placeholder mismatch: ${this.placeholderNote(source, text)}`;
|
|
124
152
|
const foreign = foreignScript(this.config.sourceLanguage, source) ? null : foreignScript(lang, text);
|
|
125
153
|
const links = hrefSignature(source) !== hrefSignature(text) ? `links changed: [${hrefSignature(source)}] -> [${hrefSignature(text)}]` : null;
|
|
126
|
-
const echoed = isUnchangedProse(source, text, this.placeholderRe) ? 'returned the source text unchanged' : null;
|
|
154
|
+
const echoed = isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe) ? 'returned the source text unchanged' : null;
|
|
127
155
|
const broken = brokenMarkup(source) ? null : brokenMarkup(text);
|
|
128
156
|
const boldMarkers = (value) => (value.match(/\*\*/g) ?? []).length % 2;
|
|
129
157
|
const brokenBold = boldMarkers(text) === 1 && boldMarkers(source) === 0 ? 'unbalanced ** markers' : null;
|
|
130
158
|
const leaked = /^(here('s| is) the translation|translation:)/i.test(text.trim()) ? 'model commentary in the output' : null;
|
|
131
|
-
const hard = placeholders ?? foreign ?? links ?? echoed ?? broken ?? brokenBold ?? droppedBullets(source, text) ?? leaked;
|
|
159
|
+
const hard = placeholders ?? unsafeAdditions(source, text) ?? foreign ?? links ?? echoed ?? broken ?? brokenBold ?? droppedBullets(source, text) ?? leaked;
|
|
132
160
|
const tags = tagCount(source) !== tagCount(text) ? `${tagCount(source)} tags in the source, got ${tagCount(text)}` : null;
|
|
133
161
|
const copied = this.englishSource && text !== source ? sourceRun(source, text) : null;
|
|
134
|
-
return { hard, soft: hard ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
|
|
162
|
+
return { hard, soft: hard ?? this.tooLong(key, text) ?? muchShorter(lang, source, text) ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
|
|
135
163
|
}
|
|
136
164
|
}
|
|
137
165
|
// ---------------------------------------------------------------------------------------
|
|
@@ -184,8 +212,25 @@ const LATIN_LANGUAGES = new Set([
|
|
|
184
212
|
]);
|
|
185
213
|
// Latin glued to a lookalike alphabet inside one word ("Вarda": Cyrillic В + Latin arda).
|
|
186
214
|
const HOMOGLYPH_WORD = /(?=\p{L}*\p{Script=Latin})(?=\p{L}*[\p{Script=Cyrillic}\p{Script=Greek}])\p{L}+/u;
|
|
215
|
+
// Common characters that exist in only one of the two Chinese scripts (pairs at the same index).
|
|
216
|
+
const SIMPLIFIED_ONLY = '们这说时会来对个为发过还让现实动门问间题体关点应开东头书长见认学页电话语读写买卖钱网设计习惯觉机帮爱车钟儿气无边进选择检样经验数据项结种类业务环节';
|
|
217
|
+
const TRADITIONAL_ONLY = '們這說時會來對個為發過還讓現實動門問間題體關點應開東頭書長見認學頁電話語讀寫買賣錢網設計習慣覺機幫愛車鐘兒氣無邊進選擇檢樣經驗數據項結種類業務環節';
|
|
218
|
+
/** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
|
|
219
|
+
export function wrongChineseScript(lang, text) {
|
|
220
|
+
if (baseLanguage(lang) !== 'zh')
|
|
221
|
+
return null;
|
|
222
|
+
const traditional = isTraditionalChinese(lang);
|
|
223
|
+
const wrong = traditional ? SIMPLIFIED_ONLY : TRADITIONAL_ONLY;
|
|
224
|
+
const found = [...new Set([...text].filter(ch => wrong.includes(ch)))];
|
|
225
|
+
return found.length >= 2
|
|
226
|
+
? `${traditional ? 'Simplified' : 'Traditional'} Chinese characters in ${lang}: ${found.slice(0, 6).join('')}`
|
|
227
|
+
: null;
|
|
228
|
+
}
|
|
187
229
|
/** Why `text` contains letters that cannot belong to `lang`, or null. */
|
|
188
230
|
export function foreignScript(lang, text) {
|
|
231
|
+
const chinese = wrongChineseScript(lang, text);
|
|
232
|
+
if (chinese)
|
|
233
|
+
return chinese;
|
|
189
234
|
const base = baseLanguage(lang);
|
|
190
235
|
const native = NATIVE_SCRIPTS[base] ?? (LATIN_LANGUAGES.has(base) ? [] : null);
|
|
191
236
|
if (native === null)
|
|
@@ -207,6 +252,63 @@ export function foreignScript(lang, text) {
|
|
|
207
252
|
return mixed ? `mixed-alphabet word: ${mixed[0]}` : null;
|
|
208
253
|
}
|
|
209
254
|
// ---------------------------------------------------------------------------------------
|
|
255
|
+
// Unsafe additions
|
|
256
|
+
const unescapeEntities = (text) => text
|
|
257
|
+
.replace(/<|�*60;|�*3c;/gi, '<')
|
|
258
|
+
.replace(/>|�*62;|�*3e;/gi, '>')
|
|
259
|
+
.replace(/"|�*34;|�*22;/gi, '"')
|
|
260
|
+
.replace(/�*39;|�*27;|'/gi, "'")
|
|
261
|
+
.replace(/:|�*58;|�*3a;/gi, ':');
|
|
262
|
+
// HTML elements, so text in angle brackets ("<minutes>", "<your name>") is not taken for markup.
|
|
263
|
+
const HTML_ELEMENTS = new Set(('a abbr address area article aside audio b base bdi bdo blockquote body br button canvas caption cite code col colgroup data datalist dd del details dfn dialog div dl dt em embed fieldset figcaption figure footer form frame frameset h1 h2 h3 h4 h5 h6 head header hr html i iframe img input ins kbd label legend li link main map mark math meta meter nav noscript object ol optgroup option output p param picture pre progress q rp rt ruby s samp script section select slot small source span strong style sub summary sup svg table tbody td template textarea tfoot th thead time title tr track u ul var video wbr animate foreignobject use image set').split(' '));
|
|
264
|
+
/** HTML element names, in lower case. Numbered <0> tags (react-i18next) are placeholders. */
|
|
265
|
+
const tagNames = (html) => new Set([...html.matchAll(/<\/?([a-z][\w-]*)/gi)].map(m => m[1].toLowerCase()).filter(tag => HTML_ELEMENTS.has(tag)));
|
|
266
|
+
const TEXT_ATTRIBUTES = new Set(['title', 'alt', 'aria-label', 'aria-description', 'placeholder']);
|
|
267
|
+
/** Every attribute as "name=value" (value without quotes), in lower case. */
|
|
268
|
+
const attributes = (html) => {
|
|
269
|
+
const found = new Set();
|
|
270
|
+
for (const [, tag, inner] of html.matchAll(/<([a-z][\w-]*)\s([^>]*)>?/gi)) {
|
|
271
|
+
// "<dakika 20)" (sw: under 20 minutes) is text, not a tag.
|
|
272
|
+
if (!HTML_ELEMENTS.has(tag.toLowerCase()))
|
|
273
|
+
continue;
|
|
274
|
+
for (const [, name, v1, v2, v3] of inner.matchAll(/([^\s"'=<>\/]+)(?:\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s"'>]+)))?/g)) {
|
|
275
|
+
const attr = name.toLowerCase();
|
|
276
|
+
// Text attributes are translated along with the visible text; only their presence counts.
|
|
277
|
+
const value = TEXT_ATTRIBUTES.has(attr) ? '*' : (v1 ?? v2 ?? v3 ?? '').trim().toLowerCase();
|
|
278
|
+
found.add(`${attr}=${value}`);
|
|
279
|
+
}
|
|
280
|
+
}
|
|
281
|
+
return found;
|
|
282
|
+
};
|
|
283
|
+
// Script URLs in a link target or Markdown link. Plain "data:" in prose is a word (pt/it "date:").
|
|
284
|
+
const DANGEROUS_URL = /(?:(?:href|src|action|formaction|xlink:href|poster|background)\s*=\s*["']?|\]\()\s*(?:javascript|vbscript|data)\s*:/gi;
|
|
285
|
+
// Formatting a translator may add for emphasis or a title (no attributes): a markup warning, not a risk.
|
|
286
|
+
const HARMLESS_TAGS = new Set(['br', 'i', 'b', 'em', 'strong', 'u', 's', 'sub', 'sup', 'small', 'mark', 'q', 'cite']);
|
|
287
|
+
/**
|
|
288
|
+
* Markup the translation adds that the source does not have: a new tag type, a new or changed
|
|
289
|
+
* attribute, an event handler or a script URL. Translations are often rendered as raw HTML
|
|
290
|
+
* (dangerouslySetInnerHTML, v-html), so such an addition is a script-injection risk, whether
|
|
291
|
+
* it comes from a model mistake or from a manipulated source string or response.
|
|
292
|
+
*/
|
|
293
|
+
export function unsafeAdditions(source, text) {
|
|
294
|
+
const [src, out] = [unescapeEntities(source), unescapeEntities(text)];
|
|
295
|
+
const srcTags = tagNames(src);
|
|
296
|
+
const newTags = [...tagNames(out)].filter(tag => !srcTags.has(tag) && !HARMLESS_TAGS.has(tag));
|
|
297
|
+
if (newTags.length > 0)
|
|
298
|
+
return `HTML tag not in the source: <${newTags.join('>, <')}>`;
|
|
299
|
+
const srcAttrs = attributes(src);
|
|
300
|
+
const newAttrs = [...attributes(out)].filter(attr => !srcAttrs.has(attr));
|
|
301
|
+
const handler = newAttrs.find(attr => /^on/.test(attr));
|
|
302
|
+
if (handler)
|
|
303
|
+
return `event handler not in the source: ${handler.split('=')[0]}`;
|
|
304
|
+
if (newAttrs.length > 0)
|
|
305
|
+
return `HTML attribute not in the source: ${newAttrs[0]}`;
|
|
306
|
+
const urls = (value) => (value.match(DANGEROUS_URL) ?? []).length;
|
|
307
|
+
if (urls(out) > urls(src))
|
|
308
|
+
return 'javascript:, vbscript: or data: URL not in the source';
|
|
309
|
+
return null;
|
|
310
|
+
}
|
|
311
|
+
// ---------------------------------------------------------------------------------------
|
|
210
312
|
// Markup
|
|
211
313
|
const unescapeHtml = (text) => text.replace(/</g, '<').replace(/>/g, '>').replace(/"/g, '"');
|
|
212
314
|
export const hrefSignature = (text) => (unescapeHtml(text).match(/href\s*=\s*"[^"]*"/g) ?? []).sort().join(' ');
|
|
@@ -259,7 +361,8 @@ const asciiDigits = (text) => text.replace(/\p{Nd}/gu, ch => {
|
|
|
259
361
|
export function yearDifference(source, text) {
|
|
260
362
|
// Decades ("the 2020s") are written in words or with suffixes in many languages.
|
|
261
363
|
const decades = new Set(asciiDigits(source).match(/(?<!\d)(?:19|20)\d0(?=['’]?s\b)/g) ?? []);
|
|
262
|
-
|
|
364
|
+
// "1,900", "1.900" and "1 900" are numbers, not years: drop thousands separators first.
|
|
365
|
+
const years = (value) => (asciiDigits(value).replace(/(\d)[,.\u00a0\u202f ](?=\d{3}(?!\d))/g, '$1').match(/(?<!\d)(?:19|20)\d\d(?!\d)/g) ?? []).filter(y => !decades.has(y)).sort().join(',');
|
|
263
366
|
const [a, b] = [years(source), years(text)];
|
|
264
367
|
return a === b ? null : `source years [${a}], got [${b}]`;
|
|
265
368
|
}
|
|
@@ -306,20 +409,25 @@ export function keptNameMissing(name, source, text) {
|
|
|
306
409
|
// ---------------------------------------------------------------------------------------
|
|
307
410
|
// Title case
|
|
308
411
|
const COMMON_NAMES = new Set(['iOS', 'Android', 'Apple', 'Google', 'iPhone', 'iPad', 'Mac', 'Windows', 'Linux', 'AI', 'API', 'URL', 'PDF', 'FAQ', 'OK']);
|
|
309
|
-
|
|
412
|
+
// Polish capitalises "you" pronouns as a sign of respect ("W Twoim planie"): correct, not Title Case.
|
|
413
|
+
const POLISH_RESPECT = /^(?:Ty|Twój|Twoja|Twoje|Twojego|Twojej|Twoim|Twoją|Twoich|Twoimi|Ciebie|Cię|Tobie|Tobą|Wy|Wasz|Wasza|Wasze|Wam|Was|Wami)$/u;
|
|
414
|
+
function midCapitals(text, source, names, lang = '') {
|
|
310
415
|
const allowed = new Set([...COMMON_NAMES, ...names.flatMap(name => name.split(/\s+/))]);
|
|
311
416
|
const isName = (word) => allowed.has(word) ||
|
|
312
417
|
new RegExp(`(?<!\\p{L})${escapeRegExp(word)}(?!\\p{L})`, 'u').test(source) ||
|
|
313
418
|
[...allowed].some(name => name.length >= 4 && word.startsWith(name));
|
|
314
419
|
return text
|
|
315
|
-
|
|
420
|
+
// Commas and brackets start segments too: list items ("SMART: Specific, Measurable") and
|
|
421
|
+
// bracketed words ("Customize (Optional)") are capitalised in many languages.
|
|
422
|
+
.split(/[.!?:—–\n•|,;()]+/)
|
|
316
423
|
.flatMap(segment => {
|
|
317
424
|
const words = segment.trim().split(/\s+/);
|
|
318
425
|
const first = words.findIndex(word => /\p{L}/u.test(word));
|
|
319
426
|
return first === -1 ? [] : words.slice(first + 1);
|
|
320
427
|
})
|
|
321
428
|
.map(word => word.replace(/^[^\p{L}]+|[^\p{L}]+$/gu, ''))
|
|
322
|
-
.filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word))
|
|
429
|
+
.filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word))
|
|
430
|
+
.filter(word => !(baseLanguage(lang) === 'pl' && POLISH_RESPECT.test(word)));
|
|
323
431
|
}
|
|
324
432
|
function isEnglishTitleCase(source) {
|
|
325
433
|
const rest = source
|
|
@@ -377,26 +485,34 @@ function sourceBoldLeadIn(source, text, names) {
|
|
|
377
485
|
for (const lead of bold(source)) {
|
|
378
486
|
if (lead.length <= 15 || !/\p{Ll}{3}/u.test(lead) || names.some(name => lead.includes(name)))
|
|
379
487
|
continue;
|
|
488
|
+
// Only capitalised words ("Apple App Store:", "Google Play"): a name, kept on purpose.
|
|
489
|
+
if (lead.split(/\s+/).every(word => !/\p{L}/u.test(word) || /^[^\p{L}]*\p{Lu}/u.test(word)))
|
|
490
|
+
continue;
|
|
380
491
|
if (translated.has(lead))
|
|
381
492
|
return lead;
|
|
382
493
|
}
|
|
383
494
|
return null;
|
|
384
495
|
}
|
|
496
|
+
const ICU_HEADER = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/g;
|
|
497
|
+
const ICU_HEADER_TEST = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/;
|
|
385
498
|
// Loanwords commonly written in Latin script inside non-Latin text.
|
|
386
499
|
const LATIN_LOANWORDS = new Set(['email', 'online', 'offline', 'emoji', 'smartphone', 'podcast', 'podcasts', 'wifi', 'blog', 'login', 'like', 'likes']);
|
|
387
500
|
/** Source-language words left inside a non-Latin-script translation ("Settings → Privacy"). */
|
|
388
|
-
function englishInNativeScript(lang, source, text, placeholderRe) {
|
|
501
|
+
function englishInNativeScript(lang, source, text, placeholderRe, allowed) {
|
|
389
502
|
const base = baseLanguage(lang);
|
|
390
503
|
// Greek writes many anglicisms in Latin script; not checked.
|
|
391
504
|
if (!NATIVE_SCRIPTS[base] || base === 'el' || base === 'sr' || text === source)
|
|
392
505
|
return null;
|
|
393
506
|
const strip = (value) => withoutUrls(value.replace(/[\w.+-]+@[\w.-]+/g, ' ').replace(/<[^>]+>/g, ' '))
|
|
507
|
+
// ICU syntax ("{count, plural, one {…} other {…}}") is code, not English text.
|
|
508
|
+
.replace(ICU_HEADER, ' ')
|
|
509
|
+
.replace(ICU_HEADER_TEST.test(value) ? /(?:^|[\s}])(?:zero|one|two|few|many|other|=\d+|[\w-]+)\s*(?=\{)/g : /$^/g, ' ')
|
|
394
510
|
.replace(placeholderRe, ' ')
|
|
395
511
|
.replace(/[((][^))]*[))]/g, ' ')
|
|
396
512
|
.replace(/["“„«「『‘'][^"”“»」』’']*["”“»」』’']/g, ' ');
|
|
397
513
|
const sourceWords = new Set(strip(source).match(/\b[a-z]{4,}\b/g) ?? []);
|
|
398
514
|
const left = [
|
|
399
|
-
...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word))),
|
|
515
|
+
...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word) && !allowed.has(word))),
|
|
400
516
|
];
|
|
401
517
|
const menu = /\b[A-Z][a-z]+ → [A-Z][a-z]+/.exec(text);
|
|
402
518
|
if (menu && source.includes(menu[0]))
|
|
@@ -425,9 +541,22 @@ const HEDGE_MARKERS = {
|
|
|
425
541
|
cs: /tendenc|obvykle|často|zpravidla|většinou|bývá|sklon|snadno|častěji|mív/iu,
|
|
426
542
|
sk: /tendenc|obvykle|často|zvyčajne|väčšinou|býva|sklon|ľahko|častejšie|zvyk|skôr/iu,
|
|
427
543
|
el: /τείν|συχνά|συνήθως|τάση|συνήθ|εύκολα|συχνότερα/iu,
|
|
428
|
-
fi: /taipu|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
|
|
544
|
+
fi: /taipu|tapana|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
|
|
429
545
|
ca: /tendeix|tendència|sol|sovint|generalment|acostum|normalment|fàcilment|freqüent/iu,
|
|
430
546
|
};
|
|
547
|
+
// Chinese, Japanese and Korean need far fewer characters than English; Thai has no spaces.
|
|
548
|
+
const COMPACT_SCRIPTS = new Set(['zh', 'ja', 'ko']);
|
|
549
|
+
/**
|
|
550
|
+
* A translation with a fraction of the source's length lost content: cut off by the model, or
|
|
551
|
+
* the source grew after it was translated. Markup and placeholders are not counted.
|
|
552
|
+
*/
|
|
553
|
+
export function muchShorter(lang, source, text) {
|
|
554
|
+
const visible = (value) => [...value.replace(/<[^>]+>|\{\{?[^{}]*\}?\}/g, '').replace(/\s+/g, ' ').trim()].length;
|
|
555
|
+
const [a, b] = [visible(source), visible(text)];
|
|
556
|
+
// Real translations from English rarely drop below 60% (Chinese/Japanese/Korean: 25%).
|
|
557
|
+
const min = COMPACT_SCRIPTS.has(baseLanguage(lang)) ? 0.2 : 0.5;
|
|
558
|
+
return a >= 200 && b < a * min ? `translation much shorter than the source (${b} of ${a} characters): content missing?` : null;
|
|
559
|
+
}
|
|
431
560
|
function droppedHedge(lang, source, text) {
|
|
432
561
|
const markers = HEDGE_MARKERS[baseLanguage(lang)];
|
|
433
562
|
return Boolean(markers) && /\btends? to\b/i.test(source) && !markers.test(text);
|