occular-ocr 0.2.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- occular_ocr-0.2.0.dist-info/METADATA +575 -0
- occular_ocr-0.2.0.dist-info/RECORD +20 -0
- occular_ocr-0.2.0.dist-info/WHEEL +5 -0
- occular_ocr-0.2.0.dist-info/entry_points.txt +2 -0
- occular_ocr-0.2.0.dist-info/licenses/LICENSE +202 -0
- occular_ocr-0.2.0.dist-info/top_level.txt +1 -0
- ocr_skel/__init__.py +274 -0
- ocr_skel/__main__.py +6 -0
- ocr_skel/_pylm.py +133 -0
- ocr_skel/_runtime.py +28 -0
- ocr_skel/cli.py +168 -0
- ocr_skel/dbnet_detector_onnx.py +110 -0
- ocr_skel/decoder_lm.py +83 -0
- ocr_skel/deskew.py +60 -0
- ocr_skel/model_files.py +115 -0
- ocr_skel/pipeline.py +352 -0
- ocr_skel/reading_order.py +84 -0
- ocr_skel/recognizer_onnx.py +139 -0
- ocr_skel/registry.py +62 -0
- ocr_skel/settings.py +56 -0
|
@@ -0,0 +1,575 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: occular-ocr
|
|
3
|
+
Version: 0.2.0
|
|
4
|
+
Summary: State-of-the-art OCR for Russian documents, with a zero-compilation install.
|
|
5
|
+
License: Apache-2.0
|
|
6
|
+
Classifier: License :: OSI Approved :: Apache Software License
|
|
7
|
+
Classifier: Programming Language :: Python :: 3
|
|
8
|
+
Classifier: Topic :: Scientific/Engineering :: Image Recognition
|
|
9
|
+
Classifier: Operating System :: OS Independent
|
|
10
|
+
Requires-Python: >=3.8
|
|
11
|
+
Description-Content-Type: text/markdown
|
|
12
|
+
License-File: LICENSE
|
|
13
|
+
Requires-Dist: numpy
|
|
14
|
+
Requires-Dist: opencv-python
|
|
15
|
+
Requires-Dist: Pillow
|
|
16
|
+
Requires-Dist: pyclipper
|
|
17
|
+
Requires-Dist: shapely
|
|
18
|
+
Requires-Dist: onnxruntime
|
|
19
|
+
Requires-Dist: pymupdf
|
|
20
|
+
Requires-Dist: huggingface_hub
|
|
21
|
+
Requires-Dist: pyctcdecode
|
|
22
|
+
Provides-Extra: gpu
|
|
23
|
+
Requires-Dist: onnxruntime-gpu; extra == "gpu"
|
|
24
|
+
Dynamic: classifier
|
|
25
|
+
Dynamic: description
|
|
26
|
+
Dynamic: description-content-type
|
|
27
|
+
Dynamic: license
|
|
28
|
+
Dynamic: license-file
|
|
29
|
+
Dynamic: provides-extra
|
|
30
|
+
Dynamic: requires-dist
|
|
31
|
+
Dynamic: requires-python
|
|
32
|
+
Dynamic: summary
|
|
33
|
+
|
|
34
|
+
<div align="center">
|
|
35
|
+
|
|
36
|
+
# Occular-OCR
|
|
37
|
+
|
|
38
|
+
**State-of-the-art OCR for Russian documents — with a zero-compilation install.**
|
|
39
|
+
|
|
40
|
+
[](LICENSE)
|
|
41
|
+
[](WEIGHTS_LICENSE.md)
|
|
42
|
+
[](#installation)
|
|
43
|
+
[](#installation)
|
|
44
|
+
|
|
45
|
+
<sub>🇬🇧 **English** · <a href="#russian">🇷🇺 Русский</a></sub>
|
|
46
|
+
|
|
47
|
+
</div>
|
|
48
|
+
|
|
49
|
+
Occular-OCR is a document OCR library purpose-built for **Russian**. It pairs a DBNet text detector with an
|
|
50
|
+
SVTR recognizer and a beam-search decoder guided by a Russian language model — and reads Russian
|
|
51
|
+
documents **markedly better than existing open-source OCR**, especially the hard stuff: forms,
|
|
52
|
+
certificates, bank statements, IDs, receipts.
|
|
53
|
+
|
|
54
|
+
It installs with a single `pip install` on Windows, Linux and macOS — **no C/C++ toolchain, no CUDA,
|
|
55
|
+
nothing to build.**
|
|
56
|
+
|
|
57
|
+
---
|
|
58
|
+
|
|
59
|
+
## See it in action
|
|
60
|
+
|
|
61
|
+
<p align="center">
|
|
62
|
+
<img src="assets/demo_passport.png" width="24%" alt="Detected lines on a specimen passport">
|
|
63
|
+
<img src="assets/demo_2ndfl.png" width="24%" alt="Detected lines on a synthetic income certificate">
|
|
64
|
+
<img src="assets/demo_book.png" width="24%" alt="Detected lines on a printed book page">
|
|
65
|
+
</p>
|
|
66
|
+
<p align="center">
|
|
67
|
+
<img src="assets/demo_schet.png" width="37%" alt="Detected lines on a synthetic VAT invoice">
|
|
68
|
+
<img src="assets/demo_invoice.png" width="37%" alt="Detected lines on a synthetic payment invoice">
|
|
69
|
+
</p>
|
|
70
|
+
|
|
71
|
+
<p align="center"><sub>Detected text lines on <b>synthetic / specimen / public-domain</b> documents — a passport, a tax certificate, a printed book page, a VAT invoice, a payment invoice.</sub></p>
|
|
72
|
+
|
|
73
|
+
## Why Occular-OCR
|
|
74
|
+
|
|
75
|
+
Occular-OCR is built for Russian and reads it dramatically better than existing open-source OCR:
|
|
76
|
+
**~35 % fewer word errors than the next-best open-source engine** in an end-to-end benchmark, and
|
|
77
|
+
ranking **#1** against large vision-language models and other leading OCR engines on a page-level
|
|
78
|
+
document benchmark. The gap is widest exactly where general-purpose OCR struggles: dense forms,
|
|
79
|
+
certificates, invoices and IDs.
|
|
80
|
+
|
|
81
|
+
- **Russian-first accuracy.** Trained and tuned on Russian document data. Where other engines guess,
|
|
82
|
+
Occular-OCR reads.
|
|
83
|
+
- **Beam + language model, on by default.** A Russian n-gram language model cuts word errors
|
|
84
|
+
~18–25 % over plain greedy decoding — with **no model retraining**, just better decoding.
|
|
85
|
+
- **Installs anywhere, no compiler.** The decoder is pure Python end to end, so you get top-tier
|
|
86
|
+
decoding quality with **zero native dependencies**.
|
|
87
|
+
- **CPU-first, GPU-ready.** Ships as ONNX and runs on CPU out of the box; set `gpu=True` to run the same models on CUDA via ONNX Runtime (`onnxruntime-gpu`).
|
|
88
|
+
- **Batteries included.** Images and PDF, folder → `.txt`, reading order, automatic deskew.
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
## Installation
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
pip install occular-ocr
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
That's it. No build tools, no CUDA. Model weights download automatically on first use and are cached
|
|
99
|
+
locally.
|
|
100
|
+
|
|
101
|
+
### GPU (optional)
|
|
102
|
+
|
|
103
|
+
GPU runs the **same ONNX models on CUDA** through ONNX Runtime's `CUDAExecutionProvider` — there are
|
|
104
|
+
no separate GPU weights. You only need the GPU runtime:
|
|
105
|
+
|
|
106
|
+
```bash
|
|
107
|
+
pip install occular-ocr[gpu] # pulls in onnxruntime-gpu
|
|
108
|
+
# equivalent to: pip install occular-ocr onnxruntime-gpu
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Then pass `gpu=True` (see below). If CUDA isn't available it falls back to CPU with a warning.
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## Quick start
|
|
116
|
+
|
|
117
|
+
```python
|
|
118
|
+
from ocr_skel import ocr
|
|
119
|
+
|
|
120
|
+
text = ocr("document.png") # an image, or a "scan.pdf"
|
|
121
|
+
print(text)
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
Line-level output with coordinates and confidence:
|
|
125
|
+
|
|
126
|
+
```python
|
|
127
|
+
from ocr_skel import ocr_detailed
|
|
128
|
+
|
|
129
|
+
for line in ocr_detailed("document.png"):
|
|
130
|
+
print(line["text"], round(line["confidence"], 2), line["quad"])
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
A whole folder → text files, from the command line:
|
|
134
|
+
|
|
135
|
+
```bash
|
|
136
|
+
python -m ocr_skel.cli ./scans ./out
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
> First run downloads the model weights automatically (and caches them). No build tools, no CUDA.
|
|
140
|
+
|
|
141
|
+
## Configuration
|
|
142
|
+
|
|
143
|
+
All behaviour is controlled by `Settings` (or the equivalent keyword arguments on `ocr`). Print the
|
|
144
|
+
current defaults at any time:
|
|
145
|
+
|
|
146
|
+
```python
|
|
147
|
+
from ocr_skel import Settings
|
|
148
|
+
print(Settings())
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
| Setting | Default | What it does |
|
|
152
|
+
|---|---|---|
|
|
153
|
+
| `num_threads` | `None` | CPU threads for inference. `None` → `min(cores, 4)`. |
|
|
154
|
+
| `gpu` | `False` | Run the ONNX models on **GPU via ONNX Runtime (CUDA)**. Needs `onnxruntime-gpu` (`pip install occular-ocr[gpu]`); falls back to CPU if unavailable. |
|
|
155
|
+
| `deskew` | `True` | Auto-correct skewed / rotated scans before detection. |
|
|
156
|
+
| `lm` | `True` | Beam search + language model (best quality). `False` → fast greedy decoding, skips the LM download. |
|
|
157
|
+
| `reading_order` | `False` | Order lines for multi-column layouts (downloads a small model on first use). |
|
|
158
|
+
| `detector` | `None` | Explicit detector name. `None` → default. |
|
|
159
|
+
| `recognizer` | `None` | Explicit recognizer name. `None` → default. |
|
|
160
|
+
|
|
161
|
+
### Full pipeline with every setting
|
|
162
|
+
|
|
163
|
+
```python
|
|
164
|
+
from ocr_skel import OCRPipeline, Settings
|
|
165
|
+
|
|
166
|
+
pipe = OCRPipeline(Settings(
|
|
167
|
+
num_threads=8, # CPU threads (None -> min(cores, 4))
|
|
168
|
+
gpu=False, # True -> ONNX Runtime CUDA (needs occular-ocr[gpu])
|
|
169
|
+
deskew=True, # auto-correct skewed scans
|
|
170
|
+
lm=True, # beam + language model; False -> greedy (faster, no LM download)
|
|
171
|
+
reading_order=False, # multi-column reading order (optional model, see below)
|
|
172
|
+
detector=None, # None = default detector
|
|
173
|
+
recognizer=None, # None = default recognizer
|
|
174
|
+
))
|
|
175
|
+
|
|
176
|
+
result = pipe.process_image("document.png")
|
|
177
|
+
# -> [{"quad": [[x, y], ...], "text": "...", "confidence": 0.97}, ...]
|
|
178
|
+
```
|
|
179
|
+
|
|
180
|
+
### PDF: render DPI and parallel workers
|
|
181
|
+
|
|
182
|
+
```python
|
|
183
|
+
pages = pipe.process_pdf(
|
|
184
|
+
"document.pdf",
|
|
185
|
+
dpi=300, # render resolution for scanned PDFs (default 300)
|
|
186
|
+
force_ocr=False, # True: OCR even PDFs that already have a text layer
|
|
187
|
+
workers=4, # parallel pages (None = auto min(cores, 4); 1 = sequential)
|
|
188
|
+
)
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
| PDF option | Default | What it does |
|
|
192
|
+
|---|---|---|
|
|
193
|
+
| `dpi` | `300` | Render resolution for scanned PDFs. Lower to `200` for speed on large batches. |
|
|
194
|
+
| `force_ocr` | `False` | OCR even PDFs that already contain a text layer. |
|
|
195
|
+
| `workers` | `None` | Parallel pages. `None` → auto `min(cores, 4)`; `1` → sequential. |
|
|
196
|
+
|
|
197
|
+
### One-liners
|
|
198
|
+
|
|
199
|
+
```python
|
|
200
|
+
from ocr_skel import ocr
|
|
201
|
+
|
|
202
|
+
ocr("document.png") # CPU (default), full quality
|
|
203
|
+
ocr("document.png", gpu=True) # GPU via ONNX Runtime CUDA (needs onnxruntime-gpu)
|
|
204
|
+
ocr("document.png", lm=False) # fast greedy mode, no LM download
|
|
205
|
+
ocr("document.png", deskew=False) # skip deskew
|
|
206
|
+
ocr("document.png", num_threads=2) # limit CPU threads
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+
### Command line
|
|
210
|
+
|
|
211
|
+
```bash
|
|
212
|
+
python -m ocr_skel.cli ./scans ./out --workers 4 --dpi 300
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
| Flag | Default | What it does |
|
|
216
|
+
|---|---|---|
|
|
217
|
+
| `--gpu` | off | Run on GPU via ONNX Runtime CUDA (needs onnxruntime-gpu). |
|
|
218
|
+
| `--dpi N` | `300` | PDF render resolution. |
|
|
219
|
+
| `--force-ocr` | off | OCR even vector PDFs. |
|
|
220
|
+
| `--workers N` | auto | Parallel workers. |
|
|
221
|
+
| `--out FILE` | — | Save structured results to a JSON file. |
|
|
222
|
+
|
|
223
|
+
### Optional: reading order for multi-column pages
|
|
224
|
+
|
|
225
|
+
Off by default. The model downloads once from the Hub.
|
|
226
|
+
|
|
227
|
+
```python
|
|
228
|
+
from ocr_skel import download_reading_order, OCRPipeline, Settings, model_info
|
|
229
|
+
|
|
230
|
+
download_reading_order() # one-time download
|
|
231
|
+
pipe = OCRPipeline(Settings(reading_order=True))
|
|
232
|
+
|
|
233
|
+
model_info() # show which weights are present locally
|
|
234
|
+
```
|
|
235
|
+
|
|
236
|
+
> 📓 Everything above is also in a runnable notebook: **[`examples.ipynb`](examples.ipynb)**.
|
|
237
|
+
|
|
238
|
+
---
|
|
239
|
+
|
|
240
|
+
## How it works
|
|
241
|
+
|
|
242
|
+
```
|
|
243
|
+
image ─▶ deskew ─▶ DBNet detector ─▶ crops ─▶ SVTR recognizer ─▶ beam + n-gram LM ─▶ text
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
The recognizer is a compact SVTR + CTC model exported to ONNX; the decoder is a CTC prefix beam search guided by a 4-gram
|
|
247
|
+
Russian language model. The entire decoding stack — language model and beam search alike — is
|
|
248
|
+
**pure Python**, which is what keeps `pip install` friction-free on every platform.
|
|
249
|
+
|
|
250
|
+
---
|
|
251
|
+
|
|
252
|
+
## Models & weights
|
|
253
|
+
|
|
254
|
+
Weights are fetched automatically on first use and cached:
|
|
255
|
+
|
|
256
|
+
| Component | What it does |
|
|
257
|
+
|---|---|
|
|
258
|
+
| DBNet text detector | finds text lines on the page |
|
|
259
|
+
| SVTR recognizer | reads text inside each line |
|
|
260
|
+
| Language model | rescoring for the beam decoder |
|
|
261
|
+
| Reading-order model *(optional)* | orders lines for multi-column layouts |
|
|
262
|
+
|
|
263
|
+
Inspect what's present locally:
|
|
264
|
+
|
|
265
|
+
```python
|
|
266
|
+
from ocr_skel import model_info
|
|
267
|
+
model_info()
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
---
|
|
271
|
+
|
|
272
|
+
## Benchmarks & methodology
|
|
273
|
+
|
|
274
|
+
Numbers above are **end-to-end** (detection **and** recognition on full pages), matched to
|
|
275
|
+
line-level ground truth by IoU, scored with word/character error rate. The Russian benchmark spans
|
|
276
|
+
30 document domains (bank, receipts, diplomas, certificates, IDs, court decisions, price tags,
|
|
277
|
+
newspapers, and more), ~15 pages each.
|
|
278
|
+
|
|
279
|
+
> **Note on ground truth.** Reference transcriptions come from a strong reference OCR system.
|
|
280
|
+
> Occular-OCR is tuned toward that transcription style, which favours it against third-party
|
|
281
|
+
> engines; the margin on Russian forms nonetheless exceeds what that bias alone explains.
|
|
282
|
+
|
|
283
|
+
---
|
|
284
|
+
|
|
285
|
+
## Licensing
|
|
286
|
+
|
|
287
|
+
- **Code — Apache License 2.0.** See [`LICENSE`](LICENSE).
|
|
288
|
+
- **Model weights — Modified AI Pubs OpenRAIL-M.** See [`WEIGHTS_LICENSE.md`](WEIGHTS_LICENSE.md).
|
|
289
|
+
**Free** for individuals, researchers, the self-employed, non-profits, and small organizations
|
|
290
|
+
(under **20 000 000 ₽** annual revenue **and** fewer than **8** employees). Larger organizations
|
|
291
|
+
need a commercial license — **300 000 ₽ / year per organization**.
|
|
292
|
+
Commercial enquiries: **user26665@gmail.com** · Telegram **[@Bodhi_b](https://t.me/Bodhi_b)**.
|
|
293
|
+
|
|
294
|
+
---
|
|
295
|
+
|
|
296
|
+
## Citation
|
|
297
|
+
|
|
298
|
+
```bibtex
|
|
299
|
+
@software{occular_ocr,
|
|
300
|
+
title = {Occular-OCR: State-of-the-art OCR for Russian documents},
|
|
301
|
+
year = {2026},
|
|
302
|
+
url = {https://github.com/<your-org>/occular-ocr}
|
|
303
|
+
}
|
|
304
|
+
```
|
|
305
|
+
|
|
306
|
+
<a id="russian"></a>
|
|
307
|
+
|
|
308
|
+
---
|
|
309
|
+
|
|
310
|
+
<div align="center">
|
|
311
|
+
|
|
312
|
+
# Occular-OCR · 🇷🇺 Русская версия
|
|
313
|
+
|
|
314
|
+
**Передовой OCR для русских документов — установка без компиляции.**
|
|
315
|
+
|
|
316
|
+
</div>
|
|
317
|
+
|
|
318
|
+
> 🇬🇧 English version is above · 🇷🇺 Ниже — то же самое по-русски.
|
|
319
|
+
|
|
320
|
+
Occular-OCR — библиотека OCR документов, созданная специально под **русский язык**. Связка:
|
|
321
|
+
DBNet-детектор текста + SVTR-распознаватель + beam-декодер с языковой моделью — читает русские
|
|
322
|
+
документы **заметно лучше существующего open-source OCR**, особенно сложное: формы, справки,
|
|
323
|
+
банковские выписки, удостоверения, чеки.
|
|
324
|
+
|
|
325
|
+
Ставится одной командой `pip install` на Windows, Linux и macOS — **без компилятора C/C++, без CUDA,
|
|
326
|
+
ничего собирать не нужно.**
|
|
327
|
+
|
|
328
|
+
---
|
|
329
|
+
|
|
330
|
+
## Демонстрация
|
|
331
|
+
|
|
332
|
+
<p align="center">
|
|
333
|
+
<img src="assets/demo_passport.png" width="24%" alt="Строки на образце паспорта">
|
|
334
|
+
<img src="assets/demo_2ndfl.png" width="24%" alt="Строки на синтетической справке о доходах">
|
|
335
|
+
<img src="assets/demo_book.png" width="24%" alt="Строки на печатной книжной странице">
|
|
336
|
+
</p>
|
|
337
|
+
<p align="center">
|
|
338
|
+
<img src="assets/demo_schet.png" width="37%" alt="Строки на синтетическом счёте-фактуре">
|
|
339
|
+
<img src="assets/demo_invoice.png" width="37%" alt="Строки на синтетическом счёте на оплату">
|
|
340
|
+
</p>
|
|
341
|
+
|
|
342
|
+
<p align="center"><sub>Найденные строки текста на <b>синтетических / образцовых / public-domain</b> документах — паспорт (образец), справка о доходах, книжная страница, счёт-фактура, счёт на оплату.</sub></p>
|
|
343
|
+
|
|
344
|
+
## Почему Occular-OCR
|
|
345
|
+
|
|
346
|
+
Occular-OCR заточен под русский и читает его значительно лучше существующего open-source OCR:
|
|
347
|
+
**~35 % меньше ошибок слов, чем ближайший open-source-движок** в сквозном бенчмарке, и **#1** против
|
|
348
|
+
больших vision-language моделей и других ведущих OCR-движков на постраничном бенчмарке. Разрыв
|
|
349
|
+
максимален там, где универсальный OCR буксует: плотные формы, справки, счета, удостоверения.
|
|
350
|
+
|
|
351
|
+
- **Точность под русский.** Обучен и настроен на русских документах. Там, где другие движки гадают,
|
|
352
|
+
Occular-OCR читает.
|
|
353
|
+
- **Beam + языковая модель по умолчанию.** Русская n-gram языковая модель снижает ошибки слов на
|
|
354
|
+
~18–25 % относительно жадного декодирования — **без дообучения модели**, просто лучше декодирование.
|
|
355
|
+
- **Ставится где угодно, без компилятора.** Декодер целиком на чистом Python — топовое качество декода
|
|
356
|
+
с **нулевыми нативными зависимостями**.
|
|
357
|
+
- **Сначала CPU, GPU — по желанию.** Поставляется как ONNX и работает на CPU из коробки; `gpu=True`
|
|
358
|
+
гоняет те же модели на CUDA через ONNX Runtime (`onnxruntime-gpu`).
|
|
359
|
+
- **Всё в комплекте.** Картинки и PDF, папка → `.txt`, порядок чтения, автоматическое выпрямление наклона.
|
|
360
|
+
|
|
361
|
+
---
|
|
362
|
+
|
|
363
|
+
## Установка
|
|
364
|
+
|
|
365
|
+
```bash
|
|
366
|
+
pip install occular-ocr
|
|
367
|
+
```
|
|
368
|
+
|
|
369
|
+
И всё. Без сборочных инструментов, без CUDA. Веса моделей скачиваются автоматически при первом запуске
|
|
370
|
+
и кэшируются.
|
|
371
|
+
|
|
372
|
+
### GPU (опционально)
|
|
373
|
+
|
|
374
|
+
GPU гоняет **те же ONNX-модели на CUDA** через `CUDAExecutionProvider` ONNX Runtime — отдельных
|
|
375
|
+
GPU-весов нет. Нужен только GPU-рантайм:
|
|
376
|
+
|
|
377
|
+
```bash
|
|
378
|
+
pip install occular-ocr[gpu] # доустанавливает onnxruntime-gpu
|
|
379
|
+
# то же самое, что: pip install occular-ocr onnxruntime-gpu
|
|
380
|
+
```
|
|
381
|
+
|
|
382
|
+
Затем передай `gpu=True` (см. ниже). Если CUDA недоступна — тихий откат на CPU с предупреждением.
|
|
383
|
+
|
|
384
|
+
---
|
|
385
|
+
|
|
386
|
+
## Быстрый старт
|
|
387
|
+
|
|
388
|
+
```python
|
|
389
|
+
from ocr_skel import ocr
|
|
390
|
+
|
|
391
|
+
text = ocr("document.png") # картинка или "scan.pdf"
|
|
392
|
+
print(text)
|
|
393
|
+
```
|
|
394
|
+
|
|
395
|
+
Построчный вывод с координатами и confidence:
|
|
396
|
+
|
|
397
|
+
```python
|
|
398
|
+
from ocr_skel import ocr_detailed
|
|
399
|
+
|
|
400
|
+
for line in ocr_detailed("document.png"):
|
|
401
|
+
print(line["text"], round(line["confidence"], 2), line["quad"])
|
|
402
|
+
```
|
|
403
|
+
|
|
404
|
+
Целая папка → текстовые файлы, из командной строки:
|
|
405
|
+
|
|
406
|
+
```bash
|
|
407
|
+
python -m ocr_skel.cli ./scans ./out
|
|
408
|
+
```
|
|
409
|
+
|
|
410
|
+
> Первый запуск сам скачает веса (и закэширует). Без сборочных инструментов, без CUDA.
|
|
411
|
+
|
|
412
|
+
## Настройки
|
|
413
|
+
|
|
414
|
+
Всё поведение задаётся через `Settings` (или эквивалентные именованные аргументы `ocr`). Посмотреть
|
|
415
|
+
текущие дефолты можно в любой момент:
|
|
416
|
+
|
|
417
|
+
```python
|
|
418
|
+
from ocr_skel import Settings
|
|
419
|
+
print(Settings())
|
|
420
|
+
```
|
|
421
|
+
|
|
422
|
+
| Настройка | По умолчанию | Что делает |
|
|
423
|
+
|---|---|---|
|
|
424
|
+
| `num_threads` | `None` | CPU-потоки для инференса. `None` → `min(ядра, 4)`. |
|
|
425
|
+
| `gpu` | `False` | Гонять ONNX-модели на **GPU через ONNX Runtime (CUDA)**. Нужен `onnxruntime-gpu` (`pip install occular-ocr[gpu]`); при отсутствии — откат на CPU. |
|
|
426
|
+
| `deskew` | `True` | Автовыпрямление наклонённых / повёрнутых сканов перед детекцией. |
|
|
427
|
+
| `lm` | `True` | Beam + языковая модель (лучшее качество). `False` → быстрое жадное декодирование, без скачивания LM. |
|
|
428
|
+
| `reading_order` | `False` | Упорядочивание строк для многоколоночных макетов (докачивает небольшую модель при первом запуске). |
|
|
429
|
+
| `detector` | `None` | Явное имя детектора. `None` → по умолчанию. |
|
|
430
|
+
| `recognizer` | `None` | Явное имя распознавателя. `None` → по умолчанию. |
|
|
431
|
+
|
|
432
|
+
### Пайплайн со всеми настройками
|
|
433
|
+
|
|
434
|
+
```python
|
|
435
|
+
from ocr_skel import OCRPipeline, Settings
|
|
436
|
+
|
|
437
|
+
pipe = OCRPipeline(Settings(
|
|
438
|
+
num_threads=8, # CPU-потоки (None -> min(ядра, 4))
|
|
439
|
+
gpu=False, # True -> ONNX Runtime CUDA (нужен occular-ocr[gpu])
|
|
440
|
+
deskew=True, # автовыпрямление наклона
|
|
441
|
+
lm=True, # beam + языковая модель; False -> жадное (быстрее, без скачивания LM)
|
|
442
|
+
reading_order=False, # порядок чтения для многоколоночных (опц. модель, см. ниже)
|
|
443
|
+
detector=None, # None = детектор по умолчанию
|
|
444
|
+
recognizer=None, # None = распознаватель по умолчанию
|
|
445
|
+
))
|
|
446
|
+
|
|
447
|
+
result = pipe.process_image("document.png")
|
|
448
|
+
# -> [{"quad": [[x, y], ...], "text": "...", "confidence": 0.97}, ...]
|
|
449
|
+
```
|
|
450
|
+
|
|
451
|
+
### PDF: DPI рендеринга и параллельные воркеры
|
|
452
|
+
|
|
453
|
+
```python
|
|
454
|
+
pages = pipe.process_pdf(
|
|
455
|
+
"document.pdf",
|
|
456
|
+
dpi=300, # разрешение рендеринга сканов (по умолчанию 300)
|
|
457
|
+
force_ocr=False, # True: OCR даже для PDF с текстовым слоем
|
|
458
|
+
workers=4, # параллельные страницы (None = авто min(ядра, 4); 1 = последовательно)
|
|
459
|
+
)
|
|
460
|
+
```
|
|
461
|
+
|
|
462
|
+
| Опция PDF | По умолчанию | Что делает |
|
|
463
|
+
|---|---|---|
|
|
464
|
+
| `dpi` | `300` | Разрешение рендеринга сканов. Снизь до `200` ради скорости на больших пачках. |
|
|
465
|
+
| `force_ocr` | `False` | OCR даже для PDF, где уже есть текстовый слой. |
|
|
466
|
+
| `workers` | `None` | Параллельные страницы. `None` → авто `min(ядра, 4)`; `1` → последовательно. |
|
|
467
|
+
|
|
468
|
+
### Однострочники
|
|
469
|
+
|
|
470
|
+
```python
|
|
471
|
+
from ocr_skel import ocr
|
|
472
|
+
|
|
473
|
+
ocr("document.png") # CPU (по умолчанию), полное качество
|
|
474
|
+
ocr("document.png", gpu=True) # GPU через ONNX Runtime CUDA (нужен onnxruntime-gpu)
|
|
475
|
+
ocr("document.png", lm=False) # быстрый жадный режим, без скачивания LM
|
|
476
|
+
ocr("document.png", deskew=False) # без выпрямления наклона
|
|
477
|
+
ocr("document.png", num_threads=2) # ограничить CPU-потоки
|
|
478
|
+
```
|
|
479
|
+
|
|
480
|
+
### Командная строка
|
|
481
|
+
|
|
482
|
+
```bash
|
|
483
|
+
python -m ocr_skel.cli ./scans ./out --workers 4 --dpi 300
|
|
484
|
+
```
|
|
485
|
+
|
|
486
|
+
| Флаг | По умолчанию | Что делает |
|
|
487
|
+
|---|---|---|
|
|
488
|
+
| `--gpu` | выкл | Гонять на GPU через ONNX Runtime CUDA (нужен onnxruntime-gpu). |
|
|
489
|
+
| `--dpi N` | `300` | Разрешение рендеринга PDF. |
|
|
490
|
+
| `--force-ocr` | выкл | OCR даже для векторных PDF. |
|
|
491
|
+
| `--workers N` | авто | Параллельные воркеры. |
|
|
492
|
+
| `--out FILE` | — | Сохранить структурированный результат в JSON. |
|
|
493
|
+
|
|
494
|
+
### Опционально: порядок чтения для многоколоночных страниц
|
|
495
|
+
|
|
496
|
+
По умолчанию выключено. Модель скачивается один раз.
|
|
497
|
+
|
|
498
|
+
```python
|
|
499
|
+
from ocr_skel import download_reading_order, OCRPipeline, Settings, model_info
|
|
500
|
+
|
|
501
|
+
download_reading_order() # разовая докачка
|
|
502
|
+
pipe = OCRPipeline(Settings(reading_order=True))
|
|
503
|
+
|
|
504
|
+
model_info() # показать, какие веса есть локально
|
|
505
|
+
```
|
|
506
|
+
|
|
507
|
+
> 📓 Всё вышеперечисленное есть и в исполняемом ноутбуке: **[`examples.ipynb`](examples.ipynb)**.
|
|
508
|
+
|
|
509
|
+
---
|
|
510
|
+
|
|
511
|
+
## Как это работает
|
|
512
|
+
|
|
513
|
+
```
|
|
514
|
+
картинка ─▶ deskew ─▶ DBNet-детектор ─▶ кропы ─▶ SVTR-распознаватель ─▶ beam + n-gram LM ─▶ текст
|
|
515
|
+
```
|
|
516
|
+
|
|
517
|
+
Распознаватель — компактная SVTR + CTC модель, экспортированная в ONNX; декодер — CTC prefix beam
|
|
518
|
+
search с 4-gram русской языковой моделью. Весь стек декодирования — и языковая модель, и beam-поиск —
|
|
519
|
+
**на чистом Python**, поэтому `pip install` проходит гладко на любой платформе.
|
|
520
|
+
|
|
521
|
+
---
|
|
522
|
+
|
|
523
|
+
## Модели и веса
|
|
524
|
+
|
|
525
|
+
Веса скачиваются автоматически при первом запуске и кэшируются:
|
|
526
|
+
|
|
527
|
+
| Компонент | Что делает |
|
|
528
|
+
|---|---|
|
|
529
|
+
| DBNet-детектор текста | находит строки текста на странице |
|
|
530
|
+
| SVTR-распознаватель | читает текст внутри каждой строки |
|
|
531
|
+
| Языковая модель | rescoring для beam-декодера |
|
|
532
|
+
| Модель порядка чтения *(опц.)* | упорядочивает строки для многоколоночных макетов |
|
|
533
|
+
|
|
534
|
+
Посмотреть, что есть локально:
|
|
535
|
+
|
|
536
|
+
```python
|
|
537
|
+
from ocr_skel import model_info
|
|
538
|
+
model_info()
|
|
539
|
+
```
|
|
540
|
+
|
|
541
|
+
---
|
|
542
|
+
|
|
543
|
+
## Бенчмарки и методология
|
|
544
|
+
|
|
545
|
+
Цифры выше — **сквозные** (детекция **и** распознавание на полных страницах), с матчингом к
|
|
546
|
+
построчному эталону по IoU, метрика — ошибка слов/символов. Русский бенчмарк охватывает 30 доменов
|
|
547
|
+
документов (банковские, чеки, дипломы, справки, удостоверения, судебные решения, ценники, газеты и
|
|
548
|
+
др.), ~15 страниц на домен.
|
|
549
|
+
|
|
550
|
+
> **Про эталон.** Эталонные транскрипции получены сильной референсной OCR-системой; Occular-OCR
|
|
551
|
+
> настроен под этот стиль транскрипции, что даёт ему фору против сторонних движков. Тем не менее на
|
|
552
|
+
> русских формах разрыв превышает то, что объясняется одним лишь этим уклоном.
|
|
553
|
+
|
|
554
|
+
---
|
|
555
|
+
|
|
556
|
+
## Лицензирование
|
|
557
|
+
|
|
558
|
+
- **Код — Apache License 2.0.** См. [`LICENSE`](LICENSE).
|
|
559
|
+
- **Веса моделей — Modified AI Pubs OpenRAIL-M.** См. [`WEIGHTS_LICENSE.md`](WEIGHTS_LICENSE.md).
|
|
560
|
+
**Бесплатно** для физлиц, исследователей, самозанятых, НКО и малых организаций (выручка до
|
|
561
|
+
**20 000 000 ₽** в год **и** менее **8** сотрудников). Крупным организациям нужна коммерческая
|
|
562
|
+
лицензия — **300 000 ₽ / год на организацию**.
|
|
563
|
+
Коммерческие вопросы: **user26665@gmail.com** · Telegram **[@Bodhi_b](https://t.me/Bodhi_b)**.
|
|
564
|
+
|
|
565
|
+
---
|
|
566
|
+
|
|
567
|
+
## Цитирование
|
|
568
|
+
|
|
569
|
+
```bibtex
|
|
570
|
+
@software{occular_ocr,
|
|
571
|
+
title = {Occular-OCR: State-of-the-art OCR for Russian documents},
|
|
572
|
+
year = {2026},
|
|
573
|
+
url = {https://github.com/<your-org>/occular-ocr}
|
|
574
|
+
}
|
|
575
|
+
```
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
occular_ocr-0.2.0.dist-info/licenses/LICENSE,sha256=z8d0m5b2O9McPEK1xHG_dWgUBT6EfBDz6wA0F7xSPTA,11358
|
|
2
|
+
ocr_skel/__init__.py,sha256=pPOfS2k0vu2dkHSSSEwjXlDZQQZhL1CZyKNYlaPfxXA,9713
|
|
3
|
+
ocr_skel/__main__.py,sha256=_5MXbPv9NY4shstjVb2cqTkRIkdC7-C9KkjkOl7Lj28,107
|
|
4
|
+
ocr_skel/_pylm.py,sha256=MNBbgsab-a_uOtGrhkw3t3cV7nDWFBvD5lBAo6YSaOM,4530
|
|
5
|
+
ocr_skel/_runtime.py,sha256=PPVnwu4tF0nl6kSSVXNmSxeOHpp4mnoXiaYC51JUErU,1524
|
|
6
|
+
ocr_skel/cli.py,sha256=Oa2qoyDiHZMyafgmAvlkD8-QN0NjMupAeoU6mIN9XiE,5174
|
|
7
|
+
ocr_skel/dbnet_detector_onnx.py,sha256=qzCuex5gfjrndXgD4y81UpH0OtiFKQY98EsVSqNxokQ,5461
|
|
8
|
+
ocr_skel/decoder_lm.py,sha256=8tnLe31HaYvaoFvPdVNbz9_pGiRFDwiNFcPAtoreplk,3998
|
|
9
|
+
ocr_skel/deskew.py,sha256=_-66VtWrOmnmxeTSYaF4av-Nc41v97PAHT9olF388jY,3415
|
|
10
|
+
ocr_skel/model_files.py,sha256=FU6vlCzhs8BVdO28Y7zzz68WhlK4B_ejEGRJIfzJVcE,6410
|
|
11
|
+
ocr_skel/pipeline.py,sha256=N-UdzVGhxrus7O4WabspPFHBjNMR0XKcmBhNl2S3gSY,16114
|
|
12
|
+
ocr_skel/reading_order.py,sha256=_HXH8ugDXFrBnsR7yBjZv57nGscW9Hzrz5pdbt2W4lI,4516
|
|
13
|
+
ocr_skel/recognizer_onnx.py,sha256=SPExiTeZUc5xDQA2xysFZfeBGjidVJe3AgPOv8y1VZk,6697
|
|
14
|
+
ocr_skel/registry.py,sha256=giD67gPp5ak0TTL3trMIKuLaZuoc0FgJDl-4NM6-XaM,2623
|
|
15
|
+
ocr_skel/settings.py,sha256=ok74_LrBrHZ2gIZGoDV45NntZpew8i6cb6PK_iE0te4,2706
|
|
16
|
+
occular_ocr-0.2.0.dist-info/METADATA,sha256=rWApmxSyVq_8XngXBWCVROpIb98Vsy7mUQ6WANJRLKc,25403
|
|
17
|
+
occular_ocr-0.2.0.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
|
|
18
|
+
occular_ocr-0.2.0.dist-info/entry_points.txt,sha256=SdSTdoXxWa6ehsU_bSpQG82Rca7jbtfIOi35n0H28g0,42
|
|
19
|
+
occular_ocr-0.2.0.dist-info/top_level.txt,sha256=givItsuFWu13xMEng9RnTAKNWH607B6KUefn0w9fM74,9
|
|
20
|
+
occular_ocr-0.2.0.dist-info/RECORD,,
|