ctrl-kd 1.1.6__tar.gz → 1.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {ctrl_kd-1.1.6/src/ctrl_kd.egg-info → ctrl_kd-1.2.0}/PKG-INFO +23 -4
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/README.md +22 -3
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/pyproject.toml +1 -1
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0/src/ctrl_kd.egg-info}/PKG-INFO +23 -4
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/src/ctrl_kd.egg-info/SOURCES.txt +1 -0
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/src/ctrlkd/__init__.py +2 -6
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/src/ctrlkd/cli.py +36 -2
- ctrl_kd-1.2.0/src/ctrlkd/convert.py +37 -0
- ctrl_kd-1.2.0/src/ctrlkd/core.py +791 -0
- ctrl_kd-1.2.0/src/ctrlkd/emit.py +566 -0
- ctrl_kd-1.2.0/src/ctrlkd/pdf.py +498 -0
- ctrl_kd-1.2.0/tests/test_ctrlkd.py +1130 -0
- ctrl_kd-1.1.6/src/ctrlkd/core.py +0 -392
- ctrl_kd-1.1.6/src/ctrlkd/emit.py +0 -241
- ctrl_kd-1.1.6/src/ctrlkd/pdf.py +0 -209
- ctrl_kd-1.1.6/tests/test_ctrlkd.py +0 -326
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/LICENSE +0 -0
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/setup.cfg +0 -0
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/src/ctrl_kd.egg-info/dependency_links.txt +0 -0
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/src/ctrl_kd.egg-info/entry_points.txt +0 -0
- {ctrl_kd-1.1.6 → ctrl_kd-1.2.0}/src/ctrl_kd.egg-info/top_level.txt +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: ctrl-kd
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.2.0
|
|
4
4
|
Summary: Convert WordStar 4-7 documents and print-to-disk files to text, Markdown, HTML, RTF, or PDF. ^KD: save and done.
|
|
5
5
|
Author: Jon Michaels
|
|
6
6
|
License: MIT
|
|
@@ -32,6 +32,8 @@ $ ctrl-kd ESSAY.WS -t html -t rtf # multiple formats
|
|
|
32
32
|
$ ctrl-kd ESSAY.WS -t pdf --mode printed # a facsimile of the 1990 printout
|
|
33
33
|
$ ctrl-kd --mode printed LETTER.WS # line-for-line, as it printed in 1990
|
|
34
34
|
$ ctrl-kd --diagnose MYSTERY.FIL # what IS this file?
|
|
35
|
+
$ ctrl-kd --comments MEMO.WS # include the author's hidden comments
|
|
36
|
+
$ ctrl-kd --no-notes PAPER.WS # body text only, no notes
|
|
35
37
|
```
|
|
36
38
|
|
|
37
39
|
## Why another converter?
|
|
@@ -67,9 +69,19 @@ against surviving period printouts of the same documents. Its rules are empirica
|
|
|
67
69
|
* **WS5+ symmetric blocks** (`0x1D`: real footnotes/endnotes, headings, page
|
|
68
70
|
breaks — machinery added in WS5) are parsed with their nested structure,
|
|
69
71
|
verified against the 86 WordStar 7 documents in Robert J. Sawyer's public
|
|
70
|
-
WordStar archive
|
|
71
|
-
|
|
72
|
-
|
|
72
|
+
WordStar archive. All **four** note kinds WordStar distinguished are read and
|
|
73
|
+
kept apart — footnote, endnote, annotation, and comment — with in-text
|
|
74
|
+
references (`[^n]` in Markdown, DPUB-ARIA anchors in HTML, real `\footnote`
|
|
75
|
+
destinations in RTF). **Comments never appear unless you ask for them**, since
|
|
76
|
+
WordStar never printed them; `--diagnose` still reports that they exist.
|
|
77
|
+
Paragraph styles become headings, and 82/86 convert with zero mojibake.
|
|
78
|
+
More WS5–7 corpora still welcome.
|
|
79
|
+
* **Page geometry** from the file's own `.pl`/`.po`/`.mt`/`.mb`, so `--mode
|
|
80
|
+
printed` uses the page the author set up rather than an assumed one — and
|
|
81
|
+
`--diagnose` says whether the size came from the file or from the default.
|
|
82
|
+
In `printed` mode footnotes are laid out the way WordStar laid them out: at
|
|
83
|
+
the foot of the page that references them, behind a twenty-dash separator,
|
|
84
|
+
split across pages with `...Continued...` when they do not fit.
|
|
73
85
|
|
|
74
86
|
## Modes
|
|
75
87
|
|
|
@@ -95,6 +107,13 @@ An output format is one function over the parsed document — register it with t
|
|
|
95
107
|
**[EXTENDING.md](EXTENDING.md)** has the IR contract, a complete worked example
|
|
96
108
|
(BBCode in ~40 lines), and a checklist.
|
|
97
109
|
|
|
110
|
+
## Siblings
|
|
111
|
+
|
|
112
|
+
**[soft-return](https://github.com/jonmichaels/soft-return)** — CtrlKD, a Swift
|
|
113
|
+
port of this engine, verified byte-for-byte against this implementation via
|
|
114
|
+
machine-generated test vectors (the two projects found six real bugs in each
|
|
115
|
+
other during the port). It grows the `sr` CLI and the Soft Return macOS app.
|
|
116
|
+
|
|
98
117
|
## Lineage
|
|
99
118
|
|
|
100
119
|
Standing on the shoulders of the tools and documentation that kept WordStar
|
|
@@ -13,6 +13,8 @@ $ ctrl-kd ESSAY.WS -t html -t rtf # multiple formats
|
|
|
13
13
|
$ ctrl-kd ESSAY.WS -t pdf --mode printed # a facsimile of the 1990 printout
|
|
14
14
|
$ ctrl-kd --mode printed LETTER.WS # line-for-line, as it printed in 1990
|
|
15
15
|
$ ctrl-kd --diagnose MYSTERY.FIL # what IS this file?
|
|
16
|
+
$ ctrl-kd --comments MEMO.WS # include the author's hidden comments
|
|
17
|
+
$ ctrl-kd --no-notes PAPER.WS # body text only, no notes
|
|
16
18
|
```
|
|
17
19
|
|
|
18
20
|
## Why another converter?
|
|
@@ -48,9 +50,19 @@ against surviving period printouts of the same documents. Its rules are empirica
|
|
|
48
50
|
* **WS5+ symmetric blocks** (`0x1D`: real footnotes/endnotes, headings, page
|
|
49
51
|
breaks — machinery added in WS5) are parsed with their nested structure,
|
|
50
52
|
verified against the 86 WordStar 7 documents in Robert J. Sawyer's public
|
|
51
|
-
WordStar archive
|
|
52
|
-
|
|
53
|
-
|
|
53
|
+
WordStar archive. All **four** note kinds WordStar distinguished are read and
|
|
54
|
+
kept apart — footnote, endnote, annotation, and comment — with in-text
|
|
55
|
+
references (`[^n]` in Markdown, DPUB-ARIA anchors in HTML, real `\footnote`
|
|
56
|
+
destinations in RTF). **Comments never appear unless you ask for them**, since
|
|
57
|
+
WordStar never printed them; `--diagnose` still reports that they exist.
|
|
58
|
+
Paragraph styles become headings, and 82/86 convert with zero mojibake.
|
|
59
|
+
More WS5–7 corpora still welcome.
|
|
60
|
+
* **Page geometry** from the file's own `.pl`/`.po`/`.mt`/`.mb`, so `--mode
|
|
61
|
+
printed` uses the page the author set up rather than an assumed one — and
|
|
62
|
+
`--diagnose` says whether the size came from the file or from the default.
|
|
63
|
+
In `printed` mode footnotes are laid out the way WordStar laid them out: at
|
|
64
|
+
the foot of the page that references them, behind a twenty-dash separator,
|
|
65
|
+
split across pages with `...Continued...` when they do not fit.
|
|
54
66
|
|
|
55
67
|
## Modes
|
|
56
68
|
|
|
@@ -76,6 +88,13 @@ An output format is one function over the parsed document — register it with t
|
|
|
76
88
|
**[EXTENDING.md](EXTENDING.md)** has the IR contract, a complete worked example
|
|
77
89
|
(BBCode in ~40 lines), and a checklist.
|
|
78
90
|
|
|
91
|
+
## Siblings
|
|
92
|
+
|
|
93
|
+
**[soft-return](https://github.com/jonmichaels/soft-return)** — CtrlKD, a Swift
|
|
94
|
+
port of this engine, verified byte-for-byte against this implementation via
|
|
95
|
+
machine-generated test vectors (the two projects found six real bugs in each
|
|
96
|
+
other during the port). It grows the `sr` CLI and the Soft Return macOS app.
|
|
97
|
+
|
|
79
98
|
## Lineage
|
|
80
99
|
|
|
81
100
|
Standing on the shoulders of the tools and documentation that kept WordStar
|
|
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
|
|
|
4
4
|
|
|
5
5
|
[project]
|
|
6
6
|
name = "ctrl-kd"
|
|
7
|
-
version = "1.
|
|
7
|
+
version = "1.2.0"
|
|
8
8
|
description = "Convert WordStar 4-7 documents and print-to-disk files to text, Markdown, HTML, RTF, or PDF. ^KD: save and done."
|
|
9
9
|
readme = "README.md"
|
|
10
10
|
license = {text = "MIT"}
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: ctrl-kd
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.2.0
|
|
4
4
|
Summary: Convert WordStar 4-7 documents and print-to-disk files to text, Markdown, HTML, RTF, or PDF. ^KD: save and done.
|
|
5
5
|
Author: Jon Michaels
|
|
6
6
|
License: MIT
|
|
@@ -32,6 +32,8 @@ $ ctrl-kd ESSAY.WS -t html -t rtf # multiple formats
|
|
|
32
32
|
$ ctrl-kd ESSAY.WS -t pdf --mode printed # a facsimile of the 1990 printout
|
|
33
33
|
$ ctrl-kd --mode printed LETTER.WS # line-for-line, as it printed in 1990
|
|
34
34
|
$ ctrl-kd --diagnose MYSTERY.FIL # what IS this file?
|
|
35
|
+
$ ctrl-kd --comments MEMO.WS # include the author's hidden comments
|
|
36
|
+
$ ctrl-kd --no-notes PAPER.WS # body text only, no notes
|
|
35
37
|
```
|
|
36
38
|
|
|
37
39
|
## Why another converter?
|
|
@@ -67,9 +69,19 @@ against surviving period printouts of the same documents. Its rules are empirica
|
|
|
67
69
|
* **WS5+ symmetric blocks** (`0x1D`: real footnotes/endnotes, headings, page
|
|
68
70
|
breaks — machinery added in WS5) are parsed with their nested structure,
|
|
69
71
|
verified against the 86 WordStar 7 documents in Robert J. Sawyer's public
|
|
70
|
-
WordStar archive
|
|
71
|
-
|
|
72
|
-
|
|
72
|
+
WordStar archive. All **four** note kinds WordStar distinguished are read and
|
|
73
|
+
kept apart — footnote, endnote, annotation, and comment — with in-text
|
|
74
|
+
references (`[^n]` in Markdown, DPUB-ARIA anchors in HTML, real `\footnote`
|
|
75
|
+
destinations in RTF). **Comments never appear unless you ask for them**, since
|
|
76
|
+
WordStar never printed them; `--diagnose` still reports that they exist.
|
|
77
|
+
Paragraph styles become headings, and 82/86 convert with zero mojibake.
|
|
78
|
+
More WS5–7 corpora still welcome.
|
|
79
|
+
* **Page geometry** from the file's own `.pl`/`.po`/`.mt`/`.mb`, so `--mode
|
|
80
|
+
printed` uses the page the author set up rather than an assumed one — and
|
|
81
|
+
`--diagnose` says whether the size came from the file or from the default.
|
|
82
|
+
In `printed` mode footnotes are laid out the way WordStar laid them out: at
|
|
83
|
+
the foot of the page that references them, behind a twenty-dash separator,
|
|
84
|
+
split across pages with `...Continued...` when they do not fit.
|
|
73
85
|
|
|
74
86
|
## Modes
|
|
75
87
|
|
|
@@ -95,6 +107,13 @@ An output format is one function over the parsed document — register it with t
|
|
|
95
107
|
**[EXTENDING.md](EXTENDING.md)** has the IR contract, a complete worked example
|
|
96
108
|
(BBCode in ~40 lines), and a checklist.
|
|
97
109
|
|
|
110
|
+
## Siblings
|
|
111
|
+
|
|
112
|
+
**[soft-return](https://github.com/jonmichaels/soft-return)** — CtrlKD, a Swift
|
|
113
|
+
port of this engine, verified byte-for-byte against this implementation via
|
|
114
|
+
machine-generated test vectors (the two projects found six real bugs in each
|
|
115
|
+
other during the port). It grows the `sr` CLI and the Soft Return macOS app.
|
|
116
|
+
|
|
98
117
|
## Lineage
|
|
99
118
|
|
|
100
119
|
Standing on the shoulders of the tools and documentation that kept WordStar
|
|
@@ -3,10 +3,6 @@ from .core import detect, parse, parse_ws, parse_printstream, Document, Block, L
|
|
|
3
3
|
from .emit import (emit_text, emit_markdown, emit_html, emit_rtf,
|
|
4
4
|
emitter, get_emitter, formats, load_plugins)
|
|
5
5
|
from .pdf import emit_pdf # registers the 'pdf' format
|
|
6
|
+
from .convert import convert, select_notes, DEFAULT_NOTE_KINDS, ALL_NOTE_KINDS
|
|
6
7
|
|
|
7
|
-
__version__ = '1.
|
|
8
|
-
|
|
9
|
-
def convert(data: bytes, to: str = 'markdown', mode: str = 'modern',
|
|
10
|
-
encoding: str = 'cp437', **options) -> str:
|
|
11
|
-
"""One-call API: bytes in, converted string out."""
|
|
12
|
-
return get_emitter(to)['fn'](parse(data, encoding=encoding), mode, **options)
|
|
8
|
+
__version__ = '1.2.0'
|
|
@@ -8,6 +8,7 @@
|
|
|
8
8
|
"""
|
|
9
9
|
import argparse, json, os, sys
|
|
10
10
|
from . import core, emit
|
|
11
|
+
from .convert import DEFAULT_NOTE_KINDS # module attr, not the re-exported convert()
|
|
11
12
|
|
|
12
13
|
def diagnose(path, data):
|
|
13
14
|
det = core.detect(data)
|
|
@@ -17,7 +18,31 @@ def diagnose(path, data):
|
|
|
17
18
|
info.update({k: doc.meta[k] for k in
|
|
18
19
|
('margin_estimate', 'dot_commands', 'unknown_codes', 'columnar')})
|
|
19
20
|
info['paragraphs'] = sum(1 for b in doc.blocks if b.kind == 'para')
|
|
20
|
-
|
|
21
|
+
# note kinds, counted separately (footnote/endnote/annotation/comment)
|
|
22
|
+
# rather than flattened, so a rescue tool can tell a file has hidden
|
|
23
|
+
# comments even when this run is only converting to plain text
|
|
24
|
+
info['notes'] = {kind: sum(1 for n in doc.notes if n.kind == kind)
|
|
25
|
+
for kind in ('footnote', 'endnote', 'annotation', 'comment')}
|
|
26
|
+
# unrecognised symmetrical-sequence types: preserved, not silently
|
|
27
|
+
# dropped, so --diagnose can report them instead of going quiet
|
|
28
|
+
info['unknown_blocks'] = [
|
|
29
|
+
{'type': f'0x{u.cmd:02x}' if u.cmd >= 0 else 'malformed',
|
|
30
|
+
'offset': u.offset, 'length': len(u.data)}
|
|
31
|
+
for u in doc.unknown_blocks]
|
|
32
|
+
# page geometry from the file's own dot commands, with provenance --
|
|
33
|
+
# a caller must be able to say "Legal (from file)" vs "Letter (default)"
|
|
34
|
+
info['page'] = doc.meta.get('page')
|
|
35
|
+
# .PT/.PSA/.PSB are WordTsar's inventions, not WordStar commands: their
|
|
36
|
+
# presence identifies who WROTE the file, not how it is encoded
|
|
37
|
+
if doc.meta.get('producer'):
|
|
38
|
+
info['producer'] = doc.meta['producer']
|
|
39
|
+
elif det['variant'] == 'printstream':
|
|
40
|
+
doc = core.parse(data)
|
|
41
|
+
# damage WordStar itself introduced at print time (comments + the
|
|
42
|
+
# ASCII/ASC256/PRVIEW/WS4 drivers truncated the rest of the line);
|
|
43
|
+
# reported so it reads as a 1990s defect, not our parse failing
|
|
44
|
+
if doc.meta.get('comment_bug'):
|
|
45
|
+
info['comment_bug'] = doc.meta['comment_bug']
|
|
21
46
|
return info
|
|
22
47
|
|
|
23
48
|
def main(argv=None):
|
|
@@ -40,11 +65,20 @@ def main(argv=None):
|
|
|
40
65
|
help='override detection')
|
|
41
66
|
ap.add_argument('--encoding', default='cp437',
|
|
42
67
|
help='byte encoding of the source (default: cp437)')
|
|
68
|
+
ap.add_argument('--no-notes', action='store_true',
|
|
69
|
+
help='omit footnotes, endnotes and annotations from the output')
|
|
70
|
+
ap.add_argument('--comments', action='store_true',
|
|
71
|
+
help="include WordStar comments, which it never printed "
|
|
72
|
+
"(author's asides, hidden since the file was written)")
|
|
43
73
|
ap.add_argument('--diagnose', action='store_true',
|
|
44
74
|
help='report what the file is (variant, margin, dot commands, '
|
|
45
75
|
'unknown codes) as JSON; no conversion')
|
|
46
76
|
a = ap.parse_args(argv)
|
|
47
77
|
formats = a.to or ['markdown']
|
|
78
|
+
if a.no_notes:
|
|
79
|
+
notes = frozenset()
|
|
80
|
+
else:
|
|
81
|
+
notes = set(DEFAULT_NOTE_KINDS) | ({'comment'} if a.comments else set())
|
|
48
82
|
if a.output and (len(a.files) > 1 or len(formats) > 1):
|
|
49
83
|
ap.error('-o works with a single input and a single format; use -d for batch')
|
|
50
84
|
|
|
@@ -69,7 +103,7 @@ def main(argv=None):
|
|
|
69
103
|
base = os.path.splitext(os.path.basename(path))[0]
|
|
70
104
|
for fmt in formats:
|
|
71
105
|
reg = emit.get_emitter(fmt)
|
|
72
|
-
out = reg['fn'](doc, a.mode, title=base)
|
|
106
|
+
out = reg['fn'](doc, a.mode, title=base, notes=notes)
|
|
73
107
|
if a.output:
|
|
74
108
|
dest = a.output
|
|
75
109
|
else:
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
"""ctrl-kd conversion API: parse + emit in one call, plus the shared note-kind
|
|
2
|
+
inclusion filter that both `convert()` and each emitter (see emit.py) apply.
|
|
3
|
+
|
|
4
|
+
WordStar 5+/7 notes come in four kinds (see core.Note): footnote, endnote,
|
|
5
|
+
annotation, comment. WordStar itself only ever PRINTED the first three --
|
|
6
|
+
comments are an author-only aside, never rendered by WordStar at all (spec-
|
|
7
|
+
documented). Text/Markdown/HTML/RTF have no notion of "don't print this" the
|
|
8
|
+
way WordStar did, so that behavior has to be modeled explicitly here instead
|
|
9
|
+
of falling out of the parse for free.
|
|
10
|
+
"""
|
|
11
|
+
from .core import parse
|
|
12
|
+
from .emit import get_emitter, DEFAULT_NOTE_KINDS, ALL_NOTE_KINDS, select_notes
|
|
13
|
+
|
|
14
|
+
__all__ = ['convert', 'select_notes', 'DEFAULT_NOTE_KINDS', 'ALL_NOTE_KINDS']
|
|
15
|
+
|
|
16
|
+
|
|
17
|
+
def convert(data: bytes, to: str = 'markdown', mode: str = 'modern',
|
|
18
|
+
encoding: str = 'cp437', notes=DEFAULT_NOTE_KINDS, **options) -> str:
|
|
19
|
+
"""One-call API: bytes in, converted string out.
|
|
20
|
+
|
|
21
|
+
`notes` selects which note kinds appear in the output: any iterable of
|
|
22
|
+
'footnote' / 'endnote' / 'annotation' / 'comment' (the strings match
|
|
23
|
+
Note.kind exactly). Default is DEFAULT_NOTE_KINDS -- footnotes, endnotes,
|
|
24
|
+
annotations, matching what WordStar itself would have printed. Comments
|
|
25
|
+
are excluded by default; pass a set that includes 'comment' (or the
|
|
26
|
+
ALL_NOTE_KINDS constant) to surface them -- e.g. a rescue tool run on
|
|
27
|
+
request to recover hidden author asides. Unrecognised kind strings are
|
|
28
|
+
silently ignored rather than raising, so a caller (CLI flag, a Swift/
|
|
29
|
+
macOS picker's checkbox state) can pass a fixed set without first
|
|
30
|
+
validating spelling against this library's kind names.
|
|
31
|
+
|
|
32
|
+
This is forwarded to the emitter as the same `notes=` keyword an emitter
|
|
33
|
+
accepts when called directly (bypassing convert()); third-party emitters
|
|
34
|
+
that don't know about it simply ignore it via **options, per the
|
|
35
|
+
EXTENDING.md contract.
|
|
36
|
+
"""
|
|
37
|
+
return get_emitter(to)['fn'](parse(data, encoding=encoding), mode, notes=notes, **options)
|