ctrl-kd 1.1.6__tar.gz → 1.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: ctrl-kd
3
- Version: 1.1.6
3
+ Version: 1.2.0
4
4
  Summary: Convert WordStar 4-7 documents and print-to-disk files to text, Markdown, HTML, RTF, or PDF. ^KD: save and done.
5
5
  Author: Jon Michaels
6
6
  License: MIT
@@ -32,6 +32,8 @@ $ ctrl-kd ESSAY.WS -t html -t rtf # multiple formats
32
32
  $ ctrl-kd ESSAY.WS -t pdf --mode printed # a facsimile of the 1990 printout
33
33
  $ ctrl-kd --mode printed LETTER.WS # line-for-line, as it printed in 1990
34
34
  $ ctrl-kd --diagnose MYSTERY.FIL # what IS this file?
35
+ $ ctrl-kd --comments MEMO.WS # include the author's hidden comments
36
+ $ ctrl-kd --no-notes PAPER.WS # body text only, no notes
35
37
  ```
36
38
 
37
39
  ## Why another converter?
@@ -67,9 +69,19 @@ against surviving period printouts of the same documents. Its rules are empirica
67
69
  * **WS5+ symmetric blocks** (`0x1D`: real footnotes/endnotes, headings, page
68
70
  breaks — machinery added in WS5) are parsed with their nested structure,
69
71
  verified against the 86 WordStar 7 documents in Robert J. Sawyer's public
70
- WordStar archive: footnotes extract with in-text references (`[^n]` in
71
- Markdown), paragraph styles become headings, and 82/86 convert with zero
72
- mojibake. More WS5–7 corpora still welcome.
72
+ WordStar archive. All **four** note kinds WordStar distinguished are read and
73
+ kept apart — footnote, endnote, annotation, and comment — with in-text
74
+ references (`[^n]` in Markdown, DPUB-ARIA anchors in HTML, real `\footnote`
75
+ destinations in RTF). **Comments never appear unless you ask for them**, since
76
+ WordStar never printed them; `--diagnose` still reports that they exist.
77
+ Paragraph styles become headings, and 82/86 convert with zero mojibake.
78
+ More WS5–7 corpora still welcome.
79
+ * **Page geometry** from the file's own `.pl`/`.po`/`.mt`/`.mb`, so `--mode
80
+ printed` uses the page the author set up rather than an assumed one — and
81
+ `--diagnose` says whether the size came from the file or from the default.
82
+ In `printed` mode footnotes are laid out the way WordStar laid them out: at
83
+ the foot of the page that references them, behind a twenty-dash separator,
84
+ split across pages with `...Continued...` when they do not fit.
73
85
 
74
86
  ## Modes
75
87
 
@@ -95,6 +107,13 @@ An output format is one function over the parsed document — register it with t
95
107
  **[EXTENDING.md](EXTENDING.md)** has the IR contract, a complete worked example
96
108
  (BBCode in ~40 lines), and a checklist.
97
109
 
110
+ ## Siblings
111
+
112
+ **[soft-return](https://github.com/jonmichaels/soft-return)** — CtrlKD, a Swift
113
+ port of this engine, verified byte-for-byte against this implementation via
114
+ machine-generated test vectors (the two projects found six real bugs in each
115
+ other during the port). It grows the `sr` CLI and the Soft Return macOS app.
116
+
98
117
  ## Lineage
99
118
 
100
119
  Standing on the shoulders of the tools and documentation that kept WordStar
@@ -13,6 +13,8 @@ $ ctrl-kd ESSAY.WS -t html -t rtf # multiple formats
13
13
  $ ctrl-kd ESSAY.WS -t pdf --mode printed # a facsimile of the 1990 printout
14
14
  $ ctrl-kd --mode printed LETTER.WS # line-for-line, as it printed in 1990
15
15
  $ ctrl-kd --diagnose MYSTERY.FIL # what IS this file?
16
+ $ ctrl-kd --comments MEMO.WS # include the author's hidden comments
17
+ $ ctrl-kd --no-notes PAPER.WS # body text only, no notes
16
18
  ```
17
19
 
18
20
  ## Why another converter?
@@ -48,9 +50,19 @@ against surviving period printouts of the same documents. Its rules are empirica
48
50
  * **WS5+ symmetric blocks** (`0x1D`: real footnotes/endnotes, headings, page
49
51
  breaks — machinery added in WS5) are parsed with their nested structure,
50
52
  verified against the 86 WordStar 7 documents in Robert J. Sawyer's public
51
- WordStar archive: footnotes extract with in-text references (`[^n]` in
52
- Markdown), paragraph styles become headings, and 82/86 convert with zero
53
- mojibake. More WS5–7 corpora still welcome.
53
+ WordStar archive. All **four** note kinds WordStar distinguished are read and
54
+ kept apart — footnote, endnote, annotation, and comment — with in-text
55
+ references (`[^n]` in Markdown, DPUB-ARIA anchors in HTML, real `\footnote`
56
+ destinations in RTF). **Comments never appear unless you ask for them**, since
57
+ WordStar never printed them; `--diagnose` still reports that they exist.
58
+ Paragraph styles become headings, and 82/86 convert with zero mojibake.
59
+ More WS5–7 corpora still welcome.
60
+ * **Page geometry** from the file's own `.pl`/`.po`/`.mt`/`.mb`, so `--mode
61
+ printed` uses the page the author set up rather than an assumed one — and
62
+ `--diagnose` says whether the size came from the file or from the default.
63
+ In `printed` mode footnotes are laid out the way WordStar laid them out: at
64
+ the foot of the page that references them, behind a twenty-dash separator,
65
+ split across pages with `...Continued...` when they do not fit.
54
66
 
55
67
  ## Modes
56
68
 
@@ -76,6 +88,13 @@ An output format is one function over the parsed document — register it with t
76
88
  **[EXTENDING.md](EXTENDING.md)** has the IR contract, a complete worked example
77
89
  (BBCode in ~40 lines), and a checklist.
78
90
 
91
+ ## Siblings
92
+
93
+ **[soft-return](https://github.com/jonmichaels/soft-return)** — CtrlKD, a Swift
94
+ port of this engine, verified byte-for-byte against this implementation via
95
+ machine-generated test vectors (the two projects found six real bugs in each
96
+ other during the port). It grows the `sr` CLI and the Soft Return macOS app.
97
+
79
98
  ## Lineage
80
99
 
81
100
  Standing on the shoulders of the tools and documentation that kept WordStar
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "ctrl-kd"
7
- version = "1.1.6"
7
+ version = "1.2.0"
8
8
  description = "Convert WordStar 4-7 documents and print-to-disk files to text, Markdown, HTML, RTF, or PDF. ^KD: save and done."
9
9
  readme = "README.md"
10
10
  license = {text = "MIT"}
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: ctrl-kd
3
- Version: 1.1.6
3
+ Version: 1.2.0
4
4
  Summary: Convert WordStar 4-7 documents and print-to-disk files to text, Markdown, HTML, RTF, or PDF. ^KD: save and done.
5
5
  Author: Jon Michaels
6
6
  License: MIT
@@ -32,6 +32,8 @@ $ ctrl-kd ESSAY.WS -t html -t rtf # multiple formats
32
32
  $ ctrl-kd ESSAY.WS -t pdf --mode printed # a facsimile of the 1990 printout
33
33
  $ ctrl-kd --mode printed LETTER.WS # line-for-line, as it printed in 1990
34
34
  $ ctrl-kd --diagnose MYSTERY.FIL # what IS this file?
35
+ $ ctrl-kd --comments MEMO.WS # include the author's hidden comments
36
+ $ ctrl-kd --no-notes PAPER.WS # body text only, no notes
35
37
  ```
36
38
 
37
39
  ## Why another converter?
@@ -67,9 +69,19 @@ against surviving period printouts of the same documents. Its rules are empirica
67
69
  * **WS5+ symmetric blocks** (`0x1D`: real footnotes/endnotes, headings, page
68
70
  breaks — machinery added in WS5) are parsed with their nested structure,
69
71
  verified against the 86 WordStar 7 documents in Robert J. Sawyer's public
70
- WordStar archive: footnotes extract with in-text references (`[^n]` in
71
- Markdown), paragraph styles become headings, and 82/86 convert with zero
72
- mojibake. More WS5–7 corpora still welcome.
72
+ WordStar archive. All **four** note kinds WordStar distinguished are read and
73
+ kept apart — footnote, endnote, annotation, and comment — with in-text
74
+ references (`[^n]` in Markdown, DPUB-ARIA anchors in HTML, real `\footnote`
75
+ destinations in RTF). **Comments never appear unless you ask for them**, since
76
+ WordStar never printed them; `--diagnose` still reports that they exist.
77
+ Paragraph styles become headings, and 82/86 convert with zero mojibake.
78
+ More WS5–7 corpora still welcome.
79
+ * **Page geometry** from the file's own `.pl`/`.po`/`.mt`/`.mb`, so `--mode
80
+ printed` uses the page the author set up rather than an assumed one — and
81
+ `--diagnose` says whether the size came from the file or from the default.
82
+ In `printed` mode footnotes are laid out the way WordStar laid them out: at
83
+ the foot of the page that references them, behind a twenty-dash separator,
84
+ split across pages with `...Continued...` when they do not fit.
73
85
 
74
86
  ## Modes
75
87
 
@@ -95,6 +107,13 @@ An output format is one function over the parsed document — register it with t
95
107
  **[EXTENDING.md](EXTENDING.md)** has the IR contract, a complete worked example
96
108
  (BBCode in ~40 lines), and a checklist.
97
109
 
110
+ ## Siblings
111
+
112
+ **[soft-return](https://github.com/jonmichaels/soft-return)** — CtrlKD, a Swift
113
+ port of this engine, verified byte-for-byte against this implementation via
114
+ machine-generated test vectors (the two projects found six real bugs in each
115
+ other during the port). It grows the `sr` CLI and the Soft Return macOS app.
116
+
98
117
  ## Lineage
99
118
 
100
119
  Standing on the shoulders of the tools and documentation that kept WordStar
@@ -8,6 +8,7 @@ src/ctrl_kd.egg-info/entry_points.txt
8
8
  src/ctrl_kd.egg-info/top_level.txt
9
9
  src/ctrlkd/__init__.py
10
10
  src/ctrlkd/cli.py
11
+ src/ctrlkd/convert.py
11
12
  src/ctrlkd/core.py
12
13
  src/ctrlkd/emit.py
13
14
  src/ctrlkd/pdf.py
@@ -3,10 +3,6 @@ from .core import detect, parse, parse_ws, parse_printstream, Document, Block, L
3
3
  from .emit import (emit_text, emit_markdown, emit_html, emit_rtf,
4
4
  emitter, get_emitter, formats, load_plugins)
5
5
  from .pdf import emit_pdf # registers the 'pdf' format
6
+ from .convert import convert, select_notes, DEFAULT_NOTE_KINDS, ALL_NOTE_KINDS
6
7
 
7
- __version__ = '1.1.6'
8
-
9
- def convert(data: bytes, to: str = 'markdown', mode: str = 'modern',
10
- encoding: str = 'cp437', **options) -> str:
11
- """One-call API: bytes in, converted string out."""
12
- return get_emitter(to)['fn'](parse(data, encoding=encoding), mode, **options)
8
+ __version__ = '1.2.0'
@@ -8,6 +8,7 @@
8
8
  """
9
9
  import argparse, json, os, sys
10
10
  from . import core, emit
11
+ from .convert import DEFAULT_NOTE_KINDS # module attr, not the re-exported convert()
11
12
 
12
13
  def diagnose(path, data):
13
14
  det = core.detect(data)
@@ -17,7 +18,31 @@ def diagnose(path, data):
17
18
  info.update({k: doc.meta[k] for k in
18
19
  ('margin_estimate', 'dot_commands', 'unknown_codes', 'columnar')})
19
20
  info['paragraphs'] = sum(1 for b in doc.blocks if b.kind == 'para')
20
- info['footnotes'] = len(doc.footnotes)
21
+ # note kinds, counted separately (footnote/endnote/annotation/comment)
22
+ # rather than flattened, so a rescue tool can tell a file has hidden
23
+ # comments even when this run is only converting to plain text
24
+ info['notes'] = {kind: sum(1 for n in doc.notes if n.kind == kind)
25
+ for kind in ('footnote', 'endnote', 'annotation', 'comment')}
26
+ # unrecognised symmetrical-sequence types: preserved, not silently
27
+ # dropped, so --diagnose can report them instead of going quiet
28
+ info['unknown_blocks'] = [
29
+ {'type': f'0x{u.cmd:02x}' if u.cmd >= 0 else 'malformed',
30
+ 'offset': u.offset, 'length': len(u.data)}
31
+ for u in doc.unknown_blocks]
32
+ # page geometry from the file's own dot commands, with provenance --
33
+ # a caller must be able to say "Legal (from file)" vs "Letter (default)"
34
+ info['page'] = doc.meta.get('page')
35
+ # .PT/.PSA/.PSB are WordTsar's inventions, not WordStar commands: their
36
+ # presence identifies who WROTE the file, not how it is encoded
37
+ if doc.meta.get('producer'):
38
+ info['producer'] = doc.meta['producer']
39
+ elif det['variant'] == 'printstream':
40
+ doc = core.parse(data)
41
+ # damage WordStar itself introduced at print time (comments + the
42
+ # ASCII/ASC256/PRVIEW/WS4 drivers truncated the rest of the line);
43
+ # reported so it reads as a 1990s defect, not our parse failing
44
+ if doc.meta.get('comment_bug'):
45
+ info['comment_bug'] = doc.meta['comment_bug']
21
46
  return info
22
47
 
23
48
  def main(argv=None):
@@ -40,11 +65,20 @@ def main(argv=None):
40
65
  help='override detection')
41
66
  ap.add_argument('--encoding', default='cp437',
42
67
  help='byte encoding of the source (default: cp437)')
68
+ ap.add_argument('--no-notes', action='store_true',
69
+ help='omit footnotes, endnotes and annotations from the output')
70
+ ap.add_argument('--comments', action='store_true',
71
+ help="include WordStar comments, which it never printed "
72
+ "(author's asides, hidden since the file was written)")
43
73
  ap.add_argument('--diagnose', action='store_true',
44
74
  help='report what the file is (variant, margin, dot commands, '
45
75
  'unknown codes) as JSON; no conversion')
46
76
  a = ap.parse_args(argv)
47
77
  formats = a.to or ['markdown']
78
+ if a.no_notes:
79
+ notes = frozenset()
80
+ else:
81
+ notes = set(DEFAULT_NOTE_KINDS) | ({'comment'} if a.comments else set())
48
82
  if a.output and (len(a.files) > 1 or len(formats) > 1):
49
83
  ap.error('-o works with a single input and a single format; use -d for batch')
50
84
 
@@ -69,7 +103,7 @@ def main(argv=None):
69
103
  base = os.path.splitext(os.path.basename(path))[0]
70
104
  for fmt in formats:
71
105
  reg = emit.get_emitter(fmt)
72
- out = reg['fn'](doc, a.mode, title=base)
106
+ out = reg['fn'](doc, a.mode, title=base, notes=notes)
73
107
  if a.output:
74
108
  dest = a.output
75
109
  else:
@@ -0,0 +1,37 @@
1
+ """ctrl-kd conversion API: parse + emit in one call, plus the shared note-kind
2
+ inclusion filter that both `convert()` and each emitter (see emit.py) apply.
3
+
4
+ WordStar 5+/7 notes come in four kinds (see core.Note): footnote, endnote,
5
+ annotation, comment. WordStar itself only ever PRINTED the first three --
6
+ comments are an author-only aside, never rendered by WordStar at all (spec-
7
+ documented). Text/Markdown/HTML/RTF have no notion of "don't print this" the
8
+ way WordStar did, so that behavior has to be modeled explicitly here instead
9
+ of falling out of the parse for free.
10
+ """
11
+ from .core import parse
12
+ from .emit import get_emitter, DEFAULT_NOTE_KINDS, ALL_NOTE_KINDS, select_notes
13
+
14
+ __all__ = ['convert', 'select_notes', 'DEFAULT_NOTE_KINDS', 'ALL_NOTE_KINDS']
15
+
16
+
17
+ def convert(data: bytes, to: str = 'markdown', mode: str = 'modern',
18
+ encoding: str = 'cp437', notes=DEFAULT_NOTE_KINDS, **options) -> str:
19
+ """One-call API: bytes in, converted string out.
20
+
21
+ `notes` selects which note kinds appear in the output: any iterable of
22
+ 'footnote' / 'endnote' / 'annotation' / 'comment' (the strings match
23
+ Note.kind exactly). Default is DEFAULT_NOTE_KINDS -- footnotes, endnotes,
24
+ annotations, matching what WordStar itself would have printed. Comments
25
+ are excluded by default; pass a set that includes 'comment' (or the
26
+ ALL_NOTE_KINDS constant) to surface them -- e.g. a rescue tool run on
27
+ request to recover hidden author asides. Unrecognised kind strings are
28
+ silently ignored rather than raising, so a caller (CLI flag, a Swift/
29
+ macOS picker's checkbox state) can pass a fixed set without first
30
+ validating spelling against this library's kind names.
31
+
32
+ This is forwarded to the emitter as the same `notes=` keyword an emitter
33
+ accepts when called directly (bypassing convert()); third-party emitters
34
+ that don't know about it simply ignore it via **options, per the
35
+ EXTENDING.md contract.
36
+ """
37
+ return get_emitter(to)['fn'](parse(data, encoding=encoding), mode, notes=notes, **options)