flowmark 0.4.1__tar.gz → 0.4.2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {flowmark-0.4.1 → flowmark-0.4.2}/PKG-INFO +20 -5
- {flowmark-0.4.1 → flowmark-0.4.2}/README.md +19 -4
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/custom_marko.py +42 -3
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/line_wrappers.py +4 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/markdown_filling.py +4 -3
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/text_wrapping.py +43 -2
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/test_wrapping.py +93 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/testdocs/testdoc.orig.md +3 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/testdocs/testdoc.out.cleaned.md +19 -20
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/testdocs/testdoc.out.plain.md +15 -14
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/testdocs/testdoc.out.semantic.md +19 -20
- {flowmark-0.4.1 → flowmark-0.4.2}/.copier-answers.yml +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/.github/workflows/ci.yml +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/.github/workflows/publish.yml +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/.gitignore +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/LICENSE +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/Makefile +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/development.md +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/devtools/lint.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/installation.md +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/poetry.lock +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/publishing.md +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/pyproject.toml +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/__init__.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/cleanups.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/cli.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/frontmatter.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/py.typed +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/sentence_split_regex.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/src/flowmark/text_filling.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/test_cleanups.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/test_filling.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/test_frontmatter.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/test_ref_docs.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/tests/test_sentences.py +0 -0
- {flowmark-0.4.1 → flowmark-0.4.2}/uv.lock +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: flowmark
|
|
3
|
-
Version: 0.4.
|
|
3
|
+
Version: 0.4.2
|
|
4
4
|
Summary: Better line wrapping and formatting for plaintext and Markdown
|
|
5
5
|
Project-URL: Repository, https://github.com/jlevy/flowmark
|
|
6
6
|
Author-email: Joshua Levy <joshua@cal.berkeley.edu>
|
|
@@ -153,8 +153,7 @@ experience.
|
|
|
153
153
|
Flowmark can be used as a library or as a CLI.
|
|
154
154
|
|
|
155
155
|
```
|
|
156
|
-
|
|
157
|
-
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [file]
|
|
156
|
+
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [--auto] [--version] [file]
|
|
158
157
|
|
|
159
158
|
Flowmark: Better line wrapping and formatting for plaintext and Markdown
|
|
160
159
|
|
|
@@ -166,9 +165,11 @@ options:
|
|
|
166
165
|
-o, --output OUTPUT Output file (use '-' for stdout)
|
|
167
166
|
-w, --width WIDTH Line width to wrap to
|
|
168
167
|
-p, --plaintext Process as plaintext (no Markdown parsing)
|
|
169
|
-
-s, --semantic Enable sentence-based line breaks (only applies to Markdown mode)
|
|
168
|
+
-s, --semantic Enable semantic (sentence-based) line breaks (only applies to Markdown mode)
|
|
170
169
|
-i, --inplace Edit the file in place (ignores --output)
|
|
171
170
|
--nobackup Do not make a backup of the original file when using --inplace
|
|
171
|
+
--auto Same as `--inplace --nobackup --semantic`, as a convenience for auto-formatting files
|
|
172
|
+
--version Show version information and exit
|
|
172
173
|
|
|
173
174
|
Flowmark provides enhanced text wrapping capabilities with special handling for
|
|
174
175
|
Markdown content. It can:
|
|
@@ -197,12 +198,26 @@ Command-line usage examples:
|
|
|
197
198
|
# Process plaintext instead of Markdown
|
|
198
199
|
flowmark --plaintext text.txt
|
|
199
200
|
|
|
200
|
-
# Use
|
|
201
|
+
# Use semantic line breaks (based on sentences, which is helpful to reduce
|
|
202
|
+
# irrelevant line wrap diffs in git history)
|
|
201
203
|
flowmark --semantic README.md
|
|
202
204
|
|
|
203
205
|
For more details, see: https://github.com/jlevy/flowmark
|
|
204
206
|
```
|
|
205
207
|
|
|
208
|
+
## Other Notes
|
|
209
|
+
|
|
210
|
+
- This enables
|
|
211
|
+
[GitHub-flavored Markdown support](https://docs.github.com/en/get-started/writing-on-github/getting-started-with-writing-and-formatting-on-github/basic-writing-and-formatting-syntax)
|
|
212
|
+
using
|
|
213
|
+
[Marko's extension](https://github.com/frostming/marko/blob/master/marko/ext/footnote.py).
|
|
214
|
+
|
|
215
|
+
- GFM-style tables are supported and also auto-formatted.
|
|
216
|
+
|
|
217
|
+
- GFM-style footnotes are supported.
|
|
218
|
+
But note these aren't actually in the GFM spec, but we follow
|
|
219
|
+
[micromark's conventions](https://github.com/frostming/marko/blob/master/marko/ext/footnote.py).
|
|
220
|
+
|
|
206
221
|
## Project Docs
|
|
207
222
|
|
|
208
223
|
For how to install uv and Python, see [installation.md](installation.md).
|
|
@@ -129,8 +129,7 @@ experience.
|
|
|
129
129
|
Flowmark can be used as a library or as a CLI.
|
|
130
130
|
|
|
131
131
|
```
|
|
132
|
-
|
|
133
|
-
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [file]
|
|
132
|
+
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [--auto] [--version] [file]
|
|
134
133
|
|
|
135
134
|
Flowmark: Better line wrapping and formatting for plaintext and Markdown
|
|
136
135
|
|
|
@@ -142,9 +141,11 @@ options:
|
|
|
142
141
|
-o, --output OUTPUT Output file (use '-' for stdout)
|
|
143
142
|
-w, --width WIDTH Line width to wrap to
|
|
144
143
|
-p, --plaintext Process as plaintext (no Markdown parsing)
|
|
145
|
-
-s, --semantic Enable sentence-based line breaks (only applies to Markdown mode)
|
|
144
|
+
-s, --semantic Enable semantic (sentence-based) line breaks (only applies to Markdown mode)
|
|
146
145
|
-i, --inplace Edit the file in place (ignores --output)
|
|
147
146
|
--nobackup Do not make a backup of the original file when using --inplace
|
|
147
|
+
--auto Same as `--inplace --nobackup --semantic`, as a convenience for auto-formatting files
|
|
148
|
+
--version Show version information and exit
|
|
148
149
|
|
|
149
150
|
Flowmark provides enhanced text wrapping capabilities with special handling for
|
|
150
151
|
Markdown content. It can:
|
|
@@ -173,12 +174,26 @@ Command-line usage examples:
|
|
|
173
174
|
# Process plaintext instead of Markdown
|
|
174
175
|
flowmark --plaintext text.txt
|
|
175
176
|
|
|
176
|
-
# Use
|
|
177
|
+
# Use semantic line breaks (based on sentences, which is helpful to reduce
|
|
178
|
+
# irrelevant line wrap diffs in git history)
|
|
177
179
|
flowmark --semantic README.md
|
|
178
180
|
|
|
179
181
|
For more details, see: https://github.com/jlevy/flowmark
|
|
180
182
|
```
|
|
181
183
|
|
|
184
|
+
## Other Notes
|
|
185
|
+
|
|
186
|
+
- This enables
|
|
187
|
+
[GitHub-flavored Markdown support](https://docs.github.com/en/get-started/writing-on-github/getting-started-with-writing-and-formatting-on-github/basic-writing-and-formatting-syntax)
|
|
188
|
+
using
|
|
189
|
+
[Marko's extension](https://github.com/frostming/marko/blob/master/marko/ext/footnote.py).
|
|
190
|
+
|
|
191
|
+
- GFM-style tables are supported and also auto-formatted.
|
|
192
|
+
|
|
193
|
+
- GFM-style footnotes are supported.
|
|
194
|
+
But note these aren't actually in the GFM spec, but we follow
|
|
195
|
+
[micromark's conventions](https://github.com/frostming/marko/blob/master/marko/ext/footnote.py).
|
|
196
|
+
|
|
182
197
|
## Project Docs
|
|
183
198
|
|
|
184
199
|
For how to install uv and Python, see [installation.md](installation.md).
|
|
@@ -7,6 +7,7 @@ from typing import Any, cast
|
|
|
7
7
|
|
|
8
8
|
from marko import Markdown, Renderer, block, inline
|
|
9
9
|
from marko.block import HTMLBlock
|
|
10
|
+
from marko.ext import footnote
|
|
10
11
|
from marko.ext.gfm import GFM
|
|
11
12
|
from marko.ext.gfm import elements as gfm_elements
|
|
12
13
|
from marko.parser import Parser
|
|
@@ -187,10 +188,18 @@ class MarkdownNormalizer(Renderer):
|
|
|
187
188
|
return result
|
|
188
189
|
|
|
189
190
|
def render_link_ref_def(self, element: block.LinkRefDef) -> str:
|
|
191
|
+
"""Render a standard link reference definition:
|
|
192
|
+
[label]: url "title"
|
|
193
|
+
"""
|
|
190
194
|
link_text = element.dest
|
|
191
195
|
if element.title:
|
|
192
|
-
|
|
193
|
-
|
|
196
|
+
# Ensure title quotes are handled correctly
|
|
197
|
+
escaped_title = element.title.replace('"', '\\"')
|
|
198
|
+
link_text += f' "{escaped_title}"'
|
|
199
|
+
result = f"{self._prefix}[{element.label}]: {link_text}\n"
|
|
200
|
+
self._prefix = self._second_prefix
|
|
201
|
+
self._suppress_item_break = True
|
|
202
|
+
return result
|
|
194
203
|
|
|
195
204
|
def render_emphasis(self, element: inline.Emphasis) -> str:
|
|
196
205
|
return f"*{self.render_children(element)}*"
|
|
@@ -243,6 +252,29 @@ class MarkdownNormalizer(Renderer):
|
|
|
243
252
|
|
|
244
253
|
# --- GFM Renderer Methods ---
|
|
245
254
|
|
|
255
|
+
def render_footnote_ref(self, element: footnote.FootnoteRef) -> str:
|
|
256
|
+
"""Render an inline footnote reference like [^label]."""
|
|
257
|
+
return f"[^{element.label}]"
|
|
258
|
+
|
|
259
|
+
def render_footnote_def(self, element: footnote.FootnoteDef) -> str:
|
|
260
|
+
"""
|
|
261
|
+
Render a GFM footnote definition, handling content wrapping.
|
|
262
|
+
Note multiline footnotes aren't very well specified but we use
|
|
263
|
+
standard 4-space indentation. See:
|
|
264
|
+
https://github.com/micromark/micromark-extension-gfm-footnote
|
|
265
|
+
"""
|
|
266
|
+
# Render label and the rest within an indented container.
|
|
267
|
+
label_part = f"[^{element.label}]: "
|
|
268
|
+
with self.container(label_part, " "):
|
|
269
|
+
content = self.render_children(element)
|
|
270
|
+
|
|
271
|
+
# Set up state for the *next* block element using the restored outer secondary prefix.
|
|
272
|
+
self._prefix = self._second_prefix
|
|
273
|
+
self._suppress_item_break = True # This definition acts as a block separator.
|
|
274
|
+
|
|
275
|
+
# Footnote defs should be separated by extra newlines.
|
|
276
|
+
return content.rstrip("\n") + "\n\n"
|
|
277
|
+
|
|
246
278
|
def render_strikethrough(self, element: gfm_elements.Strikethrough) -> str:
|
|
247
279
|
return f"~~{self.render_children(element)}~~"
|
|
248
280
|
|
|
@@ -288,12 +320,19 @@ def custom_marko(line_wrapper: LineWrapper) -> Markdown:
|
|
|
288
320
|
# Using Marko's full extension system is tricky with our customizations so simpler
|
|
289
321
|
# to do this manually.
|
|
290
322
|
custom_parser = CustomParser()
|
|
323
|
+
# Add GFM support.
|
|
291
324
|
for e in GFM.elements:
|
|
292
325
|
assert (
|
|
293
326
|
e not in custom_parser.block_elements and e not in custom_parser.inline_elements
|
|
294
327
|
)
|
|
295
328
|
custom_parser.add_element(e)
|
|
296
|
-
|
|
329
|
+
# Add GFM footnote support.
|
|
330
|
+
footnote_ext = footnote.make_extension()
|
|
331
|
+
for e in footnote_ext.elements:
|
|
332
|
+
assert (
|
|
333
|
+
e not in custom_parser.block_elements and e not in custom_parser.inline_elements
|
|
334
|
+
)
|
|
335
|
+
custom_parser.add_element(e)
|
|
297
336
|
self.parser: Parser = custom_parser
|
|
298
337
|
self.renderer: Renderer = CustomRenderer()
|
|
299
338
|
|
|
@@ -30,6 +30,7 @@ def split_sentences_no_min_length(text: str) -> list[str]:
|
|
|
30
30
|
def line_wrap_to_width(
|
|
31
31
|
width: int = DEFAULT_WRAP_WIDTH,
|
|
32
32
|
len_fn: Callable[[str], int] = DEFAULT_LEN_FUNCTION,
|
|
33
|
+
is_markdown: bool = False,
|
|
33
34
|
) -> LineWrapper:
|
|
34
35
|
"""
|
|
35
36
|
Wrap lines of text to a given width.
|
|
@@ -42,6 +43,7 @@ def line_wrap_to_width(
|
|
|
42
43
|
initial_indent=initial_indent,
|
|
43
44
|
subsequent_indent=subsequent_indent,
|
|
44
45
|
len_fn=len_fn,
|
|
46
|
+
is_markdown=is_markdown,
|
|
45
47
|
)
|
|
46
48
|
|
|
47
49
|
return line_wrapper
|
|
@@ -52,6 +54,7 @@ def line_wrap_by_sentence(
|
|
|
52
54
|
width: int = DEFAULT_WRAP_WIDTH,
|
|
53
55
|
min_line_len: int = DEFAULT_MIN_LINE_LEN,
|
|
54
56
|
len_fn: Callable[[str], int] = DEFAULT_LEN_FUNCTION,
|
|
57
|
+
is_markdown: bool = False,
|
|
55
58
|
) -> LineWrapper:
|
|
56
59
|
"""
|
|
57
60
|
Wrap lines of text to a given width but also keep sentences on their own lines.
|
|
@@ -79,6 +82,7 @@ def line_wrap_by_sentence(
|
|
|
79
82
|
width=width,
|
|
80
83
|
initial_column=current_column,
|
|
81
84
|
subsequent_offset=subsequent_indent_len,
|
|
85
|
+
is_markdown=is_markdown,
|
|
82
86
|
)
|
|
83
87
|
# If last line is shorter than min_line_len, combine with next line.
|
|
84
88
|
# Also handles if the first word doesn't fit.
|
|
@@ -99,9 +99,10 @@ def fill_markdown(
|
|
|
99
99
|
beginning of the document.
|
|
100
100
|
"""
|
|
101
101
|
if line_wrapper is None:
|
|
102
|
-
|
|
103
|
-
line_wrap_by_sentence(width=width
|
|
104
|
-
|
|
102
|
+
if semantic:
|
|
103
|
+
line_wrapper = line_wrap_by_sentence(width=width, is_markdown=True)
|
|
104
|
+
else:
|
|
105
|
+
line_wrapper = line_wrap_to_width(width=width, is_markdown=True)
|
|
105
106
|
|
|
106
107
|
# Extract frontmatter before any processing
|
|
107
108
|
frontmatter, content = split_frontmatter(markdown_text)
|
|
@@ -71,6 +71,29 @@ Split words, but not within HTML tags or Markdown links.
|
|
|
71
71
|
"""
|
|
72
72
|
|
|
73
73
|
|
|
74
|
+
# Pattern to identify words that need escaping if they start a wrapped markdown line.
|
|
75
|
+
# Matches list markers (*, +, -) bare or before a space (but not before a letter for
|
|
76
|
+
# example), blockquotes (> ), headings (#, ##, etc.).
|
|
77
|
+
_md_specials_pat = re.compile(r"^[-*+](?: |$)|> |#+$")
|
|
78
|
+
|
|
79
|
+
# Separate pattern to specifically find the numbered list cases for targeted escaping
|
|
80
|
+
_md_numeral_pat = re.compile(r"^[0-9]+[.)]$")
|
|
81
|
+
|
|
82
|
+
|
|
83
|
+
def markdown_escape_word(word: str) -> str:
|
|
84
|
+
"""
|
|
85
|
+
Prepends a backslash to a word if it matches markdown patterns
|
|
86
|
+
that need escaping at the start of a wrapped line.
|
|
87
|
+
For numbered lists (e.g., "1.", "1)"), inserts the backslash before the dot/paren.
|
|
88
|
+
"""
|
|
89
|
+
if _md_numeral_pat.match(word):
|
|
90
|
+
# Insert backslash before the `.` or `)`
|
|
91
|
+
return word[:-1] + "\\" + word[-1]
|
|
92
|
+
elif _md_specials_pat.match(word):
|
|
93
|
+
return "\\" + word
|
|
94
|
+
return word
|
|
95
|
+
|
|
96
|
+
|
|
74
97
|
def wrap_paragraph_lines(
|
|
75
98
|
text: str,
|
|
76
99
|
width: int,
|
|
@@ -80,10 +103,13 @@ def wrap_paragraph_lines(
|
|
|
80
103
|
drop_whitespace: bool = True,
|
|
81
104
|
splitter: WordSplitter = html_md_word_splitter,
|
|
82
105
|
len_fn: Callable[[str], int] = DEFAULT_LEN_FUNCTION,
|
|
106
|
+
is_markdown: bool = False,
|
|
83
107
|
) -> list[str]:
|
|
84
108
|
"""
|
|
85
109
|
Wrap a single paragraph of text, returning a list of wrapped lines.
|
|
86
110
|
Rewritten to simplify and generalize Python's textwrap.py.
|
|
111
|
+
Set `is_markdown` to True when wrapping markdown text to automatically
|
|
112
|
+
escape special markdown characters at the start of wrapped lines.
|
|
87
113
|
"""
|
|
88
114
|
if replace_whitespace:
|
|
89
115
|
text = re.sub(r"\s+", " ", text)
|
|
@@ -93,6 +119,7 @@ def wrap_paragraph_lines(
|
|
|
93
119
|
lines: list[str] = []
|
|
94
120
|
current_line: list[str] = []
|
|
95
121
|
current_width = initial_column
|
|
122
|
+
first_line = True
|
|
96
123
|
|
|
97
124
|
# Walk through words, breaking them into lines.
|
|
98
125
|
for word in words:
|
|
@@ -110,9 +137,21 @@ def wrap_paragraph_lines(
|
|
|
110
137
|
if drop_whitespace:
|
|
111
138
|
line = line.strip()
|
|
112
139
|
lines.append(line)
|
|
113
|
-
|
|
114
|
-
|
|
140
|
+
first_line = False
|
|
141
|
+
|
|
142
|
+
# Check if word needs escaping at the start of this wrapped line.
|
|
143
|
+
escaped_word = word
|
|
144
|
+
if is_markdown and not first_line:
|
|
145
|
+
escaped_word = markdown_escape_word(word)
|
|
146
|
+
|
|
147
|
+
# Recalculate width after potential escaping for the new line.
|
|
148
|
+
escaped_word_width = len_fn(escaped_word)
|
|
149
|
+
|
|
150
|
+
# Start the new line with the (potentially escaped) word
|
|
151
|
+
current_line = [escaped_word]
|
|
152
|
+
current_width = subsequent_offset + escaped_word_width
|
|
115
153
|
|
|
154
|
+
# Add the last line if necessary.
|
|
116
155
|
if current_line:
|
|
117
156
|
line = " ".join(current_line)
|
|
118
157
|
if drop_whitespace:
|
|
@@ -132,6 +171,7 @@ def wrap_paragraph(
|
|
|
132
171
|
drop_whitespace: bool = True,
|
|
133
172
|
word_splitter: WordSplitter = html_md_word_splitter,
|
|
134
173
|
len_fn: Callable[[str], int] = DEFAULT_LEN_FUNCTION,
|
|
174
|
+
is_markdown: bool = False,
|
|
135
175
|
) -> str:
|
|
136
176
|
"""
|
|
137
177
|
Wrap lines of a single paragraph of plain text, returning a new string.
|
|
@@ -146,6 +186,7 @@ def wrap_paragraph(
|
|
|
146
186
|
initial_column=initial_column + len_fn(initial_indent),
|
|
147
187
|
subsequent_offset=len_fn(subsequent_indent),
|
|
148
188
|
len_fn=len_fn,
|
|
189
|
+
is_markdown=is_markdown,
|
|
149
190
|
)
|
|
150
191
|
# Now insert indents on first and subsequent lines, if needed.
|
|
151
192
|
if initial_indent and initial_column == 0 and len(lines) > 0:
|
|
@@ -3,12 +3,105 @@ from textwrap import dedent
|
|
|
3
3
|
from flowmark.text_wrapping import (
|
|
4
4
|
_HtmlMdWordSplitter, # pyright: ignore
|
|
5
5
|
html_md_word_splitter,
|
|
6
|
+
markdown_escape_word,
|
|
6
7
|
simple_word_splitter,
|
|
7
8
|
wrap_paragraph,
|
|
8
9
|
wrap_paragraph_lines,
|
|
9
10
|
)
|
|
10
11
|
|
|
11
12
|
|
|
13
|
+
def test_markdown_escape_word_function() -> None:
|
|
14
|
+
# Cases that should be escaped
|
|
15
|
+
assert markdown_escape_word("-") == "\\-"
|
|
16
|
+
assert markdown_escape_word("+") == "\\+"
|
|
17
|
+
assert markdown_escape_word("*") == "\\*"
|
|
18
|
+
assert markdown_escape_word(">") == ">"
|
|
19
|
+
assert markdown_escape_word("#") == "\\#"
|
|
20
|
+
assert markdown_escape_word("##") == "\\##"
|
|
21
|
+
assert markdown_escape_word("1.") == "1\\."
|
|
22
|
+
assert markdown_escape_word("10.") == "10\\."
|
|
23
|
+
assert markdown_escape_word("1)") == "1\\)"
|
|
24
|
+
assert markdown_escape_word("99)") == "99\\)"
|
|
25
|
+
|
|
26
|
+
# Cases that should NOT be escaped
|
|
27
|
+
assert markdown_escape_word("word") == "word"
|
|
28
|
+
assert markdown_escape_word("-word") == "-word" # Starts with char, but not just char
|
|
29
|
+
assert markdown_escape_word("word-") == "word-" # Ends with char
|
|
30
|
+
assert markdown_escape_word("#word") == "#word"
|
|
31
|
+
assert markdown_escape_word("word#") == "word#"
|
|
32
|
+
assert markdown_escape_word("1.word") == "1.word"
|
|
33
|
+
assert markdown_escape_word("word1.") == "word1."
|
|
34
|
+
assert markdown_escape_word("1)word") == "1)word"
|
|
35
|
+
assert markdown_escape_word("word1)") == "word1)"
|
|
36
|
+
assert markdown_escape_word("<tag>") == "<tag>" # Other symbols
|
|
37
|
+
assert markdown_escape_word("[link]") == "[link]"
|
|
38
|
+
assert markdown_escape_word("1") == "1" # Just number
|
|
39
|
+
assert markdown_escape_word(".") == "." # Just dot
|
|
40
|
+
|
|
41
|
+
|
|
42
|
+
def test_wrap_paragraph_lines_markdown_escaping():
|
|
43
|
+
assert wrap_paragraph_lines(text="- word", width=10, is_markdown=True) == ["- word"]
|
|
44
|
+
|
|
45
|
+
text = "word - word * word + word > word # word ## word 1. word 2) word"
|
|
46
|
+
|
|
47
|
+
assert wrap_paragraph_lines(text=text, width=5, is_markdown=True) == [
|
|
48
|
+
"word",
|
|
49
|
+
"\\-",
|
|
50
|
+
"word",
|
|
51
|
+
"\\*",
|
|
52
|
+
"word",
|
|
53
|
+
"\\+",
|
|
54
|
+
"word",
|
|
55
|
+
">",
|
|
56
|
+
"word",
|
|
57
|
+
"\\#",
|
|
58
|
+
"word",
|
|
59
|
+
"\\##",
|
|
60
|
+
"word",
|
|
61
|
+
"1\\.",
|
|
62
|
+
"word",
|
|
63
|
+
"2\\)",
|
|
64
|
+
"word",
|
|
65
|
+
]
|
|
66
|
+
assert wrap_paragraph_lines(text=text, width=10, is_markdown=True) == [
|
|
67
|
+
"word -",
|
|
68
|
+
"word *",
|
|
69
|
+
"word +",
|
|
70
|
+
"word >",
|
|
71
|
+
"word #",
|
|
72
|
+
"word ##",
|
|
73
|
+
"word 1.",
|
|
74
|
+
"word 2)",
|
|
75
|
+
"word",
|
|
76
|
+
]
|
|
77
|
+
assert wrap_paragraph_lines(text=text, width=15, is_markdown=True) == [
|
|
78
|
+
"word - word *",
|
|
79
|
+
"word + word >",
|
|
80
|
+
"word # word ##",
|
|
81
|
+
"word 1. word 2)",
|
|
82
|
+
"word",
|
|
83
|
+
]
|
|
84
|
+
assert wrap_paragraph_lines(text=text, width=20, is_markdown=True) == [
|
|
85
|
+
"word - word * word +",
|
|
86
|
+
"word > word # word",
|
|
87
|
+
"\\## word 1. word 2)",
|
|
88
|
+
"word",
|
|
89
|
+
]
|
|
90
|
+
assert wrap_paragraph_lines(text=text, width=20, is_markdown=False) == [
|
|
91
|
+
"word - word * word +",
|
|
92
|
+
"word > word # word",
|
|
93
|
+
"## word 1. word 2)",
|
|
94
|
+
"word",
|
|
95
|
+
]
|
|
96
|
+
|
|
97
|
+
test2 = """Testing - : Is Ketamine Contraindicated in Patients with Psychiatric Disorders? - REBEL EM - more words - accessed April 24, 2025, <https://rebelem.com/is-ketamine-contraindicated-in-patients-with-psychiatric-disorders/>"""
|
|
98
|
+
assert wrap_paragraph_lines(text=test2, width=80, is_markdown=True) == [
|
|
99
|
+
"Testing - : Is Ketamine Contraindicated in Patients with Psychiatric Disorders?",
|
|
100
|
+
"\\- REBEL EM - more words - accessed April 24, 2025,",
|
|
101
|
+
"<https://rebelem.com/is-ketamine-contraindicated-in-patients-with-psychiatric-disorders/>",
|
|
102
|
+
]
|
|
103
|
+
|
|
104
|
+
|
|
12
105
|
def test_smart_splitter():
|
|
13
106
|
splitter = _HtmlMdWordSplitter()
|
|
14
107
|
|
|
@@ -569,6 +569,9 @@ Links like these underline ones come up from some exports. And let's try some li
|
|
|
569
569
|
|
|
570
570
|
[^53]: <https://www.fastcompany.com/90216464/the-29-billion-battle-to-own-how-america-sleeps>
|
|
571
571
|
|
|
572
|
+
[^217]: Testing - : Is Ketamine Contraindicated in Patients with Psychiatric Disorders? - REBEL EM - more words - accessed April 24, 2025,
|
|
573
|
+
<https://rebelem.com/is-ketamine-contraindicated-in-patients-with-psychiatric-disorders/>
|
|
574
|
+
|
|
572
575
|
[^multiline]: The distinction between “hiring” and “recruiting” isn’t
|
|
573
576
|
universally agreed upon. Some people think of hiring as a superset
|
|
574
577
|
of recruiting, some consider it to be the other way around. However
|
|
@@ -239,7 +239,7 @@ to $24,000, which means the tax benefit of buying is
|
|
|
239
239
|
* When an investor is trying to get you to agree to a term you think is unfair, you need
|
|
240
240
|
to protect your interests without sounding accusatory toward the investor: *“Sorry,
|
|
241
241
|
I’m just inexperienced, I read/was told that it’s not wise to [term they want you to
|
|
242
|
-
agree to].”* [[Paul Graham, Y Combinator]
|
|
242
|
+
agree to].”* [[Paul Graham, Y Combinator](http://paulgraham.com/fr.html)]
|
|
243
243
|
|
|
244
244
|
* When you want to test a VC’s interest to determine where to put your energy, you don’t
|
|
245
245
|
want to sound desperate or pushy: *“I know that you’re not likely to give me a strong
|
|
@@ -867,14 +867,13 @@ not complaining)[^urbanthowt.wy49lp].
|
|
|
867
867
|
❗️️️ Having multiple automatic conversion thresholds can give the investor with a higher
|
|
868
868
|
threshold leverage to block an IPO.[^210]
|
|
869
869
|
|
|
870
|
-
[^2]: Aulet, Bill. *Disciplined Entrepreneurship*: 24 Steps to a Successful Startup
|
|
871
|
-
|
|
870
|
+
[^2]: Aulet, Bill. *Disciplined Entrepreneurship*: 24 Steps to a Successful Startup (Kindle
|
|
871
|
+
Location 1220). Wiley, 2013. Kindle Edition.
|
|
872
872
|
|
|
873
873
|
[^191]: http://paulgraham.com/fr.html
|
|
874
874
|
|
|
875
|
-
[^177]: Carnegie, Dale.
|
|
876
|
-
|
|
877
|
-
Kindle Edition.
|
|
875
|
+
[^177]: Carnegie, Dale. *How To Win Friends and Influence People* (p. 35). Simon & Schuster.
|
|
876
|
+
Kindle Edition.
|
|
878
877
|
|
|
879
878
|
[^207]: [https://bostonvcblog.typepad.com/vc/2009/07/in-vc-deals-price-doesnt-matter-but-the-promote-does.html](https://bostonvcblog.typepad.com/vc/2009/07/in-vc-deals-price-doesnt-matter-but-the-promote-does.html)
|
|
880
879
|
|
|
@@ -889,29 +888,29 @@ And let's try some links with angle brackets.
|
|
|
889
888
|
|
|
890
889
|
[^axioscomth.1lioru]: <https://www.axios.com/the-rise-of-pre-seed-venture-capital-1513305959-13da61c8-15f8-441e-b016-d29902bff8bf.html>
|
|
891
890
|
|
|
892
|
-
[^carnegieda.327r3k]: Carnegie, Dale.
|
|
893
|
-
|
|
894
|
-
Kindle Edition.
|
|
891
|
+
[^carnegieda.327r3k]: Carnegie, Dale. *How To Win Friends and Influence People* (p. 35). Simon & Schuster.
|
|
892
|
+
Kindle Edition.
|
|
895
893
|
|
|
896
894
|
[^53]: <https://www.fastcompany.com/90216464/the-29-billion-battle-to-own-how-america-sleeps>
|
|
897
895
|
|
|
896
|
+
[^217]: Testing - : Is Ketamine Contraindicated in Patients with Psychiatric Disorders?
|
|
897
|
+
- REBEL EM - more words - accessed April 24, 2025,
|
|
898
|
+
<https://rebelem.com/is-ketamine-contraindicated-in-patients-with-psychiatric-disorders/>
|
|
899
|
+
|
|
898
900
|
[^multiline]: The distinction between “hiring” and “recruiting” isn’t universally agreed
|
|
899
|
-
upon.
|
|
900
|
-
|
|
901
|
-
|
|
902
|
-
|
|
903
|
-
|
|
901
|
+
upon. Some people think of hiring as a superset of recruiting, some consider it to be
|
|
902
|
+
the other way around.
|
|
903
|
+
However you think of it, both recruiting and hiring involve selling candidates on
|
|
904
|
+
the value proposition of a company and ensuring the alignment of interests between
|
|
905
|
+
the two parties.
|
|
904
906
|
|
|
905
907
|
[^multiparagraph]: This is an even longer footnote...
|
|
906
908
|
|
|
907
|
-
|
|
908
|
-
Paragraph 1.
|
|
909
|
+
Paragraph 1.
|
|
909
910
|
|
|
910
|
-
> And even a
|
|
911
|
-
> block quote.
|
|
911
|
+
> And even a block quote.
|
|
912
912
|
|
|
913
|
-
Paragraph 3.
|
|
914
|
-
```
|
|
913
|
+
Paragraph 3.
|
|
915
914
|
|
|
916
915
|
And by contrast here a bare link is like this https://www.google.com/
|
|
917
916
|
|
|
@@ -225,7 +225,7 @@ to $24,000, which means the tax benefit of buying is
|
|
|
225
225
|
* When an investor is trying to get you to agree to a term you think is unfair, you need
|
|
226
226
|
to protect your interests without sounding accusatory toward the investor: *“Sorry,
|
|
227
227
|
I’m just inexperienced, I read/was told that it’s not wise to [term they want you to
|
|
228
|
-
agree to].”* [[Paul Graham, Y Combinator]
|
|
228
|
+
agree to].”* [[Paul Graham, Y Combinator](http://paulgraham.com/fr.html)]
|
|
229
229
|
|
|
230
230
|
* When you want to test a VC’s interest to determine where to put your energy, you don’t
|
|
231
231
|
want to sound desperate or pushy: *“I know that you’re not likely to give me a strong
|
|
@@ -830,12 +830,12 @@ not complaining)[^urbanthowt.wy49lp].
|
|
|
830
830
|
threshold leverage to block an IPO.[^210]
|
|
831
831
|
|
|
832
832
|
[^2]: Aulet, Bill. *Disciplined Entrepreneurship*: 24 Steps to a Successful Startup
|
|
833
|
-
(Kindle Location 1220). Wiley, 2013. Kindle Edition.
|
|
833
|
+
(Kindle Location 1220). Wiley, 2013. Kindle Edition.
|
|
834
834
|
|
|
835
835
|
[^191]: http://paulgraham.com/fr.html
|
|
836
836
|
|
|
837
837
|
[^177]: Carnegie, Dale. *How To Win Friends and Influence People* (p. 35). Simon &
|
|
838
|
-
Schuster. Kindle Edition.
|
|
838
|
+
Schuster. Kindle Edition.
|
|
839
839
|
|
|
840
840
|
[^207]: [https://bostonvcblog.typepad.com/vc/2009/07/in-vc-deals-price-doesnt-matter-but-the-promote-does.html](https://bostonvcblog.typepad.com/vc/2009/07/in-vc-deals-price-doesnt-matter-but-the-promote-does.html)
|
|
841
841
|
|
|
@@ -851,26 +851,27 @@ angle brackets.
|
|
|
851
851
|
[^axioscomth.1lioru]: <https://www.axios.com/the-rise-of-pre-seed-venture-capital-1513305959-13da61c8-15f8-441e-b016-d29902bff8bf.html>
|
|
852
852
|
|
|
853
853
|
[^carnegieda.327r3k]: Carnegie, Dale. *How To Win Friends and Influence People* (p. 35).
|
|
854
|
-
Simon & Schuster. Kindle Edition.
|
|
854
|
+
Simon & Schuster. Kindle Edition.
|
|
855
855
|
|
|
856
856
|
[^53]: <https://www.fastcompany.com/90216464/the-29-billion-battle-to-own-how-america-sleeps>
|
|
857
857
|
|
|
858
|
+
[^217]: Testing - : Is Ketamine Contraindicated in Patients with Psychiatric Disorders?
|
|
859
|
+
\- REBEL EM - more words - accessed April 24, 2025,
|
|
860
|
+
<https://rebelem.com/is-ketamine-contraindicated-in-patients-with-psychiatric-disorders/>
|
|
861
|
+
|
|
858
862
|
[^multiline]: The distinction between “hiring” and “recruiting” isn’t universally agreed
|
|
859
|
-
upon. Some people think of hiring as a superset of recruiting, some consider it to
|
|
860
|
-
the other way around. However you think of it, both recruiting and hiring involve
|
|
861
|
-
selling candidates on the value proposition of a company and ensuring the alignment
|
|
862
|
-
interests between the two parties.
|
|
863
|
+
upon. Some people think of hiring as a superset of recruiting, some consider it to
|
|
864
|
+
be the other way around. However you think of it, both recruiting and hiring involve
|
|
865
|
+
selling candidates on the value proposition of a company and ensuring the alignment
|
|
866
|
+
of interests between the two parties.
|
|
863
867
|
|
|
864
868
|
[^multiparagraph]: This is an even longer footnote...
|
|
865
869
|
|
|
866
|
-
|
|
867
|
-
Paragraph 1.
|
|
870
|
+
Paragraph 1.
|
|
868
871
|
|
|
869
|
-
> And even a
|
|
870
|
-
> block quote.
|
|
872
|
+
> And even a block quote.
|
|
871
873
|
|
|
872
|
-
Paragraph 3.
|
|
873
|
-
```
|
|
874
|
+
Paragraph 3.
|
|
874
875
|
|
|
875
876
|
And by contrast here a bare link is like this https://www.google.com/
|
|
876
877
|
|
|
@@ -239,7 +239,7 @@ to $24,000, which means the tax benefit of buying is
|
|
|
239
239
|
* When an investor is trying to get you to agree to a term you think is unfair, you need
|
|
240
240
|
to protect your interests without sounding accusatory toward the investor: *“Sorry,
|
|
241
241
|
I’m just inexperienced, I read/was told that it’s not wise to [term they want you to
|
|
242
|
-
agree to].”* [[Paul Graham, Y Combinator]
|
|
242
|
+
agree to].”* [[Paul Graham, Y Combinator](http://paulgraham.com/fr.html)]
|
|
243
243
|
|
|
244
244
|
* When you want to test a VC’s interest to determine where to put your energy, you don’t
|
|
245
245
|
want to sound desperate or pushy: *“I know that you’re not likely to give me a strong
|
|
@@ -867,14 +867,13 @@ not complaining)[^urbanthowt.wy49lp].
|
|
|
867
867
|
❗️️️ Having multiple automatic conversion thresholds can give the investor with a higher
|
|
868
868
|
threshold leverage to block an IPO.[^210]
|
|
869
869
|
|
|
870
|
-
[^2]: Aulet, Bill. *Disciplined Entrepreneurship*: 24 Steps to a Successful Startup
|
|
871
|
-
|
|
870
|
+
[^2]: Aulet, Bill. *Disciplined Entrepreneurship*: 24 Steps to a Successful Startup (Kindle
|
|
871
|
+
Location 1220). Wiley, 2013. Kindle Edition.
|
|
872
872
|
|
|
873
873
|
[^191]: http://paulgraham.com/fr.html
|
|
874
874
|
|
|
875
|
-
[^177]: Carnegie, Dale.
|
|
876
|
-
|
|
877
|
-
Kindle Edition.
|
|
875
|
+
[^177]: Carnegie, Dale. *How To Win Friends and Influence People* (p. 35). Simon & Schuster.
|
|
876
|
+
Kindle Edition.
|
|
878
877
|
|
|
879
878
|
[^207]: [https://bostonvcblog.typepad.com/vc/2009/07/in-vc-deals-price-doesnt-matter-but-the-promote-does.html](https://bostonvcblog.typepad.com/vc/2009/07/in-vc-deals-price-doesnt-matter-but-the-promote-does.html)
|
|
880
879
|
|
|
@@ -889,29 +888,29 @@ And let's try some links with angle brackets.
|
|
|
889
888
|
|
|
890
889
|
[^axioscomth.1lioru]: <https://www.axios.com/the-rise-of-pre-seed-venture-capital-1513305959-13da61c8-15f8-441e-b016-d29902bff8bf.html>
|
|
891
890
|
|
|
892
|
-
[^carnegieda.327r3k]: Carnegie, Dale.
|
|
893
|
-
|
|
894
|
-
Kindle Edition.
|
|
891
|
+
[^carnegieda.327r3k]: Carnegie, Dale. *How To Win Friends and Influence People* (p. 35). Simon & Schuster.
|
|
892
|
+
Kindle Edition.
|
|
895
893
|
|
|
896
894
|
[^53]: <https://www.fastcompany.com/90216464/the-29-billion-battle-to-own-how-america-sleeps>
|
|
897
895
|
|
|
896
|
+
[^217]: Testing - : Is Ketamine Contraindicated in Patients with Psychiatric Disorders?
|
|
897
|
+
- REBEL EM - more words - accessed April 24, 2025,
|
|
898
|
+
<https://rebelem.com/is-ketamine-contraindicated-in-patients-with-psychiatric-disorders/>
|
|
899
|
+
|
|
898
900
|
[^multiline]: The distinction between “hiring” and “recruiting” isn’t universally agreed
|
|
899
|
-
upon.
|
|
900
|
-
|
|
901
|
-
|
|
902
|
-
|
|
903
|
-
|
|
901
|
+
upon. Some people think of hiring as a superset of recruiting, some consider it to be
|
|
902
|
+
the other way around.
|
|
903
|
+
However you think of it, both recruiting and hiring involve selling candidates on
|
|
904
|
+
the value proposition of a company and ensuring the alignment of interests between
|
|
905
|
+
the two parties.
|
|
904
906
|
|
|
905
907
|
[^multiparagraph]: This is an even longer footnote...
|
|
906
908
|
|
|
907
|
-
|
|
908
|
-
Paragraph 1.
|
|
909
|
+
Paragraph 1.
|
|
909
910
|
|
|
910
|
-
> And even a
|
|
911
|
-
> block quote.
|
|
911
|
+
> And even a block quote.
|
|
912
912
|
|
|
913
|
-
Paragraph 3.
|
|
914
|
-
```
|
|
913
|
+
Paragraph 3.
|
|
915
914
|
|
|
916
915
|
And by contrast here a bare link is like this https://www.google.com/
|
|
917
916
|
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|