flowmark 0.2.0__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- flowmark-0.3.0/PKG-INFO +178 -0
- flowmark-0.3.0/README.md +158 -0
- {flowmark-0.2.0 → flowmark-0.3.0}/pyproject.toml +2 -2
- {flowmark-0.2.0 → flowmark-0.3.0}/src/flowmark/cli.py +26 -15
- flowmark-0.3.0/src/flowmark/frontmatter.py +44 -0
- {flowmark-0.2.0 → flowmark-0.3.0}/src/flowmark/markdown_filling.py +47 -10
- flowmark-0.2.0/PKG-INFO +0 -124
- flowmark-0.2.0/README.md +0 -104
- {flowmark-0.2.0 → flowmark-0.3.0}/LICENSE +0 -0
- {flowmark-0.2.0 → flowmark-0.3.0}/src/flowmark/__init__.py +0 -0
- {flowmark-0.2.0 → flowmark-0.3.0}/src/flowmark/line_wrappers.py +0 -0
- {flowmark-0.2.0 → flowmark-0.3.0}/src/flowmark/sentence_split_regex.py +0 -0
- {flowmark-0.2.0 → flowmark-0.3.0}/src/flowmark/text_filling.py +0 -0
- {flowmark-0.2.0 → flowmark-0.3.0}/src/flowmark/text_wrapping.py +0 -0
flowmark-0.3.0/PKG-INFO
ADDED
|
@@ -0,0 +1,178 @@
|
|
|
1
|
+
Metadata-Version: 2.3
|
|
2
|
+
Name: flowmark
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: Better line wrapping and formatting for plaintext and Markdown
|
|
5
|
+
License: MIT
|
|
6
|
+
Author: Joshua Levy
|
|
7
|
+
Author-email: joshua@cal.berkeley.edu
|
|
8
|
+
Requires-Python: >=3.10,<4.0
|
|
9
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
10
|
+
Classifier: Programming Language :: Python :: 3
|
|
11
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
15
|
+
Requires-Dist: marko (>=2.1.2,<3.0.0)
|
|
16
|
+
Requires-Dist: regex (>=2024.11.6,<2025.0.0)
|
|
17
|
+
Requires-Dist: strif (>=2.0.0,<3.0.0)
|
|
18
|
+
Description-Content-Type: text/markdown
|
|
19
|
+
|
|
20
|
+
# flowmark
|
|
21
|
+
|
|
22
|
+
Flowmark is a new Python implementation of **text and Markdown line wrapping and
|
|
23
|
+
filling**.
|
|
24
|
+
|
|
25
|
+
In addition, it adds optional **support for Markdown** and offers **Markdown
|
|
26
|
+
auto-formatting and normalization**. This is much like
|
|
27
|
+
[markdownfmt](https://github.com/shurcooL/markdownfmt) or
|
|
28
|
+
[prettier's Markdown support](https://prettier.io/blog/2017/11/07/1.8.0) but is pure
|
|
29
|
+
Python and has (in my humble opinion) better options and defaults.
|
|
30
|
+
|
|
31
|
+
Use cases:
|
|
32
|
+
|
|
33
|
+
- As a **command line formatter** to format text or Markdown files using the `flowmark`
|
|
34
|
+
command.
|
|
35
|
+
|
|
36
|
+
- To **autoformat Markdown on save in VSCode/Cursor** or any other editor that supports
|
|
37
|
+
running a command on save.
|
|
38
|
+
Flowmark uses a readable format that makes diffs easy to read and use on GitHub.
|
|
39
|
+
It also normalizes all Markdown syntax variations (such as different header or
|
|
40
|
+
formatting styles). This can be especially useful for documentation and editing
|
|
41
|
+
workflows where clean diffs and minimal merge conflicts on GitHub are important.
|
|
42
|
+
|
|
43
|
+
- As a **library to autoformat Markdown**. For example, it is great to normalize the
|
|
44
|
+
outputs from LLMs to be consistent, or to run on the inputs and outputs of LLM
|
|
45
|
+
transformations that edit text, so that the resulting diffs are clean.
|
|
46
|
+
Having this as a simple Python library makes this easy in AI-related document
|
|
47
|
+
pipelines.
|
|
48
|
+
|
|
49
|
+
- As a **drop-in replacement library for Python's default
|
|
50
|
+
[`textwrap`](https://docs.python.org/3/library/textwrap.html)** but with more options.
|
|
51
|
+
It simplifies and generalizes that library, offering better control over **initial and
|
|
52
|
+
subsequent indentation** and **when to split words and lines**, e.g. using a word
|
|
53
|
+
splitter that won't break lines within HTML tags.
|
|
54
|
+
|
|
55
|
+
- Flowmark has the option to to use **semantic line breaks** (using a heuristic to break
|
|
56
|
+
lines on sentences sentences when that is reasonable), which is an underrated feature
|
|
57
|
+
that can **make diffs on GitHub much more readable**. The the change may seem subtle
|
|
58
|
+
but avoids having paragraphs reflow for very small edits, which does a lot to
|
|
59
|
+
**minimize merge conflicts**. An example of what sentence-guided wrapping looks like,
|
|
60
|
+
see the
|
|
61
|
+
[Markdown source](https://github.com/jlevy/flowmark/blob/main/README.md?plain=1) of
|
|
62
|
+
this readme file.)
|
|
63
|
+
|
|
64
|
+
- Very, very simple and fast **regex-based sentence splitting**. This should work fine
|
|
65
|
+
for English and many other latin/Cyrillic languages but it hasn't been tested on CJK.
|
|
66
|
+
|
|
67
|
+
It aims to be small and simple and have only a few dependencies, currently only
|
|
68
|
+
[`marko`](https://github.com/frostming/marko),
|
|
69
|
+
[`regex`](https://pypi.org/project/regex/), and
|
|
70
|
+
[`strif`](https://github.com/jlevy/strif).
|
|
71
|
+
|
|
72
|
+
Because **YAML frontmatter** is common on Markdown files, the Markdown autoformat
|
|
73
|
+
preserves all frontmatter (content between `---` delimiters at the front of a file).
|
|
74
|
+
|
|
75
|
+
This is a new and simple package.
|
|
76
|
+
Previously I'd implemented something very similar with
|
|
77
|
+
[for Atom](https://github.com/jlevy/atom-flowmark) but Atom is no more and there seems
|
|
78
|
+
to be greater need for this in Python now.
|
|
79
|
+
From many years experience, I found the Markdown formatting conventions enforced by the
|
|
80
|
+
Atom Flowmark plugin worked well for editing and publishing large or collaboratively
|
|
81
|
+
edited documents.
|
|
82
|
+
|
|
83
|
+
## Installation
|
|
84
|
+
|
|
85
|
+
The simplest way to use the tool is to use [pipx](https://github.com/pypa/pipx):
|
|
86
|
+
|
|
87
|
+
```shell
|
|
88
|
+
pipx install flowmark
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
Then
|
|
92
|
+
|
|
93
|
+
```
|
|
94
|
+
flowmark --help
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
To use as a library, use pip/poetry/uv to install
|
|
98
|
+
[`flowmark`](https://pypi.org/project/flowmark/).
|
|
99
|
+
|
|
100
|
+
## Use in VSCode/Cursor
|
|
101
|
+
|
|
102
|
+
You can use Flowmark to auto-format Markdown on save in VSCode or Cursor.
|
|
103
|
+
Install the "Run on Save" (`emeraldwalk.runonsave`) extension.
|
|
104
|
+
Then add to your `settings.json`:
|
|
105
|
+
|
|
106
|
+
```json
|
|
107
|
+
"emeraldwalk.runonsave": {
|
|
108
|
+
"commands": [
|
|
109
|
+
{
|
|
110
|
+
"match": "\\.md$",
|
|
111
|
+
"cmd": "flowmark --auto ${file}"
|
|
112
|
+
}
|
|
113
|
+
]
|
|
114
|
+
}
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
The `--auto` option is just the same as `--inplace --nobackup --semantic`.
|
|
118
|
+
|
|
119
|
+
## Usage
|
|
120
|
+
|
|
121
|
+
Flowmark can be used as a library or as a CLI.
|
|
122
|
+
|
|
123
|
+
```
|
|
124
|
+
$ flowmark --help
|
|
125
|
+
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [file]
|
|
126
|
+
|
|
127
|
+
Flowmark: Better line wrapping and formatting for plaintext and Markdown
|
|
128
|
+
|
|
129
|
+
positional arguments:
|
|
130
|
+
file Input file (use '-' for stdin)
|
|
131
|
+
|
|
132
|
+
options:
|
|
133
|
+
-h, --help show this help message and exit
|
|
134
|
+
-o, --output OUTPUT Output file (use '-' for stdout)
|
|
135
|
+
-w, --width WIDTH Line width to wrap to
|
|
136
|
+
-p, --plaintext Process as plaintext (no Markdown parsing)
|
|
137
|
+
-s, --semantic Enable sentence-based line breaks (only applies to Markdown mode)
|
|
138
|
+
-i, --inplace Edit the file in place (ignores --output)
|
|
139
|
+
--nobackup Do not make a backup of the original file when using --inplace
|
|
140
|
+
|
|
141
|
+
Flowmark provides enhanced text wrapping capabilities with special handling for
|
|
142
|
+
Markdown content. It can:
|
|
143
|
+
|
|
144
|
+
- Format Markdown with proper line wrapping while preserving structure
|
|
145
|
+
and normalizing Markdown formatting
|
|
146
|
+
|
|
147
|
+
- Optionally break lines at sentence boundaries for better diff readability
|
|
148
|
+
|
|
149
|
+
- Process plaintext with HTML-aware word splitting
|
|
150
|
+
|
|
151
|
+
It is both a library and a command-line tool.
|
|
152
|
+
|
|
153
|
+
Command-line usage examples:
|
|
154
|
+
|
|
155
|
+
# Format a Markdown file to stdout
|
|
156
|
+
flowmark README.md
|
|
157
|
+
|
|
158
|
+
# Format a Markdown file and save to a new file
|
|
159
|
+
flowmark README.md -o README_formatted.md
|
|
160
|
+
|
|
161
|
+
# Edit a file in-place (with or without making a backup)
|
|
162
|
+
flowmark --inplace README.md
|
|
163
|
+
flowmark --inplace --nobackup README.md
|
|
164
|
+
|
|
165
|
+
# Process plaintext instead of Markdown
|
|
166
|
+
flowmark --plaintext text.txt
|
|
167
|
+
|
|
168
|
+
# Use sentences to guide line breaks (good for many purposes git history and diffs)
|
|
169
|
+
flowmark --semantic README.md
|
|
170
|
+
|
|
171
|
+
For more details, see: https://github.com/jlevy/flowmark
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
* * *
|
|
175
|
+
|
|
176
|
+
*This project was built from
|
|
177
|
+
[simple-modern-poetry](https://github.com/jlevy/simple-modern-poetry).*
|
|
178
|
+
|
flowmark-0.3.0/README.md
ADDED
|
@@ -0,0 +1,158 @@
|
|
|
1
|
+
# flowmark
|
|
2
|
+
|
|
3
|
+
Flowmark is a new Python implementation of **text and Markdown line wrapping and
|
|
4
|
+
filling**.
|
|
5
|
+
|
|
6
|
+
In addition, it adds optional **support for Markdown** and offers **Markdown
|
|
7
|
+
auto-formatting and normalization**. This is much like
|
|
8
|
+
[markdownfmt](https://github.com/shurcooL/markdownfmt) or
|
|
9
|
+
[prettier's Markdown support](https://prettier.io/blog/2017/11/07/1.8.0) but is pure
|
|
10
|
+
Python and has (in my humble opinion) better options and defaults.
|
|
11
|
+
|
|
12
|
+
Use cases:
|
|
13
|
+
|
|
14
|
+
- As a **command line formatter** to format text or Markdown files using the `flowmark`
|
|
15
|
+
command.
|
|
16
|
+
|
|
17
|
+
- To **autoformat Markdown on save in VSCode/Cursor** or any other editor that supports
|
|
18
|
+
running a command on save.
|
|
19
|
+
Flowmark uses a readable format that makes diffs easy to read and use on GitHub.
|
|
20
|
+
It also normalizes all Markdown syntax variations (such as different header or
|
|
21
|
+
formatting styles). This can be especially useful for documentation and editing
|
|
22
|
+
workflows where clean diffs and minimal merge conflicts on GitHub are important.
|
|
23
|
+
|
|
24
|
+
- As a **library to autoformat Markdown**. For example, it is great to normalize the
|
|
25
|
+
outputs from LLMs to be consistent, or to run on the inputs and outputs of LLM
|
|
26
|
+
transformations that edit text, so that the resulting diffs are clean.
|
|
27
|
+
Having this as a simple Python library makes this easy in AI-related document
|
|
28
|
+
pipelines.
|
|
29
|
+
|
|
30
|
+
- As a **drop-in replacement library for Python's default
|
|
31
|
+
[`textwrap`](https://docs.python.org/3/library/textwrap.html)** but with more options.
|
|
32
|
+
It simplifies and generalizes that library, offering better control over **initial and
|
|
33
|
+
subsequent indentation** and **when to split words and lines**, e.g. using a word
|
|
34
|
+
splitter that won't break lines within HTML tags.
|
|
35
|
+
|
|
36
|
+
- Flowmark has the option to to use **semantic line breaks** (using a heuristic to break
|
|
37
|
+
lines on sentences sentences when that is reasonable), which is an underrated feature
|
|
38
|
+
that can **make diffs on GitHub much more readable**. The the change may seem subtle
|
|
39
|
+
but avoids having paragraphs reflow for very small edits, which does a lot to
|
|
40
|
+
**minimize merge conflicts**. An example of what sentence-guided wrapping looks like,
|
|
41
|
+
see the
|
|
42
|
+
[Markdown source](https://github.com/jlevy/flowmark/blob/main/README.md?plain=1) of
|
|
43
|
+
this readme file.)
|
|
44
|
+
|
|
45
|
+
- Very, very simple and fast **regex-based sentence splitting**. This should work fine
|
|
46
|
+
for English and many other latin/Cyrillic languages but it hasn't been tested on CJK.
|
|
47
|
+
|
|
48
|
+
It aims to be small and simple and have only a few dependencies, currently only
|
|
49
|
+
[`marko`](https://github.com/frostming/marko),
|
|
50
|
+
[`regex`](https://pypi.org/project/regex/), and
|
|
51
|
+
[`strif`](https://github.com/jlevy/strif).
|
|
52
|
+
|
|
53
|
+
Because **YAML frontmatter** is common on Markdown files, the Markdown autoformat
|
|
54
|
+
preserves all frontmatter (content between `---` delimiters at the front of a file).
|
|
55
|
+
|
|
56
|
+
This is a new and simple package.
|
|
57
|
+
Previously I'd implemented something very similar with
|
|
58
|
+
[for Atom](https://github.com/jlevy/atom-flowmark) but Atom is no more and there seems
|
|
59
|
+
to be greater need for this in Python now.
|
|
60
|
+
From many years experience, I found the Markdown formatting conventions enforced by the
|
|
61
|
+
Atom Flowmark plugin worked well for editing and publishing large or collaboratively
|
|
62
|
+
edited documents.
|
|
63
|
+
|
|
64
|
+
## Installation
|
|
65
|
+
|
|
66
|
+
The simplest way to use the tool is to use [pipx](https://github.com/pypa/pipx):
|
|
67
|
+
|
|
68
|
+
```shell
|
|
69
|
+
pipx install flowmark
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
Then
|
|
73
|
+
|
|
74
|
+
```
|
|
75
|
+
flowmark --help
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
To use as a library, use pip/poetry/uv to install
|
|
79
|
+
[`flowmark`](https://pypi.org/project/flowmark/).
|
|
80
|
+
|
|
81
|
+
## Use in VSCode/Cursor
|
|
82
|
+
|
|
83
|
+
You can use Flowmark to auto-format Markdown on save in VSCode or Cursor.
|
|
84
|
+
Install the "Run on Save" (`emeraldwalk.runonsave`) extension.
|
|
85
|
+
Then add to your `settings.json`:
|
|
86
|
+
|
|
87
|
+
```json
|
|
88
|
+
"emeraldwalk.runonsave": {
|
|
89
|
+
"commands": [
|
|
90
|
+
{
|
|
91
|
+
"match": "\\.md$",
|
|
92
|
+
"cmd": "flowmark --auto ${file}"
|
|
93
|
+
}
|
|
94
|
+
]
|
|
95
|
+
}
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
The `--auto` option is just the same as `--inplace --nobackup --semantic`.
|
|
99
|
+
|
|
100
|
+
## Usage
|
|
101
|
+
|
|
102
|
+
Flowmark can be used as a library or as a CLI.
|
|
103
|
+
|
|
104
|
+
```
|
|
105
|
+
$ flowmark --help
|
|
106
|
+
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [file]
|
|
107
|
+
|
|
108
|
+
Flowmark: Better line wrapping and formatting for plaintext and Markdown
|
|
109
|
+
|
|
110
|
+
positional arguments:
|
|
111
|
+
file Input file (use '-' for stdin)
|
|
112
|
+
|
|
113
|
+
options:
|
|
114
|
+
-h, --help show this help message and exit
|
|
115
|
+
-o, --output OUTPUT Output file (use '-' for stdout)
|
|
116
|
+
-w, --width WIDTH Line width to wrap to
|
|
117
|
+
-p, --plaintext Process as plaintext (no Markdown parsing)
|
|
118
|
+
-s, --semantic Enable sentence-based line breaks (only applies to Markdown mode)
|
|
119
|
+
-i, --inplace Edit the file in place (ignores --output)
|
|
120
|
+
--nobackup Do not make a backup of the original file when using --inplace
|
|
121
|
+
|
|
122
|
+
Flowmark provides enhanced text wrapping capabilities with special handling for
|
|
123
|
+
Markdown content. It can:
|
|
124
|
+
|
|
125
|
+
- Format Markdown with proper line wrapping while preserving structure
|
|
126
|
+
and normalizing Markdown formatting
|
|
127
|
+
|
|
128
|
+
- Optionally break lines at sentence boundaries for better diff readability
|
|
129
|
+
|
|
130
|
+
- Process plaintext with HTML-aware word splitting
|
|
131
|
+
|
|
132
|
+
It is both a library and a command-line tool.
|
|
133
|
+
|
|
134
|
+
Command-line usage examples:
|
|
135
|
+
|
|
136
|
+
# Format a Markdown file to stdout
|
|
137
|
+
flowmark README.md
|
|
138
|
+
|
|
139
|
+
# Format a Markdown file and save to a new file
|
|
140
|
+
flowmark README.md -o README_formatted.md
|
|
141
|
+
|
|
142
|
+
# Edit a file in-place (with or without making a backup)
|
|
143
|
+
flowmark --inplace README.md
|
|
144
|
+
flowmark --inplace --nobackup README.md
|
|
145
|
+
|
|
146
|
+
# Process plaintext instead of Markdown
|
|
147
|
+
flowmark --plaintext text.txt
|
|
148
|
+
|
|
149
|
+
# Use sentences to guide line breaks (good for many purposes git history and diffs)
|
|
150
|
+
flowmark --semantic README.md
|
|
151
|
+
|
|
152
|
+
For more details, see: https://github.com/jlevy/flowmark
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
* * *
|
|
156
|
+
|
|
157
|
+
*This project was built from
|
|
158
|
+
[simple-modern-poetry](https://github.com/jlevy/simple-modern-poetry).*
|
|
@@ -8,7 +8,7 @@ readme = "README.md"
|
|
|
8
8
|
license = "MIT"
|
|
9
9
|
requires-python = ">=3.10,<4.0"
|
|
10
10
|
dynamic = [ "dependencies", ]
|
|
11
|
-
version = "0.
|
|
11
|
+
version = "0.3.0"
|
|
12
12
|
|
|
13
13
|
[tool.poetry]
|
|
14
14
|
requires-poetry = ">=2.0"
|
|
@@ -88,7 +88,7 @@ disable_error_code = [
|
|
|
88
88
|
|
|
89
89
|
[tool.codespell]
|
|
90
90
|
# ignore-words-list = "foo,bar"
|
|
91
|
-
|
|
91
|
+
skip = "tests/testdocs/*.md"
|
|
92
92
|
|
|
93
93
|
[tool.pytest.ini_options]
|
|
94
94
|
python_files = ["*.py"]
|
|
@@ -29,8 +29,9 @@ Command-line usage examples:
|
|
|
29
29
|
# Process plaintext instead of Markdown
|
|
30
30
|
flowmark --plaintext text.txt
|
|
31
31
|
|
|
32
|
-
# Use
|
|
33
|
-
|
|
32
|
+
# Use semantic line breaks (based on sentences, which is helpful to reduce
|
|
33
|
+
# irrelevant line wrap diffs in git history)
|
|
34
|
+
flowmark --semantic README.md
|
|
34
35
|
|
|
35
36
|
For more details, see: https://github.com/jlevy/flowmark
|
|
36
37
|
"""
|
|
@@ -53,7 +54,7 @@ class Options:
|
|
|
53
54
|
output: str
|
|
54
55
|
width: int
|
|
55
56
|
plaintext: bool
|
|
56
|
-
|
|
57
|
+
semantic: bool
|
|
57
58
|
inplace: bool
|
|
58
59
|
nobackup: bool
|
|
59
60
|
|
|
@@ -91,10 +92,10 @@ def _parse_args(args: Optional[List[str]] = None) -> Options:
|
|
|
91
92
|
)
|
|
92
93
|
parser.add_argument(
|
|
93
94
|
"-s",
|
|
94
|
-
"--
|
|
95
|
+
"--semantic",
|
|
95
96
|
action="store_true",
|
|
96
|
-
default=
|
|
97
|
-
help="Enable sentence-based line breaks (only applies to Markdown mode)",
|
|
97
|
+
default=False,
|
|
98
|
+
help="Enable semantic (sentence-based) line breaks (only applies to Markdown mode)",
|
|
98
99
|
)
|
|
99
100
|
parser.add_argument(
|
|
100
101
|
"-i", "--inplace", action="store_true", help="Edit the file in place (ignores --output)"
|
|
@@ -104,16 +105,26 @@ def _parse_args(args: Optional[List[str]] = None) -> Options:
|
|
|
104
105
|
action="store_true",
|
|
105
106
|
help="Do not make a backup of the original file when using --inplace",
|
|
106
107
|
)
|
|
107
|
-
|
|
108
|
+
parser.add_argument(
|
|
109
|
+
"--auto",
|
|
110
|
+
action="store_true",
|
|
111
|
+
help="Same as `--inplace --nobackup --semantic`, as a convenience for auto-formatting files",
|
|
112
|
+
)
|
|
113
|
+
opts = parser.parse_args(args)
|
|
114
|
+
|
|
115
|
+
if opts.auto:
|
|
116
|
+
opts.inplace = True
|
|
117
|
+
opts.nobackup = True
|
|
118
|
+
opts.semantic = True
|
|
108
119
|
|
|
109
120
|
return Options(
|
|
110
|
-
file=
|
|
111
|
-
output=
|
|
112
|
-
width=
|
|
113
|
-
plaintext=
|
|
114
|
-
|
|
115
|
-
inplace=
|
|
116
|
-
nobackup=
|
|
121
|
+
file=opts.file,
|
|
122
|
+
output=opts.output,
|
|
123
|
+
width=opts.width,
|
|
124
|
+
plaintext=opts.plaintext,
|
|
125
|
+
semantic=opts.semantic,
|
|
126
|
+
inplace=opts.inplace,
|
|
127
|
+
nobackup=opts.nobackup,
|
|
117
128
|
)
|
|
118
129
|
|
|
119
130
|
|
|
@@ -149,7 +160,7 @@ def main(args: Optional[List[str]] = None) -> int:
|
|
|
149
160
|
result = fill_markdown(
|
|
150
161
|
text,
|
|
151
162
|
width=options.width,
|
|
152
|
-
|
|
163
|
+
semantic=options.semantic,
|
|
153
164
|
dedent_input=True,
|
|
154
165
|
)
|
|
155
166
|
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
from typing import Tuple
|
|
2
|
+
|
|
3
|
+
|
|
4
|
+
def split_frontmatter(text: str) -> Tuple[str, str]:
|
|
5
|
+
"""
|
|
6
|
+
Split a text document into frontmatter and content parts.
|
|
7
|
+
|
|
8
|
+
Checks if the string starts with YAML frontmatter, delimited by `---`
|
|
9
|
+
lines. If so, returns the frontmatter, including the `---` lines, and the
|
|
10
|
+
rest of the document. If no frontmatter is found, returns an empty string
|
|
11
|
+
and the original text.
|
|
12
|
+
"""
|
|
13
|
+
lines = text.splitlines()
|
|
14
|
+
|
|
15
|
+
# Skip empty lines at the beginning
|
|
16
|
+
start_idx = 0
|
|
17
|
+
while start_idx < len(lines) and lines[start_idx].strip() == "":
|
|
18
|
+
start_idx += 1
|
|
19
|
+
|
|
20
|
+
# If no content or doesn't start with '---', return empty frontmatter
|
|
21
|
+
if start_idx >= len(lines) or lines[start_idx].strip() != "---":
|
|
22
|
+
return "", text
|
|
23
|
+
|
|
24
|
+
# Look for the closing '---'
|
|
25
|
+
end_idx = start_idx + 1
|
|
26
|
+
while end_idx < len(lines):
|
|
27
|
+
if lines[end_idx].strip() == "---":
|
|
28
|
+
# Found the closing delimiter - extract frontmatter and content
|
|
29
|
+
frontmatter = "\n".join(lines[start_idx : end_idx + 1]) + "\n"
|
|
30
|
+
content = "\n".join(lines[end_idx + 1 :])
|
|
31
|
+
return frontmatter, content
|
|
32
|
+
end_idx += 1
|
|
33
|
+
|
|
34
|
+
# If no closing delimiter found, everything is considered frontmatter
|
|
35
|
+
# and nothing should change
|
|
36
|
+
return text, ""
|
|
37
|
+
|
|
38
|
+
|
|
39
|
+
def has_frontmatter(text: str) -> bool:
|
|
40
|
+
"""
|
|
41
|
+
Check if the text starts with YAML frontmatter.
|
|
42
|
+
"""
|
|
43
|
+
frontmatter, _ = split_frontmatter(text)
|
|
44
|
+
return frontmatter != ""
|
|
@@ -20,6 +20,7 @@ from marko.parser import Parser
|
|
|
20
20
|
from marko.renderer import Renderer
|
|
21
21
|
from marko.source import Source
|
|
22
22
|
|
|
23
|
+
from flowmark.frontmatter import split_frontmatter
|
|
23
24
|
from flowmark.line_wrappers import line_wrap_by_sentence, line_wrap_to_width, LineWrapper
|
|
24
25
|
from flowmark.sentence_split_regex import split_sentences_regex
|
|
25
26
|
from flowmark.text_filling import DEFAULT_WRAP_WIDTH
|
|
@@ -127,16 +128,33 @@ class _MarkdownNormalizer(Renderer):
|
|
|
127
128
|
self._prefix = self._second_prefix
|
|
128
129
|
return wrapped_text + "\n"
|
|
129
130
|
|
|
131
|
+
def _has_multiple_paragraphs(self, item: object) -> bool:
|
|
132
|
+
"""
|
|
133
|
+
Check if a list item contains multiple paragraphs.
|
|
134
|
+
"""
|
|
135
|
+
list_item = cast(block.ListItem, item)
|
|
136
|
+
paragraphs = [c for c in list_item.children if isinstance(c, block.Paragraph)]
|
|
137
|
+
return len(paragraphs) > 1
|
|
138
|
+
|
|
130
139
|
def render_list(self, element: block.List) -> str:
|
|
131
140
|
result: List[str] = []
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
141
|
+
|
|
142
|
+
for i, child in enumerate(element.children):
|
|
143
|
+
# Configure the appropriate prefix based on list type
|
|
144
|
+
if element.ordered:
|
|
145
|
+
num = i + element.start
|
|
146
|
+
prefix = f"{num}. "
|
|
147
|
+
subsequent_indent = " " * (len(str(num)) + 2)
|
|
148
|
+
else:
|
|
149
|
+
prefix = f"{element.bullet} "
|
|
150
|
+
subsequent_indent = " "
|
|
151
|
+
|
|
152
|
+
with self.container(prefix, subsequent_indent):
|
|
153
|
+
# Add an extra newline before multi-paragraph list items (except the first)
|
|
154
|
+
if i > 0 and self._has_multiple_paragraphs(child):
|
|
155
|
+
result.append(self._second_prefix.strip() + "\n")
|
|
156
|
+
|
|
157
|
+
result.append(self.render(child))
|
|
140
158
|
|
|
141
159
|
self._prefix = self._second_prefix
|
|
142
160
|
return "".join(result)
|
|
@@ -150,6 +168,7 @@ class _MarkdownNormalizer(Renderer):
|
|
|
150
168
|
# Add the newline between paragraphs. Normally this would be an empty line but
|
|
151
169
|
# within a quote block it would be the secondary prefix, like `> `.
|
|
152
170
|
result += self._second_prefix.strip() + "\n"
|
|
171
|
+
|
|
153
172
|
result += self.render_children(element)
|
|
154
173
|
return result
|
|
155
174
|
|
|
@@ -271,7 +290,7 @@ def fill_markdown(
|
|
|
271
290
|
markdown_text: str,
|
|
272
291
|
dedent_input: bool = True,
|
|
273
292
|
width: int = DEFAULT_WRAP_WIDTH,
|
|
274
|
-
|
|
293
|
+
semantic: bool = False,
|
|
275
294
|
line_wrapper: Optional[LineWrapper] = None,
|
|
276
295
|
) -> str:
|
|
277
296
|
"""
|
|
@@ -285,12 +304,25 @@ def fill_markdown(
|
|
|
285
304
|
|
|
286
305
|
Optionally also dedents and strips the input, so it can be used
|
|
287
306
|
on docstrings.
|
|
307
|
+
|
|
308
|
+
With `semantic` enabled, the line breaks are wrapped approximately
|
|
309
|
+
by sentence boundaries, to make diffs more readable.
|
|
310
|
+
|
|
311
|
+
Preserves YAML frontmatter (delimited by --- lines) if present at the
|
|
312
|
+
beginning of the document.
|
|
288
313
|
"""
|
|
289
314
|
if line_wrapper is None:
|
|
290
315
|
line_wrapper = (
|
|
291
|
-
line_wrap_by_sentence(width=width) if
|
|
316
|
+
line_wrap_by_sentence(width=width) if semantic else line_wrap_to_width(width=width)
|
|
292
317
|
)
|
|
293
318
|
|
|
319
|
+
# Extract frontmatter before any processing
|
|
320
|
+
frontmatter, content = split_frontmatter(markdown_text)
|
|
321
|
+
|
|
322
|
+
# Only format the content part if there's frontmatter
|
|
323
|
+
if frontmatter:
|
|
324
|
+
markdown_text = content
|
|
325
|
+
|
|
294
326
|
if dedent_input:
|
|
295
327
|
markdown_text = dedent(markdown_text).strip()
|
|
296
328
|
|
|
@@ -303,4 +335,9 @@ def fill_markdown(
|
|
|
303
335
|
parser = CustomParser()
|
|
304
336
|
parsed = parser.parse(markdown_text)
|
|
305
337
|
result = _MarkdownNormalizer(line_wrapper).render(parsed)
|
|
338
|
+
|
|
339
|
+
# Reattach frontmatter if it was present
|
|
340
|
+
if frontmatter:
|
|
341
|
+
result = frontmatter + result
|
|
342
|
+
|
|
306
343
|
return result
|
flowmark-0.2.0/PKG-INFO
DELETED
|
@@ -1,124 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.3
|
|
2
|
-
Name: flowmark
|
|
3
|
-
Version: 0.2.0
|
|
4
|
-
Summary: Better line wrapping and formatting for plaintext and Markdown
|
|
5
|
-
License: MIT
|
|
6
|
-
Author: Joshua Levy
|
|
7
|
-
Author-email: joshua@cal.berkeley.edu
|
|
8
|
-
Requires-Python: >=3.10,<4.0
|
|
9
|
-
Classifier: License :: OSI Approved :: MIT License
|
|
10
|
-
Classifier: Programming Language :: Python :: 3
|
|
11
|
-
Classifier: Programming Language :: Python :: 3.10
|
|
12
|
-
Classifier: Programming Language :: Python :: 3.11
|
|
13
|
-
Classifier: Programming Language :: Python :: 3.12
|
|
14
|
-
Classifier: Programming Language :: Python :: 3.13
|
|
15
|
-
Requires-Dist: marko (>=2.1.2,<3.0.0)
|
|
16
|
-
Requires-Dist: regex (>=2024.11.6,<2025.0.0)
|
|
17
|
-
Requires-Dist: strif (>=2.0.0,<3.0.0)
|
|
18
|
-
Description-Content-Type: text/markdown
|
|
19
|
-
|
|
20
|
-
# flowmark
|
|
21
|
-
|
|
22
|
-
Flowmark is a new Python implementation of text line wrapping and filling.
|
|
23
|
-
|
|
24
|
-
It simplifies and generalizes Python's
|
|
25
|
-
[`textwrap`](https://docs.python.org/3/library/textwrap.html) with a few more
|
|
26
|
-
capabilities:
|
|
27
|
-
|
|
28
|
-
- Full customizability of initial and subsequent indentation strings
|
|
29
|
-
|
|
30
|
-
- Control over when to split words, by default using a word splitter that won't break
|
|
31
|
-
lines within HTML tags
|
|
32
|
-
|
|
33
|
-
In addition, it adds optional support for Markdown and offers Markdown auto-formatting,
|
|
34
|
-
like [markdownfmt](https://github.com/shurcooL/markdownfmt), also with controllable line
|
|
35
|
-
wrapping options.
|
|
36
|
-
|
|
37
|
-
One key use case is to normalize Markdown in a standard, readable way that makes diffs
|
|
38
|
-
easy to read and use on GitHub.
|
|
39
|
-
This can be useful for documentation workflows and also to compare LLM outputs that are
|
|
40
|
-
Markdown.
|
|
41
|
-
|
|
42
|
-
Finally, it has options to use heuristics to split on sentences, which can make diffs
|
|
43
|
-
much more readable. (For an example of this, look at the
|
|
44
|
-
[Markdown source](https://github.com/jlevy/flowmark/blob/main/README.md?plain=1) of this
|
|
45
|
-
readme file.)
|
|
46
|
-
|
|
47
|
-
It aims to be small and simple and have only a few dependencies, currently only
|
|
48
|
-
[`marko`](https://github.com/frostming/marko) and
|
|
49
|
-
[`regex`](https://pypi.org/project/regex/).
|
|
50
|
-
|
|
51
|
-
This is a new and simple package (previously I'd implemented something like this
|
|
52
|
-
[for Atom](https://github.com/jlevy/atom-flowmark)) but I plan to add more support for
|
|
53
|
-
command line usage and VSCode/Cursor auto-formatting in the future.
|
|
54
|
-
|
|
55
|
-
## Installation
|
|
56
|
-
|
|
57
|
-
The simplest way to use the tool is to use pipx:
|
|
58
|
-
|
|
59
|
-
```shell
|
|
60
|
-
pipx install flowmark
|
|
61
|
-
```
|
|
62
|
-
|
|
63
|
-
To use as a library, use pip or poetry to install `flowmark`.
|
|
64
|
-
|
|
65
|
-
## Usage
|
|
66
|
-
|
|
67
|
-
Flowmark can be used as a library or as a CLI.
|
|
68
|
-
|
|
69
|
-
```
|
|
70
|
-
$ flowmark --help
|
|
71
|
-
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [file]
|
|
72
|
-
|
|
73
|
-
Flowmark: Better line wrapping and formatting for plaintext and Markdown
|
|
74
|
-
|
|
75
|
-
positional arguments:
|
|
76
|
-
file Input file (use '-' for stdin)
|
|
77
|
-
|
|
78
|
-
options:
|
|
79
|
-
-h, --help show this help message and exit
|
|
80
|
-
-o, --output OUTPUT Output file (use '-' for stdout)
|
|
81
|
-
-w, --width WIDTH Line width to wrap to
|
|
82
|
-
-p, --plaintext Process as plaintext (no Markdown parsing)
|
|
83
|
-
-s, --sentences Enable sentence-based line breaks (only applies to Markdown mode)
|
|
84
|
-
-i, --inplace Edit the file in place (ignores --output)
|
|
85
|
-
--nobackup Do not make a backup of the original file when using --inplace
|
|
86
|
-
|
|
87
|
-
Flowmark provides enhanced text wrapping capabilities with special handling for
|
|
88
|
-
Markdown content. It can:
|
|
89
|
-
|
|
90
|
-
- Format Markdown with proper line wrapping while preserving structure
|
|
91
|
-
and normalizing Markdown formatting
|
|
92
|
-
|
|
93
|
-
- Optionally break lines at sentence boundaries for better diff readability
|
|
94
|
-
|
|
95
|
-
- Process plaintext with HTML-aware word splitting
|
|
96
|
-
|
|
97
|
-
It is both a library and a command-line tool.
|
|
98
|
-
|
|
99
|
-
Command-line usage examples:
|
|
100
|
-
|
|
101
|
-
# Format a Markdown file to stdout
|
|
102
|
-
flowmark README.md
|
|
103
|
-
|
|
104
|
-
# Format a Markdown file and save to a new file
|
|
105
|
-
flowmark README.md -o README_formatted.md
|
|
106
|
-
|
|
107
|
-
# Edit a file in-place (with or without making a backup)
|
|
108
|
-
flowmark --inplace README.md
|
|
109
|
-
flowmark --inplace --nobackup README.md
|
|
110
|
-
|
|
111
|
-
# Process plaintext instead of Markdown
|
|
112
|
-
flowmark --plaintext text.txt
|
|
113
|
-
|
|
114
|
-
# Use sentences to guide line breaks (good for many purposes git history and diffs)
|
|
115
|
-
flowmark --sentences README.md
|
|
116
|
-
|
|
117
|
-
For more details, see: https://github.com/jlevy/flowmark
|
|
118
|
-
```
|
|
119
|
-
|
|
120
|
-
* * *
|
|
121
|
-
|
|
122
|
-
*This project was built from
|
|
123
|
-
[simple-modern-poetry](https://github.com/jlevy/simple-modern-poetry).*
|
|
124
|
-
|
flowmark-0.2.0/README.md
DELETED
|
@@ -1,104 +0,0 @@
|
|
|
1
|
-
# flowmark
|
|
2
|
-
|
|
3
|
-
Flowmark is a new Python implementation of text line wrapping and filling.
|
|
4
|
-
|
|
5
|
-
It simplifies and generalizes Python's
|
|
6
|
-
[`textwrap`](https://docs.python.org/3/library/textwrap.html) with a few more
|
|
7
|
-
capabilities:
|
|
8
|
-
|
|
9
|
-
- Full customizability of initial and subsequent indentation strings
|
|
10
|
-
|
|
11
|
-
- Control over when to split words, by default using a word splitter that won't break
|
|
12
|
-
lines within HTML tags
|
|
13
|
-
|
|
14
|
-
In addition, it adds optional support for Markdown and offers Markdown auto-formatting,
|
|
15
|
-
like [markdownfmt](https://github.com/shurcooL/markdownfmt), also with controllable line
|
|
16
|
-
wrapping options.
|
|
17
|
-
|
|
18
|
-
One key use case is to normalize Markdown in a standard, readable way that makes diffs
|
|
19
|
-
easy to read and use on GitHub.
|
|
20
|
-
This can be useful for documentation workflows and also to compare LLM outputs that are
|
|
21
|
-
Markdown.
|
|
22
|
-
|
|
23
|
-
Finally, it has options to use heuristics to split on sentences, which can make diffs
|
|
24
|
-
much more readable. (For an example of this, look at the
|
|
25
|
-
[Markdown source](https://github.com/jlevy/flowmark/blob/main/README.md?plain=1) of this
|
|
26
|
-
readme file.)
|
|
27
|
-
|
|
28
|
-
It aims to be small and simple and have only a few dependencies, currently only
|
|
29
|
-
[`marko`](https://github.com/frostming/marko) and
|
|
30
|
-
[`regex`](https://pypi.org/project/regex/).
|
|
31
|
-
|
|
32
|
-
This is a new and simple package (previously I'd implemented something like this
|
|
33
|
-
[for Atom](https://github.com/jlevy/atom-flowmark)) but I plan to add more support for
|
|
34
|
-
command line usage and VSCode/Cursor auto-formatting in the future.
|
|
35
|
-
|
|
36
|
-
## Installation
|
|
37
|
-
|
|
38
|
-
The simplest way to use the tool is to use pipx:
|
|
39
|
-
|
|
40
|
-
```shell
|
|
41
|
-
pipx install flowmark
|
|
42
|
-
```
|
|
43
|
-
|
|
44
|
-
To use as a library, use pip or poetry to install `flowmark`.
|
|
45
|
-
|
|
46
|
-
## Usage
|
|
47
|
-
|
|
48
|
-
Flowmark can be used as a library or as a CLI.
|
|
49
|
-
|
|
50
|
-
```
|
|
51
|
-
$ flowmark --help
|
|
52
|
-
usage: flowmark [-h] [-o OUTPUT] [-w WIDTH] [-p] [-s] [-i] [--nobackup] [file]
|
|
53
|
-
|
|
54
|
-
Flowmark: Better line wrapping and formatting for plaintext and Markdown
|
|
55
|
-
|
|
56
|
-
positional arguments:
|
|
57
|
-
file Input file (use '-' for stdin)
|
|
58
|
-
|
|
59
|
-
options:
|
|
60
|
-
-h, --help show this help message and exit
|
|
61
|
-
-o, --output OUTPUT Output file (use '-' for stdout)
|
|
62
|
-
-w, --width WIDTH Line width to wrap to
|
|
63
|
-
-p, --plaintext Process as plaintext (no Markdown parsing)
|
|
64
|
-
-s, --sentences Enable sentence-based line breaks (only applies to Markdown mode)
|
|
65
|
-
-i, --inplace Edit the file in place (ignores --output)
|
|
66
|
-
--nobackup Do not make a backup of the original file when using --inplace
|
|
67
|
-
|
|
68
|
-
Flowmark provides enhanced text wrapping capabilities with special handling for
|
|
69
|
-
Markdown content. It can:
|
|
70
|
-
|
|
71
|
-
- Format Markdown with proper line wrapping while preserving structure
|
|
72
|
-
and normalizing Markdown formatting
|
|
73
|
-
|
|
74
|
-
- Optionally break lines at sentence boundaries for better diff readability
|
|
75
|
-
|
|
76
|
-
- Process plaintext with HTML-aware word splitting
|
|
77
|
-
|
|
78
|
-
It is both a library and a command-line tool.
|
|
79
|
-
|
|
80
|
-
Command-line usage examples:
|
|
81
|
-
|
|
82
|
-
# Format a Markdown file to stdout
|
|
83
|
-
flowmark README.md
|
|
84
|
-
|
|
85
|
-
# Format a Markdown file and save to a new file
|
|
86
|
-
flowmark README.md -o README_formatted.md
|
|
87
|
-
|
|
88
|
-
# Edit a file in-place (with or without making a backup)
|
|
89
|
-
flowmark --inplace README.md
|
|
90
|
-
flowmark --inplace --nobackup README.md
|
|
91
|
-
|
|
92
|
-
# Process plaintext instead of Markdown
|
|
93
|
-
flowmark --plaintext text.txt
|
|
94
|
-
|
|
95
|
-
# Use sentences to guide line breaks (good for many purposes git history and diffs)
|
|
96
|
-
flowmark --sentences README.md
|
|
97
|
-
|
|
98
|
-
For more details, see: https://github.com/jlevy/flowmark
|
|
99
|
-
```
|
|
100
|
-
|
|
101
|
-
* * *
|
|
102
|
-
|
|
103
|
-
*This project was built from
|
|
104
|
-
[simple-modern-poetry](https://github.com/jlevy/simple-modern-poetry).*
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|