spanmark 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,436 @@
1
+ Metadata-Version: 2.5
2
+ Name: spanmark
3
+ Version: 0.1.0
4
+ Summary: A lightweight span-annotation widget for interactive Python environments
5
+ Project-URL: Homepage, https://github.com/pdhall99/spanmark
6
+ Project-URL: Documentation, https://github.com/pdhall99/spanmark/blob/main/README.md
7
+ Project-URL: Source, https://github.com/pdhall99/spanmark
8
+ Project-URL: Issues, https://github.com/pdhall99/spanmark/issues
9
+ Project-URL: Changelog, https://github.com/pdhall99/spanmark/blob/main/docs/CHANGELOG.md
10
+ Author-email: PD Hall <20580126+pdhall99@users.noreply.github.com>
11
+ License-Expression: MIT
12
+ License-File: LICENSE
13
+ Keywords: annotation,anywidget,data labeling,data labelling,human in the loop,information extraction,jupyter,named entity recognition,ner,nlp,notebook,sequence labeling,sequence labelling,span annotation,text annotation,text labeling,text labelling,widget
14
+ Classifier: Development Status :: 4 - Beta
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Operating System :: OS Independent
17
+ Classifier: Programming Language :: Python
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3 :: Only
20
+ Classifier: Programming Language :: Python :: 3.10
21
+ Classifier: Programming Language :: Python :: 3.11
22
+ Classifier: Programming Language :: Python :: 3.12
23
+ Classifier: Programming Language :: Python :: 3.13
24
+ Classifier: Programming Language :: Python :: 3.14
25
+ Classifier: Typing :: Typed
26
+ Requires-Python: >=3.10
27
+ Requires-Dist: anywidget
28
+ Requires-Dist: filelock>=3.16
29
+ Requires-Dist: ipython
30
+ Requires-Dist: traitlets
31
+ Requires-Dist: typing-extensions
32
+ Description-Content-Type: text/markdown
33
+
34
+ # spanmark
35
+
36
+ A lightweight span-annotation widget for interactive Python environments.
37
+
38
+ Add span annotations to text datasets in Jupyter Notebook, JupyterLab, VS Code, Google Colab, marimo, or other [anywidget](https://anywidget.dev/)-compatible environments.
39
+
40
+ ![spanmark widget showing named-entity spans and document decisions](https://github.com/pdhall99/spanmark/blob/main/assets/screenshot.png)
41
+
42
+ ## Table of contents
43
+
44
+ - [Installation](#installation)
45
+ - [Usage](#usage)
46
+ - [Use cases](#use-cases)
47
+ - [Related tools](#related-tools)
48
+ - [Compatibility and versioning](#compatibility-and-versioning)
49
+ - [Acknowledgements](#acknowledgements)
50
+ - [Contributing](#contributing)
51
+ - [License](#license)
52
+
53
+ ## Installation
54
+
55
+ `spanmark` requires Python 3.10 or later (see [compatibility and versioning](#compatibility-and-versioning)).
56
+
57
+ Install the latest release from [PyPI](https://pypi.org/project/spanmark/):
58
+
59
+ ```shell
60
+ python -m pip install spanmark
61
+ ```
62
+
63
+ _JupyterLab environments_:
64
+ In most setups, installing `spanmark` in the notebook's Python environment is enough.
65
+ If JupyterLab itself is installed in a different environment from the notebook kernel, install `anywidget` in the JupyterLab environment as well so its frontend extension is available.
66
+
67
+ ## Usage
68
+
69
+ ### Quickstart
70
+
71
+ Prepare a dataset as a sequence of documents.
72
+ Each document must have:
73
+
74
+ - `id` – a string identifier, unique in the dataset
75
+ - `text` – the string to be annotated.
76
+
77
+ Each document may have `suggestions`, pre-annotated spans, which can be edited during annotation.
78
+
79
+ ```python
80
+ documents = [
81
+ {
82
+ "id": "example-001",
83
+ "text": "Alice joined Acme Corp in London.",
84
+ "meta": {"split": "train"},
85
+ "suggestions": [
86
+ {"start": 0, "end": 5, "label": "PERSON", "score": 0.99},
87
+ {"start": 13, "end": 22, "label": "ORG", "score": 0.94},
88
+ {"start": 26, "end": 32, "label": "LOCATION", "score": 0.97},
89
+ ],
90
+ },
91
+ {
92
+ "id": "example-002",
93
+ "text": "Bob moved from Paris to Edinburgh.",
94
+ "suggestions": [
95
+ {"start": 0, "end": 3, "label": "PERSON", "score": 0.96},
96
+ {"start": 15, "end": 20, "label": "LOCATION", "score": 0.91},
97
+ {"start": 24, "end": 33, "label": "LOCATION", "score": 0.89},
98
+ ],
99
+ },
100
+ ]
101
+ ```
102
+
103
+ Make an `AnnotationSession` with the input data, the allowed span labels, and output JSONL path:
104
+
105
+ ```python
106
+ from spanmark import AnnotationSession
107
+
108
+ session = AnnotationSession(
109
+ documents,
110
+ labels=["PERSON", "ORG", "LOCATION"],
111
+ output_path="annotations.jsonl",
112
+ )
113
+ ```
114
+
115
+ Display the widget and annotate the documents:
116
+
117
+ ```python
118
+ session.display()
119
+ ```
120
+
121
+ Cleanly end the annotation session:
122
+
123
+ ```python
124
+ session.close()
125
+ ```
126
+
127
+ The completed annotations and decisions are written to the output JSONL, `annotations.jsonl`.
128
+
129
+ ### Terminology
130
+
131
+ - **dataset** — an ordered collection of documents
132
+ - **document** — one item to annotate, identified by `id`
133
+ - **text** — the string within a document to which spans refer
134
+ - **span** — a labelled right-open character range `[start, end)` over the text
135
+ - **suggestion** — an optional pre-annotated span, usually produced by a model
136
+ - **annotation** — the persisted annotation state for a document
137
+ - **decision** — one of Accept, Reject, or Ignore for the current document
138
+ - **flag** — an independent marker indicating that a document should be reviewed later
139
+
140
+ ### Character-based spans
141
+
142
+ spanmark spans are labelled right-open intervals of Unicode code-point offsets into the document text: `[start, end)`, where `start` is inclusive and `end` is exclusive.
143
+ For a persisted span, the covered text is therefore exactly:
144
+
145
+ ```python
146
+ span_text = text[span["start"] : span["end"]]
147
+ ```
148
+
149
+ ### Overlapping spans
150
+
151
+ The default annotation mode rejects overlapping spans.
152
+ To allow overlapping or nested spans, set `allow_overlaps=True`:
153
+
154
+ ```python
155
+ session = AnnotationSession(
156
+ documents,
157
+ labels=["PERSON", "TITLE", "ORG"],
158
+ output_path="annotations.jsonl",
159
+ allow_overlaps=True,
160
+ )
161
+ ```
162
+
163
+ ### Workflow
164
+
165
+ spanmark supports the following workflow:
166
+
167
+ 1. Prepare input data
168
+ 2. Annotate
169
+ 3. Review flagged documents
170
+ 4. Use the output JSONL
171
+
172
+ #### 1. Prepare input data
173
+
174
+ Prepare a sequence of documents:
175
+
176
+ ```python
177
+ {
178
+ "id": "doc-1",
179
+ "text": "Alice joined Acme.",
180
+ "meta": {
181
+ "source": "demo"
182
+ },
183
+ "suggestions": [
184
+ {
185
+ "start": 0,
186
+ "end": 5,
187
+ "label": "PERSON",
188
+ "score": 0.97
189
+ }
190
+ ]
191
+ }
192
+ ```
193
+
194
+ The document schema is:
195
+
196
+ | Field | Required | Type | Meaning |
197
+ | --------------- | ---------: | ------------------------------------ | ------------------------------ |
198
+ | `id` | yes | string | Unique document identifier |
199
+ | `text` | yes | string | Text to annotate |
200
+ | _`meta`_ | _no_ | _object or null_ | _Arbitrary user metadata_ |
201
+ | _`suggestions`_ | _no_ | _list of suggestion spans_ | _Optional pre-annotated spans_ |
202
+ | _other fields_ | _no_ | _any JSON value_ | _Preserved unchanged_ |
203
+
204
+ The suggestion span schema is:
205
+
206
+ | Field | Required | Type | Description |
207
+ | --------- | ---------: | --------- | ----------------------------------------------- |
208
+ | `start` | yes | integer | Start character offset, inclusive |
209
+ | `end` | yes | integer | End character offset, exclusive |
210
+ | `label` | yes | string | Span label - one of the session labels |
211
+ | _`score`_ | _no_ | _number_ | _Optional model confidence score_ |
212
+ | _`id`_ | _no_ | _string_ | _Optional upstream suggestion identifier_ |
213
+
214
+ spanmark spans are labelled right-open intervals of Unicode code-point offsets into the document text: `[start, end)`, where `start` is inclusive and `end` is exclusive.
215
+ JSONL input files must be UTF-8 encoded.
216
+ spanmark writes output JSONL files as UTF-8.
217
+
218
+ Documents can be passed directly or loaded from JSONL:
219
+
220
+ ```python
221
+ session = AnnotationSession.from_jsonl(
222
+ "documents.jsonl",
223
+ labels=["PERSON", "ORG", "LOCATION"],
224
+ output_path="annotations.jsonl",
225
+ )
226
+ ```
227
+
228
+ #### 2. Annotate
229
+
230
+ spanmark shows one document at a time, including any suggested spans.
231
+ Add, remove, or relabel spans, then select a document decision to advance to the next undecided document:
232
+
233
+ - **Accept** — keep the document with the current spans.
234
+ - **Reject** — exclude the document because it fails the annotation task criteria.
235
+ - **Ignore** — exclude the document without making a task judgement, for example because it is malformed, ambiguous, or cannot be annotated reliably.
236
+
237
+ You can optionally add a Flag before selecting a decision:
238
+
239
+ - **Flag** — mark the document to be revisited later.
240
+
241
+ The Flag can be used, for example, to indicate that an annotation is uncertain by selecting it alongside **Ignore**.
242
+ A span edit invalidates an existing decision because the decision certifies the exact current span set.
243
+
244
+ The main keyboard shortcuts are:
245
+
246
+ | Shortcut | Action |
247
+ | --- | --- |
248
+ | `A` | Accept |
249
+ | `X` | Reject |
250
+ | `Space` | Ignore |
251
+ | `F` | Toggle Flag |
252
+ | `1`–`9` | Choose a label or relabel the selected span |
253
+ | `Ctrl/Cmd+Z` | Undo the current span edit |
254
+ | `Backspace` / Mac `Delete` | Undo the previous submitted decision |
255
+ | `Delete` / Mac `Fn+Delete` | Remove the selected span |
256
+ | `Esc` | Clear the span selection |
257
+
258
+ #### 3. Review flagged documents
259
+
260
+ When all documents have a decision, the completion screen will show.
261
+ Pressing **Review flagged** begins a flagged review session.
262
+
263
+ Flagged documents are reviewed in dataset order.
264
+ Existing spans and the existing decision are loaded as normal, and you can edit spans or change the decision.
265
+ Changing the Accept/Reject/Ignore decision during flagged review autosaves but keeps you on the same document.
266
+ Clearing **Flag** means the review item is resolved and automatically advances to the next flagged document.
267
+
268
+ Span edits still invalidate the previous decision because a decision certifies the exact current span set.
269
+ If you edit spans during flagged review, submit a new Accept, Reject, or Ignore decision before clearing Flag.
270
+
271
+ If you close a completed session partway through flagged review, the unresolved flags remain in the data.
272
+ Reopening that JSONL returns to the completion screen, where **Review flagged** continues with the remaining flagged documents.
273
+ If the JSONL still contains any undecided documents, spanmark instead resumes the main pass at the first undecided document and flagged review remains unavailable until the pass is complete.
274
+
275
+ #### 4. Use the output JSONL
276
+
277
+ The output JSONL contains one document record per line, for example:
278
+
279
+ ```json
280
+ {
281
+ "id": "doc-1",
282
+ "text": "Alice joined Acme.",
283
+ "suggestions": [
284
+ {"start": 0, "end": 5, "label": "PERSON", "score": 0.97}
285
+ ],
286
+ "annotation": {
287
+ "spans": [
288
+ {"start": 0, "end": 5, "label": "PERSON", "source": "suggestion", "score": 0.97},
289
+ {"start": 13, "end": 17, "label": "ORG", "source": "user"}
290
+ ],
291
+ "answer": "accept",
292
+ "flagged": false
293
+ }
294
+ }
295
+ ```
296
+
297
+ The document schema is as for the input with the addition of the `annotation` field.
298
+ The `annotation` object has the following schema:
299
+
300
+ | Field | Required | Type | Meaning |
301
+ | --------- | -------: | ----------------------- | ------------------------------------------------------ |
302
+ | `spans` | yes | list of spans or `null` | Current annotated spans; `null` means untouched |
303
+ | `answer` | yes | string or `null` | `accept`, `reject`, `ignore`, or `null` if undecided |
304
+ | `flagged` | yes | boolean | Whether the document is flagged for review |
305
+
306
+ Each persisted annotation span has the following schema:
307
+
308
+ | Field | Required | Type | Meaning |
309
+ | ---------- | ---------: | ---------- | ----------------------------------------------------------- |
310
+ | `start` | yes | integer | Start character offset, inclusive |
311
+ | `end` | yes | integer | End character offset, exclusive |
312
+ | `label` | yes | string | Span label |
313
+ | `source` | yes | string | `suggestion` for unchanged suggestions; otherwise `user` |
314
+ | _`score`_ | _no_ | _number_ | _Optional confidence score, preserved when present_ |
315
+
316
+ spanmark spans are labelled right-open intervals of Unicode code-point offsets into the document text: `[start, end)`, where `start` is inclusive and `end` is exclusive.
317
+
318
+ `annotation.spans: null` means the document has not yet been changed or decided, so the widget initializes its editable spans from `suggestions`.
319
+ Once the annotator edits, flags, accepts, rejects, or ignores the document, `annotation.spans` becomes a list.
320
+ An empty list is therefore meaningful: it can represent a deliberate zero-span annotation and will not be repopulated from suggestions on resume.
321
+
322
+ Persisted spans have one of two `source` values:
323
+
324
+ - `"suggestion"` for an unchanged model suggestion
325
+ - `"user"` for a span created or edited by the annotator
326
+
327
+ ### Annotating in multiple sessions
328
+
329
+ At the first session for a given dataset, the selected `output_path` must not already exist.
330
+ spanmark writes a complete annotated copy of the input data there.
331
+
332
+ To resume annotation in another session, use `AnnotationSession.from_jsonl` and set the previous annotated output as both the input and the output:
333
+
334
+ ```python
335
+ session = AnnotationSession.from_jsonl(
336
+ "annotations.jsonl",
337
+ labels=["PERSON", "ORG", "LOCATION"],
338
+ output_path="annotations.jsonl",
339
+ )
340
+ ```
341
+
342
+ spanmark starts at the first undecided document.
343
+ If all documents are decided, it shows the completion screen.
344
+
345
+ If a kernel or process stops between checkpoints, spanmark may leave a hidden sibling file named like `.annotations.jsonl.spanmark-autosave`.
346
+ Reopening the JSONL as both input and output automatically recovers the latest complete autosaved state and folds it back into the JSONL.
347
+ A clean close removes the autosave file, so the resulting JSONL is self-contained.
348
+
349
+ After reopening a session, previous-decision Undo is reconstructed in dataset order.
350
+ That matches the normal one-way first pass, but it is not an exact record of arbitrary action order from an earlier session.
351
+
352
+ The JSONL reader builds a byte-offset index and loads source records on demand.
353
+ Normal annotation changes append only the complete current `annotation` state for the changed document to the hidden autosave overlay.
354
+ Full JSONL work is reserved for explicit `save()`, clean close, workflow completion, and occasional internal checkpoints when the overlay grows large.
355
+
356
+ ### Autosave and locking
357
+
358
+ A live `AnnotationSession` owns its `output_path` exclusively.
359
+ spanmark acquires a cross-platform lock at session creation and holds it until `session.close()`.
360
+ Trying to open a second live session against the same output raises an error instead of risking last-writer-wins data loss.
361
+
362
+ If an output path already exists but is not the JSONL file being opened for resume, spanmark refuses to overwrite it.
363
+
364
+ Each annotation change is written as one compact JSON object to the hidden autosave overlay, then flushed and `fsync`ed before the in-memory state is committed.
365
+ Each entry contains the complete latest annotation state for one document, so recovery is last-writer-wins by document rather than an edit-event replay.
366
+
367
+ JSONL checkpoints still use a unique temporary file in the same directory, `fsync` it, then atomically replace `output_path`.
368
+ On POSIX systems spanmark also performs a best-effort directory `fsync`.
369
+ The autosave overlay records a content fingerprint of the JSONL checkpoint it belongs to; spanmark refuses ambiguous recovery if that checkpoint was modified externally.
370
+
371
+ To force a complete JSONL checkpoint, call:
372
+
373
+ ```python
374
+ session.save()
375
+ ```
376
+
377
+ Normal session close and workflow completion also checkpoint.
378
+ If the process is interrupted before that happens, keep the hidden autosave file beside the JSONL so spanmark can recover it on the next open.
379
+
380
+ ## Use cases
381
+
382
+ spanmark helps you to make a human-verified dataset of labelled, character-based spans over document text.
383
+ Such datasets are needed for the training and evaluation of **span recognition** tasks such as
384
+
385
+ - **Named entity recognition (NER)** — people, organizations, locations, products, dates, and other entity mentions
386
+ - **PII and sensitive-data annotation** — names, addresses, account identifiers, phone numbers, email addresses, and spans for redaction datasets
387
+ - **Keyphrase, terminology, and concept extraction** — domain terms in technical, legal, biomedical, financial, or product text
388
+ - **Slot and field extraction** — destinations, dates, quantities, order numbers, product names, and similar values in conversational or transactional text
389
+ - **Event-trigger and mention detection** — the exact text that expresses an event or concept
390
+ - **Model correction and human-in-the-loop review** — preload model suggestions, then accept, remove, relabel, or supplement them
391
+
392
+ ## Related tools
393
+
394
+ - [Label Studio](https://labelstud.io/) is a general-purpose data labeling across text and other modalities
395
+ - [doccano](https://github.com/doccano/doccano) is an open-source collaborative text annotation for tasks including text classification and named entity recognition
396
+ - [INCEpTION](https://inception-project.github.io/) is a collaborative text annotation with configurable span, relation, and chain layers plus curation workflows
397
+ - [brat](https://brat.nlplab.org/) is a web-based structured text annotation for spans, relations, events, and attributes
398
+ - [Prodigy](https://prodi.gy/) is a commercial NLP annotation software with named-entity, overlapping-span, classification, relation, and model-assisted workflows
399
+
400
+ ## Compatibility and versioning
401
+
402
+ ### Package versioning
403
+
404
+ This project follows [Semantic Versioning](https://semver.org/).
405
+ Releases have version numbers of the form `MAJOR.MINOR.PATCH`:
406
+
407
+ - **MAJOR** releases may contain backwards-incompatible changes to the public API
408
+ - **MINOR** releases may add functionality and deprecate public APIs while remaining backwards compatible
409
+ - **PATCH** releases contain backwards-compatible fixes
410
+
411
+ APIs explicitly documented as experimental are not covered by the same backwards-compatibility guarantees.
412
+
413
+ For releases before `1.0.0`, the public API should be considered under development and may change between minor releases, as permitted by Semantic Versioning.
414
+
415
+ ### Python-version compatibility
416
+
417
+ This project supports Python feature releases from their official final release until their official end-of-life (EOL).
418
+
419
+ Support for a new Python feature release is generally introduced in the first minor release of this project following the upstream Python release.
420
+ Python feature releases may be dropped once they reach the end of their upstream support cycle.
421
+ The currently supported Python versions are declared in the package metadata.
422
+
423
+ Dropping an EOL Python version is considered a change to the supported runtime environment rather than a backwards-incompatible change to this project's public API, and therefore does not by itself require a new major release.
424
+
425
+ ## Acknowledgements
426
+
427
+ This project is developed with assistance from AI coding tools.
428
+ All code included in the project is reviewed and tested by [@pdhall99](https://github.com/pdhall99), who takes responsibility for its quality and maintenance.
429
+
430
+ ## Contributing
431
+
432
+ See the [contributor guide](https://github.com/pdhall99/spanmark/blob/main/docs/CONTRIBUTING.md).
433
+
434
+ ## License
435
+
436
+ [MIT © PD Hall](LICENSE)
@@ -0,0 +1,14 @@
1
+ spanmark/__init__.py,sha256=3wXVke30C9f5Or547vIL9qJ0aHoJqwA7DRzwD2yt3cc,360
2
+ spanmark/_autosave.py,sha256=4xHa09p7vac4am_pkn94lXeXyUDhPESje4NVJQPlnfc,7572
3
+ spanmark/_model.py,sha256=oNhqBHQS1MVMk8uYjYrBnS2m7mJ0Sj17v_twN6h__l4,8395
4
+ spanmark/_session.py,sha256=OW_hp3KACFy7_8YB6pGddNfAMHMBzRflqevU_JQRkqY,32917
5
+ spanmark/_source.py,sha256=CG218bdmJkQPxHp70YkVvNMEv29XDl8QeaRvWhSUMM8,5779
6
+ spanmark/_storage.py,sha256=RxwT5Q8vBDXJLnjgN1ISkmWw32ax-mNr4iXGovkG4DU,2548
7
+ spanmark/_version.py,sha256=n_5vdJsPNu7wZ57LGuRL585uvll-hiuvZUBWzdG0RQU,520
8
+ spanmark/_widget.py,sha256=EtO1zLf1aa5ga7oRUNr9tbRHsX-CcWmVle_d6iyRdGA,1438
9
+ spanmark/static/widget.css,sha256=_L4Uz7GQy9iC_lF0J7NnD0BkXJQY1qFgqrMsaGFDKUw,6469
10
+ spanmark/static/widget.js,sha256=jaFMPTzJsOifTM77XsKCI0ZeizdllQrb__H4gZR37ZA,28501
11
+ spanmark-0.1.0.dist-info/METADATA,sha256=Q4e4G11eEIQgn9lUvIUNQh1DfP52cOVyE72CpYHflCI,19599
12
+ spanmark-0.1.0.dist-info/WHEEL,sha256=zOwg4jB6zX2kU910N-cMawjivD6tO8NEWvE12je1bVk,87
13
+ spanmark-0.1.0.dist-info/licenses/LICENSE,sha256=OoaHuf_DLafMGytJruBa9ke4Gbv5DmhkfONcuV7E4so,1064
14
+ spanmark-0.1.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: hatchling 1.32.0
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 PD Hall
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.