sslabdata 3.1.0__tar.gz → 4.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. {sslabdata-3.1.0/sslabdata.egg-info → sslabdata-4.0.0}/PKG-INFO +33 -5
  2. {sslabdata-3.1.0 → sslabdata-4.0.0}/README.md +32 -4
  3. {sslabdata-3.1.0 → sslabdata-4.0.0}/SPEC.md +26 -19
  4. {sslabdata-3.1.0 → sslabdata-4.0.0}/pyproject.toml +3 -3
  5. sslabdata-4.0.0/schema/v6/output.schema.json +481 -0
  6. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/__init__.py +3 -2
  7. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/diagnostics.py +2 -0
  8. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/models.py +18 -2
  9. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/parsers/bibtex.py +91 -4
  10. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/templates/init/bib/publications.bib +4 -1
  11. {sslabdata-3.1.0 → sslabdata-4.0.0/sslabdata.egg-info}/PKG-INFO +33 -5
  12. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata.egg-info/SOURCES.txt +1 -0
  13. {sslabdata-3.1.0 → sslabdata-4.0.0}/LICENSE +0 -0
  14. {sslabdata-3.1.0 → sslabdata-4.0.0}/MANIFEST.in +0 -0
  15. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/__init__.py +0 -0
  16. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/input/v1/collaborators.schema.json +0 -0
  17. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/input/v1/lab.schema.json +0 -0
  18. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/input/v1/people.schema.json +0 -0
  19. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/input/v1/projects.schema.json +0 -0
  20. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/v3/output.schema.json +0 -0
  21. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/v4/output.schema.json +0 -0
  22. {sslabdata-3.1.0 → sslabdata-4.0.0}/schema/v5/output.schema.json +0 -0
  23. {sslabdata-3.1.0 → sslabdata-4.0.0}/setup.cfg +0 -0
  24. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/assembler.py +0 -0
  25. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/cli.py +0 -0
  26. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/config.py +0 -0
  27. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/exporters.py +0 -0
  28. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/loaders.py +0 -0
  29. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/parsers/__init__.py +0 -0
  30. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/parsers/latex.py +0 -0
  31. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/resolver.py +0 -0
  32. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/templates/init/collaborators.yaml +0 -0
  33. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/templates/init/lab.yaml +0 -0
  34. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/templates/init/people.yaml +0 -0
  35. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata/templates/init/projects.yaml +0 -0
  36. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata.egg-info/dependency_links.txt +0 -0
  37. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata.egg-info/entry_points.txt +0 -0
  38. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata.egg-info/requires.txt +0 -0
  39. {sslabdata-3.1.0 → sslabdata-4.0.0}/sslabdata.egg-info/top_level.txt +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: sslabdata
3
- Version: 3.1.0
3
+ Version: 4.0.0
4
4
  Summary: Renderer-agnostic academic lab data assembler: BibTeX + YAML → structured data
5
5
  Author: Siddhartha Srinivasa
6
6
  License-Expression: MIT
@@ -37,6 +37,11 @@ Dynamic: license-file
37
37
 
38
38
  # sslabdata
39
39
 
40
+ [![PyPI](https://img.shields.io/pypi/v/sslabdata.svg)](https://pypi.org/project/sslabdata/)
41
+ [![Python](https://img.shields.io/pypi/pyversions/sslabdata.svg)](https://pypi.org/project/sslabdata/)
42
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://github.com/siddhss5/sslabdata/blob/main/LICENSE)
43
+ [![Tests](https://github.com/siddhss5/sslabdata/actions/workflows/test.yml/badge.svg?branch=main)](https://github.com/siddhss5/sslabdata/actions/workflows/test.yml)
44
+
40
45
  sslabdata compiles BibTeX and a little YAML into one schema-specified document —
41
46
  works, people, projects and the links between them — that any website, CV or
42
47
  script can read.
@@ -56,10 +61,11 @@ sslabdata --config lab.yaml --output lab.yml
56
61
  order the lists are in, which fields are derived, when the version changes.
57
62
  - [`CHANGELOG.md`](https://github.com/siddhss5/sslabdata/blob/main/CHANGELOG.md) — what changed at each release, and what it
58
63
  replaced.
59
- - [`schema/v5/output.schema.json`](https://github.com/siddhss5/sslabdata/blob/main/schema/v5/output.schema.json) — the
64
+ - [`schema/v6/output.schema.json`](https://github.com/siddhss5/sslabdata/blob/main/schema/v6/output.schema.json) — the
60
65
  document's JSON Schema. Published versions are immutable and live at their
61
- own paths; [`schema/v3/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v3/output.schema.json) and
62
- [`schema/v4/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v4/output.schema.json) are still there.
66
+ own paths; [`schema/v3/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v3/output.schema.json),
67
+ [`schema/v4/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v4/output.schema.json) and
68
+ [`schema/v5/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v5/output.schema.json) are still there.
63
69
  - [`tests/COVERAGE.md`](https://github.com/siddhss5/sslabdata/blob/main/tests/COVERAGE.md) — every input case sslabdata
64
70
  supports, and every case it does not, with the fixture and test for each.
65
71
 
@@ -182,6 +188,7 @@ nothing else:
182
188
  | `doi`, `isbn`, `issn`, `eprint` + `archivePrefix` (or `eprinttype`) | `identifiers`, an open map from scheme to a list of identifiers, plus the links built from them. An `eprint`'s scheme is the repository `archivePrefix` or `eprinttype` named, lower-cased, so that field needs no property of its own — and an `eprint` in a repository other than arXiv gets no arXiv link |
183
189
  | `abstract` | `abstract` |
184
190
  | `note` | `note` |
191
+ | `award` | `awards`, a list of `{name, year}` (see below). `note` is never read for awards |
185
192
  | `url` | A link of kind `video` when its host is YouTube or Vimeo (or a subdomain of either), otherwise of kind `url` |
186
193
  | `video` | A link of kind `video`, whatever its host, so `url` can hold the work's website |
187
194
  | `pdf` | The work's one link of kind `pdf`. An entry without it gets `pdf_base_url` plus its citation key, when `pdf_base_url` is set |
@@ -192,6 +199,26 @@ The entry is also re-serialized into a `bibtex` field, so fields sslabdata does
192
199
  not interpret are still carried. It is a re-serialization, not a copy
193
200
  ([`SPEC.md` §5](https://github.com/siddhss5/sslabdata/blob/main/SPEC.md#5-input-versus-derived)).
194
201
 
202
+ ### Awards
203
+
204
+ A paper's awards go in its `award` field. Several are separated by `and`, as
205
+ the names in `author` are, and braces keep an `and` inside one name. An award
206
+ may start with the year it was given, as `YYYY:` and a space, when that is not
207
+ the paper's year:
208
+
209
+ ```bibtex
210
+ award = {2026: Test of Time Award and {Best Systems and Software Paper Award}}
211
+ ```
212
+
213
+ Each award becomes `{name, year}` in the work's `awards`, in the order
214
+ written, with its name converted from LaTeX as `title` is. An award without a
215
+ year takes the work's `year`, or `null` when the work has none. `awards` is
216
+ `[]` for a work with no `award` field. An empty award, or a prefix that looks
217
+ like a year and is not four digits, a colon and a space, is reported
218
+ (`BIB-AWARD-EMPTY`, `BIB-AWARD-YEAR-MALFORMED`); a malformed prefix stays in
219
+ the name. Awards a person holds, such as fellowships, are not part of the
220
+ document.
221
+
195
222
  ### The `project` tag
196
223
 
197
224
  sslabdata adds one custom BibTeX field, `project`, to link a paper to a research
@@ -205,6 +232,7 @@ project:
205
232
  year = {2024},
206
233
  eprint = {2406.99812},
207
234
  archivePrefix = {arXiv},
235
+ award = {Best Paper Award},
208
236
  project = {homebot}
209
237
  }
210
238
  ```
@@ -349,7 +377,7 @@ pip install jsonschema
349
377
  python -c "
350
378
  import json, yaml, jsonschema
351
379
  from importlib.resources import files
352
- schema = json.loads(files('sslabdata.schema').joinpath('v5/output.schema.json').read_text())
380
+ schema = json.loads(files('sslabdata.schema').joinpath('v6/output.schema.json').read_text())
353
381
  jsonschema.Draft202012Validator(schema).validate(yaml.safe_load(open('lab.yml')))
354
382
  print('valid')
355
383
  "
@@ -1,5 +1,10 @@
1
1
  # sslabdata
2
2
 
3
+ [![PyPI](https://img.shields.io/pypi/v/sslabdata.svg)](https://pypi.org/project/sslabdata/)
4
+ [![Python](https://img.shields.io/pypi/pyversions/sslabdata.svg)](https://pypi.org/project/sslabdata/)
5
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://github.com/siddhss5/sslabdata/blob/main/LICENSE)
6
+ [![Tests](https://github.com/siddhss5/sslabdata/actions/workflows/test.yml/badge.svg?branch=main)](https://github.com/siddhss5/sslabdata/actions/workflows/test.yml)
7
+
3
8
  sslabdata compiles BibTeX and a little YAML into one schema-specified document —
4
9
  works, people, projects and the links between them — that any website, CV or
5
10
  script can read.
@@ -19,10 +24,11 @@ sslabdata --config lab.yaml --output lab.yml
19
24
  order the lists are in, which fields are derived, when the version changes.
20
25
  - [`CHANGELOG.md`](https://github.com/siddhss5/sslabdata/blob/main/CHANGELOG.md) — what changed at each release, and what it
21
26
  replaced.
22
- - [`schema/v5/output.schema.json`](https://github.com/siddhss5/sslabdata/blob/main/schema/v5/output.schema.json) — the
27
+ - [`schema/v6/output.schema.json`](https://github.com/siddhss5/sslabdata/blob/main/schema/v6/output.schema.json) — the
23
28
  document's JSON Schema. Published versions are immutable and live at their
24
- own paths; [`schema/v3/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v3/output.schema.json) and
25
- [`schema/v4/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v4/output.schema.json) are still there.
29
+ own paths; [`schema/v3/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v3/output.schema.json),
30
+ [`schema/v4/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v4/output.schema.json) and
31
+ [`schema/v5/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v5/output.schema.json) are still there.
26
32
  - [`tests/COVERAGE.md`](https://github.com/siddhss5/sslabdata/blob/main/tests/COVERAGE.md) — every input case sslabdata
27
33
  supports, and every case it does not, with the fixture and test for each.
28
34
 
@@ -145,6 +151,7 @@ nothing else:
145
151
  | `doi`, `isbn`, `issn`, `eprint` + `archivePrefix` (or `eprinttype`) | `identifiers`, an open map from scheme to a list of identifiers, plus the links built from them. An `eprint`'s scheme is the repository `archivePrefix` or `eprinttype` named, lower-cased, so that field needs no property of its own — and an `eprint` in a repository other than arXiv gets no arXiv link |
146
152
  | `abstract` | `abstract` |
147
153
  | `note` | `note` |
154
+ | `award` | `awards`, a list of `{name, year}` (see below). `note` is never read for awards |
148
155
  | `url` | A link of kind `video` when its host is YouTube or Vimeo (or a subdomain of either), otherwise of kind `url` |
149
156
  | `video` | A link of kind `video`, whatever its host, so `url` can hold the work's website |
150
157
  | `pdf` | The work's one link of kind `pdf`. An entry without it gets `pdf_base_url` plus its citation key, when `pdf_base_url` is set |
@@ -155,6 +162,26 @@ The entry is also re-serialized into a `bibtex` field, so fields sslabdata does
155
162
  not interpret are still carried. It is a re-serialization, not a copy
156
163
  ([`SPEC.md` §5](https://github.com/siddhss5/sslabdata/blob/main/SPEC.md#5-input-versus-derived)).
157
164
 
165
+ ### Awards
166
+
167
+ A paper's awards go in its `award` field. Several are separated by `and`, as
168
+ the names in `author` are, and braces keep an `and` inside one name. An award
169
+ may start with the year it was given, as `YYYY:` and a space, when that is not
170
+ the paper's year:
171
+
172
+ ```bibtex
173
+ award = {2026: Test of Time Award and {Best Systems and Software Paper Award}}
174
+ ```
175
+
176
+ Each award becomes `{name, year}` in the work's `awards`, in the order
177
+ written, with its name converted from LaTeX as `title` is. An award without a
178
+ year takes the work's `year`, or `null` when the work has none. `awards` is
179
+ `[]` for a work with no `award` field. An empty award, or a prefix that looks
180
+ like a year and is not four digits, a colon and a space, is reported
181
+ (`BIB-AWARD-EMPTY`, `BIB-AWARD-YEAR-MALFORMED`); a malformed prefix stays in
182
+ the name. Awards a person holds, such as fellowships, are not part of the
183
+ document.
184
+
158
185
  ### The `project` tag
159
186
 
160
187
  sslabdata adds one custom BibTeX field, `project`, to link a paper to a research
@@ -168,6 +195,7 @@ project:
168
195
  year = {2024},
169
196
  eprint = {2406.99812},
170
197
  archivePrefix = {arXiv},
198
+ award = {Best Paper Award},
171
199
  project = {homebot}
172
200
  }
173
201
  ```
@@ -312,7 +340,7 @@ pip install jsonschema
312
340
  python -c "
313
341
  import json, yaml, jsonschema
314
342
  from importlib.resources import files
315
- schema = json.loads(files('sslabdata.schema').joinpath('v5/output.schema.json').read_text())
343
+ schema = json.loads(files('sslabdata.schema').joinpath('v6/output.schema.json').read_text())
316
344
  jsonschema.Draft202012Validator(schema).validate(yaml.safe_load(open('lab.yml')))
317
345
  print('valid')
318
346
  "
@@ -7,7 +7,7 @@ them.
7
7
  This file states the parts of that contract a JSON Schema cannot express:
8
8
  what the strings in the document are, what order the lists are in, what an
9
9
  absent key means, which fields are computed, when the version changes, and
10
- how a repeated `@string` macro resolves. `schema/v5/output.schema.json`
10
+ how a repeated `@string` macro resolves. `schema/v6/output.schema.json`
11
11
  states the rest.
12
12
 
13
13
  Everything here is normative unless it carries a `Target` note. A `Target`
@@ -16,8 +16,8 @@ names the issue that will make it true. Until that issue lands, the rule is
16
16
  the intent and the note is the fact. What changed at each release, and what
17
17
  it replaced, is in [`CHANGELOG.md`](CHANGELOG.md), not here.
18
18
 
19
- - Applies to: `schema_version` 5 (`sslabdata.models.SCHEMA_VERSION`), package
20
- version 3.0.0 (`sslabdata.__version__`).
19
+ - Applies to: `schema_version` 6 (`sslabdata.models.SCHEMA_VERSION`), package
20
+ version 4.0.0 (`sslabdata.__version__`).
21
21
 
22
22
  ### How this file cites the code
23
23
 
@@ -25,7 +25,7 @@ Every rule below is grounded in a named part of the code rather than a line
25
25
  number, because line numbers rot silently: a function such as
26
26
  `sslabdata.parsers.bibtex.parse_all_works()`, a method such as
27
27
  `Person.to_dict()`, a module-level constant such as `TEXT_FIELDS`, a JSON
28
- Pointer into `schema/v5/output.schema.json` such as `/$defs/person/required`, or
28
+ Pointer into `schema/v6/output.schema.json` such as `/$defs/person/required`, or
29
29
  a `tests/COVERAGE.md` row key such as `config.people_file.missing`. A bare
30
30
  statement is cited by its enclosing function.
31
31
 
@@ -201,7 +201,7 @@ without depending on English wording. Codes obey three rules:
201
201
  | **Fatal at load** | `Error loading configuration: <CODE> …` (`Error: <CODE> …` for `CONFIG-NOT-FOUND`) on standard error; exits `1` before anything is compiled, so there is no report. | The same. | `CONFIG-BIB-FILE-ABSOLUTE`, `CONFIG-BIB-FILE-OUTSIDE-BIB-DIR`, `CONFIG-NOT-A-MAPPING`, `CONFIG-KEY-MISSING`, `CONFIG-TYPE-INVALID`, `CONFIG-VALUE-NOT-JSON`, `CONFIG-KEY-REPEATED`, `CONFIG-NOT-FOUND`, `CONFIG-UNREADABLE` |
202
202
  | **Fatal** | Listed under `Bibliography errors` and counted; exits `1`. | Written to standard error unprefixed; exits `1`, and `--output` writes nothing. | `BIB-CROSSREF-UNSUPPORTED`, `BIB-ENCODING-INVALID`, `CONFIG-FILE-NOT-FOUND`, `CONFIG-PATH-WRONG-KIND`, `PEOPLE-YAML-INVALID`, `PEOPLE-NOT-A-LIST`, `PEOPLE-FIELD-MISSING`, `PROJECTS-YAML-INVALID`, `PROJECTS-NOT-A-LIST`, `PROJECTS-FIELD-MISSING`, `COLLABORATORS-YAML-INVALID`, `COLLABORATORS-NOT-A-LIST`, `COLLABORATORS-FIELD-MISSING`, `RECORD-KEY-REPEATED`, `OUTPUT-WRITE-FAILED`, `INIT-FILE-EXISTS`, `INIT-PATH-WRONG-KIND`, `INIT-PATH-OUTSIDE-DIR`, `INIT-WRITE-FAILED` |
203
203
  | **Validation error** | Listed under `Bibliography errors` and counted; exits `1`. | Prefixed `Warning: ` on standard error; the run continues and exits `0`. | `BIB-DUPLICATE-KEY`, `RESOLVE-PROJECT-UNKNOWN`, `PEOPLE-ID-DUPLICATE`, `PROJECTS-ID-DUPLICATE` |
204
- | **Warning** | Listed under `Warnings`; not counted, and does not change the exit code. | Prefixed `Warning: ` on standard error; the run continues. | `BIB-YEAR-MISSING`, `BIB-YEAR-INVALID`, `BIB-DOI-INVALID`, `BIB-OTHERS-NOT-LAST`, `BIB-STRING-UNDEFINED`, `BIB-SYNTAX-ERROR`, `BIB-BRACE-MISMATCH`, `BIB-COMMENTED-COMMAND-READ`, `BIB-VENUE-MISSING`, `BIB-ENTRY-TYPE-UNSUPPORTED`, `LATEX-COMMAND-UNKNOWN`, `ID-GROUPING-SPANS-SPELLINGS`, `ID-GROUPING-INITIALS-AMBIGUOUS`, `RESOLVE-AMBIGUOUS-NAME`, `RESOLVE-SUGGESTION`, `RESOLVE-COLLABORATOR-ALIAS-IS-MEMBER`, `PEOPLE-ALIAS-AMBIGUOUS`, `PEOPLE-ROLE-INVALID`, `PEOPLE-STATUS-INVALID`, `PROJECTS-STATUS-INVALID`, `CONFIG-LAB-NAME-MISSING`, `CONFIG-KEY-UNKNOWN`, `RECORD-KEY-UNKNOWN`, `RECORD-TYPE-INVALID`, `CONFIG-BIB-FILES-MISSING`, `BIB-PARSER-MESSAGE`, `LATEX-CONVERSION-FAILED`, `LINK-SCHEME-UNSUPPORTED`, `TEXT-CONTROL-CHARACTER`, `BIB-WRITE-BACK-FAILED`, `BIB-STRING-REDEFINED`, `ID-GROUPING-AMBIGUOUS-DECLARED`, `RESOLVE-UNRESOLVED-NAME` |
204
+ | **Warning** | Listed under `Warnings`; not counted, and does not change the exit code. | Prefixed `Warning: ` on standard error; the run continues. | `BIB-YEAR-MISSING`, `BIB-YEAR-INVALID`, `BIB-DOI-INVALID`, `BIB-AWARD-EMPTY`, `BIB-AWARD-YEAR-MALFORMED`, `BIB-OTHERS-NOT-LAST`, `BIB-STRING-UNDEFINED`, `BIB-SYNTAX-ERROR`, `BIB-BRACE-MISMATCH`, `BIB-COMMENTED-COMMAND-READ`, `BIB-VENUE-MISSING`, `BIB-ENTRY-TYPE-UNSUPPORTED`, `LATEX-COMMAND-UNKNOWN`, `ID-GROUPING-SPANS-SPELLINGS`, `ID-GROUPING-INITIALS-AMBIGUOUS`, `RESOLVE-AMBIGUOUS-NAME`, `RESOLVE-SUGGESTION`, `RESOLVE-COLLABORATOR-ALIAS-IS-MEMBER`, `PEOPLE-ALIAS-AMBIGUOUS`, `PEOPLE-ROLE-INVALID`, `PEOPLE-STATUS-INVALID`, `PROJECTS-STATUS-INVALID`, `CONFIG-LAB-NAME-MISSING`, `CONFIG-KEY-UNKNOWN`, `RECORD-KEY-UNKNOWN`, `RECORD-TYPE-INVALID`, `CONFIG-BIB-FILES-MISSING`, `BIB-PARSER-MESSAGE`, `LATEX-CONVERSION-FAILED`, `LINK-SCHEME-UNSUPPORTED`, `TEXT-CONTROL-CHARACTER`, `BIB-WRITE-BACK-FAILED`, `BIB-STRING-REDEFINED`, `ID-GROUPING-AMBIGUOUS-DECLARED`, `RESOLVE-UNRESOLVED-NAME` |
205
205
 
206
206
  The same code always carries the same class. What varies with the mode is
207
207
  how the run reacts to it, which is why the class is not in the code, and
@@ -249,6 +249,8 @@ Codes in use:
249
249
  | `CONFIG-LAB-NAME-MISSING` | The `lab` header declares no `name`. A `lab` that is not a mapping at all is a different condition and is not reported under this code. A warning. |
250
250
  | `BIB-YEAR-INVALID` | An entry's `year` is present but is not an unsigned run of the ASCII digits `0`–`9` (`sslabdata.parsers.bibtex.YEAR_DIGITS`), such as `in press`, or `-5`, `+2020`, `2_020` and full-width `2020`, which Python's `int()` would read as numbers. The work is emitted with `year: null` and sorts last, as one with no year does. A warning. |
251
251
  | `BIB-DOI-INVALID` | An entry's `doi` is a DOI resolver URL with nothing after it, such as `https://doi.org/` (`sslabdata.parsers.bibtex.DOI_RESOLVERS`), so it names no DOI. Located at `<file>:<key>:doi`, naming the value. The work gets no DOI identifier and no link of kind `doi`, rather than an empty one; the entry is kept. A `doi` that is empty or blank is read as absent, and is not reported. A warning. |
252
+ | `BIB-AWARD-EMPTY` | An entry's `award` field names no award: the field is empty or whitespace alone, or one of its awards is empty — two `and`s with nothing between them, an `and` that opens or closes the list, or a name that is empty once converted from LaTeX, such as `{}` — or is a year prefix with no name after it, such as `2026:` (`sslabdata.parsers.bibtex.parse_awards()`). Located at `<file>:<key>:award`; the prose names the award's place in the list, and the award as written when it has a year. The empty award is dropped and the others are kept, so an empty field gives `awards: []`. A warning. |
253
+ | `BIB-AWARD-YEAR-MALFORMED` | An award in an entry's `award` field starts with what reads as a year but is not the prefix `YYYY:` — four ASCII digits, a colon and whitespace (`sslabdata.parsers.bibtex.AWARD_YEAR`): digits of any kind and a colon, such as `26: …`, `2026:…` with no space, or full-width `2026: …`, or four digits and a space with no colon, `2026 …` (`AWARD_YEAR_LIKE`). Located at `<file>:<key>:award`, naming the award as written and the part that is not a prefix. No year is read from it: **the text is kept as part of the name**, and the award takes the work's `year`. A prefix inside braces, `{2026}: …`, is text the author protected and is not reported. A warning. |
252
254
  | `LINK-SCHEME-UNSUPPORTED` | A link in `work.links` whose URL scheme is not `http`, `https` or `mailto` (`sslabdata.parsers.bibtex.LINK_SCHEMES`), such as `javascript:`, `data:` or `vbscript:`, which a renderer escaping the document (§2) should refuse to turn into a link. The scheme is the one Python's URL parser reads (`urllib.parse.urlsplit()`, in `sslabdata.parsers.bibtex.url_scheme()`) once surrounding whitespace is off, compared without regard to case; the parser drops a tab or line break anywhere in the URL first, as a browser does, so `java<tab>script:` is `javascript`. A URL with no scheme is a relative path and is not reported; that includes a protocol-relative `//host/path`, which takes the scheme of the page that links it. A scheme of one letter is a Windows drive, read by Windows rules as `sslabdata.config` reads one (`PureWindowsPath`), and not a scheme: `C:/papers`, `C:\papers` and the drive-relative `c:papers` are local paths and are not reported. No registered scheme is one letter long. Every link is checked, whether the entry wrote it (`url`, `video`, `pdf`) or sslabdata built it (`doi`, `arxiv`, and `pdf` from `pdf_base_url`). Located at `<file>:<key>:<field>`, the field the link came from, or at `<file>:<key>:` for the PDF link `pdf_base_url` builds, which no field of the entry wrote; the prose names the link's kind, which for `url` can be `video`, and the URL. Once per link. A person's `website` and `photo`, a project's `website` and `image`, and everything under `lab` are fields, not links, and are not checked. **The link is emitted unchanged**: this code only reports it. A warning. |
253
255
  | `BIB-OTHERS-NOT-LAST` | An `author` or `editor` list has `and others` somewhere other than at its end, the one place BibTeX reads it as "et al.". It names nobody there either, so it is dropped as a terminal one is, and the names around it are kept in their order, with `position` counting only them. Located at `<file>:<key>:author` or `:editor`, once per list however many times it occurs. A terminal `and others` is not reported. A warning. |
254
256
  | `BIB-STRING-UNDEFINED` | A field value names an `@string` macro that nothing defined earlier in the same file. Located at the entry and field that use it, and naming the macro. It is read as empty, as BibTeX reads it, and the entry and its neighbours are kept. A macro used inside another `@string` definition is located at the file alone. A warning. |
@@ -359,7 +361,7 @@ What each mode puts in the array:
359
361
 
360
362
  **The Python API is convenience only.** Public: the names in `sslabdata.__all__`
361
363
  — `assemble`, `AssemblyResult`, `AssemblyError`, the models `LabData`,
362
- `Work`, `Author`, `Contributor`, `Venue`, `Link`, `Person`, `Project`,
364
+ `Work`, `Award`, `Author`, `Contributor`, `Venue`, `Link`, `Person`, `Project`,
363
365
  `Collaborator`, the config loader `LabDataConfig` with `BibFile`, and the exporters
364
366
  `export_to_yaml` and `export_to_json`, and the exception
365
367
  `ConfigurationError`.
@@ -521,7 +523,8 @@ the URL, and in math it is escaped to `\&` or `\%`. Exactly the fields in
521
523
  `publisher`, `address`, `organization`, `archivePrefix` and `eprinttype` —
522
524
  applied in `entry_fields()`. Name
523
525
  parts are converted the same way, in `person_name_parts()`, for authors and
524
- editors alike.
526
+ editors alike, and so is each award's name, once `award` is split into its
527
+ awards and each award's year prefix is read (`parse_awards()`).
525
528
 
526
529
  Being converted is not the same as being emitted. Of these fields, `title`,
527
530
  `abstract`, `note`, `type`, `series`, `publisher`, `address` and
@@ -581,6 +584,7 @@ rule does not apply to the input itself — only to whatever it produces.
581
584
  | `pdf` | Becomes the work's one link of kind `pdf`, with `origin: input`, in place of the one `pdf_base_url` would give (`build_links()`). Empty or whitespace-only is read as absent. |
582
585
  | `author` | Parsed into the `authors` list (`parse_author_list()`); the name parts are converted under heading 1. |
583
586
  | `editor` | Parsed into the `editors` list (`parse_editor_list()`), resolved by the same machinery, and excluded from `person.work_ids`, from a project's people and from `collaborators`. |
587
+ | `award` | Parsed into the `awards` list (`parse_awards()`), and emitted nowhere else: `note` is never read for awards, and an award in `note` stays there as text. The field is split into awards as BibTeX splits `author` into names: on `and`, in any case, with whitespace on both sides, outside braces, which are counted as BibTeX counts them, so `{Systems and Software Award}` is one award. An `and` that opens or closes the list, or follows another, separates an empty award, which is dropped and reported (`BIB-AWARD-EMPTY`). Each award may then start with the year it was given, four ASCII digits, a colon and whitespace, `2026: Test of Time Award` (`AWARD_YEAR`), read before the name is converted from LaTeX under heading 1 and trimmed; it is that award's `year`, which need not be the work's. An award without one takes the work's `year`, or `null` when the work has none. A prefix that reads as a year and is not one, `26:` or `2026 …`, is kept in the name and reported (`BIB-AWARD-YEAR-MALFORMED`). |
584
588
  | `year` | Emitted as the integer `year` — not a string — or `null` with a `BIB-YEAR-MISSING` diagnostic when the entry supplied none, or with `BIB-YEAR-INVALID` when it is not written in the digits `0`–`9` alone. It drives the works order (§3). |
585
589
  | `crossref` | **Rejected, on presence rather than on value.** An entry carrying the field is an error under `BIB-CROSSREF-UNSUPPORTED`, whatever is inside it: an empty `crossref = {}` is a field the entry carries, and letting it through would put the silent path back under a different spelling. The entry is not emitted and the run fails in every mode (`parse_all_works()`). No field of any entry is filled in from any other entry. |
586
590
  | `journal`, `booktitle`, `school`, `institution` | Converted under heading 1, then consumed by `build_venue()` into `venue.name`, with the `venue.kind` each implies. |
@@ -657,6 +661,7 @@ accident.
657
661
  | `works` | `year` **descending**, with works that have no year **last**. Ties keep *read order* (below). The sort is the final statement of `sslabdata.parsers.bibtex.parse_all_works()`. |
658
662
  | `work.authors` | The order the `author` field wrote them (`sslabdata.parsers.bibtex.parse_author_list()`). A terminal `and others` is BibTeX's "et al." and is dropped rather than emitted as an author; one anywhere else is dropped too, and reported as `BIB-OTHERS-NOT-LAST`. `position` is that order, 1-based, and counts only the names that reach the document. |
659
663
  | `work.editors` | The order the `editor` field wrote them, read the same way (`parse_editor_list()`). |
664
+ | `work.awards` | The order the `award` field wrote them, with each empty award dropped (`parse_awards()`). |
660
665
  | `work.project_ids` | The order the `project` field wrote them, comma-separated, whitespace trimmed, empty entries dropped (`sslabdata.parsers.bibtex.parse_project_ids()`). |
661
666
  | `people` | The order of `people_file`. sslabdata does not sort people (`sslabdata.loaders.load_people()`, called by `sslabdata.assembler.assemble()`). |
662
667
  | `projects` | The order of `projects_file`, likewise (`sslabdata.loaders.load_projects()`). |
@@ -752,7 +757,7 @@ unconditionally, and by `LabData.to_dict()`, which always emits `lab`.
752
757
  itself and each entity type as **closed**: `/additionalProperties` and each
753
758
  of `/$defs/authorship`, `/$defs/editorship`, `/$defs/work`, `/$defs/person`,
754
759
  `/$defs/project`, `/$defs/collaborator`, `/$defs/venue`, `/$defs/link`,
755
- `/$defs/resolution`, and the `source` and `generator` objects, set
760
+ `/$defs/resolution`, `/$defs/award`, and the `source` and `generator` objects, set
756
761
  `additionalProperties: false`. Within those, a property that is not declared
757
762
  cannot appear, and "absent" always means a declared property with no value.
758
763
 
@@ -783,6 +788,7 @@ sslabdata's own output as input, and a wrong derivation becomes permanent.
783
788
  | `work.source.file` | Input — the `name` of the `bib_files` entry the file was listed under, **never an absolute path**. A *relative* directory is fine and is passed through as written: `sub/journal.bib` is a name under `bib_dir`; one that leaves `bib_dir` is rejected under `CONFIG-BIB-FILE-OUTSIDE-BIB-DIR`. The guarantee is kept by rejecting the input rather than by rewriting it, which would quietly discard that directory, and is enforced under `CONFIG-BIB-FILE-ABSOLUTE` (§1, *`sslabdata.ConfigurationError`*). |
784
789
  | `work.entry_type` | Input — the BibTeX entry type, **lowercased** by `entry_fields()`. Of it and `bib_id`, it is the only one that is case-folded. |
785
790
  | `work.title`, `abstract`, `note` | Input — BibTeX fields, converted from LaTeX to text (§2). `note` additionally has trailing `.` and whitespace trimmed (`sslabdata.parsers.bibtex.extract_note()`). |
791
+ | `work.awards` | Input — the BibTeX `award` field, split into awards, each `{name, year}`: `name` converted from LaTeX, and `year` the year the award was given, from its `YYYY:` prefix, else the work's `year`, else `null` (`sslabdata.parsers.bibtex.parse_awards()`, §2). `[]` when the entry writes none. These are a work's awards only; an award a person holds, such as a fellowship, is not in the document. |
786
792
  | `work.year` | Input — the BibTeX `year`, as an integer; `null` when the entry supplied none, with a diagnostic (`entry_year()`). |
787
793
  | `work.category` | Input — the `category` of the `bib_files` entry the file was listed under, not anything in the `.bib` file (`sslabdata.config.BibFile`, read by `parse_all_works()`). |
788
794
  | `work.venue` | **Derived** — the first of `journal`, `booktitle`, `school` and `institution` the entry wrote, as `name`, with the `kind` that field and the entry type imply; a preprint's repository when the entry has only an `eprint`; `null` when it names no container (`sslabdata.parsers.bibtex.build_venue()`). See below. |
@@ -1114,13 +1120,14 @@ meaning and guarantees, and each of them is a bump.
1114
1120
  has been published is never edited. Version `N`'s schema stays reachable, byte
1115
1121
  for byte, at its own path after version `N+1` ships, so a consumer pinned to
1116
1122
  `N` keeps a stable target. `schema/v3/output.schema.json`,
1117
- `schema/v4/output.schema.json` and `schema/v5/output.schema.json` are those
1118
- paths, and `tests/COVERAGE.md` row `output.versioned_schema` asserts that the
1119
- older two are unchanged byte for byte and still say `3` and `4`.
1120
-
1121
- **The `$id` is a pinned tag URL.** v5's `$id` is
1122
- `https://raw.githubusercontent.com/siddhss5/sslabdata/schema-v5/schema/v5/output.schema.json`.
1123
- The rule that makes it a contract rather than a guess: **the `schema-v5` tag
1123
+ `schema/v4/output.schema.json`, `schema/v5/output.schema.json` and
1124
+ `schema/v6/output.schema.json` are those paths, and `tests/COVERAGE.md` row
1125
+ `output.versioned_schema` asserts that the older three are unchanged byte for
1126
+ byte and still say `3`, `4` and `5`.
1127
+
1128
+ **The `$id` is a pinned tag URL.** v6's `$id` is
1129
+ `https://raw.githubusercontent.com/siddhss5/sslabdata/schema-v6/schema/v6/output.schema.json`.
1130
+ The rule that makes it a contract rather than a guess: **the `schema-v6` tag
1124
1131
  is created when this version ships and is never moved.** A branch URL such as
1125
1132
  `blob/main` is not usable — it serves an HTML page rather than the schema, so
1126
1133
  no consumer can ever have resolved v3's `$id` — and this repository publishes
@@ -1135,8 +1142,8 @@ rule.
1135
1142
  descriptions inside those files, name the repository `labdata`, and the files
1136
1143
  are left byte for byte as published rather than rewritten. v4's raw `$id`
1137
1144
  resolves, because GitHub redirects `labdata` to `sslabdata`. v3's `blob/main`
1138
- `$id` does not resolve, as above. v5's `$id`, title and descriptions say
1139
- `sslabdata`.
1145
+ `$id` does not resolve, as above. The `$id`s, titles and descriptions of v5
1146
+ and v6 say `sslabdata`.
1140
1147
 
1141
1148
  **The input schemas are versioned apart from the document.**
1142
1149
  `schema/input/v1/` holds a JSON Schema for each input file: `lab.schema.json`,
@@ -1243,7 +1250,7 @@ So each rejected type is rejected for a stated reason, and each has a home:
1243
1250
  | Teaching and courses | A prose page in your site repository, or the institution's course catalogue. |
1244
1251
  | Press and media coverage | A typed link on the work it covers: `links` is an open map from kind to links, so #27 adds the kind without a version bump. |
1245
1252
  | Galleries, photos and videos | Your site repository; a video already reaches the document as a link of kind `video`. |
1246
- | Awards and honours | An attribute of the work, today `work.note`; #27 moves it out of `note`. |
1253
+ | Awards and honours | A paper's awards are an attribute of the work, `work.awards`, read from its `award` field (§5). An award a person holds, such as a fellowship or an endowed chair, belongs in your site repository. |
1247
1254
  | Funding and grants | Your site repository. Nothing in the document depends on it. |
1248
1255
  | Software and datasets | Not a separate collection — they are kinds of *work*, added by #31. |
1249
1256
  | Alumni | Not a collection — a `status` on a person (`sslabdata.models.Person.status`). |
@@ -1260,7 +1267,7 @@ vocabularies, not over arbitrary content.
1260
1267
 
1261
1268
  ## 9. What this file is not
1262
1269
 
1263
- It does not list the document's fields; `schema/v5/output.schema.json` does.
1270
+ It does not list the document's fields; `schema/v6/output.schema.json` does.
1264
1271
  It does not describe renderers such as
1265
1272
  [sslabdata-site](https://github.com/siddhss5/sslabdata-site), which are
1266
1273
  downstream consumers in their own repositories. It does not describe the input formats `lab.yaml`,
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "sslabdata"
3
- version = "3.1.0"
3
+ version = "4.0.0"
4
4
  description = "Renderer-agnostic academic lab data assembler: BibTeX + YAML → structured data"
5
5
  readme = "README.md"
6
6
  requires-python = ">=3.10"
@@ -58,7 +58,7 @@ requires = ["setuptools>=77"]
58
58
  build-backend = "setuptools.build_meta"
59
59
 
60
60
  # The current output schema and input schemas are public wheel data, installed
61
- # at sslabdata/schema/v5/output.schema.json and
61
+ # at sslabdata/schema/v6/output.schema.json and
62
62
  # sslabdata/schema/input/v1/*.schema.json. The package directory is mapped onto
63
63
  # the repository's schema/ directory so each frozen file has one copy and one
64
64
  # path; earlier versions stay in the sdist only. schema/__init__.py makes the
@@ -71,7 +71,7 @@ packages = ["sslabdata", "sslabdata.parsers", "sslabdata.schema"]
71
71
  "sslabdata.schema" = "schema"
72
72
 
73
73
  [tool.setuptools.package-data]
74
- "sslabdata.schema" = ["v5/output.schema.json", "input/v1/*.schema.json"]
74
+ "sslabdata.schema" = ["v6/output.schema.json", "input/v1/*.schema.json"]
75
75
  # The starting point `sslabdata init` copies.
76
76
  "sslabdata" = ["templates/init/*.yaml", "templates/init/bib/*.bib"]
77
77
 
@@ -0,0 +1,481 @@
1
+ {
2
+ "$schema": "https://json-schema.org/draft/2020-12/schema",
3
+ "$id": "https://raw.githubusercontent.com/siddhss5/sslabdata/schema-v6/schema/v6/output.schema.json",
4
+ "title": "sslabdata output",
5
+ "description": "The lab.yml / lab.json file written by `sslabdata --output`. YAML and JSON exports hold the same data. Published schemas are immutable: this one is served from the `schema-v6` tag, which is created when this version ships and is never moved.",
6
+ "type": "object",
7
+ "required": [
8
+ "schema_version", "generator", "lab", "works", "people", "projects",
9
+ "collaborators"
10
+ ],
11
+ "additionalProperties": false,
12
+ "properties": {
13
+ "schema_version": {
14
+ "description": "Version of this output format. Bumped when a change could break a consumer.",
15
+ "const": 6
16
+ },
17
+ "generator": {
18
+ "description": "What compiled this document. It carries no timestamp: a build timestamp would make every run differ, and git records when.",
19
+ "type": "object",
20
+ "required": ["name", "version", "schema_version"],
21
+ "additionalProperties": false,
22
+ "properties": {
23
+ "name": { "type": "string", "minLength": 1 },
24
+ "version": {
25
+ "description": "The sslabdata package version. Independent of schema_version; neither can be derived from the other.",
26
+ "type": "string",
27
+ "minLength": 1
28
+ },
29
+ "schema_version": { "const": 6 }
30
+ }
31
+ },
32
+ "lab": {
33
+ "description": "The lab section of lab.yaml, copied through unchanged and always emitted, so an empty header is distinguishable from no header. The properties below are declared for documentation only: this object stays open, and a consumer must probe for the keys it wants rather than expect a fixed set.",
34
+ "type": "object",
35
+ "additionalProperties": true,
36
+ "properties": {
37
+ "name": { "type": "string" },
38
+ "description": { "type": "string" },
39
+ "institution": { "type": "string" },
40
+ "department": { "type": "string" },
41
+ "website": { "type": "string" },
42
+ "email": { "type": "string" },
43
+ "address": { "type": "string" },
44
+ "logo": { "type": "string" },
45
+ "links": { "type": "object" }
46
+ }
47
+ },
48
+ "works": {
49
+ "description": "Everything the lab produced that carries a citation key. Ordered by year descending, works with no year last; ties keep read order.",
50
+ "type": "array",
51
+ "items": { "$ref": "#/$defs/work" }
52
+ },
53
+ "people": {
54
+ "type": "array",
55
+ "items": { "$ref": "#/$defs/person" }
56
+ },
57
+ "projects": {
58
+ "type": "array",
59
+ "items": { "$ref": "#/$defs/project" }
60
+ },
61
+ "collaborators": {
62
+ "description": "A derived grouping over the authorships that resolved to nobody in people.yaml. Not a list of people: see /$defs/collaborator.",
63
+ "type": "array",
64
+ "items": { "$ref": "#/$defs/collaborator" }
65
+ }
66
+ },
67
+ "$defs": {
68
+ "optionalString": { "type": ["string", "null"] },
69
+ "idList": {
70
+ "type": "array",
71
+ "items": { "type": "string", "minLength": 1 }
72
+ },
73
+ "derived": {
74
+ "description": "Reserved for sslabdata. Every key inside it is sslabdata's; consumers tolerate keys they do not know and must not write their own. A key graduating out of it into a named property is still a major change. It is not a general extension mechanism.",
75
+ "type": "object",
76
+ "additionalProperties": true
77
+ },
78
+ "venue": {
79
+ "description": "Where a work appeared. The one place sslabdata normalises across entry types, collapsing `journal`, `booktitle`, `school` and `institution` to one name plus a kind.",
80
+ "type": "object",
81
+ "required": ["kind", "name"],
82
+ "additionalProperties": false,
83
+ "properties": {
84
+ "kind": {
85
+ "description": "An open string, deliberately not an enum so that new work types need no version bump. Documented values: `journal`, `conference`, `book`, `institution`, `repository`, `other`.",
86
+ "type": "string",
87
+ "minLength": 1
88
+ },
89
+ "name": { "type": "string", "minLength": 1 }
90
+ }
91
+ },
92
+ "link": {
93
+ "description": "One URL a work can be reached at. A link that fails verification is kept and labelled, never deleted, so a consumer can tell a broken link from an absent one.",
94
+ "type": "object",
95
+ "required": ["url", "label", "origin", "verification"],
96
+ "additionalProperties": false,
97
+ "properties": {
98
+ "url": { "type": "string", "minLength": 1 },
99
+ "label": { "$ref": "#/$defs/optionalString" },
100
+ "origin": {
101
+ "description": "Where the link came from. An open string; documented values are `input` (the entry's own field), `sidecar`, `enrichment`, `inferred` and `derived` (sslabdata built it). Only `input` means the entry supplied it.",
102
+ "type": "string",
103
+ "minLength": 1
104
+ },
105
+ "verification": {
106
+ "type": "object",
107
+ "required": ["status"],
108
+ "additionalProperties": false,
109
+ "properties": {
110
+ "status": {
111
+ "description": "`unchecked` when nothing has looked, `verified` when a local filesystem check or a committed cache found it, `missing` when one looked and it was not there. A build never fetches.",
112
+ "enum": ["unchecked", "verified", "missing"]
113
+ }
114
+ }
115
+ }
116
+ }
117
+ },
118
+ "linkMap": {
119
+ "description": "An open map from link kind to a list of links, so that two code repositories or two videos can both be carried. Documented kinds: `pdf`, `doi`, `arxiv`, `url`, `video`. Unordered; consumers filter.",
120
+ "type": "object",
121
+ "additionalProperties": {
122
+ "type": "array",
123
+ "items": { "$ref": "#/$defs/link" }
124
+ }
125
+ },
126
+ "identifierMap": {
127
+ "description": "An open map from identifier scheme to a unique list of identifiers, for example {\"doi\": [\"10.5555/...\"], \"arxiv\": [\"2501.01234\"]}. Schemes are documented, not enumerated. A list because ISBN and ISSN genuinely repeat. Unordered; consumers filter.",
128
+ "type": "object",
129
+ "additionalProperties": {
130
+ "type": "array",
131
+ "items": { "type": "string", "minLength": 1 },
132
+ "uniqueItems": true
133
+ }
134
+ },
135
+ "resolution": {
136
+ "description": "Whether this contributor was matched to a person, and how. Both are open strings: #24 adds `ambiguous` to the status and #25 fills the method with an override or an ORCID, and neither is breaking.",
137
+ "type": "object",
138
+ "required": ["status", "method"],
139
+ "additionalProperties": false,
140
+ "properties": {
141
+ "status": {
142
+ "description": "Documented values: `resolved`, `unresolved`.",
143
+ "type": "string",
144
+ "minLength": 1
145
+ },
146
+ "method": {
147
+ "description": "Documented values: `exact`, `fuzzy`, or null when nothing matched.",
148
+ "$ref": "#/$defs/optionalString"
149
+ }
150
+ }
151
+ },
152
+ "nameParts": {
153
+ "description": "The name properties every contributor record carries: the readable form, the parts BibTeX split it into, its position in the list, the person it resolved to, and how. Referenced in place, so each record below still declares every key it allows -- `additionalProperties: false` sees only its own `properties`.",
154
+ "type": "object",
155
+ "required": [
156
+ "name", "position", "person_id", "given", "von", "family", "suffix",
157
+ "literal", "resolution", "derived"
158
+ ],
159
+ "properties": {
160
+ "name": {
161
+ "description": "The structured parts joined in reading order. A readable form of the input name, not a citation form: it does not abbreviate, expand or normalise.",
162
+ "type": "string",
163
+ "minLength": 1
164
+ },
165
+ "position": {
166
+ "description": "1-based position in this work's list. With the work's bib_id it addresses the record.",
167
+ "type": "integer",
168
+ "minimum": 1
169
+ },
170
+ "person_id": {
171
+ "description": "The id of the matching person in people.yaml, or null. It means nothing else.",
172
+ "type": ["string", "null"]
173
+ },
174
+ "given": {
175
+ "description": "Given and middle names, as the entry supplied them: an entry writing `Brown, B.` yields `B.`, which is correct rather than a gap. Null for a name written as one brace-protected unit.",
176
+ "$ref": "#/$defs/optionalString"
177
+ },
178
+ "von": {
179
+ "description": "Surname particles such as `van den`. Null when the name has none.",
180
+ "$ref": "#/$defs/optionalString"
181
+ },
182
+ "family": {
183
+ "description": "The surname, without particles or suffix. Null for a name written as one brace-protected unit.",
184
+ "$ref": "#/$defs/optionalString"
185
+ },
186
+ "suffix": {
187
+ "description": "A lineage suffix such as `Jr.`. Null when the name has none.",
188
+ "$ref": "#/$defs/optionalString"
189
+ },
190
+ "literal": {
191
+ "description": "A name written as one brace-protected unit, such as `{Example Robotics Consortium}`. Null for a parsed name, where the four parts above apply instead. Brace protection means `do not parse this`, so it does not assert that the name is an organisation.",
192
+ "$ref": "#/$defs/optionalString"
193
+ },
194
+ "resolution": { "$ref": "#/$defs/resolution" },
195
+ "derived": { "$ref": "#/$defs/derived" }
196
+ }
197
+ },
198
+ "authorship": {
199
+ "description": "One authorship of a work, addressed by (work.bib_id, position). The authorship is the primary contributor record, so two people who write their names identically are never merged at this level. It references exactly one contributor: `person_id` for a person in people.yaml, `collaborator_key` for the grouping an unresolved authorship fell into.",
200
+ "$ref": "#/$defs/nameParts",
201
+ "type": "object",
202
+ "required": ["collaborator_key", "equal_contribution"],
203
+ "additionalProperties": false,
204
+ "properties": {
205
+ "name": { "type": "string" },
206
+ "position": { "type": "integer" },
207
+ "person_id": { "type": ["string", "null"] },
208
+ "given": { "$ref": "#/$defs/optionalString" },
209
+ "von": { "$ref": "#/$defs/optionalString" },
210
+ "family": { "$ref": "#/$defs/optionalString" },
211
+ "suffix": { "$ref": "#/$defs/optionalString" },
212
+ "literal": { "$ref": "#/$defs/optionalString" },
213
+ "resolution": { "$ref": "#/$defs/resolution" },
214
+ "derived": { "$ref": "#/$defs/derived" },
215
+ "collaborator_key": {
216
+ "description": "The key of the collaborator grouping this authorship fell into, or null when it resolved to a person.",
217
+ "$ref": "#/$defs/optionalString"
218
+ },
219
+ "equal_contribution": {
220
+ "description": "True when the BibTeX entry marked this author with a `*`. The marker itself is not part of the name.",
221
+ "type": "boolean"
222
+ }
223
+ },
224
+ "oneOf": [
225
+ {
226
+ "properties": {
227
+ "person_id": { "type": "string", "minLength": 1 },
228
+ "collaborator_key": { "type": "null" }
229
+ }
230
+ },
231
+ {
232
+ "properties": {
233
+ "person_id": { "type": "null" },
234
+ "collaborator_key": { "type": "string", "minLength": 1 }
235
+ }
236
+ }
237
+ ]
238
+ },
239
+ "editorship": {
240
+ "description": "One editor of a work. Editing a volume is not an authorship: an editor is excluded from person.work_ids, from a project's people and from collaborators, so one that matched nobody is simply `person_id: null` and references no grouping.",
241
+ "$ref": "#/$defs/nameParts",
242
+ "type": "object",
243
+ "additionalProperties": false,
244
+ "properties": {
245
+ "name": { "type": "string" },
246
+ "position": { "type": "integer" },
247
+ "person_id": { "type": ["string", "null"] },
248
+ "given": { "$ref": "#/$defs/optionalString" },
249
+ "von": { "$ref": "#/$defs/optionalString" },
250
+ "family": { "$ref": "#/$defs/optionalString" },
251
+ "suffix": { "$ref": "#/$defs/optionalString" },
252
+ "literal": { "$ref": "#/$defs/optionalString" },
253
+ "resolution": { "$ref": "#/$defs/resolution" },
254
+ "derived": { "$ref": "#/$defs/derived" }
255
+ }
256
+ },
257
+ "award": {
258
+ "description": "One award a work received.",
259
+ "type": "object",
260
+ "required": ["name", "year"],
261
+ "additionalProperties": false,
262
+ "properties": {
263
+ "name": {
264
+ "description": "The award's name, as plain text.",
265
+ "type": "string",
266
+ "minLength": 1
267
+ },
268
+ "year": {
269
+ "description": "The year the award was given, which need not be the work's: the year its `YYYY:` prefix wrote, else the work's `year`, else null.",
270
+ "type": ["integer", "null"],
271
+ "minimum": 0
272
+ }
273
+ }
274
+ },
275
+ "work": {
276
+ "type": "object",
277
+ "required": [
278
+ "bib_id", "source", "title", "authors", "editors", "year", "venue",
279
+ "volume", "number", "pages", "series", "edition", "publisher",
280
+ "address", "organization", "chapter", "month", "howpublished", "type",
281
+ "category", "entry_type", "abstract", "note", "awards", "identifiers",
282
+ "links", "project_ids", "bibtex", "derived"
283
+ ],
284
+ "additionalProperties": false,
285
+ "properties": {
286
+ "bib_id": {
287
+ "description": "The BibTeX citation key, as written.",
288
+ "type": "string",
289
+ "minLength": 1
290
+ },
291
+ "source": {
292
+ "description": "Where this work was read from. Per-field provenance stays out of the document; #30 uses `derived`.",
293
+ "type": "object",
294
+ "required": ["file", "key"],
295
+ "additionalProperties": false,
296
+ "properties": {
297
+ "file": {
298
+ "description": "The configured `bib_files[].name`, never an absolute path.",
299
+ "type": "string"
300
+ },
301
+ "key": {
302
+ "description": "The citation key as written, which is `bib_id` again.",
303
+ "type": "string",
304
+ "minLength": 1
305
+ }
306
+ }
307
+ },
308
+ "title": { "type": "string" },
309
+ "authors": { "type": "array", "items": { "$ref": "#/$defs/authorship" } },
310
+ "editors": { "type": "array", "items": { "$ref": "#/$defs/editorship" } },
311
+ "year": {
312
+ "description": "Null when the entry supplied none, which is reported as a diagnostic. A work with no year sorts last.",
313
+ "type": ["integer", "null"],
314
+ "minimum": 0
315
+ },
316
+ "venue": {
317
+ "description": "The container this work appeared in, or null when the entry names none.",
318
+ "oneOf": [{ "$ref": "#/$defs/venue" }, { "type": "null" }]
319
+ },
320
+ "volume": { "$ref": "#/$defs/optionalString" },
321
+ "number": {
322
+ "description": "BibTeX's `number`, with BibTeX's meaning: an issue number for an article and a report number for a technical report. Reinterpreting it is not sslabdata's job; `venue.kind` gives a consumer the branch it needs.",
323
+ "$ref": "#/$defs/optionalString"
324
+ },
325
+ "pages": { "$ref": "#/$defs/optionalString" },
326
+ "series": { "$ref": "#/$defs/optionalString" },
327
+ "edition": { "$ref": "#/$defs/optionalString" },
328
+ "publisher": { "$ref": "#/$defs/optionalString" },
329
+ "address": { "$ref": "#/$defs/optionalString" },
330
+ "organization": { "$ref": "#/$defs/optionalString" },
331
+ "chapter": { "$ref": "#/$defs/optionalString" },
332
+ "month": { "$ref": "#/$defs/optionalString" },
333
+ "howpublished": { "$ref": "#/$defs/optionalString" },
334
+ "type": {
335
+ "description": "BibTeX's `type` field: a report's own label, or a section label for a chapter. Not the entry type, which is `entry_type`.",
336
+ "$ref": "#/$defs/optionalString"
337
+ },
338
+ "category": {
339
+ "description": "The category of the bib_files entry the file was listed under, not anything in the .bib file.",
340
+ "type": "string"
341
+ },
342
+ "entry_type": {
343
+ "description": "The BibTeX entry type, lower-cased.",
344
+ "type": "string",
345
+ "minLength": 1
346
+ },
347
+ "abstract": { "$ref": "#/$defs/optionalString" },
348
+ "note": { "$ref": "#/$defs/optionalString" },
349
+ "awards": {
350
+ "description": "The awards this work received, in the order the entry's `award` field wrote them; `[]` when it names none. Awards held by a person are not here.",
351
+ "type": "array",
352
+ "items": { "$ref": "#/$defs/award" }
353
+ },
354
+ "identifiers": { "$ref": "#/$defs/identifierMap" },
355
+ "links": { "$ref": "#/$defs/linkMap" },
356
+ "project_ids": { "$ref": "#/$defs/idList" },
357
+ "bibtex": {
358
+ "description": "The entry re-serialized as BibTeX, for a reader to copy, and null when it could not be written back out. It is a re-serialization of the entry's data and explicitly not a source of properties.",
359
+ "$ref": "#/$defs/optionalString"
360
+ },
361
+ "derived": { "$ref": "#/$defs/derived" }
362
+ }
363
+ },
364
+ "person": {
365
+ "type": "object",
366
+ "required": [
367
+ "id", "name", "role", "status", "photo", "email", "website",
368
+ "co_advisor", "start_year", "end_year", "degree", "thesis_title",
369
+ "current_position", "work_ids", "derived"
370
+ ],
371
+ "additionalProperties": false,
372
+ "properties": {
373
+ "id": { "type": "string", "minLength": 1 },
374
+ "name": { "type": "string", "minLength": 1 },
375
+ "role": { "$ref": "#/$defs/optionalString" },
376
+ "status": { "type": "string" },
377
+ "photo": { "$ref": "#/$defs/optionalString" },
378
+ "email": { "$ref": "#/$defs/optionalString" },
379
+ "website": { "$ref": "#/$defs/optionalString" },
380
+ "co_advisor": { "$ref": "#/$defs/optionalString" },
381
+ "start_year": { "type": ["integer", "null"] },
382
+ "end_year": { "type": ["integer", "null"] },
383
+ "degree": { "$ref": "#/$defs/optionalString" },
384
+ "thesis_title": { "$ref": "#/$defs/optionalString" },
385
+ "current_position": { "$ref": "#/$defs/optionalString" },
386
+ "work_ids": {
387
+ "description": "The works this person authored. Editors are not authorships and are not listed.",
388
+ "$ref": "#/$defs/idList"
389
+ },
390
+ "derived": { "$ref": "#/$defs/derived" }
391
+ }
392
+ },
393
+ "project": {
394
+ "type": "object",
395
+ "required": [
396
+ "id", "title", "description", "website", "image", "status",
397
+ "work_ids", "people_ids", "derived"
398
+ ],
399
+ "additionalProperties": false,
400
+ "properties": {
401
+ "id": { "type": "string", "minLength": 1 },
402
+ "title": { "type": "string" },
403
+ "description": { "$ref": "#/$defs/optionalString" },
404
+ "website": { "$ref": "#/$defs/optionalString" },
405
+ "image": {
406
+ "description": "A URL or a site path, the same kind of value as a person's `photo`. Plain text: deciding which URLs are safe to render is the renderer's job.",
407
+ "$ref": "#/$defs/optionalString"
408
+ },
409
+ "status": { "type": "string" },
410
+ "work_ids": { "$ref": "#/$defs/idList" },
411
+ "people_ids": { "$ref": "#/$defs/idList" },
412
+ "derived": { "$ref": "#/$defs/derived" }
413
+ }
414
+ },
415
+ "collaborator": {
416
+ "description": "A grouping over unresolved authorships, not an identity. A consumer that distrusts the grouping can ignore it and work from `authorships` instead; that property is what makes the grouping safe to publish.",
417
+ "type": "object",
418
+ "required": [
419
+ "key", "grouped_by", "name_kind", "name", "given", "von", "family",
420
+ "suffix", "literal", "name_variants", "authorships", "work_ids",
421
+ "last_year", "derived"
422
+ ],
423
+ "additionalProperties": false,
424
+ "properties": {
425
+ "key": {
426
+ "description": "A lookup key, explicitly not an assertion about a human: a readable slug of the normalised name plus an always-present short digest, so adding an unrelated collaborator can never change an existing key.",
427
+ "type": "string",
428
+ "minLength": 1
429
+ },
430
+ "grouped_by": {
431
+ "description": "The policy that built the key. An open string; `normalized_name`, or `declared` for a grouping the collaborators file declares.",
432
+ "type": "string",
433
+ "minLength": 1
434
+ },
435
+ "name_kind": {
436
+ "description": "`personal` or `literal`. Not `person`/`organization`: brace protection in BibTeX means `do not parse this`, which covers organisations but also mononyms, so the document must not assert corporate-ness.",
437
+ "enum": ["personal", "literal"]
438
+ },
439
+ "name": {
440
+ "description": "The readable form of the first spelling this key grouped, in document order.",
441
+ "type": "string",
442
+ "minLength": 1
443
+ },
444
+ "given": { "$ref": "#/$defs/optionalString" },
445
+ "von": { "$ref": "#/$defs/optionalString" },
446
+ "family": { "$ref": "#/$defs/optionalString" },
447
+ "suffix": { "$ref": "#/$defs/optionalString" },
448
+ "literal": { "$ref": "#/$defs/optionalString" },
449
+ "name_variants": {
450
+ "description": "The distinct source spellings that normalised to this key, ascending.",
451
+ "type": "array",
452
+ "items": { "type": "string", "minLength": 1 },
453
+ "uniqueItems": true
454
+ },
455
+ "authorships": {
456
+ "description": "The occurrences this key grouped, in document order.",
457
+ "type": "array",
458
+ "items": {
459
+ "type": "object",
460
+ "required": ["work_id", "position"],
461
+ "additionalProperties": false,
462
+ "properties": {
463
+ "work_id": { "type": "string", "minLength": 1 },
464
+ "position": { "type": "integer", "minimum": 1 }
465
+ }
466
+ },
467
+ "minItems": 1
468
+ },
469
+ "work_ids": {
470
+ "description": "The distinct works this key grouped, in document order. One work listing two authorships under this key appears once; `authorships` lists both.",
471
+ "$ref": "#/$defs/idList"
472
+ },
473
+ "last_year": {
474
+ "description": "The greatest year of any work contributing an occurrence, or null when none of them has one.",
475
+ "type": ["integer", "null"]
476
+ },
477
+ "derived": { "$ref": "#/$defs/derived" }
478
+ }
479
+ }
480
+ }
481
+ }
@@ -12,7 +12,7 @@ MIT License - see LICENSE file for details.
12
12
 
13
13
  from .config import ConfigurationError, LabDataConfig, BibFile
14
14
  from .models import (
15
- LabData, Work, Author, Contributor, Venue, Link, Person, Project,
15
+ LabData, Work, Award, Author, Contributor, Venue, Link, Person, Project,
16
16
  Collaborator,
17
17
  )
18
18
  from .assembler import assemble, AssemblyError, AssemblyResult
@@ -24,6 +24,7 @@ __all__ = [
24
24
  "ConfigurationError",
25
25
  "LabData",
26
26
  "Work",
27
+ "Award",
27
28
  "Author",
28
29
  "Contributor",
29
30
  "Venue",
@@ -37,4 +38,4 @@ __all__ = [
37
38
  "export_to_yaml",
38
39
  "export_to_json",
39
40
  ]
40
- __version__ = "3.1.0"
41
+ __version__ = "4.0.0"
@@ -89,6 +89,8 @@ CLASSES: Dict[str, str] = {
89
89
  "BIB-YEAR-MISSING": WARNING,
90
90
  "BIB-YEAR-INVALID": WARNING,
91
91
  "BIB-DOI-INVALID": WARNING,
92
+ "BIB-AWARD-EMPTY": WARNING,
93
+ "BIB-AWARD-YEAR-MALFORMED": WARNING,
92
94
  "BIB-OTHERS-NOT-LAST": WARNING,
93
95
  "BIB-STRING-UNDEFINED": WARNING,
94
96
  "BIB-STRING-REDEFINED": WARNING,
@@ -18,9 +18,9 @@ from typing import Dict, List, Optional
18
18
  from .config import json_lab, reject_absolute_name
19
19
 
20
20
 
21
- # The document's schema version (schema/v5/output.schema.json). When it
21
+ # The document's schema version (schema/v6/output.schema.json). When it
22
22
  # changes is SPEC.md §6.
23
- SCHEMA_VERSION = 5
23
+ SCHEMA_VERSION = 6
24
24
 
25
25
  GENERATOR_NAME = "sslabdata"
26
26
 
@@ -56,6 +56,17 @@ class Link:
56
56
  }
57
57
 
58
58
 
59
+ @dataclass
60
+ class Award:
61
+ """One award a work received: its name, and the year it was given, which
62
+ need not be the work's (SPEC.md §5)."""
63
+ name: str
64
+ year: Optional[int] = None
65
+
66
+ def to_dict(self) -> dict:
67
+ return {'name': self.name, 'year': self.year}
68
+
69
+
59
70
  @dataclass
60
71
  class Contributor:
61
72
  """One person named on a work: the parts of the name, and who it resolved
@@ -152,6 +163,10 @@ class Work:
152
163
  bibtex: Optional[str] = None
153
164
  derived: Dict[str, object] = field(default_factory=dict)
154
165
 
166
+ # Declared last so it takes no other field's position in a positional
167
+ # call; to_dict() emits it beside `note`.
168
+ awards: List[Award] = field(default_factory=list)
169
+
155
170
  def to_dict(self) -> dict:
156
171
  """Convert to dictionary for serialization.
157
172
 
@@ -184,6 +199,7 @@ class Work:
184
199
  'entry_type': self.entry_type,
185
200
  'abstract': self.abstract,
186
201
  'note': self.note,
202
+ 'awards': [award.to_dict() for award in self.awards],
187
203
  'identifiers': {scheme: list(values)
188
204
  for scheme, values in self.identifiers.items()},
189
205
  'links': {kind: [link.to_dict() for link in links]
@@ -21,6 +21,7 @@ from typing import Callable, Dict, List, Optional, Tuple
21
21
  from urllib.parse import urlsplit
22
22
 
23
23
  import pybtex.errors
24
+ from pybtex.bibtex.utils import split_tex_string
24
25
  from pybtex.database import BibliographyData, Entry, Person
25
26
  from pybtex.exceptions import PybtexError
26
27
  from pybtex.database.output.bibtex import Writer as BibTeXWriter
@@ -34,7 +35,7 @@ from ..config import (
34
35
  CONTROL_CHARACTER, control_message, nfc, without_control_characters,
35
36
  )
36
37
  from ..diagnostics import Diagnostic, diagnostic
37
- from ..models import Author, Contributor, Link, Venue, Work
38
+ from ..models import Author, Award, Contributor, Link, Venue, Work
38
39
 
39
40
 
40
41
  # The fields converted from LaTeX to plain text (SPEC.md §2). The repository
@@ -1085,6 +1086,86 @@ def extract_note(entry: dict) -> Optional[str]:
1085
1086
  return note
1086
1087
 
1087
1088
 
1089
+ # The award list is split as BibTeX splits `author` (SPEC.md §5): on `and`, in
1090
+ # any case, with whitespace on both sides, outside braces. The whitespace
1091
+ # before it is part of the separator and the whitespace after it is not, so
1092
+ # `and and` is two separators with an empty award between them. The value is
1093
+ # padded with a space at each end before it is split, so an `and` that opens
1094
+ # or closes the list separates an empty award too.
1095
+ AWARD_SEPARATOR = r"[ \t\r\n][Aa][Nn][Dd](?=[ \t\r\n])"
1096
+
1097
+ # The year an award was given, written before it: four ASCII digits, a colon
1098
+ # and whitespace. The whitespace is not needed when nothing follows, so
1099
+ # `2026:` alone is a year with no name rather than a malformed prefix.
1100
+ AWARD_YEAR = re.compile(r"([0-9]{4}):(?:\s+|$)")
1101
+
1102
+ # What an author may have meant as a year and did not write as one: digits of
1103
+ # any kind and a colon, or four digits and a space with no colon between them.
1104
+ AWARD_YEAR_LIKE = re.compile(r"\d+\s*:|\d{4}\s")
1105
+
1106
+ AWARD_EMPTY = "BIB-AWARD-EMPTY"
1107
+ AWARD_YEAR_MALFORMED = "BIB-AWARD-YEAR-MALFORMED"
1108
+
1109
+
1110
+ def split_awards(value: str) -> List[str]:
1111
+ """One `award` field value split into its awards, each as written and
1112
+ trimmed, with an empty string for each empty award.
1113
+
1114
+ pybtex's own brace-aware splitter does the work, so braces protect an
1115
+ `and` exactly as they do in a name list.
1116
+ """
1117
+ return split_tex_string(f" {value} ", AWARD_SEPARATOR)
1118
+
1119
+
1120
+ def parse_awards(entry: dict, work_year: Optional[int], source: str, report,
1121
+ on_unknown) -> List[Award]:
1122
+ """The work's awards, from its `award` field, in the order written
1123
+ (SPEC.md §5).
1124
+
1125
+ Each award may start with the year it was given, `YYYY:`, read before
1126
+ its name is converted from LaTeX; without one it takes ``work_year``. An
1127
+ empty award is dropped and reported, and so is a year with no name after
1128
+ it. A prefix that is not a year is reported and kept in the name.
1129
+ """
1130
+ raw = entry.get("award")
1131
+ if raw is None:
1132
+ return []
1133
+ key = entry.get("ID")
1134
+ if not raw.strip():
1135
+ report(diagnostic(AWARD_EMPTY, source, key, "award",
1136
+ "the award field names no award; the work has none"))
1137
+ return []
1138
+ items = split_awards(raw)
1139
+ awards: List[Award] = []
1140
+ for position, item in enumerate(items, 1):
1141
+ where = f"award {position} of {len(items)}"
1142
+ year = work_year
1143
+ text = item
1144
+ prefix = AWARD_YEAR.match(item)
1145
+ malformed = None if prefix else AWARD_YEAR_LIKE.match(item)
1146
+ if prefix:
1147
+ year = int(prefix.group(1))
1148
+ text = item[prefix.end():]
1149
+ elif malformed:
1150
+ report(diagnostic(
1151
+ AWARD_YEAR_MALFORMED, source, key, "award",
1152
+ f"{where}, '{item}', starts with "
1153
+ f"'{malformed.group().strip()}', which is not a year "
1154
+ "prefix: that is four digits 0-9, a colon and a space. It is "
1155
+ "kept as part of the name, and the award takes the work's "
1156
+ "year"))
1157
+ name = _convert(text, on_unknown).strip()
1158
+ if not name:
1159
+ written = f", '{item}', has a year and no name" if prefix else \
1160
+ " is empty"
1161
+ report(diagnostic(AWARD_EMPTY, source, key, "award",
1162
+ f"{where}{written}; it is dropped, and the "
1163
+ "other awards are kept"))
1164
+ continue
1165
+ awards.append(Award(name=name, year=year))
1166
+ return awards
1167
+
1168
+
1088
1169
  def parse_project_ids(entry: dict) -> List[str]:
1089
1170
  """Parse the project field from a BibTeX entry."""
1090
1171
  project_field = entry.get("project", "").strip()
@@ -1165,7 +1246,7 @@ def entry_to_work(
1165
1246
  check_entry_type(fields, source, report)
1166
1247
  check_others(entry, bib_id, source, report)
1167
1248
 
1168
- return Work(
1249
+ work = Work(
1169
1250
  bib_id=bib_id,
1170
1251
  title=fields.get("title", ""),
1171
1252
  authors=parse_author_list(entry, unknown_in("author")),
@@ -1182,11 +1263,17 @@ def entry_to_work(
1182
1263
  source=source, report=report),
1183
1264
  project_ids=parse_project_ids(fields),
1184
1265
  bibtex=format_bibtex(bib_id, entry, source, report),
1185
- # Empty, as its default is; named so that the mapping below is
1186
- # type-checked against the flat fields alone.
1266
+ # Empty, as their defaults are; named so that the mapping below is
1267
+ # type-checked against the flat fields alone. The awards are read
1268
+ # below, once the work's year is.
1187
1269
  derived={},
1270
+ awards=[],
1188
1271
  **{name: fields.get(name) for name in FLAT_FIELDS},
1189
1272
  )
1273
+ # An award without a year of its own takes the work's.
1274
+ work.awards = parse_awards(fields, work.year, source, report,
1275
+ unknown_in("award"))
1276
+ return work
1190
1277
 
1191
1278
 
1192
1279
  def _encoding_error(path: str, error: UnicodeDecodeError) -> Diagnostic:
@@ -1,5 +1,7 @@
1
1
  % A fictional work to start from: replace it with your own entries.
2
- % `project` names a project id from projects.yaml.
2
+ % `project` names a project id from projects.yaml. `award` lists the work's
3
+ % awards, separated by `and`; start one with `YYYY: ` when it was given in
4
+ % another year than the work's.
3
5
 
4
6
  @article{adams2024tidy,
5
7
  title = {Learning to Tidy Up from Demonstrations},
@@ -7,5 +9,6 @@
7
9
  journal = {Journal of Example Robotics},
8
10
  year = {2024},
9
11
  url = {https://example.org/papers/adams2024tidy},
12
+ award = {Best Paper Award},
10
13
  project = {homebot}
11
14
  }
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: sslabdata
3
- Version: 3.1.0
3
+ Version: 4.0.0
4
4
  Summary: Renderer-agnostic academic lab data assembler: BibTeX + YAML → structured data
5
5
  Author: Siddhartha Srinivasa
6
6
  License-Expression: MIT
@@ -37,6 +37,11 @@ Dynamic: license-file
37
37
 
38
38
  # sslabdata
39
39
 
40
+ [![PyPI](https://img.shields.io/pypi/v/sslabdata.svg)](https://pypi.org/project/sslabdata/)
41
+ [![Python](https://img.shields.io/pypi/pyversions/sslabdata.svg)](https://pypi.org/project/sslabdata/)
42
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://github.com/siddhss5/sslabdata/blob/main/LICENSE)
43
+ [![Tests](https://github.com/siddhss5/sslabdata/actions/workflows/test.yml/badge.svg?branch=main)](https://github.com/siddhss5/sslabdata/actions/workflows/test.yml)
44
+
40
45
  sslabdata compiles BibTeX and a little YAML into one schema-specified document —
41
46
  works, people, projects and the links between them — that any website, CV or
42
47
  script can read.
@@ -56,10 +61,11 @@ sslabdata --config lab.yaml --output lab.yml
56
61
  order the lists are in, which fields are derived, when the version changes.
57
62
  - [`CHANGELOG.md`](https://github.com/siddhss5/sslabdata/blob/main/CHANGELOG.md) — what changed at each release, and what it
58
63
  replaced.
59
- - [`schema/v5/output.schema.json`](https://github.com/siddhss5/sslabdata/blob/main/schema/v5/output.schema.json) — the
64
+ - [`schema/v6/output.schema.json`](https://github.com/siddhss5/sslabdata/blob/main/schema/v6/output.schema.json) — the
60
65
  document's JSON Schema. Published versions are immutable and live at their
61
- own paths; [`schema/v3/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v3/output.schema.json) and
62
- [`schema/v4/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v4/output.schema.json) are still there.
66
+ own paths; [`schema/v3/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v3/output.schema.json),
67
+ [`schema/v4/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v4/output.schema.json) and
68
+ [`schema/v5/`](https://github.com/siddhss5/sslabdata/blob/main/schema/v5/output.schema.json) are still there.
63
69
  - [`tests/COVERAGE.md`](https://github.com/siddhss5/sslabdata/blob/main/tests/COVERAGE.md) — every input case sslabdata
64
70
  supports, and every case it does not, with the fixture and test for each.
65
71
 
@@ -182,6 +188,7 @@ nothing else:
182
188
  | `doi`, `isbn`, `issn`, `eprint` + `archivePrefix` (or `eprinttype`) | `identifiers`, an open map from scheme to a list of identifiers, plus the links built from them. An `eprint`'s scheme is the repository `archivePrefix` or `eprinttype` named, lower-cased, so that field needs no property of its own — and an `eprint` in a repository other than arXiv gets no arXiv link |
183
189
  | `abstract` | `abstract` |
184
190
  | `note` | `note` |
191
+ | `award` | `awards`, a list of `{name, year}` (see below). `note` is never read for awards |
185
192
  | `url` | A link of kind `video` when its host is YouTube or Vimeo (or a subdomain of either), otherwise of kind `url` |
186
193
  | `video` | A link of kind `video`, whatever its host, so `url` can hold the work's website |
187
194
  | `pdf` | The work's one link of kind `pdf`. An entry without it gets `pdf_base_url` plus its citation key, when `pdf_base_url` is set |
@@ -192,6 +199,26 @@ The entry is also re-serialized into a `bibtex` field, so fields sslabdata does
192
199
  not interpret are still carried. It is a re-serialization, not a copy
193
200
  ([`SPEC.md` §5](https://github.com/siddhss5/sslabdata/blob/main/SPEC.md#5-input-versus-derived)).
194
201
 
202
+ ### Awards
203
+
204
+ A paper's awards go in its `award` field. Several are separated by `and`, as
205
+ the names in `author` are, and braces keep an `and` inside one name. An award
206
+ may start with the year it was given, as `YYYY:` and a space, when that is not
207
+ the paper's year:
208
+
209
+ ```bibtex
210
+ award = {2026: Test of Time Award and {Best Systems and Software Paper Award}}
211
+ ```
212
+
213
+ Each award becomes `{name, year}` in the work's `awards`, in the order
214
+ written, with its name converted from LaTeX as `title` is. An award without a
215
+ year takes the work's `year`, or `null` when the work has none. `awards` is
216
+ `[]` for a work with no `award` field. An empty award, or a prefix that looks
217
+ like a year and is not four digits, a colon and a space, is reported
218
+ (`BIB-AWARD-EMPTY`, `BIB-AWARD-YEAR-MALFORMED`); a malformed prefix stays in
219
+ the name. Awards a person holds, such as fellowships, are not part of the
220
+ document.
221
+
195
222
  ### The `project` tag
196
223
 
197
224
  sslabdata adds one custom BibTeX field, `project`, to link a paper to a research
@@ -205,6 +232,7 @@ project:
205
232
  year = {2024},
206
233
  eprint = {2406.99812},
207
234
  archivePrefix = {arXiv},
235
+ award = {Best Paper Award},
208
236
  project = {homebot}
209
237
  }
210
238
  ```
@@ -349,7 +377,7 @@ pip install jsonschema
349
377
  python -c "
350
378
  import json, yaml, jsonschema
351
379
  from importlib.resources import files
352
- schema = json.loads(files('sslabdata.schema').joinpath('v5/output.schema.json').read_text())
380
+ schema = json.loads(files('sslabdata.schema').joinpath('v6/output.schema.json').read_text())
353
381
  jsonschema.Draft202012Validator(schema).validate(yaml.safe_load(open('lab.yml')))
354
382
  print('valid')
355
383
  "
@@ -11,6 +11,7 @@ schema/input/v1/projects.schema.json
11
11
  schema/v3/output.schema.json
12
12
  schema/v4/output.schema.json
13
13
  schema/v5/output.schema.json
14
+ schema/v6/output.schema.json
14
15
  sslabdata/__init__.py
15
16
  sslabdata/assembler.py
16
17
  sslabdata/cli.py
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes