bioai-evidence-validator 0.4.1__py3-none-any.whl → 0.7.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,9 +1,9 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: bioai-evidence-validator
3
- Version: 0.4.1
3
+ Version: 0.7.0
4
4
  Summary: Standards-aligned evidence policy validation for AI-assisted biological curation
5
5
  Project-URL: Homepage, https://github.com/NingyuSUN/bioai-evidence-validator
6
- Project-URL: Documentation, https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/docs
6
+ Project-URL: Documentation, https://ningyusun.github.io/bioai-evidence-validator/
7
7
  Project-URL: Changelog, https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CHANGELOG.md
8
8
  Project-URL: Issues, https://github.com/NingyuSUN/bioai-evidence-validator/issues
9
9
  Author: Ningyu Sun
@@ -25,7 +25,13 @@ Requires-Dist: jsonschema<5,>=4.23
25
25
  Requires-Dist: linkml<2,>=1.8
26
26
  Requires-Dist: pyyaml<7,>=6.0
27
27
  Provides-Extra: dev
28
+ Requires-Dist: mypy<3,>=2.3; extra == 'dev'
29
+ Requires-Dist: openpyxl<4,>=3.1; extra == 'dev'
30
+ Requires-Dist: pytest-cov<8,>=7; extra == 'dev'
28
31
  Requires-Dist: pytest<9,>=8; extra == 'dev'
32
+ Requires-Dist: ruff<0.17,>=0.16; extra == 'dev'
33
+ Requires-Dist: types-jsonschema; extra == 'dev'
34
+ Requires-Dist: types-pyyaml; extra == 'dev'
29
35
  Description-Content-Type: text/markdown
30
36
 
31
37
  # BioAI Evidence Validator
@@ -34,6 +40,8 @@ Description-Content-Type: text/markdown
34
40
  [![PyPI](https://img.shields.io/pypi/v/bioai-evidence-validator)](https://pypi.org/project/bioai-evidence-validator/)
35
41
  [![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue)](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/pyproject.toml)
36
42
  [![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue)](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/LICENSE)
43
+ [![Open in Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/NingyuSUN/bioai-evidence-validator/blob/main/examples/quickstart.ipynb)
44
+ [![Docs](https://img.shields.io/badge/docs-site-blue)](https://ningyusun.github.io/bioai-evidence-validator/)
37
45
 
38
46
  **Stop AI-extracted biological claims from entering your knowledge base or
39
47
  training set before their evidence is good enough for that use.**
@@ -49,6 +57,10 @@ each requested use.
49
57
  pip install bioai-evidence-validator
50
58
  ```
51
59
 
60
+ Or try it in the browser, nothing to install:
61
+ [quickstart notebook on Colab](https://colab.research.google.com/github/NingyuSUN/bioai-evidence-validator/blob/main/examples/quickstart.ipynb).
62
+ Full documentation: **[ningyusun.github.io/bioai-evidence-validator](https://ningyusun.github.io/bioai-evidence-validator/)**.
63
+
52
64
  ## 30-second example
53
65
 
54
66
  The two records below are identical except for one field: how the supporting
@@ -99,19 +111,54 @@ record is trustworthy enough for a particular purpose.
99
111
  | Provenance consistency (source hashes, resolved references, scope) | — | ✅ |
100
112
  | Human adjudications bound to a specific statement and use | — | ✅ |
101
113
  | Machine-readable audit report with hashes of input, schema and profile | — | ✅ |
114
+ | Mapped to ECO, Biolink and GA4GH VA-Spec ([standards alignment](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/STANDARDS.md)) | — | ✅ |
102
115
 
103
- On the real-data benchmark below, schema-only checks admitted **160/160**
104
- injected faults; the full validator admitted **0/160**.
116
+ On both real-data benchmarks below, schema-only checks admitted **160/160**
117
+ injected faults; the full validator admitted **0/160**. On ClinVar, 1★ and 2★ variants
118
+ the validator holds back were **about 4–5× more likely** to be reclassified or put in
119
+ conflict three years later.
105
120
 
106
121
  ## Use it
107
122
 
123
+ ### Write a draft, not a full record
124
+
125
+ A full record spells out identifiers, evidence lines and hashes. A draft states each fact
126
+ once; `bioevidence build` derives the rest and hashes local source files:
127
+
128
+ ```yaml
129
+ profile: literature-claim
130
+ uses: [research_summary]
131
+ statement:
132
+ subject: {id: "SYN:GENE_A", label: Synthetic gene A, type: gene}
133
+ predicate: associated_with
134
+ object: {id: "SYN:PHENOTYPE_A", label: Synthetic phenotype A, type: phenotype}
135
+ scope: ["taxon:synthetic"]
136
+ sources:
137
+ - {id: paper, title: Synthetic paper, type: publication, version: v1,
138
+ retrieved_at: "2026-09-21T00:00:00Z", file: synthetic_paper.txt}
139
+ evidence:
140
+ - {source: paper, locator: Table 2, type: publication_result,
141
+ method: llm_extraction, scope: ["taxon:synthetic"]}
142
+ ```
143
+
144
+ ```bash
145
+ bioevidence build examples/drafts/llm_claim.yaml --output record.json
146
+ bioevidence validate record.json --profile literature-claim # exit 2: review required
147
+ ```
148
+
149
+ Building never fills in scope, extraction method, retrieval time or review decisions for
150
+ you. For LLM pipelines, `bioevidence draft-schema --profile literature-claim` prints a JSON
151
+ Schema for structured output. See the [draft format](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/DRAFTS.md).
152
+
108
153
  ### Command line
109
154
 
110
155
  ```bash
111
156
  bioevidence profiles # list built-in profiles and their use contracts
157
+ bioevidence build draft.yaml --output record.json # expand a compact draft
112
158
  bioevidence validate record.json --profile literature-claim
113
159
  bioevidence validate record.json --profile my_profile.yaml --output report.json
114
- bioevidence generate-schema --output record.schema.json # JSON Schema for the input format
160
+ bioevidence draft-schema --profile literature-claim # JSON Schema for drafts (e.g. LLM output)
161
+ bioevidence generate-schema --output record.schema.json # JSON Schema for full records
115
162
  ```
116
163
 
117
164
  Exit codes: **0** admitted, **1** rejected, **2** review required, **3** input or configuration error.
@@ -119,12 +166,9 @@ Exit codes: **0** admitted, **1** rejected, **2** review required, **3** input o
119
166
  ### Python
120
167
 
121
168
  ```python
122
- import json
123
- from pathlib import Path
124
-
125
- from bioevidence_validator.engine import validate_record
169
+ from bioevidence_validator import build_record, load_draft, validate_record
126
170
 
127
- record = json.loads(Path("record.json").read_text(encoding="utf-8"))
171
+ record = build_record(load_draft("draft.yaml"), base_dir=".") # or load a full record JSON
128
172
  report = validate_record(record, profile="literature-claim")
129
173
 
130
174
  for decision in report["use_decisions"]:
@@ -133,6 +177,30 @@ for decision in report["use_decisions"]:
133
177
 
134
178
  `profile` accepts a built-in name or a path to your own YAML profile.
135
179
 
180
+ ### Check records in CI
181
+
182
+ Validate every record or draft in a pull request, with a summary table and inline annotations:
183
+
184
+ ```yaml
185
+ # .github/workflows/evidence.yml
186
+ on: pull_request
187
+ jobs:
188
+ evidence:
189
+ runs-on: ubuntu-latest
190
+ steps:
191
+ - uses: actions/checkout@v4
192
+ - uses: NingyuSUN/bioai-evidence-validator@v0.7.0
193
+ with:
194
+ files: records/**/*.yaml # whitespace-separated globs
195
+ format: draft # or: record (default)
196
+ profile: literature-claim # built-in name or path to your profile YAML
197
+ fail-on: review # or: rejected
198
+ ```
199
+
200
+ With `fail-on: review` (default) the job fails unless every file is admitted; with
201
+ `fail-on: rejected` it fails only on rejected files or files that cannot be read. The
202
+ action's outputs `admitted`, `review_required`, `rejected` and `error` hold the counts.
203
+
136
204
  ## How it works
137
205
 
138
206
  ```mermaid
@@ -193,7 +261,36 @@ Use the [annotation templates](https://github.com/NingyuSUN/bioai-evidence-valid
193
261
  source version and intended use. Document reviewer roles and whether labels are
194
262
  single-reviewed or independently reviewed by multiple people.
195
263
 
196
- ## Real-data case
264
+ ## Real-data cases
265
+
266
+ ### ClinVar germline classifications, three years later
267
+
268
+ The [ClinVar case](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/examples/clinvar_germline/README.md)
269
+ turns every 2023-09 lab submission into evidence, validates 5,026 sampled variants with a
270
+ ClinVar-style profile, and checks what happened to them by 2026-09.
271
+
272
+ ![ClinVar benchmark: share of 2023 pathogenic classifications reclassified or conflicting by 2026, with and without a dissenting submission, and false admissions under controlled faults and the trust boundary](https://raw.githubusercontent.com/NingyuSUN/bioai-evidence-validator/main/docs/assets/clinvar_germline_benchmark.svg)
273
+
274
+ - **Policy reproduction:** the profile matches NCBI's own 2023-09 review status on 98–99% of
275
+ decisions; every disagreement is listed with its cause.
276
+ - **Where the validator is stricter, classifications were less stable.** It sends any P/LP
277
+ variant with a dissenting submission to review, even when ClinVar's aggregate does not.
278
+ Across all 218,920 germline P/LP variants, the share later reclassified or put in conflict:
279
+
280
+ | 2023-09 ClinVar review status | No dissenting submission | With one (validator: review) |
281
+ |---|---:|---:|
282
+ | 1★ single submitter | 2.63% (2.54–2.71) | **13.77%** (11.46–16.47) |
283
+ | 2★ multiple submitters | 2.19% (2.05–2.33) | **8.32%** (6.51–10.59) |
284
+ | 3★ expert panel | 0.07% (0.03–0.16) | **1.30%** (0.63–2.66) |
285
+
286
+ Wilson 95% intervals. Stability is not correctness, and this is observational; see the
287
+ case's interpretation limits. Not for clinical use.
288
+
289
+ ```bash
290
+ uv run python examples/clinvar_germline/run.py --output artifacts/clinvar
291
+ ```
292
+
293
+ ### VBO canine name mapping
197
294
 
198
295
  [VBO canine name mapping](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/examples/vbo_canine/README.md) uses a frozen public ontology:
199
296
  72 real-name cases, 160 controlled errors, and 16 separately reported trust-boundary
@@ -204,7 +301,7 @@ per-required-evidence-type validation. Source-derived labels are not expert anno
204
301
  uv run python examples/vbo_canine/run.py --output artifacts/vbo-canine
205
302
  ```
206
303
 
207
- ### Benchmark results (v0.4.1)
304
+ #### Benchmark results (v0.4.1)
208
305
 
209
306
  ![VBO canine benchmark comparing false admissions across three validation methods](https://raw.githubusercontent.com/NingyuSUN/bioai-evidence-validator/main/docs/assets/vbo_canine_benchmark.svg)
210
307
 
@@ -214,20 +311,31 @@ On 72 real-source name mappings, the full validator admitted all 48 unambiguous
214
311
 
215
312
  ## Scope
216
313
 
217
- The VBO case uses attributed public data; other fixtures are synthetic.
314
+ The VBO and ClinVar cases use attributed public data; other fixtures are synthetic.
218
315
  Admission means **the supplied record meets the selected
219
316
  profile**, not that a biological claim is true. The toolkit does not retrieve papers,
220
- verify reviewer identities, train models, or measure prediction accuracy. The generic core compares supplied hashes; the VBO importer also hashes its local source
221
- projection. External source truth and cohort independence require upstream verification.
317
+ verify reviewer identities, train models, or measure prediction accuracy. The generic core compares supplied hashes; the VBO and ClinVar importers also hash their local source
318
+ projections. External source truth and cohort independence require upstream verification.
319
+ Neither benchmark has independent expert annotation yet; a blinded [expert-review kit](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/evaluation/clinvar_review) for the ClinVar case is ready for reviewers.
222
320
 
223
321
  ## Versions and branches
224
322
 
225
- `main` is the domain-neutral framework (0.4.1). The complete canine implementation
226
- and SQLite adapter from 0.3 live on the
227
- [`canine-breed` branch](https://github.com/NingyuSUN/bioai-evidence-validator/tree/canine-breed);
323
+ `main` is the domain-neutral framework (0.7.0). The complete canine implementation
324
+ and SQLite adapter from 0.3 are preserved at the
325
+ [`canine-0.3` tag](https://github.com/NingyuSUN/bioai-evidence-validator/tree/canine-0.3)
326
+ (also the `canine-breed` branch);
228
327
  see the [0.4 migration guide](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/MIGRATION-0.4.md) and
229
328
  [changelog](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CHANGELOG.md).
230
329
 
330
+ ## Contributing
331
+
332
+ Bug reports, domain profiles and new benchmarks are welcome. See
333
+ [CONTRIBUTING.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CONTRIBUTING.md),
334
+ the [community profiles](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/community/profiles)
335
+ and issues labelled [`good first issue`](https://github.com/NingyuSUN/bioai-evidence-validator/labels/good%20first%20issue).
336
+ Report security problems privately as described in
337
+ [SECURITY.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/SECURITY.md).
338
+
231
339
  ## Citing
232
340
 
233
341
  If you use this toolkit in research, please cite it using the metadata in
@@ -235,6 +343,8 @@ If you use this toolkit in research, please cite it using the metadata in
235
343
  (GitHub's "Cite this repository" button generates APA and BibTeX).
236
344
 
237
345
  [Create a profile](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/PROFILES.md) ·
346
+ [Draft format](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/DRAFTS.md) ·
347
+ [Standards alignment](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/STANDARDS.md) ·
238
348
  [Engineering contract](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/ENGINEERING.md) ·
239
349
  [Design case study](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/CASE_STUDY.md) ·
240
350
  [Architecture decision](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/ADR-002-domain-neutral-main.md) ·
@@ -0,0 +1,15 @@
1
+ bioevidence_validator/__init__.py,sha256=PsEzz-ifHjfLOJX7KkUqMFNbniIQMNPKRFxgNdoF3JI,312
2
+ bioevidence_validator/cli.py,sha256=7TxCuCN0jcbFwUnim0QiIARJC3XPQcsYN5qBcT7EFF0,9965
3
+ bioevidence_validator/config.py,sha256=Ahc9eRPNR_ArqnbnJkdZsPFq9yERTYied0oGwUkANYI,1110
4
+ bioevidence_validator/draft.py,sha256=2ge9l1BA44e-_ukO_Nnz9UZAV_RS8yqj3Ujr86eFxtY,13631
5
+ bioevidence_validator/engine.py,sha256=OSH9Zd-UqCVqYRU6eLkyoavRWAJGzz8Q6mcCNkhW48I,14989
6
+ bioevidence_validator/review.py,sha256=yp9DI2TafTiZqZmNTejtUb9oHeeXPEjZHXx3fe2qCGI,20293
7
+ bioevidence_validator/profiles/dataset-label.yaml,sha256=4utz9YcFFF2UnaLjIb0ATb-27hF8Smw7zmv0sBTaQdM,843
8
+ bioevidence_validator/profiles/general.yaml,sha256=i9fJCWovJZ2e2d4rbn0Bx7mrhumNb2zjS9TOarkBWZw,778
9
+ bioevidence_validator/profiles/literature-claim.yaml,sha256=7l82gX3oGGyYFq9B29-1qGXeL4LRcMhJi8Y9fieVE0c,612
10
+ bioevidence_validator/schema/bioevidence_core.yaml,sha256=c2YIBOhuT6v7j4MDsMScImdyLZCMy-3X--ExCB_IYu8,7514
11
+ bioai_evidence_validator-0.7.0.dist-info/METADATA,sha256=mt7dqba6hjGQjYBUoBRTS_09bvyD7GU3x1c_RmK4amg,18818
12
+ bioai_evidence_validator-0.7.0.dist-info/WHEEL,sha256=W3fkpkm7-wf9vBI5Z-7s0eWkeM-spu78I8Neb98DeEg,87
13
+ bioai_evidence_validator-0.7.0.dist-info/entry_points.txt,sha256=6oufzT2x1nIhTqM5B6y4aWubExKAoFb_wXXTdz3mBQQ,63
14
+ bioai_evidence_validator-0.7.0.dist-info/licenses/LICENSE,sha256=z8d0m5b2O9McPEK1xHG_dWgUBT6EfBDz6wA0F7xSPTA,11358
15
+ bioai_evidence_validator-0.7.0.dist-info/RECORD,,
@@ -1,3 +1,8 @@
1
1
  """Evidence validation primitives and domain policy runners."""
2
2
 
3
- __version__ = "0.4.1"
3
+ __version__ = "0.7.0"
4
+
5
+ from .draft import build_record, draft_json_schema, load_draft # noqa: E402
6
+ from .engine import validate_record # noqa: E402
7
+
8
+ __all__ = ["__version__", "build_record", "draft_json_schema", "load_draft", "validate_record"]
@@ -9,7 +9,9 @@ from pathlib import Path
9
9
 
10
10
  import yaml
11
11
 
12
- from .engine import default_schema_path, generate_json_schema, validate_record, profile_path, list_profiles
12
+ from . import review as review_module
13
+ from .draft import build_record, draft_json_schema, load_draft
14
+ from .engine import default_schema_path, generate_json_schema, list_profiles, profile_path, validate_record
13
15
 
14
16
 
15
17
  def parser() -> argparse.ArgumentParser:
@@ -20,10 +22,46 @@ def parser() -> argparse.ArgumentParser:
20
22
  validate.add_argument("--output", type=Path)
21
23
  validate.add_argument("--schema", type=Path, default=default_schema_path())
22
24
  validate.add_argument("--profile", default="general", help="Built-in profile name or YAML file path")
25
+ build = commands.add_parser("build", help="Expand a compact YAML/JSON draft into a full record")
26
+ build.add_argument("draft", type=Path)
27
+ build.add_argument("--output", type=Path)
28
+ draft_schema = commands.add_parser("draft-schema", help="JSON Schema for drafts under one profile")
29
+ draft_schema.add_argument("--profile", default="general", help="Built-in profile name or YAML file path")
30
+ draft_schema.add_argument("--output", type=Path)
23
31
  generate = commands.add_parser("generate-schema", help="Generate JSON Schema from LinkML")
24
32
  generate.add_argument("--output", type=Path, required=True)
25
33
  generate.add_argument("--schema", type=Path, default=default_schema_path())
26
34
  commands.add_parser("profiles", help="List built-in profiles and their use contracts")
35
+
36
+ review = commands.add_parser("review", help="Independent human review: agreement, adjudication, scoring")
37
+ steps = review.add_subparsers(dest="step", required=True)
38
+ check = steps.add_parser("check", help="Check annotations (and adjudications) against the protocol format")
39
+ check.add_argument("annotations", type=Path)
40
+ check.add_argument("--adjudications", type=Path)
41
+ agree = steps.add_parser("agreement", help="Inter-reviewer agreement per use, before adjudication")
42
+ agree.add_argument("annotations", type=Path)
43
+ agree.add_argument("--output", type=Path)
44
+ sheet = steps.add_parser("adjudication-sheet", help="Blank adjudication rows for every disagreement")
45
+ sheet.add_argument("annotations", type=Path)
46
+ sheet.add_argument("--output", type=Path, required=True)
47
+ sheet.add_argument("--min-reviewers", type=int, default=2)
48
+ scoring = steps.add_parser("score", help="Compare predictions with the resolved reference labels")
49
+ scoring.add_argument("--annotations", type=Path, required=True)
50
+ scoring.add_argument("--adjudications", type=Path, required=True)
51
+ scoring.add_argument("--predictions", type=Path, required=True)
52
+ scoring.add_argument("--split", default="test", choices=sorted(review_module.SPLITS))
53
+ scoring.add_argument("--min-reviewers", type=int, default=2)
54
+ scoring.add_argument("--output", type=Path)
55
+ frozen = steps.add_parser("freeze", help="Write a frozen manifest for a completed reference set")
56
+ frozen.add_argument("--manifest", type=Path, required=True, help="Draft manifest to complete")
57
+ frozen.add_argument("--annotations", type=Path, required=True)
58
+ frozen.add_argument("--adjudications", type=Path, required=True)
59
+ frozen.add_argument("--file", type=Path, action="append", default=[], help="Extra file to hash (repeatable)")
60
+ frozen.add_argument("--dataset-id", required=True)
61
+ frozen.add_argument("--version", required=True)
62
+ frozen.add_argument("--frozen-at", required=True, help="ISO 8601 time of the freeze")
63
+ frozen.add_argument("--min-reviewers", type=int, default=2)
64
+ frozen.add_argument("--output", type=Path, required=True)
27
65
  return result
28
66
 
29
67
 
@@ -81,6 +119,24 @@ def _run(args) -> int:
81
119
  print(json.dumps(list_profiles(), indent=2))
82
120
  return 0
83
121
 
122
+ if args.command == "review":
123
+ return _review(args)
124
+
125
+ if args.command in {"build", "draft-schema"}:
126
+ if args.command == "build":
127
+ if args.output and args.output.resolve() == args.draft.resolve():
128
+ raise ValueError("Output must not overwrite the draft")
129
+ draft = load_draft(args.draft)
130
+ _check_depth(draft)
131
+ result = build_record(draft, base_dir=args.draft.parent)
132
+ else:
133
+ result = draft_json_schema(args.profile)
134
+ if args.output:
135
+ _write_json(args.output, result)
136
+ else:
137
+ print(json.dumps(result, indent=2))
138
+ return 0
139
+
84
140
  selected_profile = profile_path(args.profile)
85
141
  protected = [args.input, args.schema, selected_profile, default_schema_path()]
86
142
  if args.output and args.output.resolve() in {p.resolve() for p in protected}:
@@ -101,6 +157,48 @@ def _run(args) -> int:
101
157
  return 0
102
158
 
103
159
 
160
+ def _review(args) -> int:
161
+ rv = review_module
162
+ annotations = rv.load_annotations(args.annotations)
163
+ if args.step == "check":
164
+ adjudications = rv.load_adjudications(args.adjudications, annotations) if args.adjudications else []
165
+ units = {(a["case_id"], a["requested_use"]) for a in annotations}
166
+ print(f"OK: {len(annotations)} annotations, {len(units)} case-use units, "
167
+ f"{len({a['reviewer_id'] for a in annotations})} reviewers, {len(adjudications)} adjudications")
168
+ return 0
169
+ if args.step == "agreement":
170
+ report = rv.agreement(annotations)
171
+ if args.output:
172
+ _write_json(args.output, report)
173
+ print(rv.render_agreement(report), end="")
174
+ return 0
175
+ if args.step == "adjudication-sheet":
176
+ if args.output.exists():
177
+ raise ValueError(f"{args.output} exists; refusing to overwrite adjudication work")
178
+ rows, incomplete = rv.adjudication_sheet(annotations, min_reviewers=args.min_reviewers)
179
+ rv.write_csv(args.output, rows, rv.ADJUDICATION_COLUMNS)
180
+ print(f"{len(rows)} disagreement(s) to adjudicate; {len(incomplete)} unit(s) with too few reviews")
181
+ return 0
182
+ adjudications = rv.load_adjudications(args.adjudications, annotations)
183
+ final = rv.resolve(annotations, adjudications, min_reviewers=args.min_reviewers)
184
+ if args.step == "score":
185
+ report = rv.score(final, rv.load_predictions(args.predictions), split=args.split)
186
+ if args.output:
187
+ _write_json(args.output, report)
188
+ print(rv.render_score(report), end="")
189
+ return 0
190
+ manifest = json.loads(args.manifest.read_text(encoding="utf-8"))
191
+ files = [args.annotations, args.adjudications, *args.file]
192
+ if args.output.resolve() in {f.resolve() for f in [args.manifest, *files]}:
193
+ raise ValueError("Output must not overwrite an input")
194
+ frozen = rv.freeze(manifest, final, annotations, files, dataset_id=args.dataset_id,
195
+ version=args.version, frozen_at=args.frozen_at)
196
+ _write_json(args.output, frozen)
197
+ print(f"Frozen {frozen['reviewed_case_use_count']} case-use labels from "
198
+ f"{frozen['independent_human_reviewer_count']} reviewers ({frozen['reference_type']})")
199
+ return 0
200
+
201
+
104
202
  def main(argv: list[str] | None = None) -> int:
105
203
  args = parser().parse_args(argv)
106
204
  try:
@@ -0,0 +1,253 @@
1
+ """Compact drafts: a short, strict authoring format that expands to a full evidence record.
2
+
3
+ A draft names each fact once. Building it derives identifiers, groups evidence items into
4
+ evidence lines by direction, and hashes local source files. It never fills in facts the
5
+ caller did not state: scopes, extraction methods, retrieval times and reviewer decisions
6
+ must be written explicitly, and the resulting record still goes through full validation.
7
+ """
8
+ from __future__ import annotations
9
+
10
+ import datetime as dt
11
+ import hashlib
12
+ import json
13
+ from pathlib import Path
14
+ from typing import Any
15
+
16
+ from .config import load_mapping, nonblank
17
+ from .engine import load_profile, profile_path
18
+
19
+ DRAFT_KEYS = {"required": {"profile", "uses", "statement", "sources", "evidence"},
20
+ "optional": {"id", "status", "reviews"}}
21
+ STATEMENT_KEYS = {"required": {"subject", "predicate", "object", "scope"}, "optional": set()}
22
+ ENTITY_KEYS = {"required": {"id", "label", "type"}, "optional": set()}
23
+ SOURCE_KEYS = {"required": {"id", "title", "type", "version", "retrieved_at"},
24
+ "optional": {"uri", "sha256", "file"}}
25
+ EVIDENCE_KEYS = {"required": {"source", "locator", "type", "method", "scope"},
26
+ "optional": {"text", "direction"}}
27
+ REVIEW_KEYS = {"required": {"reviewer", "decision", "uses", "rationale", "decided_at"}, "optional": set()}
28
+ REVIEWER_KEYS = {"required": {"id", "type"}, "optional": {"name"}}
29
+
30
+ STATUSES = ["proposed", "accepted", "rejected", "superseded"]
31
+ SOURCE_TYPES = ["ontology_snapshot", "registry_snapshot", "dataset_snapshot", "publication", "web_page", "local_file"]
32
+ METHODS = ["deterministic_parser", "manual_curation", "normalized_string_match", "llm_extraction"]
33
+ DIRECTIONS = ["supports", "contradicts", "neutral"]
34
+ DECISIONS = ["accept", "reject", "defer"]
35
+ AGENT_TYPES = ["human", "software", "organization"]
36
+
37
+
38
+ def _fields(value: Any, keys: dict, path: str) -> dict:
39
+ if not isinstance(value, dict):
40
+ raise ValueError(f"{path} must be a mapping")
41
+ missing = keys["required"] - value.keys()
42
+ unknown = value.keys() - keys["required"] - keys["optional"]
43
+ if missing:
44
+ raise ValueError(f"{path} is missing: {', '.join(sorted(missing))}")
45
+ if unknown:
46
+ raise ValueError(f"{path} has unknown fields: {', '.join(sorted(map(str, unknown)))}")
47
+ return value
48
+
49
+
50
+ def _text(value: Any, path: str, choices: list[str] | None = None) -> str:
51
+ if isinstance(value, (dt.date, dt.datetime)):
52
+ value = value.isoformat() # YAML loads unquoted timestamps as datetime objects.
53
+ if not nonblank(value):
54
+ raise ValueError(f"{path} must be a nonblank string (quote numbers and versions in YAML)")
55
+ if choices is not None and value not in choices:
56
+ raise ValueError(f"{path} must be one of: {', '.join(choices)}")
57
+ return value
58
+
59
+
60
+ def _texts(value: Any, path: str) -> list[str]:
61
+ if not isinstance(value, list) or not value:
62
+ raise ValueError(f"{path} must be a nonempty list")
63
+ items = [_text(item, f"{path}[{index}]") for index, item in enumerate(value)]
64
+ if len(set(items)) != len(items):
65
+ raise ValueError(f"{path} must not repeat values")
66
+ return items
67
+
68
+
69
+ def _rows(value: Any, path: str, required: bool = True) -> list:
70
+ if not isinstance(value, list) or (required and not value):
71
+ raise ValueError(f"{path} must be a {'nonempty ' if required else ''}list")
72
+ return value
73
+
74
+
75
+ def _entity(value: Any, path: str) -> dict:
76
+ entity = _fields(value, ENTITY_KEYS, path)
77
+ return {"id": _text(entity["id"], path + ".id"), "label": _text(entity["label"], path + ".label"),
78
+ "entity_type": _text(entity["type"], path + ".type")}
79
+
80
+
81
+ def _source(value: Any, path: str, base_dir: Path) -> tuple[str, dict]:
82
+ source = _fields(value, SOURCE_KEYS, path)
83
+ artifact = {"title": _text(source["title"], path + ".title"),
84
+ "source_type": _text(source["type"], path + ".type", SOURCE_TYPES)}
85
+ if "uri" in source:
86
+ artifact["uri"] = _text(source["uri"], path + ".uri")
87
+ artifact.update({"version": _text(source["version"], path + ".version"),
88
+ "retrieved_at": _text(source["retrieved_at"], path + ".retrieved_at")})
89
+ observed = None
90
+ if "file" in source:
91
+ file = base_dir / _text(source["file"], path + ".file")
92
+ if not file.is_file():
93
+ raise ValueError(f"{path}.file does not exist: {file}")
94
+ observed = hashlib.sha256(file.read_bytes()).hexdigest()
95
+ # A stated hash is the frozen reference; a file hash is what was observed now.
96
+ # With both, validation reports a mismatch (BEV002) instead of trusting either.
97
+ stated = _text(source["sha256"], path + ".sha256").lower() if "sha256" in source else None
98
+ frozen = stated or observed
99
+ if frozen is None:
100
+ raise ValueError(f"{path} needs sha256, file, or both")
101
+ artifact["sha256"] = frozen
102
+ if stated is not None and observed is not None:
103
+ artifact["observed_sha256"] = observed
104
+ return _text(source["id"], path + ".id"), artifact
105
+
106
+
107
+ def load_draft(path: Path) -> dict:
108
+ """Read a YAML or JSON draft with duplicate-key rejection."""
109
+ return load_mapping(Path(path).read_bytes(), "Draft")
110
+
111
+
112
+ def build_record(draft: dict, *, base_dir: Path | str = ".") -> dict[str, Any]:
113
+ """Expand a compact draft into a full record. Raises ValueError on any malformed field."""
114
+ draft = _fields(draft, DRAFT_KEYS, "draft")
115
+ base_dir = Path(base_dir)
116
+ profile = _text(draft["profile"], "draft.profile")
117
+ uses = _texts(draft["uses"], "draft.uses")
118
+ statement = _fields(draft["statement"], STATEMENT_KEYS, "draft.statement")
119
+ scope = _texts(statement["scope"], "draft.statement.scope")
120
+
121
+ record_id = _text(draft["id"], "draft.id") if "id" in draft else None
122
+ if record_id is None:
123
+ canonical = json.dumps(draft, sort_keys=True, separators=(",", ":"), default=str).encode()
124
+ record_id = "bioev:draft-" + hashlib.sha256(canonical).hexdigest()[:16]
125
+
126
+ sources, local_ids = [], {}
127
+ for index, row in enumerate(_rows(draft["sources"], "draft.sources")):
128
+ name, artifact = _source(row, f"draft.sources[{index}]", base_dir)
129
+ if name in local_ids:
130
+ raise ValueError(f"draft.sources[{index}].id repeats {name!r}")
131
+ local_ids[name] = f"{record_id}/source/{name}"
132
+ sources.append({"id": local_ids[name], **artifact})
133
+
134
+ items: list[dict[str, Any]] = []
135
+ lines: dict[str, list[str]] = {}
136
+ for index, row in enumerate(_rows(draft["evidence"], "draft.evidence")):
137
+ path = f"draft.evidence[{index}]"
138
+ evidence = _fields(row, EVIDENCE_KEYS, path)
139
+ source = _text(evidence["source"], path + ".source")
140
+ if source not in local_ids:
141
+ raise ValueError(f"{path}.source {source!r} is not a declared source id")
142
+ item: dict[str, Any] = {"id": f"{record_id}/evidence/{index + 1}", "source_artifact_id": local_ids[source],
143
+ "locator": _text(evidence["locator"], path + ".locator")}
144
+ if "text" in evidence:
145
+ item["extracted_text"] = _text(evidence["text"], path + ".text")
146
+ item.update({"evidence_type": _text(evidence["type"], path + ".type"),
147
+ "extraction_method": _text(evidence["method"], path + ".method", METHODS),
148
+ "scope": _texts(evidence["scope"], path + ".scope")})
149
+ items.append(item)
150
+ direction = _text(evidence.get("direction", "supports"), path + ".direction", DIRECTIONS)
151
+ lines.setdefault(direction, []).append(item["id"])
152
+
153
+ statement_id = f"{record_id}/statement"
154
+ adjudications = []
155
+ for index, row in enumerate(_rows(draft.get("reviews", []), "draft.reviews", required=False)):
156
+ path = f"draft.reviews[{index}]"
157
+ review = _fields(row, REVIEW_KEYS, path)
158
+ reviewer = _fields(review["reviewer"], REVIEWER_KEYS, path + ".reviewer")
159
+ agent = {"id": _text(reviewer["id"], path + ".reviewer.id")}
160
+ if "name" in reviewer:
161
+ agent["name"] = _text(reviewer["name"], path + ".reviewer.name")
162
+ agent["agent_type"] = _text(reviewer["type"], path + ".reviewer.type", AGENT_TYPES)
163
+ adjudications.append({"id": f"{record_id}/review/{index + 1}", "statement_id": statement_id,
164
+ "applies_to_uses": _texts(review["uses"], path + ".uses"),
165
+ "decision": _text(review["decision"], path + ".decision", DECISIONS),
166
+ "reviewer": agent, "rationale": _text(review["rationale"], path + ".rationale"),
167
+ "decided_at": _text(review["decided_at"], path + ".decided_at")})
168
+
169
+ return {
170
+ "record_id": record_id,
171
+ "profile_id": profile,
172
+ "statement": {
173
+ "id": statement_id,
174
+ "subject": _entity(statement["subject"], "draft.statement.subject"),
175
+ "predicate": _text(statement["predicate"], "draft.statement.predicate"),
176
+ "object": _entity(statement["object"], "draft.statement.object"),
177
+ "scope": scope,
178
+ "evidence_lines": [{"id": f"{record_id}/line/{direction}", "direction": direction,
179
+ "evidence_item_ids": ids} for direction, ids in lines.items()],
180
+ "statement_status": _text(draft.get("status", "proposed"), "draft.status", STATUSES),
181
+ },
182
+ "source_artifacts": sources,
183
+ "evidence_items": items,
184
+ "adjudications": adjudications,
185
+ "requested_uses": uses,
186
+ }
187
+
188
+
189
+ def _object(keys: dict, properties: dict, description: str | None = None) -> dict:
190
+ if properties.keys() != keys["required"] | keys["optional"]:
191
+ raise RuntimeError("Draft schema and parser fields diverged")
192
+ schema = {"type": "object", "additionalProperties": False,
193
+ "required": sorted(keys["required"]), "properties": properties}
194
+ return {"description": description, **schema} if description else schema
195
+
196
+
197
+ def _choice(values: list[str], description: str | None = None) -> dict:
198
+ if not values:
199
+ return {"type": "string", "minLength": 1, **({"description": description} if description else {})}
200
+ return {"type": "string", "enum": list(values), **({"description": description} if description else {})}
201
+
202
+
203
+ def draft_json_schema(profile: str | Path = "general") -> dict[str, Any]:
204
+ """JSON Schema for drafts under one profile, e.g. for LLM structured output.
205
+
206
+ Profile allowlists become enums. The schema guides authoring only; build_record and
207
+ validation remain authoritative.
208
+ """
209
+ contract = load_profile(profile_path(profile).read_bytes())
210
+ evidence_types = sorted({kind for use in contract["uses"].values() for kind in use["required_evidence_types"]})
211
+ text = {"type": "string", "minLength": 1}
212
+ tokens = {"type": "array", "minItems": 1, "uniqueItems": True, "items": text}
213
+
214
+ def entity(types):
215
+ return _object(ENTITY_KEYS, {"id": {**text, "description": "Stable identifier, e.g. a CURIE."},
216
+ "label": text, "type": _choice(types)})
217
+
218
+ return {
219
+ "$schema": "https://json-schema.org/draft/2020-12/schema",
220
+ "title": f"BioAI evidence draft ({contract['id']} profile {contract['version']})",
221
+ **_object(DRAFT_KEYS, {
222
+ "id": {**text, "description": "Optional record id; derived from the draft content when omitted."},
223
+ "profile": {"const": contract["id"]},
224
+ "uses": {"type": "array", "minItems": 1, "uniqueItems": True,
225
+ "items": _choice(sorted(contract["uses"]))},
226
+ "status": _choice(STATUSES, "Defaults to proposed."),
227
+ "statement": _object(STATEMENT_KEYS, {
228
+ "subject": entity(contract["subject_types"]),
229
+ "predicate": _choice(contract["predicates"]),
230
+ "object": entity(contract["object_types"]),
231
+ "scope": {**tokens, "description": "Explicit scope tokens, e.g. taxon:9606."}}),
232
+ "sources": {"type": "array", "minItems": 1, "items": _object(SOURCE_KEYS, {
233
+ "id": {**text, "description": "Local name referenced by evidence[].source."},
234
+ "title": text, "type": _choice(SOURCE_TYPES), "uri": text, "version": text,
235
+ "retrieved_at": {"type": "string", "format": "date-time"},
236
+ "sha256": {"type": "string", "pattern": "^[0-9a-fA-F]{64}$",
237
+ "description": "Frozen source hash."},
238
+ "file": {**text, "description": "Local file hashed at build time, relative to the draft."}},
239
+ "Provide sha256, file, or both.")},
240
+ "evidence": {"type": "array", "minItems": 1, "items": _object(EVIDENCE_KEYS, {
241
+ "source": text, "locator": {**text, "description": "Where in the source, e.g. 'Table 2'."},
242
+ "text": text, "type": _choice(evidence_types),
243
+ "method": _choice(METHODS, "Set by the pipeline that produced this item, not by a model."),
244
+ "scope": {**tokens, "description": "Scope the source actually supports."},
245
+ "direction": _choice(DIRECTIONS, "Defaults to supports.")})},
246
+ "reviews": {"type": "array", "items": _object(REVIEW_KEYS, {
247
+ "reviewer": _object(REVIEWER_KEYS, {"id": text, "name": text, "type": _choice(AGENT_TYPES)}),
248
+ "decision": _choice(DECISIONS),
249
+ "uses": {"type": "array", "minItems": 1, "uniqueItems": True,
250
+ "items": _choice(sorted(contract["uses"]))},
251
+ "rationale": text, "decided_at": {"type": "string", "format": "date-time"}})},
252
+ }),
253
+ }
@@ -3,7 +3,7 @@ from __future__ import annotations
3
3
  import hashlib
4
4
  import json
5
5
  from dataclasses import asdict, dataclass
6
- from datetime import datetime, timezone
6
+ from datetime import UTC, datetime
7
7
  from pathlib import Path
8
8
  from typing import Any
9
9
 
@@ -255,7 +255,7 @@ class RecordValidator:
255
255
  "profile_sha256": self.profile_sha256,
256
256
  "schema_version": self.schema_version, "schema_sha256": self.schema_sha256,
257
257
  "schema_sources": [dict(source) for source in self.schema_sources],
258
- "input_sha256": sha256_bytes(canonical), "validated_at": datetime.now(timezone.utc).isoformat(),
258
+ "input_sha256": sha256_bytes(canonical), "validated_at": datetime.now(UTC).isoformat(),
259
259
  "schema_valid": not any(f.rule_id == "SCHEMA" for f in findings),
260
260
  "overall_status": overall, "findings": [asdict(f) for f in findings], "use_decisions": decisions,
261
261
  }
@@ -0,0 +1,388 @@
1
+ """Independent human review: check annotations, measure agreement, adjudicate, score and freeze.
2
+
3
+ Implements the workflow in docs/GOLD_STANDARD.md for the CSV formats in evaluation/gold_standard/.
4
+ Reference labels live here, never in the engine: they evaluate decisions, they do not make them.
5
+ """
6
+ from __future__ import annotations
7
+
8
+ import csv
9
+ import datetime as dt
10
+ import hashlib
11
+ import json
12
+ import math
13
+ import random
14
+ import re
15
+ from collections import Counter, defaultdict
16
+ from collections.abc import Sequence
17
+ from pathlib import Path
18
+ from typing import Any
19
+
20
+ ANNOTATION_COLUMNS = [
21
+ "annotation_id", "case_id", "record_sha256", "profile_id", "profile_sha256", "requested_use", "group_id",
22
+ "split", "reviewer_id", "reviewer_qualification", "annotated_at", "mapping_label", "mapping_rationale",
23
+ "admission_label", "admission_rationale", "evidence_refs_json",
24
+ ]
25
+ OPTIONAL_ANNOTATION_COLUMNS = ["later_information_seen", "minutes_spent"]
26
+ ADJUDICATION_COLUMNS = [
27
+ "adjudication_id", "case_id", "record_sha256", "profile_sha256", "requested_use", "annotation_ids_json",
28
+ "mapping_label", "admission_label", "adjudicator_id", "adjudicated_at", "mapping_rationale",
29
+ "admission_rationale", "evidence_refs_json",
30
+ ]
31
+ PREDICTION_COLUMNS = ["case_id", "requested_use", "record_sha256", "profile_sha256", "method", "predicted_status"]
32
+ MAPPING_LABELS = ["correct", "incorrect", "uncertain"]
33
+ ADMISSION_LABELS = ["admitted", "review_required", "rejected"]
34
+ PREDICTED = ADMISSION_LABELS + ["not_admitted"] # binary baselines predict admitted / not_admitted
35
+ SPLITS = {"development", "test"}
36
+ HEX64 = re.compile(r"^[0-9a-f]{64}$")
37
+
38
+ Unit = tuple[str, str] # (case_id, requested_use)
39
+
40
+
41
+ def _read_csv(path: Path, required: list[str], optional: Sequence[str] = ()) -> list[dict[str, str]]:
42
+ with Path(path).open(encoding="utf-8-sig", newline="") as handle: # tolerate Excel's BOM
43
+ reader = csv.DictReader(handle)
44
+ header = reader.fieldnames or []
45
+ missing = [c for c in required if c not in header]
46
+ unknown = [c for c in header if c not in required and c not in optional]
47
+ if missing or unknown:
48
+ raise ValueError(f"{path}: missing columns {missing}, unknown columns {unknown}")
49
+ return [dict(row) for row in reader]
50
+
51
+
52
+ def _iso(value: str, where: str) -> None:
53
+ try:
54
+ dt.datetime.fromisoformat(value.replace("Z", "+00:00"))
55
+ except ValueError:
56
+ raise ValueError(f"{where}: not an ISO 8601 time: {value!r}") from None
57
+
58
+
59
+ def _json_list(value: str, where: str) -> list:
60
+ try:
61
+ parsed = json.loads(value)
62
+ except json.JSONDecodeError:
63
+ raise ValueError(f"{where}: not valid JSON: {value!r}") from None
64
+ if not isinstance(parsed, list) or not all(isinstance(x, str) and x.strip() for x in parsed):
65
+ raise ValueError(f"{where}: must be a JSON array of nonblank strings")
66
+ return parsed
67
+
68
+
69
+ def load_annotations(path: Path) -> list[dict[str, str]]:
70
+ """Read and strictly check an annotations CSV; raise ValueError listing the first problems."""
71
+ rows = _read_csv(path, ANNOTATION_COLUMNS, OPTIONAL_ANNOTATION_COLUMNS)
72
+ errors: list[str] = []
73
+ seen_ids: set[str] = set()
74
+ seen_reviews: set[tuple[str, str, str]] = set()
75
+ bindings: dict[Unit, tuple[str, ...]] = {}
76
+ for number, row in enumerate(rows, start=2): # line 1 is the header
77
+ where = f"{path}:{number}"
78
+ blank = [c for c in ANNOTATION_COLUMNS
79
+ if c not in ("mapping_rationale", "admission_rationale") and not (row[c] or "").strip()]
80
+ if blank:
81
+ errors.append(f"{where}: blank {blank}")
82
+ continue
83
+ if row["annotation_id"] in seen_ids:
84
+ errors.append(f"{where}: duplicate annotation_id {row['annotation_id']!r}")
85
+ seen_ids.add(row["annotation_id"])
86
+ review = (row["case_id"], row["requested_use"], row["reviewer_id"])
87
+ if review in seen_reviews:
88
+ errors.append(f"{where}: reviewer {row['reviewer_id']!r} labelled {review[:2]} twice")
89
+ seen_reviews.add(review)
90
+ if row["mapping_label"] not in MAPPING_LABELS:
91
+ errors.append(f"{where}: mapping_label must be one of {MAPPING_LABELS}")
92
+ if row["admission_label"] not in ADMISSION_LABELS:
93
+ errors.append(f"{where}: admission_label must be one of {ADMISSION_LABELS}")
94
+ if row["split"] not in SPLITS:
95
+ errors.append(f"{where}: split must be one of {sorted(SPLITS)}")
96
+ for column in ("record_sha256", "profile_sha256"):
97
+ if not HEX64.match(row[column]):
98
+ errors.append(f"{where}: {column} must be 64 lowercase hex characters")
99
+ try:
100
+ _iso(row["annotated_at"], where)
101
+ _json_list(row["evidence_refs_json"], where)
102
+ except ValueError as exc:
103
+ errors.append(str(exc))
104
+ binding = (row["record_sha256"], row["profile_sha256"], row["profile_id"], row["group_id"], row["split"])
105
+ if bindings.setdefault(review[:2], binding) != binding:
106
+ errors.append(f"{where}: reviewers of {review[:2]} saw a different record, profile, group or split")
107
+ if errors:
108
+ raise ValueError("Invalid annotations:\n" + "\n".join(errors[:20]))
109
+ return rows
110
+
111
+
112
+ def load_adjudications(path: Path, annotations: list[dict[str, str]]) -> list[dict[str, str]]:
113
+ rows = _read_csv(path, ADJUDICATION_COLUMNS)
114
+ by_id = {a["annotation_id"]: a for a in annotations}
115
+ errors, units = [], set()
116
+ for number, row in enumerate(rows, start=2):
117
+ where = f"{path}:{number}"
118
+ blank = [c for c in ADJUDICATION_COLUMNS if not (row[c] or "").strip()]
119
+ if blank:
120
+ errors.append(f"{where}: blank {blank}")
121
+ continue
122
+ unit = (row["case_id"], row["requested_use"])
123
+ if unit in units:
124
+ errors.append(f"{where}: {unit} adjudicated twice")
125
+ units.add(unit)
126
+ if row["mapping_label"] not in MAPPING_LABELS or row["admission_label"] not in ADMISSION_LABELS:
127
+ errors.append(f"{where}: labels must be from {MAPPING_LABELS} and {ADMISSION_LABELS}")
128
+ try:
129
+ _iso(row["adjudicated_at"], where)
130
+ _json_list(row["evidence_refs_json"], where)
131
+ linked = _json_list(row["annotation_ids_json"], where)
132
+ except ValueError as exc:
133
+ errors.append(str(exc))
134
+ continue
135
+ for annotation_id in linked:
136
+ original = by_id.get(annotation_id)
137
+ if original is None or (original["case_id"], original["requested_use"]) != unit:
138
+ errors.append(f"{where}: {annotation_id!r} is not an annotation of {unit}")
139
+ elif (original["record_sha256"], original["profile_sha256"]) != (row["record_sha256"], row["profile_sha256"]):
140
+ errors.append(f"{where}: record or profile hash differs from annotation {annotation_id!r}")
141
+ if errors:
142
+ raise ValueError("Invalid adjudications:\n" + "\n".join(errors[:20]))
143
+ return rows
144
+
145
+
146
+ def _units(annotations: list[dict[str, str]], field: str, use: str | None = None) -> dict[Unit, dict[str, str]]:
147
+ units: dict[Unit, dict[str, str]] = defaultdict(dict)
148
+ for row in annotations:
149
+ if use is None or row["requested_use"] == use:
150
+ units[(row["case_id"], row["requested_use"])][row["reviewer_id"]] = row[field]
151
+ return units
152
+
153
+
154
+ def krippendorff_alpha_nominal(units: list[list[str]]) -> float | None:
155
+ """Krippendorff's alpha for nominal data; units with fewer than two values are not pairable.
156
+
157
+ Returns None when there is no variation to measure agreement against (expected disagreement is 0).
158
+ """
159
+ coincidence: dict[tuple[str, str], float] = defaultdict(float)
160
+ for values in units:
161
+ if len(values) < 2:
162
+ continue
163
+ for i, c in enumerate(values):
164
+ for j, k in enumerate(values):
165
+ if i != j:
166
+ coincidence[(c, k)] += 1 / (len(values) - 1)
167
+ totals: dict[str, float] = defaultdict(float)
168
+ for (c, _), weight in coincidence.items():
169
+ totals[c] += weight
170
+ n = sum(totals.values())
171
+ observed = sum(w for (c, k), w in coincidence.items() if c != k)
172
+ expected = (n * n - sum(t * t for t in totals.values())) / (n - 1) if n > 1 else 0.0
173
+ return None if expected == 0 else 1 - observed / expected
174
+
175
+
176
+ def cohen_kappa(pairs: list[tuple[str, str]]) -> float | None:
177
+ if not pairs:
178
+ return None
179
+ n = len(pairs)
180
+ observed = sum(a == b for a, b in pairs) / n
181
+ first, second = Counter(a for a, _ in pairs), Counter(b for _, b in pairs)
182
+ expected = sum(first[c] * second[c] for c in set(first) | set(second)) / (n * n)
183
+ return None if expected == 1 else (observed - expected) / (1 - expected)
184
+
185
+
186
+ def wilson(events: int, n: int, z: float = 1.959963984540054) -> list[float] | None:
187
+ if not n:
188
+ return None
189
+ p = events / n
190
+ centre = (p + z * z / (2 * n)) / (1 + z * z / n)
191
+ half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
192
+ return [round(max(0.0, centre - half), 6), round(min(1.0, centre + half), 6)]
193
+
194
+
195
+ def _bootstrap_alpha(units: list[list[str]], resamples: int, seed: int) -> list[float] | None:
196
+ rng = random.Random(seed)
197
+ values = []
198
+ for _ in range(resamples):
199
+ alpha = krippendorff_alpha_nominal([units[rng.randrange(len(units))] for _ in units])
200
+ if alpha is not None:
201
+ values.append(alpha)
202
+ if len(values) < resamples * 0.9:
203
+ return None # too many resamples without variation for a meaningful interval
204
+ values.sort()
205
+ last = len(values) - 1
206
+ return [round(values[int(0.025 * last)], 4), round(values[math.ceil(0.975 * last)], 4)]
207
+
208
+
209
+ def agreement(annotations: list[dict[str, str]], *, resamples: int = 2000, seed: int = 20260925) -> dict[str, Any]:
210
+ """Inter-reviewer agreement per requested use, before adjudication."""
211
+ reviewers = sorted({row["reviewer_id"] for row in annotations})
212
+ result: dict[str, Any] = {"reviewers": reviewers, "uses": {}}
213
+ for use in sorted({row["requested_use"] for row in annotations}):
214
+ per_field = {}
215
+ for field in ("admission_label", "mapping_label"):
216
+ units = _units(annotations, field, use)
217
+ multi = [list(v.values()) for v in units.values() if len(v) >= 2]
218
+ alpha = krippendorff_alpha_nominal(multi)
219
+ entry: dict[str, Any] = {
220
+ "units": len(units), "units_with_2plus_reviews": len(multi),
221
+ "percent_agreement": round(sum(len(set(v)) == 1 for v in multi) / len(multi), 4) if multi else None,
222
+ "krippendorff_alpha": None if alpha is None else round(alpha, 4),
223
+ "alpha_95_bootstrap": _bootstrap_alpha(multi, resamples, seed) if multi else None,
224
+ }
225
+ if len(reviewers) == 2:
226
+ a, b = reviewers
227
+ pairs = [(v[a], v[b]) for v in units.values() if a in v and b in v]
228
+ kappa = cohen_kappa(pairs)
229
+ entry["cohen_kappa"] = None if kappa is None else round(kappa, 4)
230
+ entry["confusion"] = {f"{x}|{y}": n for (x, y), n in sorted(Counter(pairs).items())}
231
+ per_field[field] = entry
232
+ result["uses"][use] = per_field
233
+ return result
234
+
235
+
236
+ def adjudication_sheet(annotations: list[dict[str, str]], *, min_reviewers: int = 2) -> tuple[list[dict], list[Unit]]:
237
+ """Blank adjudication rows for every disagreement, plus units that still lack enough reviews."""
238
+ grouped: dict[Unit, list[dict[str, str]]] = defaultdict(list)
239
+ for row in annotations:
240
+ grouped[(row["case_id"], row["requested_use"])].append(row)
241
+ rows, incomplete = [], []
242
+ for unit, reviews in sorted(grouped.items()):
243
+ if len(reviews) < min_reviewers:
244
+ incomplete.append(unit)
245
+ continue
246
+ if len({r["admission_label"] for r in reviews}) > 1 or len({r["mapping_label"] for r in reviews}) > 1:
247
+ first = reviews[0]
248
+ rows.append({"adjudication_id": f"adj:{unit[0]}:{unit[1]}", "case_id": unit[0],
249
+ "record_sha256": first["record_sha256"], "profile_sha256": first["profile_sha256"],
250
+ "requested_use": unit[1],
251
+ "annotation_ids_json": json.dumps(sorted(r["annotation_id"] for r in reviews)),
252
+ **{c: "" for c in ADJUDICATION_COLUMNS[6:]}})
253
+ return rows, incomplete
254
+
255
+
256
+ def resolve(annotations: list[dict[str, str]], adjudications: list[dict[str, str]], *,
257
+ min_reviewers: int = 2) -> dict[Unit, dict[str, str]]:
258
+ """Final reference labels: unanimous reviews, or the adjudicated label where reviewers disagreed."""
259
+ grouped: dict[Unit, list[dict[str, str]]] = defaultdict(list)
260
+ for row in annotations:
261
+ grouped[(row["case_id"], row["requested_use"])].append(row)
262
+ decided = {(a["case_id"], a["requested_use"]): a for a in adjudications}
263
+ final, problems = {}, []
264
+ for unit, reviews in sorted(grouped.items()):
265
+ first = reviews[0]
266
+ base = {"record_sha256": first["record_sha256"], "profile_sha256": first["profile_sha256"],
267
+ "split": first["split"], "group_id": first["group_id"], "reviews": str(len(reviews))}
268
+ if unit in decided:
269
+ final[unit] = {**base, "admission_label": decided[unit]["admission_label"],
270
+ "mapping_label": decided[unit]["mapping_label"], "resolution": "adjudicated"}
271
+ elif len(reviews) < min_reviewers:
272
+ problems.append(f"{unit}: {len(reviews)} review(s), {min_reviewers} required")
273
+ elif len({r["admission_label"] for r in reviews}) == 1 and len({r["mapping_label"] for r in reviews}) == 1:
274
+ final[unit] = {**base, "admission_label": first["admission_label"],
275
+ "mapping_label": first["mapping_label"], "resolution": "unanimous"}
276
+ else:
277
+ problems.append(f"{unit}: reviewers disagree and there is no adjudication")
278
+ unknown = sorted(set(decided) - set(grouped))
279
+ problems += [f"{unit}: adjudicated but never reviewed" for unit in unknown]
280
+ if problems:
281
+ raise ValueError("Unresolved reference labels:\n" + "\n".join(problems[:20]))
282
+ return final
283
+
284
+
285
+ def load_predictions(path: Path) -> list[dict[str, str]]:
286
+ rows = _read_csv(path, PREDICTION_COLUMNS, ["subset"])
287
+ seen = set()
288
+ for number, row in enumerate(rows, start=2):
289
+ key = (row["case_id"], row["requested_use"], row["method"])
290
+ if key in seen:
291
+ raise ValueError(f"{path}:{number}: duplicate prediction for {key}")
292
+ seen.add(key)
293
+ if row["predicted_status"] not in PREDICTED:
294
+ raise ValueError(f"{path}:{number}: predicted_status must be one of {PREDICTED}")
295
+ return rows
296
+
297
+
298
+ def score(final: dict[Unit, dict[str, str]], predictions: list[dict[str, str]], *, split: str = "test") -> dict[str, Any]:
299
+ """Compare predictions with reference labels on one split; hashes must match the reviewed records."""
300
+ selected = {u: v for u, v in final.items() if v["split"] == split}
301
+ if not selected:
302
+ raise ValueError(f"No reference labels in split {split!r}")
303
+ rows: dict[tuple[str, str], list[tuple[str, str]]] = defaultdict(list) # (method, subset) -> (gold, predicted)
304
+ for p in predictions:
305
+ unit = (p["case_id"], p["requested_use"])
306
+ if unit not in final:
307
+ continue
308
+ gold = final[unit]
309
+ if (p["record_sha256"], p["profile_sha256"]) != (gold["record_sha256"], gold["profile_sha256"]):
310
+ raise ValueError(f"{unit}: prediction was made on a different record or profile than was reviewed")
311
+ if unit in selected:
312
+ for subset in {"all", p.get("subset") or "all"}:
313
+ rows[(p["method"], subset)].append((gold["admission_label"], p["predicted_status"]))
314
+ predicted = {(p["method"], p["case_id"], p["requested_use"]) for p in predictions}
315
+ missing = [(m, u) for m in sorted({p["method"] for p in predictions}) for u in sorted(selected)
316
+ if (m, *u) not in predicted]
317
+ if missing:
318
+ raise ValueError(f"Missing predictions for {len(missing)} reviewed unit(s), e.g. {missing[0]}")
319
+ report: dict[str, Any] = {"split": split, "reference_units": len(selected), "methods": {}}
320
+ for (method, subset), pairs in sorted(rows.items()):
321
+ positives = [p for g, p in pairs if g == "admitted"]
322
+ negatives = [p for g, p in pairs if g != "admitted"]
323
+ false_admissions = sum(p == "admitted" for p in negatives)
324
+ false_blocks = sum(p != "admitted" for p in positives)
325
+ report["methods"].setdefault(method, {})[subset] = {
326
+ "n": len(pairs), "reference_admitted": len(positives), "reference_not_admitted": len(negatives),
327
+ "false_admissions": false_admissions, "false_admission_rate_wilson_95": wilson(false_admissions, len(negatives)),
328
+ "false_blocks": false_blocks, "false_block_rate_wilson_95": wilson(false_blocks, len(positives)),
329
+ "admission_agreement": sum((g == "admitted") == (p == "admitted") for g, p in pairs),
330
+ "exact_matches": sum(g == p for g, p in pairs),
331
+ "confusion": {f"{g}|{p}": n for (g, p), n in sorted(Counter(pairs).items())},
332
+ }
333
+ return report
334
+
335
+
336
+ def sha256_file(path: Path) -> str:
337
+ return hashlib.sha256(Path(path).read_bytes()).hexdigest()
338
+
339
+
340
+ def freeze(manifest: dict[str, Any], final: dict[Unit, dict[str, str]], annotations: list[dict[str, str]],
341
+ files: list[Path], *, dataset_id: str, version: str, frozen_at: str) -> dict[str, Any]:
342
+ """Fill a manifest for a completed reference set. Hashes identify files; they do not prove review happened."""
343
+ _iso(frozen_at, "frozen_at")
344
+ reviewers = {row["reviewer_id"] for row in annotations}
345
+ single = all(v["reviews"] == "1" for v in final.values())
346
+ return {**manifest, "dataset_id": dataset_id, "version": version, "status": "frozen",
347
+ "reference_type": "single-reviewer reference set" if single else "independently reviewed reference set",
348
+ "reviewed_case_use_count": len(final), "independent_human_reviewer_count": len(reviewers),
349
+ "resolution_counts": dict(sorted(Counter(v["resolution"] for v in final.values()).items())),
350
+ "split_groups": sorted({v["group_id"] for v in final.values()}),
351
+ "files": {Path(f).name: sha256_file(f) for f in files}, "frozen_at": frozen_at,
352
+ "gold_standard_metrics_available": True}
353
+
354
+
355
+ def _fmt(value: Any) -> str:
356
+ return "–" if value is None else str(value)
357
+
358
+
359
+ def render_agreement(report: dict[str, Any]) -> str:
360
+ lines = [f"Reviewers: {', '.join(report['reviewers'])}", "",
361
+ "| Use | Label | Units (2+ reviews) | % agreement | Krippendorff α (95% bootstrap) | Cohen κ |",
362
+ "|---|---|---:|---:|---|---:|"]
363
+ for use, fields in report["uses"].items():
364
+ for field, e in fields.items():
365
+ ci = e["alpha_95_bootstrap"]
366
+ alpha = _fmt(e["krippendorff_alpha"]) + (f" ({ci[0]}–{ci[1]})" if ci else "")
367
+ lines.append(f"| {use} | {field} | {e['units_with_2plus_reviews']} | {_fmt(e['percent_agreement'])} | "
368
+ f"{alpha} | {_fmt(e.get('cohen_kappa'))} |")
369
+ return "\n".join(lines) + "\n"
370
+
371
+
372
+ def render_score(report: dict[str, Any]) -> str:
373
+ lines = [f"Split: {report['split']} · reference units: {report['reference_units']}", "",
374
+ "| Method | Subset | n | False admissions | False blocks | Admission agreement | Exact |",
375
+ "|---|---|---:|---:|---:|---:|---:|"]
376
+ for method, subsets in report["methods"].items():
377
+ for subset, m in subsets.items():
378
+ lines.append(f"| {method} | {subset} | {m['n']} | {m['false_admissions']}/{m['reference_not_admitted']} | "
379
+ f"{m['false_blocks']}/{m['reference_admitted']} | {m['admission_agreement']}/{m['n']} | "
380
+ f"{m['exact_matches']}/{m['n']} |")
381
+ return "\n".join(lines) + "\n"
382
+
383
+
384
+ def write_csv(path: Path, rows: list[dict[str, str]], columns: list[str]) -> None:
385
+ with Path(path).open("w", encoding="utf-8", newline="") as handle:
386
+ writer = csv.DictWriter(handle, fieldnames=columns, lineterminator="\n")
387
+ writer.writeheader()
388
+ writer.writerows(rows)
@@ -10,6 +10,8 @@ prefixes:
10
10
  linkml: https://w3id.org/linkml/
11
11
  prov: http://www.w3.org/ns/prov#
12
12
  sepio: http://purl.obolibrary.org/obo/SEPIO_
13
+ ECO: http://purl.obolibrary.org/obo/ECO_
14
+ biolink: https://w3id.org/biolink/vocab/
13
15
  default_prefix: bioev
14
16
  default_range: string
15
17
  imports:
@@ -33,12 +35,24 @@ enums:
33
35
  supports: null
34
36
  contradicts: null
35
37
  neutral: null
38
+ # How an evidence item was produced. Meanings are the closest ECO terms (release
39
+ # 2026-07-10); see docs/STANDARDS.md for Biolink agent types and caveats. Kept as a
40
+ # comment so the compiled JSON Schema, and every report's schema_sha256, are unchanged.
36
41
  ExtractionMethod:
37
42
  permissible_values:
38
- deterministic_parser: null
39
- manual_curation: null
40
- normalized_string_match: null
41
- llm_extraction: null
43
+ deterministic_parser:
44
+ description: Rule-based import of an existing source record.
45
+ meaning: ECO:0000313
46
+ manual_curation:
47
+ description: A person read the source and recorded the evidence.
48
+ meaning: ECO:0000352
49
+ normalized_string_match:
50
+ description: Matching normalized strings, e.g. names to identifiers.
51
+ meaning: ECO:0008021
52
+ llm_extraction:
53
+ description: A language model extracted the evidence from text. ECO has no
54
+ LLM-specific term; the machine-learning term is the closest.
55
+ meaning: ECO:0008004
42
56
  StatementStatus:
43
57
  permissible_values:
44
58
  proposed: null
@@ -1,13 +0,0 @@
1
- bioevidence_validator/__init__.py,sha256=w-OLV_MserELKh9U9FrLM8ZICJqNoWNOlqMwglL1lIk,87
2
- bioevidence_validator/cli.py,sha256=fknYldtjN50OvguGMcp_kmAky5PxZ0i5WRUF4HEQVkA,4326
3
- bioevidence_validator/config.py,sha256=Ahc9eRPNR_ArqnbnJkdZsPFq9yERTYied0oGwUkANYI,1110
4
- bioevidence_validator/engine.py,sha256=fm1gIOySsLiBbI3SklbEStiMD93RrhR3lA0vqdLJJXQ,15003
5
- bioevidence_validator/profiles/dataset-label.yaml,sha256=4utz9YcFFF2UnaLjIb0ATb-27hF8Smw7zmv0sBTaQdM,843
6
- bioevidence_validator/profiles/general.yaml,sha256=i9fJCWovJZ2e2d4rbn0Bx7mrhumNb2zjS9TOarkBWZw,778
7
- bioevidence_validator/profiles/literature-claim.yaml,sha256=7l82gX3oGGyYFq9B29-1qGXeL4LRcMhJi8Y9fieVE0c,612
8
- bioevidence_validator/schema/bioevidence_core.yaml,sha256=18MMngua0jlu4VnAbpb7SfI3numnCdjO986bZP6-Xn4,6700
9
- bioai_evidence_validator-0.4.1.dist-info/METADATA,sha256=JkN3mmJrD-cx86fBjdrjrTaxadrM1SdzbTrtHv81uO0,12563
10
- bioai_evidence_validator-0.4.1.dist-info/WHEEL,sha256=W3fkpkm7-wf9vBI5Z-7s0eWkeM-spu78I8Neb98DeEg,87
11
- bioai_evidence_validator-0.4.1.dist-info/entry_points.txt,sha256=6oufzT2x1nIhTqM5B6y4aWubExKAoFb_wXXTdz3mBQQ,63
12
- bioai_evidence_validator-0.4.1.dist-info/licenses/LICENSE,sha256=z8d0m5b2O9McPEK1xHG_dWgUBT6EfBDz6wA0F7xSPTA,11358
13
- bioai_evidence_validator-0.4.1.dist-info/RECORD,,