bioai-evidence-validator 0.5.0__tar.gz → 0.7.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (114) hide show
  1. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.gitattributes +1 -0
  2. bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/bug_report.yml +53 -0
  3. bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/config.yml +8 -0
  4. bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/feature_request.yml +28 -0
  5. bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/profile_proposal.yml +42 -0
  6. bioai_evidence_validator-0.7.0/.github/pull_request_template.md +18 -0
  7. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.github/workflows/ci.yml +23 -2
  8. bioai_evidence_validator-0.7.0/.github/workflows/docs.yml +43 -0
  9. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.gitignore +4 -0
  10. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/CHANGELOG.md +33 -0
  11. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/CITATION.cff +1 -1
  12. bioai_evidence_validator-0.7.0/CODE_OF_CONDUCT.md +16 -0
  13. bioai_evidence_validator-0.7.0/CONTRIBUTING.md +76 -0
  14. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/PKG-INFO +65 -13
  15. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/README.md +57 -11
  16. bioai_evidence_validator-0.7.0/SECURITY.md +35 -0
  17. bioai_evidence_validator-0.7.0/community/profiles/README.md +37 -0
  18. bioai_evidence_validator-0.7.0/community/profiles/_template/README.md +18 -0
  19. bioai_evidence_validator-0.7.0/community/profiles/_template/cases/curated_assay.yaml +14 -0
  20. bioai_evidence_validator-0.7.0/community/profiles/_template/cases/llm_only.yaml +14 -0
  21. bioai_evidence_validator-0.7.0/community/profiles/_template/cases/network_without_review.yaml +14 -0
  22. bioai_evidence_validator-0.7.0/community/profiles/_template/cases/synthetic_study.txt +2 -0
  23. bioai_evidence_validator-0.7.0/community/profiles/_template/expected.yaml +14 -0
  24. bioai_evidence_validator-0.7.0/community/profiles/_template/profile.yaml +19 -0
  25. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/ENGINEERING.md +19 -1
  26. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/GOLD_STANDARD.md +10 -5
  27. bioai_evidence_validator-0.7.0/docs/STANDARDS.md +79 -0
  28. bioai_evidence_validator-0.7.0/docs/assets/clinvar_germline_benchmark.svg +4101 -0
  29. bioai_evidence_validator-0.7.0/docs/index.md +58 -0
  30. bioai_evidence_validator-0.7.0/evaluation/clinvar_review/README.md +125 -0
  31. bioai_evidence_validator-0.7.0/evaluation/clinvar_review/RUBRIC.md +74 -0
  32. bioai_evidence_validator-0.7.0/evaluation/clinvar_review/import_sheets.py +126 -0
  33. bioai_evidence_validator-0.7.0/evaluation/clinvar_review/prepare_packets.py +221 -0
  34. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/README.md +13 -3
  35. bioai_evidence_validator-0.7.0/examples/clinvar_germline/README.md +178 -0
  36. bioai_evidence_validator-0.7.0/examples/clinvar_germline/pipeline.py +231 -0
  37. bioai_evidence_validator-0.7.0/examples/clinvar_germline/plot.py +135 -0
  38. bioai_evidence_validator-0.7.0/examples/clinvar_germline/prepare_source.py +200 -0
  39. bioai_evidence_validator-0.7.0/examples/clinvar_germline/profile.yaml +24 -0
  40. bioai_evidence_validator-0.7.0/examples/clinvar_germline/results/divergences.csv +246 -0
  41. bioai_evidence_validator-0.7.0/examples/clinvar_germline/results/summary.json +556 -0
  42. bioai_evidence_validator-0.7.0/examples/clinvar_germline/results/summary.md +48 -0
  43. bioai_evidence_validator-0.7.0/examples/clinvar_germline/run.py +179 -0
  44. bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/README.md +33 -0
  45. bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/clinvar-sample.jsonl.gz +0 -0
  46. bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/manifest.json +39 -0
  47. bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/population_outcomes.json +121 -0
  48. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/pipeline.py +1 -1
  49. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/summary.json +1 -1
  50. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/run.py +2 -1
  51. bioai_evidence_validator-0.7.0/mkdocs.yml +62 -0
  52. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/pyproject.toml +41 -3
  53. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/__init__.py +1 -1
  54. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/cli.py +77 -1
  55. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/draft.py +9 -6
  56. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/engine.py +2 -2
  57. bioai_evidence_validator-0.7.0/src/bioevidence_validator/review.py +388 -0
  58. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/schema/bioevidence_core.yaml +18 -4
  59. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_cli.py +1 -2
  60. bioai_evidence_validator-0.7.0/tests/test_clinvar_case.py +121 -0
  61. bioai_evidence_validator-0.7.0/tests/test_clinvar_review_kit.py +138 -0
  62. bioai_evidence_validator-0.7.0/tests/test_community_profiles.py +39 -0
  63. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_engine.py +2 -1
  64. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_evidence_quality.py +3 -1
  65. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_fail_closed.py +2 -0
  66. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_profiles.py +2 -2
  67. bioai_evidence_validator-0.7.0/tests/test_repository_files.py +21 -0
  68. bioai_evidence_validator-0.7.0/tests/test_review.py +230 -0
  69. bioai_evidence_validator-0.7.0/tests/test_standards.py +34 -0
  70. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_vbo_case.py +3 -1
  71. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tools/check_distribution.py +10 -1
  72. bioai_evidence_validator-0.7.0/tools/mkdocs_hooks.py +73 -0
  73. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/uv.lock +508 -3
  74. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.github/workflows/release.yml +0 -0
  75. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/LICENSE +0 -0
  76. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/action.yml +0 -0
  77. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/ADR-001-canine-breed-first.md +0 -0
  78. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/ADR-002-domain-neutral-main.md +0 -0
  79. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/CASE_STUDY.md +0 -0
  80. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/DRAFTS.md +0 -0
  81. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/MIGRATION-0.4.md +0 -0
  82. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/PROFILES.md +0 -0
  83. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/assets/vbo_canine_benchmark.svg +0 -0
  84. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/adjudications.template.csv +0 -0
  85. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/annotations.template.csv +0 -0
  86. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/manifest.template.json +0 -0
  87. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/custom_profile/assay.yaml +0 -0
  88. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/custom_profile/assay_record.json +0 -0
  89. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/dataset_label/curated_sample_label.json +0 -0
  90. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/dataset_label/missing_sample_link.json +0 -0
  91. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/drafts/llm_claim.yaml +0 -0
  92. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/drafts/reviewed_claim.yaml +0 -0
  93. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/drafts/synthetic_paper.txt +0 -0
  94. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/general/curated_assertion.json +0 -0
  95. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/literature_claim/curated_association.json +0 -0
  96. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/literature_claim/llm_only.json +0 -0
  97. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/quickstart.ipynb +0 -0
  98. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/README.md +0 -0
  99. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/prepare_source.py +0 -0
  100. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/profile.yaml +0 -0
  101. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/reference_cases.json +0 -0
  102. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/decisions.jsonl +0 -0
  103. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/review_queue.csv +0 -0
  104. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/summary.md +0 -0
  105. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/sources/README.md +0 -0
  106. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/sources/manifest.json +0 -0
  107. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/sources/vbo-dogs.json +0 -0
  108. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/config.py +0 -0
  109. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/profiles/dataset-label.yaml +0 -0
  110. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/profiles/general.yaml +0 -0
  111. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/profiles/literature-claim.yaml +0 -0
  112. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_draft.py +0 -0
  113. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_github_action.py +0 -0
  114. {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tools/github_action.py +0 -0
@@ -1 +1,2 @@
1
1
  * text=auto eol=lf
2
+ *.gz binary
@@ -0,0 +1,53 @@
1
+ name: Bug report
2
+ description: The validator, CLI or GitHub Action does something it should not.
3
+ labels: [bug]
4
+ body:
5
+ - type: markdown
6
+ attributes:
7
+ value: |
8
+ Thanks for reporting. For security problems (for example a record admitted that a
9
+ documented rule should reject), use [private reporting](https://github.com/NingyuSUN/bioai-evidence-validator/security/advisories/new) instead.
10
+ **Do not paste real patient data**; a synthetic record that shows the behavior is enough.
11
+ - type: input
12
+ id: version
13
+ attributes:
14
+ label: Version
15
+ description: Output of `pip show bioai-evidence-validator` or `bioevidence --help` header, or the commit.
16
+ placeholder: "0.7.0"
17
+ validations:
18
+ required: true
19
+ - type: textarea
20
+ id: command
21
+ attributes:
22
+ label: Command or code
23
+ description: The exact command or Python call, including `--profile`.
24
+ render: shell
25
+ validations:
26
+ required: true
27
+ - type: textarea
28
+ id: input
29
+ attributes:
30
+ label: Smallest record, draft or profile that reproduces it
31
+ render: yaml
32
+ validations:
33
+ required: true
34
+ - type: textarea
35
+ id: expected
36
+ attributes:
37
+ label: Expected result
38
+ description: Which status, reason code or exit code did you expect, and which rule says so?
39
+ validations:
40
+ required: true
41
+ - type: textarea
42
+ id: actual
43
+ attributes:
44
+ label: Actual result
45
+ description: The report's `overall_status`, `findings` and exit code, or the error message.
46
+ render: json
47
+ validations:
48
+ required: true
49
+ - type: input
50
+ id: environment
51
+ attributes:
52
+ label: Environment
53
+ placeholder: "Python 3.12, Ubuntu 24.04"
@@ -0,0 +1,8 @@
1
+ blank_issues_enabled: true
2
+ contact_links:
3
+ - name: Security vulnerability
4
+ url: https://github.com/NingyuSUN/bioai-evidence-validator/security/advisories/new
5
+ about: Report privately; please do not open a public issue.
6
+ - name: Documentation
7
+ url: https://github.com/NingyuSUN/bioai-evidence-validator#readme
8
+ about: Usage, draft format, profiles and benchmarks.
@@ -0,0 +1,28 @@
1
+ name: Feature request
2
+ description: Propose a change to the engine, CLI, formats or integrations.
3
+ labels: [enhancement]
4
+ body:
5
+ - type: textarea
6
+ id: problem
7
+ attributes:
8
+ label: Problem
9
+ description: What are you trying to do, and what gets in the way today?
10
+ validations:
11
+ required: true
12
+ - type: textarea
13
+ id: proposal
14
+ attributes:
15
+ label: Proposal
16
+ description: The behavior you would like. For new rules, describe a record that should pass and one that should not.
17
+ validations:
18
+ required: true
19
+ - type: textarea
20
+ id: failure
21
+ attributes:
22
+ label: Failure mode
23
+ description: If this goes wrong, does it admit too much or block too much? How would a user notice?
24
+ - type: textarea
25
+ id: alternatives
26
+ attributes:
27
+ label: Alternatives considered
28
+ description: For example, doing it in a profile or an importer instead of the engine.
@@ -0,0 +1,42 @@
1
+ name: Domain profile proposal
2
+ description: Propose an admission profile for a biological domain.
3
+ labels: [profile]
4
+ body:
5
+ - type: markdown
6
+ attributes:
7
+ value: |
8
+ Profiles decide which evidence each intended use requires; they do not change the engine.
9
+ See [docs/PROFILES.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/PROFILES.md)
10
+ and the [community profile template](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/community/profiles/_template).
11
+ - type: input
12
+ id: domain
13
+ attributes:
14
+ label: Domain and assertion
15
+ placeholder: "Protein–protein interactions: protein A interacts_with protein B"
16
+ validations:
17
+ required: true
18
+ - type: textarea
19
+ id: uses
20
+ attributes:
21
+ label: Intended uses and their evidence requirements
22
+ description: For each use (e.g. research summary, knowledge-base admission, training data), which evidence types are required, and is human acceptance required?
23
+ validations:
24
+ required: true
25
+ - type: textarea
26
+ id: basis
27
+ attributes:
28
+ label: Basis for the policy
29
+ description: Community guideline, database policy or published standard the requirements follow (with links), or state that it is your own proposal.
30
+ validations:
31
+ required: true
32
+ - type: textarea
33
+ id: cases
34
+ attributes:
35
+ label: Example cases
36
+ description: At least one claim that should be admitted and one that should not, with the reason.
37
+ - type: checkboxes
38
+ id: pr
39
+ attributes:
40
+ label: Contribution
41
+ options:
42
+ - label: I am willing to open a pull request with the profile and example cases.
@@ -0,0 +1,18 @@
1
+ ## What and why
2
+
3
+ <!-- One purpose per pull request. Link the issue: "Closes #123". -->
4
+
5
+ ## Admission boundary
6
+
7
+ <!-- Does this change what gets admitted, sent to review or rejected? If yes, say which
8
+ records move and why, and point to the tests that pin the new boundary. -->
9
+
10
+ - [ ] No change to admission decisions
11
+ - [ ] Changes admission decisions (explained above; tests on both sides of the boundary)
12
+
13
+ ## Checklist
14
+
15
+ - [ ] `uv run --frozen ruff check .`, `uv run --frozen mypy` and `uv run --frozen pytest --cov` pass
16
+ - [ ] Committed benchmark results regenerated if report contents changed
17
+ - [ ] `CHANGELOG.md` updated if users will notice
18
+ - [ ] New third-party data has its license, attribution and exact version recorded
@@ -6,6 +6,23 @@ on:
6
6
  permissions:
7
7
  contents: read
8
8
  jobs:
9
+ lint:
10
+ timeout-minutes: 10
11
+ runs-on: ubuntu-latest
12
+ steps:
13
+ - uses: actions/checkout@v4
14
+ - uses: actions/setup-python@v5
15
+ with:
16
+ python-version: '3.12'
17
+ - name: Install locked tooling
18
+ run: python -m pip install uv==0.12.17
19
+ - name: Install locked project
20
+ run: uv sync --frozen --extra dev
21
+ - name: Lint
22
+ run: uv run --frozen ruff check --output-format github .
23
+ - name: Type-check
24
+ run: uv run --frozen mypy
25
+
9
26
  verify:
10
27
  timeout-minutes: 15
11
28
  strategy:
@@ -29,8 +46,12 @@ jobs:
29
46
  run: python -m pip install uv==0.12.17
30
47
  - name: Install locked project
31
48
  run: uv sync --frozen --extra dev --python ${{ matrix.python }}
32
- - name: Run tests
33
- run: uv run --frozen pytest
49
+ - name: Run tests with coverage (minimum in pyproject.toml)
50
+ run: uv run --frozen pytest --cov --cov-report=term
51
+ - name: Coverage summary
52
+ if: always()
53
+ shell: bash
54
+ run: uv run --frozen coverage report --format=markdown >> "$GITHUB_STEP_SUMMARY" || true
34
55
  - name: Build distributable
35
56
  run: uv build
36
57
  - name: Verify installed wheel and CLI outside editable source
@@ -0,0 +1,43 @@
1
+ name: Documentation
2
+ on:
3
+ push:
4
+ branches: [main]
5
+ pull_request:
6
+ workflow_dispatch:
7
+ permissions:
8
+ contents: read
9
+ concurrency:
10
+ group: docs-${{ github.ref }}
11
+ cancel-in-progress: true
12
+ jobs:
13
+ build:
14
+ timeout-minutes: 10
15
+ runs-on: ubuntu-latest
16
+ steps:
17
+ - uses: actions/checkout@v4
18
+ - uses: actions/setup-python@v5
19
+ with:
20
+ python-version: '3.12'
21
+ # Pinned: MkDocs 2.0 drops the plugin and theme system this site uses.
22
+ - name: Install MkDocs
23
+ run: python -m pip install mkdocs==1.6.1 mkdocs-material==9.7.7
24
+ - name: Build with strict link checking
25
+ run: mkdocs build --strict --site-dir _site
26
+ - if: github.event_name != 'pull_request'
27
+ uses: actions/upload-pages-artifact@v3
28
+ with:
29
+ path: _site
30
+
31
+ deploy:
32
+ if: github.event_name != 'pull_request'
33
+ needs: build
34
+ runs-on: ubuntu-latest
35
+ permissions:
36
+ pages: write
37
+ id-token: write
38
+ environment:
39
+ name: github-pages
40
+ url: ${{ steps.deployment.outputs.page_url }}
41
+ steps:
42
+ - id: deployment
43
+ uses: actions/deploy-pages@v4
@@ -7,3 +7,7 @@ dist/
7
7
  build/
8
8
  artifacts/
9
9
  validation_report.json
10
+ _site/
11
+ site/
12
+ .coverage
13
+ htmlcov/
@@ -1,5 +1,38 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.7.0 — Expert review, quality checks, community and documentation site
4
+
5
+ - Add `bioevidence review` (`check`, `agreement`, `adjudication-sheet`, `score`, `freeze`)
6
+ implementing the gold-standard protocol: strict annotation checks, Krippendorff's α with
7
+ bootstrap interval and Cohen's κ (both cross-checked against reference implementations),
8
+ adjudication of disagreements, scoring on the test split with hash-bound joins, and frozen
9
+ manifests. It never produces labels.
10
+ - Add a blinded ClinVar expert-review kit (`evaluation/clinvar_review/`): all 95 variants where
11
+ the validator and NCBI disagree plus 95 stratum-matched controls, an Excel workbook with
12
+ dropdowns, a reviewer rubric, a private key, and an importer into the protocol format.
13
+ - CI runs ruff and mypy, and enforces a 95% test-coverage minimum (currently 98%).
14
+ - Add CONTRIBUTING, CODE_OF_CONDUCT (Contributor Covenant 2.1), SECURITY, issue forms and a
15
+ pull request template.
16
+ - Add `community/profiles/`: contributed domain profiles whose example cases are built,
17
+ validated and checked against expected outcomes in CI; `_template/` to copy.
18
+ - Add a documentation site (MkDocs, strict link checking) published to GitHub Pages.
19
+ - The canine 0.3 implementation is also preserved at the `canine-0.3` tag.
20
+ - No change to validation decisions or report format; both benchmarks reproduce 0.6.0 exactly
21
+ apart from `validator_version`.
22
+
23
+ ## 0.6.0 — ClinVar case and standards alignment
24
+
25
+ - Add a second real-data case: ClinVar germline classifications. 5,026 sampled variants
26
+ from the 2023-09 release are built from per-submission evidence and validated with a
27
+ ClinVar-style profile; decisions are compared with NCBI's own 2023-09 review status and
28
+ with each classification's 2026-09 outcome, plus controlled faults and a trust-boundary
29
+ cohort. Frozen, hash-pinned sample; rebuild script verifies the three upstream files.
30
+ - Add `docs/STANDARDS.md`, mapping the record model to ECO, Biolink 4.4.4, GA4GH VA-Spec
31
+ 1.0.1 and PROV-O, with mapping strength and caveats.
32
+ - Annotate `ExtractionMethod` values with ECO meanings in the LinkML schema. The compiled
33
+ JSON Schema, and therefore every report's `schema_sha256`, is unchanged.
34
+ - The VBO benchmark summary changes only `validator_version`.
35
+
3
36
  ## 0.5.0 — Drafts, LLM draft schema and GitHub Action
4
37
 
5
38
  - Add compact YAML/JSON drafts: `bioevidence build` and `build_record()` derive identifiers,
@@ -9,7 +9,7 @@ abstract: >-
9
9
  authors:
10
10
  - family-names: Sun
11
11
  given-names: Ningyu
12
- version: 0.5.0
12
+ version: 0.7.0
13
13
  license: Apache-2.0
14
14
  repository-code: "https://github.com/NingyuSUN/bioai-evidence-validator"
15
15
  keywords:
@@ -0,0 +1,16 @@
1
+ # Code of conduct
2
+
3
+ This project adopts the
4
+ [Contributor Covenant, version 2.1](https://www.contributor-covenant.org/version/2/1/code_of_conduct/)
5
+ as its code of conduct. It applies to all project spaces (issues, pull requests,
6
+ discussions and reviews) and to anyone representing the project elsewhere.
7
+
8
+ In short: be respectful and constructive, assume good faith, critique work rather than
9
+ people, and remember that contributors bring expertise from many fields, from biocuration
10
+ to software engineering.
11
+
12
+ ## Reporting
13
+
14
+ Report unacceptable behavior to the maintainer, Ningyu Sun, at <woshiwosunny@gmail.com>.
15
+ Reports are handled confidentially, and the enforcement guidelines of the Contributor
16
+ Covenant 2.1 apply.
@@ -0,0 +1,76 @@
1
+ # Contributing
2
+
3
+ Thanks for helping make AI-assisted biocuration safer. Contributions of every size are
4
+ welcome: a typo fix, a bug report with a failing record, a new domain profile, or a new
5
+ real-data benchmark.
6
+
7
+ By participating you agree to follow the [code of conduct](CODE_OF_CONDUCT.md).
8
+
9
+ ## Ways to contribute
10
+
11
+ | You have… | Start here |
12
+ |---|---|
13
+ | A record the validator judges wrongly | [Bug report](https://github.com/NingyuSUN/bioai-evidence-validator/issues/new?template=bug_report.yml); attach the smallest record or draft that reproduces it |
14
+ | An admission policy for your domain | [Profile proposal](https://github.com/NingyuSUN/bioai-evidence-validator/issues/new?template=profile_proposal.yml), then a pull request to [`community/profiles/`](community/profiles/README.md) |
15
+ | An idea for the engine, CLI or formats | [Feature request](https://github.com/NingyuSUN/bioai-evidence-validator/issues/new?template=feature_request.yml) first, so we can agree on the contract before code |
16
+ | A security problem | Do **not** open an issue; see [SECURITY.md](SECURITY.md) |
17
+
18
+ Issues labelled [`good first issue`](https://github.com/NingyuSUN/bioai-evidence-validator/labels/good%20first%20issue)
19
+ are scoped to be finished in an afternoon.
20
+
21
+ ## Development setup
22
+
23
+ Python 3.11+ and [uv](https://docs.astral.sh/uv/):
24
+
25
+ ```bash
26
+ git clone https://github.com/NingyuSUN/bioai-evidence-validator.git
27
+ cd bioai-evidence-validator
28
+ uv sync --frozen --extra dev
29
+ ```
30
+
31
+ Before opening a pull request, run what CI runs:
32
+
33
+ ```bash
34
+ uv run --frozen ruff check .
35
+ uv run --frozen mypy
36
+ uv run --frozen pytest --cov
37
+ ```
38
+
39
+ CI also runs the tests on Linux (Python 3.11–3.13) and Windows, builds the wheel and
40
+ smoke-tests it outside the source tree, and runs the GitHub Action. Coverage must stay at
41
+ or above the minimum in `pyproject.toml`.
42
+
43
+ ## Ground rules for changes
44
+
45
+ This project's value is that it **fails closed** and **says exactly what it checked**.
46
+ Changes are reviewed against that:
47
+
48
+ - **No silent weakening.** A change that admits something previously rejected needs an
49
+ explicit reason in the pull request and a test that shows the new boundary.
50
+ - **Every rule has a test on both sides**: a record that passes and a minimal one that fails.
51
+ - **Reports stay reproducible.** If a change alters report contents, the committed benchmark
52
+ results must be regenerated in the same pull request, and the diff explained.
53
+ - **Domain logic stays out of the engine.** New domains are profiles and importers, not
54
+ special cases in `engine.py`.
55
+ - **Honest limits.** Benchmarks state what their labels are (source-derived, authored, or
56
+ independently reviewed) and what they do not measure.
57
+ - Match the surrounding style; ruff enforces correctness rules, not formatting.
58
+
59
+ ## Contributing a domain profile
60
+
61
+ Community profiles live in [`community/profiles/`](community/profiles/README.md). Each one is
62
+ a folder with a profile, a short README and example cases whose expected outcomes are
63
+ checked by the test suite. Copy `community/profiles/_template/` to get started.
64
+
65
+ ## Pull requests
66
+
67
+ - Keep each pull request to one purpose; link the issue it resolves.
68
+ - Update `CHANGELOG.md` under a new heading if users will notice the change.
69
+ - New third-party data needs its license, attribution and exact source version recorded,
70
+ as in `examples/*/sources/README.md`.
71
+
72
+ ## Releases
73
+
74
+ Maintainers release by bumping the version in `pyproject.toml` and
75
+ `src/bioevidence_validator/__init__.py`, adding a `CHANGELOG.md` section, and pushing a
76
+ `vX.Y.Z` tag. The release workflow tests, publishes to PyPI and creates the GitHub release.
@@ -1,9 +1,9 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: bioai-evidence-validator
3
- Version: 0.5.0
3
+ Version: 0.7.0
4
4
  Summary: Standards-aligned evidence policy validation for AI-assisted biological curation
5
5
  Project-URL: Homepage, https://github.com/NingyuSUN/bioai-evidence-validator
6
- Project-URL: Documentation, https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/docs
6
+ Project-URL: Documentation, https://ningyusun.github.io/bioai-evidence-validator/
7
7
  Project-URL: Changelog, https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CHANGELOG.md
8
8
  Project-URL: Issues, https://github.com/NingyuSUN/bioai-evidence-validator/issues
9
9
  Author: Ningyu Sun
@@ -25,7 +25,13 @@ Requires-Dist: jsonschema<5,>=4.23
25
25
  Requires-Dist: linkml<2,>=1.8
26
26
  Requires-Dist: pyyaml<7,>=6.0
27
27
  Provides-Extra: dev
28
+ Requires-Dist: mypy<3,>=2.3; extra == 'dev'
29
+ Requires-Dist: openpyxl<4,>=3.1; extra == 'dev'
30
+ Requires-Dist: pytest-cov<8,>=7; extra == 'dev'
28
31
  Requires-Dist: pytest<9,>=8; extra == 'dev'
32
+ Requires-Dist: ruff<0.17,>=0.16; extra == 'dev'
33
+ Requires-Dist: types-jsonschema; extra == 'dev'
34
+ Requires-Dist: types-pyyaml; extra == 'dev'
29
35
  Description-Content-Type: text/markdown
30
36
 
31
37
  # BioAI Evidence Validator
@@ -35,6 +41,7 @@ Description-Content-Type: text/markdown
35
41
  [![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue)](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/pyproject.toml)
36
42
  [![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue)](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/LICENSE)
37
43
  [![Open in Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/NingyuSUN/bioai-evidence-validator/blob/main/examples/quickstart.ipynb)
44
+ [![Docs](https://img.shields.io/badge/docs-site-blue)](https://ningyusun.github.io/bioai-evidence-validator/)
38
45
 
39
46
  **Stop AI-extracted biological claims from entering your knowledge base or
40
47
  training set before their evidence is good enough for that use.**
@@ -52,6 +59,7 @@ pip install bioai-evidence-validator
52
59
 
53
60
  Or try it in the browser, nothing to install:
54
61
  [quickstart notebook on Colab](https://colab.research.google.com/github/NingyuSUN/bioai-evidence-validator/blob/main/examples/quickstart.ipynb).
62
+ Full documentation: **[ningyusun.github.io/bioai-evidence-validator](https://ningyusun.github.io/bioai-evidence-validator/)**.
55
63
 
56
64
  ## 30-second example
57
65
 
@@ -103,9 +111,12 @@ record is trustworthy enough for a particular purpose.
103
111
  | Provenance consistency (source hashes, resolved references, scope) | — | ✅ |
104
112
  | Human adjudications bound to a specific statement and use | — | ✅ |
105
113
  | Machine-readable audit report with hashes of input, schema and profile | — | ✅ |
114
+ | Mapped to ECO, Biolink and GA4GH VA-Spec ([standards alignment](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/STANDARDS.md)) | — | ✅ |
106
115
 
107
- On the real-data benchmark below, schema-only checks admitted **160/160**
108
- injected faults; the full validator admitted **0/160**.
116
+ On both real-data benchmarks below, schema-only checks admitted **160/160**
117
+ injected faults; the full validator admitted **0/160**. On ClinVar, 1★ and 2★ variants
118
+ the validator holds back were **about 4–5× more likely** to be reclassified or put in
119
+ conflict three years later.
109
120
 
110
121
  ## Use it
111
122
 
@@ -178,7 +189,7 @@ jobs:
178
189
  runs-on: ubuntu-latest
179
190
  steps:
180
191
  - uses: actions/checkout@v4
181
- - uses: NingyuSUN/bioai-evidence-validator@v0.5.0
192
+ - uses: NingyuSUN/bioai-evidence-validator@v0.7.0
182
193
  with:
183
194
  files: records/**/*.yaml # whitespace-separated globs
184
195
  format: draft # or: record (default)
@@ -250,7 +261,36 @@ Use the [annotation templates](https://github.com/NingyuSUN/bioai-evidence-valid
250
261
  source version and intended use. Document reviewer roles and whether labels are
251
262
  single-reviewed or independently reviewed by multiple people.
252
263
 
253
- ## Real-data case
264
+ ## Real-data cases
265
+
266
+ ### ClinVar germline classifications, three years later
267
+
268
+ The [ClinVar case](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/examples/clinvar_germline/README.md)
269
+ turns every 2023-09 lab submission into evidence, validates 5,026 sampled variants with a
270
+ ClinVar-style profile, and checks what happened to them by 2026-09.
271
+
272
+ ![ClinVar benchmark: share of 2023 pathogenic classifications reclassified or conflicting by 2026, with and without a dissenting submission, and false admissions under controlled faults and the trust boundary](https://raw.githubusercontent.com/NingyuSUN/bioai-evidence-validator/main/docs/assets/clinvar_germline_benchmark.svg)
273
+
274
+ - **Policy reproduction:** the profile matches NCBI's own 2023-09 review status on 98–99% of
275
+ decisions; every disagreement is listed with its cause.
276
+ - **Where the validator is stricter, classifications were less stable.** It sends any P/LP
277
+ variant with a dissenting submission to review, even when ClinVar's aggregate does not.
278
+ Across all 218,920 germline P/LP variants, the share later reclassified or put in conflict:
279
+
280
+ | 2023-09 ClinVar review status | No dissenting submission | With one (validator: review) |
281
+ |---|---:|---:|
282
+ | 1★ single submitter | 2.63% (2.54–2.71) | **13.77%** (11.46–16.47) |
283
+ | 2★ multiple submitters | 2.19% (2.05–2.33) | **8.32%** (6.51–10.59) |
284
+ | 3★ expert panel | 0.07% (0.03–0.16) | **1.30%** (0.63–2.66) |
285
+
286
+ Wilson 95% intervals. Stability is not correctness, and this is observational; see the
287
+ case's interpretation limits. Not for clinical use.
288
+
289
+ ```bash
290
+ uv run python examples/clinvar_germline/run.py --output artifacts/clinvar
291
+ ```
292
+
293
+ ### VBO canine name mapping
254
294
 
255
295
  [VBO canine name mapping](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/examples/vbo_canine/README.md) uses a frozen public ontology:
256
296
  72 real-name cases, 160 controlled errors, and 16 separately reported trust-boundary
@@ -261,7 +301,7 @@ per-required-evidence-type validation. Source-derived labels are not expert anno
261
301
  uv run python examples/vbo_canine/run.py --output artifacts/vbo-canine
262
302
  ```
263
303
 
264
- ### Benchmark results (v0.4.1)
304
+ #### Benchmark results (v0.4.1)
265
305
 
266
306
  ![VBO canine benchmark comparing false admissions across three validation methods](https://raw.githubusercontent.com/NingyuSUN/bioai-evidence-validator/main/docs/assets/vbo_canine_benchmark.svg)
267
307
 
@@ -271,20 +311,31 @@ On 72 real-source name mappings, the full validator admitted all 48 unambiguous
271
311
 
272
312
  ## Scope
273
313
 
274
- The VBO case uses attributed public data; other fixtures are synthetic.
314
+ The VBO and ClinVar cases use attributed public data; other fixtures are synthetic.
275
315
  Admission means **the supplied record meets the selected
276
316
  profile**, not that a biological claim is true. The toolkit does not retrieve papers,
277
- verify reviewer identities, train models, or measure prediction accuracy. The generic core compares supplied hashes; the VBO importer also hashes its local source
278
- projection. External source truth and cohort independence require upstream verification.
317
+ verify reviewer identities, train models, or measure prediction accuracy. The generic core compares supplied hashes; the VBO and ClinVar importers also hash their local source
318
+ projections. External source truth and cohort independence require upstream verification.
319
+ Neither benchmark has independent expert annotation yet; a blinded [expert-review kit](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/evaluation/clinvar_review) for the ClinVar case is ready for reviewers.
279
320
 
280
321
  ## Versions and branches
281
322
 
282
- `main` is the domain-neutral framework (0.5.0). The complete canine implementation
283
- and SQLite adapter from 0.3 live on the
284
- [`canine-breed` branch](https://github.com/NingyuSUN/bioai-evidence-validator/tree/canine-breed);
323
+ `main` is the domain-neutral framework (0.7.0). The complete canine implementation
324
+ and SQLite adapter from 0.3 are preserved at the
325
+ [`canine-0.3` tag](https://github.com/NingyuSUN/bioai-evidence-validator/tree/canine-0.3)
326
+ (also the `canine-breed` branch);
285
327
  see the [0.4 migration guide](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/MIGRATION-0.4.md) and
286
328
  [changelog](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CHANGELOG.md).
287
329
 
330
+ ## Contributing
331
+
332
+ Bug reports, domain profiles and new benchmarks are welcome. See
333
+ [CONTRIBUTING.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CONTRIBUTING.md),
334
+ the [community profiles](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/community/profiles)
335
+ and issues labelled [`good first issue`](https://github.com/NingyuSUN/bioai-evidence-validator/labels/good%20first%20issue).
336
+ Report security problems privately as described in
337
+ [SECURITY.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/SECURITY.md).
338
+
288
339
  ## Citing
289
340
 
290
341
  If you use this toolkit in research, please cite it using the metadata in
@@ -293,6 +344,7 @@ If you use this toolkit in research, please cite it using the metadata in
293
344
 
294
345
  [Create a profile](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/PROFILES.md) ·
295
346
  [Draft format](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/DRAFTS.md) ·
347
+ [Standards alignment](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/STANDARDS.md) ·
296
348
  [Engineering contract](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/ENGINEERING.md) ·
297
349
  [Design case study](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/CASE_STUDY.md) ·
298
350
  [Architecture decision](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/ADR-002-domain-neutral-main.md) ·