bioai-evidence-validator 0.5.0__tar.gz → 0.7.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.gitattributes +1 -0
- bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/bug_report.yml +53 -0
- bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/config.yml +8 -0
- bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/feature_request.yml +28 -0
- bioai_evidence_validator-0.7.0/.github/ISSUE_TEMPLATE/profile_proposal.yml +42 -0
- bioai_evidence_validator-0.7.0/.github/pull_request_template.md +18 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.github/workflows/ci.yml +23 -2
- bioai_evidence_validator-0.7.0/.github/workflows/docs.yml +43 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.gitignore +4 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/CHANGELOG.md +33 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/CITATION.cff +1 -1
- bioai_evidence_validator-0.7.0/CODE_OF_CONDUCT.md +16 -0
- bioai_evidence_validator-0.7.0/CONTRIBUTING.md +76 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/PKG-INFO +65 -13
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/README.md +57 -11
- bioai_evidence_validator-0.7.0/SECURITY.md +35 -0
- bioai_evidence_validator-0.7.0/community/profiles/README.md +37 -0
- bioai_evidence_validator-0.7.0/community/profiles/_template/README.md +18 -0
- bioai_evidence_validator-0.7.0/community/profiles/_template/cases/curated_assay.yaml +14 -0
- bioai_evidence_validator-0.7.0/community/profiles/_template/cases/llm_only.yaml +14 -0
- bioai_evidence_validator-0.7.0/community/profiles/_template/cases/network_without_review.yaml +14 -0
- bioai_evidence_validator-0.7.0/community/profiles/_template/cases/synthetic_study.txt +2 -0
- bioai_evidence_validator-0.7.0/community/profiles/_template/expected.yaml +14 -0
- bioai_evidence_validator-0.7.0/community/profiles/_template/profile.yaml +19 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/ENGINEERING.md +19 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/GOLD_STANDARD.md +10 -5
- bioai_evidence_validator-0.7.0/docs/STANDARDS.md +79 -0
- bioai_evidence_validator-0.7.0/docs/assets/clinvar_germline_benchmark.svg +4101 -0
- bioai_evidence_validator-0.7.0/docs/index.md +58 -0
- bioai_evidence_validator-0.7.0/evaluation/clinvar_review/README.md +125 -0
- bioai_evidence_validator-0.7.0/evaluation/clinvar_review/RUBRIC.md +74 -0
- bioai_evidence_validator-0.7.0/evaluation/clinvar_review/import_sheets.py +126 -0
- bioai_evidence_validator-0.7.0/evaluation/clinvar_review/prepare_packets.py +221 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/README.md +13 -3
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/README.md +178 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/pipeline.py +231 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/plot.py +135 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/prepare_source.py +200 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/profile.yaml +24 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/results/divergences.csv +246 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/results/summary.json +556 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/results/summary.md +48 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/run.py +179 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/README.md +33 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/clinvar-sample.jsonl.gz +0 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/manifest.json +39 -0
- bioai_evidence_validator-0.7.0/examples/clinvar_germline/sources/population_outcomes.json +121 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/pipeline.py +1 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/summary.json +1 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/run.py +2 -1
- bioai_evidence_validator-0.7.0/mkdocs.yml +62 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/pyproject.toml +41 -3
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/__init__.py +1 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/cli.py +77 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/draft.py +9 -6
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/engine.py +2 -2
- bioai_evidence_validator-0.7.0/src/bioevidence_validator/review.py +388 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/schema/bioevidence_core.yaml +18 -4
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_cli.py +1 -2
- bioai_evidence_validator-0.7.0/tests/test_clinvar_case.py +121 -0
- bioai_evidence_validator-0.7.0/tests/test_clinvar_review_kit.py +138 -0
- bioai_evidence_validator-0.7.0/tests/test_community_profiles.py +39 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_engine.py +2 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_evidence_quality.py +3 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_fail_closed.py +2 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_profiles.py +2 -2
- bioai_evidence_validator-0.7.0/tests/test_repository_files.py +21 -0
- bioai_evidence_validator-0.7.0/tests/test_review.py +230 -0
- bioai_evidence_validator-0.7.0/tests/test_standards.py +34 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_vbo_case.py +3 -1
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tools/check_distribution.py +10 -1
- bioai_evidence_validator-0.7.0/tools/mkdocs_hooks.py +73 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/uv.lock +508 -3
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/.github/workflows/release.yml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/LICENSE +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/action.yml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/ADR-001-canine-breed-first.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/ADR-002-domain-neutral-main.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/CASE_STUDY.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/DRAFTS.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/MIGRATION-0.4.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/PROFILES.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/docs/assets/vbo_canine_benchmark.svg +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/adjudications.template.csv +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/annotations.template.csv +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/evaluation/gold_standard/manifest.template.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/custom_profile/assay.yaml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/custom_profile/assay_record.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/dataset_label/curated_sample_label.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/dataset_label/missing_sample_link.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/drafts/llm_claim.yaml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/drafts/reviewed_claim.yaml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/drafts/synthetic_paper.txt +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/general/curated_assertion.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/literature_claim/curated_association.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/literature_claim/llm_only.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/quickstart.ipynb +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/README.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/prepare_source.py +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/profile.yaml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/reference_cases.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/decisions.jsonl +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/review_queue.csv +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/results/summary.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/sources/README.md +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/sources/manifest.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/examples/vbo_canine/sources/vbo-dogs.json +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/config.py +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/profiles/dataset-label.yaml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/profiles/general.yaml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/src/bioevidence_validator/profiles/literature-claim.yaml +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_draft.py +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tests/test_github_action.py +0 -0
- {bioai_evidence_validator-0.5.0 → bioai_evidence_validator-0.7.0}/tools/github_action.py +0 -0
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
name: Bug report
|
|
2
|
+
description: The validator, CLI or GitHub Action does something it should not.
|
|
3
|
+
labels: [bug]
|
|
4
|
+
body:
|
|
5
|
+
- type: markdown
|
|
6
|
+
attributes:
|
|
7
|
+
value: |
|
|
8
|
+
Thanks for reporting. For security problems (for example a record admitted that a
|
|
9
|
+
documented rule should reject), use [private reporting](https://github.com/NingyuSUN/bioai-evidence-validator/security/advisories/new) instead.
|
|
10
|
+
**Do not paste real patient data**; a synthetic record that shows the behavior is enough.
|
|
11
|
+
- type: input
|
|
12
|
+
id: version
|
|
13
|
+
attributes:
|
|
14
|
+
label: Version
|
|
15
|
+
description: Output of `pip show bioai-evidence-validator` or `bioevidence --help` header, or the commit.
|
|
16
|
+
placeholder: "0.7.0"
|
|
17
|
+
validations:
|
|
18
|
+
required: true
|
|
19
|
+
- type: textarea
|
|
20
|
+
id: command
|
|
21
|
+
attributes:
|
|
22
|
+
label: Command or code
|
|
23
|
+
description: The exact command or Python call, including `--profile`.
|
|
24
|
+
render: shell
|
|
25
|
+
validations:
|
|
26
|
+
required: true
|
|
27
|
+
- type: textarea
|
|
28
|
+
id: input
|
|
29
|
+
attributes:
|
|
30
|
+
label: Smallest record, draft or profile that reproduces it
|
|
31
|
+
render: yaml
|
|
32
|
+
validations:
|
|
33
|
+
required: true
|
|
34
|
+
- type: textarea
|
|
35
|
+
id: expected
|
|
36
|
+
attributes:
|
|
37
|
+
label: Expected result
|
|
38
|
+
description: Which status, reason code or exit code did you expect, and which rule says so?
|
|
39
|
+
validations:
|
|
40
|
+
required: true
|
|
41
|
+
- type: textarea
|
|
42
|
+
id: actual
|
|
43
|
+
attributes:
|
|
44
|
+
label: Actual result
|
|
45
|
+
description: The report's `overall_status`, `findings` and exit code, or the error message.
|
|
46
|
+
render: json
|
|
47
|
+
validations:
|
|
48
|
+
required: true
|
|
49
|
+
- type: input
|
|
50
|
+
id: environment
|
|
51
|
+
attributes:
|
|
52
|
+
label: Environment
|
|
53
|
+
placeholder: "Python 3.12, Ubuntu 24.04"
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
blank_issues_enabled: true
|
|
2
|
+
contact_links:
|
|
3
|
+
- name: Security vulnerability
|
|
4
|
+
url: https://github.com/NingyuSUN/bioai-evidence-validator/security/advisories/new
|
|
5
|
+
about: Report privately; please do not open a public issue.
|
|
6
|
+
- name: Documentation
|
|
7
|
+
url: https://github.com/NingyuSUN/bioai-evidence-validator#readme
|
|
8
|
+
about: Usage, draft format, profiles and benchmarks.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
name: Feature request
|
|
2
|
+
description: Propose a change to the engine, CLI, formats or integrations.
|
|
3
|
+
labels: [enhancement]
|
|
4
|
+
body:
|
|
5
|
+
- type: textarea
|
|
6
|
+
id: problem
|
|
7
|
+
attributes:
|
|
8
|
+
label: Problem
|
|
9
|
+
description: What are you trying to do, and what gets in the way today?
|
|
10
|
+
validations:
|
|
11
|
+
required: true
|
|
12
|
+
- type: textarea
|
|
13
|
+
id: proposal
|
|
14
|
+
attributes:
|
|
15
|
+
label: Proposal
|
|
16
|
+
description: The behavior you would like. For new rules, describe a record that should pass and one that should not.
|
|
17
|
+
validations:
|
|
18
|
+
required: true
|
|
19
|
+
- type: textarea
|
|
20
|
+
id: failure
|
|
21
|
+
attributes:
|
|
22
|
+
label: Failure mode
|
|
23
|
+
description: If this goes wrong, does it admit too much or block too much? How would a user notice?
|
|
24
|
+
- type: textarea
|
|
25
|
+
id: alternatives
|
|
26
|
+
attributes:
|
|
27
|
+
label: Alternatives considered
|
|
28
|
+
description: For example, doing it in a profile or an importer instead of the engine.
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
name: Domain profile proposal
|
|
2
|
+
description: Propose an admission profile for a biological domain.
|
|
3
|
+
labels: [profile]
|
|
4
|
+
body:
|
|
5
|
+
- type: markdown
|
|
6
|
+
attributes:
|
|
7
|
+
value: |
|
|
8
|
+
Profiles decide which evidence each intended use requires; they do not change the engine.
|
|
9
|
+
See [docs/PROFILES.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/PROFILES.md)
|
|
10
|
+
and the [community profile template](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/community/profiles/_template).
|
|
11
|
+
- type: input
|
|
12
|
+
id: domain
|
|
13
|
+
attributes:
|
|
14
|
+
label: Domain and assertion
|
|
15
|
+
placeholder: "Protein–protein interactions: protein A interacts_with protein B"
|
|
16
|
+
validations:
|
|
17
|
+
required: true
|
|
18
|
+
- type: textarea
|
|
19
|
+
id: uses
|
|
20
|
+
attributes:
|
|
21
|
+
label: Intended uses and their evidence requirements
|
|
22
|
+
description: For each use (e.g. research summary, knowledge-base admission, training data), which evidence types are required, and is human acceptance required?
|
|
23
|
+
validations:
|
|
24
|
+
required: true
|
|
25
|
+
- type: textarea
|
|
26
|
+
id: basis
|
|
27
|
+
attributes:
|
|
28
|
+
label: Basis for the policy
|
|
29
|
+
description: Community guideline, database policy or published standard the requirements follow (with links), or state that it is your own proposal.
|
|
30
|
+
validations:
|
|
31
|
+
required: true
|
|
32
|
+
- type: textarea
|
|
33
|
+
id: cases
|
|
34
|
+
attributes:
|
|
35
|
+
label: Example cases
|
|
36
|
+
description: At least one claim that should be admitted and one that should not, with the reason.
|
|
37
|
+
- type: checkboxes
|
|
38
|
+
id: pr
|
|
39
|
+
attributes:
|
|
40
|
+
label: Contribution
|
|
41
|
+
options:
|
|
42
|
+
- label: I am willing to open a pull request with the profile and example cases.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
## What and why
|
|
2
|
+
|
|
3
|
+
<!-- One purpose per pull request. Link the issue: "Closes #123". -->
|
|
4
|
+
|
|
5
|
+
## Admission boundary
|
|
6
|
+
|
|
7
|
+
<!-- Does this change what gets admitted, sent to review or rejected? If yes, say which
|
|
8
|
+
records move and why, and point to the tests that pin the new boundary. -->
|
|
9
|
+
|
|
10
|
+
- [ ] No change to admission decisions
|
|
11
|
+
- [ ] Changes admission decisions (explained above; tests on both sides of the boundary)
|
|
12
|
+
|
|
13
|
+
## Checklist
|
|
14
|
+
|
|
15
|
+
- [ ] `uv run --frozen ruff check .`, `uv run --frozen mypy` and `uv run --frozen pytest --cov` pass
|
|
16
|
+
- [ ] Committed benchmark results regenerated if report contents changed
|
|
17
|
+
- [ ] `CHANGELOG.md` updated if users will notice
|
|
18
|
+
- [ ] New third-party data has its license, attribution and exact version recorded
|
|
@@ -6,6 +6,23 @@ on:
|
|
|
6
6
|
permissions:
|
|
7
7
|
contents: read
|
|
8
8
|
jobs:
|
|
9
|
+
lint:
|
|
10
|
+
timeout-minutes: 10
|
|
11
|
+
runs-on: ubuntu-latest
|
|
12
|
+
steps:
|
|
13
|
+
- uses: actions/checkout@v4
|
|
14
|
+
- uses: actions/setup-python@v5
|
|
15
|
+
with:
|
|
16
|
+
python-version: '3.12'
|
|
17
|
+
- name: Install locked tooling
|
|
18
|
+
run: python -m pip install uv==0.12.17
|
|
19
|
+
- name: Install locked project
|
|
20
|
+
run: uv sync --frozen --extra dev
|
|
21
|
+
- name: Lint
|
|
22
|
+
run: uv run --frozen ruff check --output-format github .
|
|
23
|
+
- name: Type-check
|
|
24
|
+
run: uv run --frozen mypy
|
|
25
|
+
|
|
9
26
|
verify:
|
|
10
27
|
timeout-minutes: 15
|
|
11
28
|
strategy:
|
|
@@ -29,8 +46,12 @@ jobs:
|
|
|
29
46
|
run: python -m pip install uv==0.12.17
|
|
30
47
|
- name: Install locked project
|
|
31
48
|
run: uv sync --frozen --extra dev --python ${{ matrix.python }}
|
|
32
|
-
- name: Run tests
|
|
33
|
-
run: uv run --frozen pytest
|
|
49
|
+
- name: Run tests with coverage (minimum in pyproject.toml)
|
|
50
|
+
run: uv run --frozen pytest --cov --cov-report=term
|
|
51
|
+
- name: Coverage summary
|
|
52
|
+
if: always()
|
|
53
|
+
shell: bash
|
|
54
|
+
run: uv run --frozen coverage report --format=markdown >> "$GITHUB_STEP_SUMMARY" || true
|
|
34
55
|
- name: Build distributable
|
|
35
56
|
run: uv build
|
|
36
57
|
- name: Verify installed wheel and CLI outside editable source
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
name: Documentation
|
|
2
|
+
on:
|
|
3
|
+
push:
|
|
4
|
+
branches: [main]
|
|
5
|
+
pull_request:
|
|
6
|
+
workflow_dispatch:
|
|
7
|
+
permissions:
|
|
8
|
+
contents: read
|
|
9
|
+
concurrency:
|
|
10
|
+
group: docs-${{ github.ref }}
|
|
11
|
+
cancel-in-progress: true
|
|
12
|
+
jobs:
|
|
13
|
+
build:
|
|
14
|
+
timeout-minutes: 10
|
|
15
|
+
runs-on: ubuntu-latest
|
|
16
|
+
steps:
|
|
17
|
+
- uses: actions/checkout@v4
|
|
18
|
+
- uses: actions/setup-python@v5
|
|
19
|
+
with:
|
|
20
|
+
python-version: '3.12'
|
|
21
|
+
# Pinned: MkDocs 2.0 drops the plugin and theme system this site uses.
|
|
22
|
+
- name: Install MkDocs
|
|
23
|
+
run: python -m pip install mkdocs==1.6.1 mkdocs-material==9.7.7
|
|
24
|
+
- name: Build with strict link checking
|
|
25
|
+
run: mkdocs build --strict --site-dir _site
|
|
26
|
+
- if: github.event_name != 'pull_request'
|
|
27
|
+
uses: actions/upload-pages-artifact@v3
|
|
28
|
+
with:
|
|
29
|
+
path: _site
|
|
30
|
+
|
|
31
|
+
deploy:
|
|
32
|
+
if: github.event_name != 'pull_request'
|
|
33
|
+
needs: build
|
|
34
|
+
runs-on: ubuntu-latest
|
|
35
|
+
permissions:
|
|
36
|
+
pages: write
|
|
37
|
+
id-token: write
|
|
38
|
+
environment:
|
|
39
|
+
name: github-pages
|
|
40
|
+
url: ${{ steps.deployment.outputs.page_url }}
|
|
41
|
+
steps:
|
|
42
|
+
- id: deployment
|
|
43
|
+
uses: actions/deploy-pages@v4
|
|
@@ -1,5 +1,38 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.7.0 — Expert review, quality checks, community and documentation site
|
|
4
|
+
|
|
5
|
+
- Add `bioevidence review` (`check`, `agreement`, `adjudication-sheet`, `score`, `freeze`)
|
|
6
|
+
implementing the gold-standard protocol: strict annotation checks, Krippendorff's α with
|
|
7
|
+
bootstrap interval and Cohen's κ (both cross-checked against reference implementations),
|
|
8
|
+
adjudication of disagreements, scoring on the test split with hash-bound joins, and frozen
|
|
9
|
+
manifests. It never produces labels.
|
|
10
|
+
- Add a blinded ClinVar expert-review kit (`evaluation/clinvar_review/`): all 95 variants where
|
|
11
|
+
the validator and NCBI disagree plus 95 stratum-matched controls, an Excel workbook with
|
|
12
|
+
dropdowns, a reviewer rubric, a private key, and an importer into the protocol format.
|
|
13
|
+
- CI runs ruff and mypy, and enforces a 95% test-coverage minimum (currently 98%).
|
|
14
|
+
- Add CONTRIBUTING, CODE_OF_CONDUCT (Contributor Covenant 2.1), SECURITY, issue forms and a
|
|
15
|
+
pull request template.
|
|
16
|
+
- Add `community/profiles/`: contributed domain profiles whose example cases are built,
|
|
17
|
+
validated and checked against expected outcomes in CI; `_template/` to copy.
|
|
18
|
+
- Add a documentation site (MkDocs, strict link checking) published to GitHub Pages.
|
|
19
|
+
- The canine 0.3 implementation is also preserved at the `canine-0.3` tag.
|
|
20
|
+
- No change to validation decisions or report format; both benchmarks reproduce 0.6.0 exactly
|
|
21
|
+
apart from `validator_version`.
|
|
22
|
+
|
|
23
|
+
## 0.6.0 — ClinVar case and standards alignment
|
|
24
|
+
|
|
25
|
+
- Add a second real-data case: ClinVar germline classifications. 5,026 sampled variants
|
|
26
|
+
from the 2023-09 release are built from per-submission evidence and validated with a
|
|
27
|
+
ClinVar-style profile; decisions are compared with NCBI's own 2023-09 review status and
|
|
28
|
+
with each classification's 2026-09 outcome, plus controlled faults and a trust-boundary
|
|
29
|
+
cohort. Frozen, hash-pinned sample; rebuild script verifies the three upstream files.
|
|
30
|
+
- Add `docs/STANDARDS.md`, mapping the record model to ECO, Biolink 4.4.4, GA4GH VA-Spec
|
|
31
|
+
1.0.1 and PROV-O, with mapping strength and caveats.
|
|
32
|
+
- Annotate `ExtractionMethod` values with ECO meanings in the LinkML schema. The compiled
|
|
33
|
+
JSON Schema, and therefore every report's `schema_sha256`, is unchanged.
|
|
34
|
+
- The VBO benchmark summary changes only `validator_version`.
|
|
35
|
+
|
|
3
36
|
## 0.5.0 — Drafts, LLM draft schema and GitHub Action
|
|
4
37
|
|
|
5
38
|
- Add compact YAML/JSON drafts: `bioevidence build` and `build_record()` derive identifiers,
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
# Code of conduct
|
|
2
|
+
|
|
3
|
+
This project adopts the
|
|
4
|
+
[Contributor Covenant, version 2.1](https://www.contributor-covenant.org/version/2/1/code_of_conduct/)
|
|
5
|
+
as its code of conduct. It applies to all project spaces (issues, pull requests,
|
|
6
|
+
discussions and reviews) and to anyone representing the project elsewhere.
|
|
7
|
+
|
|
8
|
+
In short: be respectful and constructive, assume good faith, critique work rather than
|
|
9
|
+
people, and remember that contributors bring expertise from many fields, from biocuration
|
|
10
|
+
to software engineering.
|
|
11
|
+
|
|
12
|
+
## Reporting
|
|
13
|
+
|
|
14
|
+
Report unacceptable behavior to the maintainer, Ningyu Sun, at <woshiwosunny@gmail.com>.
|
|
15
|
+
Reports are handled confidentially, and the enforcement guidelines of the Contributor
|
|
16
|
+
Covenant 2.1 apply.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
Thanks for helping make AI-assisted biocuration safer. Contributions of every size are
|
|
4
|
+
welcome: a typo fix, a bug report with a failing record, a new domain profile, or a new
|
|
5
|
+
real-data benchmark.
|
|
6
|
+
|
|
7
|
+
By participating you agree to follow the [code of conduct](CODE_OF_CONDUCT.md).
|
|
8
|
+
|
|
9
|
+
## Ways to contribute
|
|
10
|
+
|
|
11
|
+
| You have… | Start here |
|
|
12
|
+
|---|---|
|
|
13
|
+
| A record the validator judges wrongly | [Bug report](https://github.com/NingyuSUN/bioai-evidence-validator/issues/new?template=bug_report.yml); attach the smallest record or draft that reproduces it |
|
|
14
|
+
| An admission policy for your domain | [Profile proposal](https://github.com/NingyuSUN/bioai-evidence-validator/issues/new?template=profile_proposal.yml), then a pull request to [`community/profiles/`](community/profiles/README.md) |
|
|
15
|
+
| An idea for the engine, CLI or formats | [Feature request](https://github.com/NingyuSUN/bioai-evidence-validator/issues/new?template=feature_request.yml) first, so we can agree on the contract before code |
|
|
16
|
+
| A security problem | Do **not** open an issue; see [SECURITY.md](SECURITY.md) |
|
|
17
|
+
|
|
18
|
+
Issues labelled [`good first issue`](https://github.com/NingyuSUN/bioai-evidence-validator/labels/good%20first%20issue)
|
|
19
|
+
are scoped to be finished in an afternoon.
|
|
20
|
+
|
|
21
|
+
## Development setup
|
|
22
|
+
|
|
23
|
+
Python 3.11+ and [uv](https://docs.astral.sh/uv/):
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
git clone https://github.com/NingyuSUN/bioai-evidence-validator.git
|
|
27
|
+
cd bioai-evidence-validator
|
|
28
|
+
uv sync --frozen --extra dev
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Before opening a pull request, run what CI runs:
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
uv run --frozen ruff check .
|
|
35
|
+
uv run --frozen mypy
|
|
36
|
+
uv run --frozen pytest --cov
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
CI also runs the tests on Linux (Python 3.11–3.13) and Windows, builds the wheel and
|
|
40
|
+
smoke-tests it outside the source tree, and runs the GitHub Action. Coverage must stay at
|
|
41
|
+
or above the minimum in `pyproject.toml`.
|
|
42
|
+
|
|
43
|
+
## Ground rules for changes
|
|
44
|
+
|
|
45
|
+
This project's value is that it **fails closed** and **says exactly what it checked**.
|
|
46
|
+
Changes are reviewed against that:
|
|
47
|
+
|
|
48
|
+
- **No silent weakening.** A change that admits something previously rejected needs an
|
|
49
|
+
explicit reason in the pull request and a test that shows the new boundary.
|
|
50
|
+
- **Every rule has a test on both sides**: a record that passes and a minimal one that fails.
|
|
51
|
+
- **Reports stay reproducible.** If a change alters report contents, the committed benchmark
|
|
52
|
+
results must be regenerated in the same pull request, and the diff explained.
|
|
53
|
+
- **Domain logic stays out of the engine.** New domains are profiles and importers, not
|
|
54
|
+
special cases in `engine.py`.
|
|
55
|
+
- **Honest limits.** Benchmarks state what their labels are (source-derived, authored, or
|
|
56
|
+
independently reviewed) and what they do not measure.
|
|
57
|
+
- Match the surrounding style; ruff enforces correctness rules, not formatting.
|
|
58
|
+
|
|
59
|
+
## Contributing a domain profile
|
|
60
|
+
|
|
61
|
+
Community profiles live in [`community/profiles/`](community/profiles/README.md). Each one is
|
|
62
|
+
a folder with a profile, a short README and example cases whose expected outcomes are
|
|
63
|
+
checked by the test suite. Copy `community/profiles/_template/` to get started.
|
|
64
|
+
|
|
65
|
+
## Pull requests
|
|
66
|
+
|
|
67
|
+
- Keep each pull request to one purpose; link the issue it resolves.
|
|
68
|
+
- Update `CHANGELOG.md` under a new heading if users will notice the change.
|
|
69
|
+
- New third-party data needs its license, attribution and exact source version recorded,
|
|
70
|
+
as in `examples/*/sources/README.md`.
|
|
71
|
+
|
|
72
|
+
## Releases
|
|
73
|
+
|
|
74
|
+
Maintainers release by bumping the version in `pyproject.toml` and
|
|
75
|
+
`src/bioevidence_validator/__init__.py`, adding a `CHANGELOG.md` section, and pushing a
|
|
76
|
+
`vX.Y.Z` tag. The release workflow tests, publishes to PyPI and creates the GitHub release.
|
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: bioai-evidence-validator
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.7.0
|
|
4
4
|
Summary: Standards-aligned evidence policy validation for AI-assisted biological curation
|
|
5
5
|
Project-URL: Homepage, https://github.com/NingyuSUN/bioai-evidence-validator
|
|
6
|
-
Project-URL: Documentation, https://github.
|
|
6
|
+
Project-URL: Documentation, https://ningyusun.github.io/bioai-evidence-validator/
|
|
7
7
|
Project-URL: Changelog, https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CHANGELOG.md
|
|
8
8
|
Project-URL: Issues, https://github.com/NingyuSUN/bioai-evidence-validator/issues
|
|
9
9
|
Author: Ningyu Sun
|
|
@@ -25,7 +25,13 @@ Requires-Dist: jsonschema<5,>=4.23
|
|
|
25
25
|
Requires-Dist: linkml<2,>=1.8
|
|
26
26
|
Requires-Dist: pyyaml<7,>=6.0
|
|
27
27
|
Provides-Extra: dev
|
|
28
|
+
Requires-Dist: mypy<3,>=2.3; extra == 'dev'
|
|
29
|
+
Requires-Dist: openpyxl<4,>=3.1; extra == 'dev'
|
|
30
|
+
Requires-Dist: pytest-cov<8,>=7; extra == 'dev'
|
|
28
31
|
Requires-Dist: pytest<9,>=8; extra == 'dev'
|
|
32
|
+
Requires-Dist: ruff<0.17,>=0.16; extra == 'dev'
|
|
33
|
+
Requires-Dist: types-jsonschema; extra == 'dev'
|
|
34
|
+
Requires-Dist: types-pyyaml; extra == 'dev'
|
|
29
35
|
Description-Content-Type: text/markdown
|
|
30
36
|
|
|
31
37
|
# BioAI Evidence Validator
|
|
@@ -35,6 +41,7 @@ Description-Content-Type: text/markdown
|
|
|
35
41
|
[](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/pyproject.toml)
|
|
36
42
|
[](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/LICENSE)
|
|
37
43
|
[](https://colab.research.google.com/github/NingyuSUN/bioai-evidence-validator/blob/main/examples/quickstart.ipynb)
|
|
44
|
+
[](https://ningyusun.github.io/bioai-evidence-validator/)
|
|
38
45
|
|
|
39
46
|
**Stop AI-extracted biological claims from entering your knowledge base or
|
|
40
47
|
training set before their evidence is good enough for that use.**
|
|
@@ -52,6 +59,7 @@ pip install bioai-evidence-validator
|
|
|
52
59
|
|
|
53
60
|
Or try it in the browser, nothing to install:
|
|
54
61
|
[quickstart notebook on Colab](https://colab.research.google.com/github/NingyuSUN/bioai-evidence-validator/blob/main/examples/quickstart.ipynb).
|
|
62
|
+
Full documentation: **[ningyusun.github.io/bioai-evidence-validator](https://ningyusun.github.io/bioai-evidence-validator/)**.
|
|
55
63
|
|
|
56
64
|
## 30-second example
|
|
57
65
|
|
|
@@ -103,9 +111,12 @@ record is trustworthy enough for a particular purpose.
|
|
|
103
111
|
| Provenance consistency (source hashes, resolved references, scope) | — | ✅ |
|
|
104
112
|
| Human adjudications bound to a specific statement and use | — | ✅ |
|
|
105
113
|
| Machine-readable audit report with hashes of input, schema and profile | — | ✅ |
|
|
114
|
+
| Mapped to ECO, Biolink and GA4GH VA-Spec ([standards alignment](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/STANDARDS.md)) | — | ✅ |
|
|
106
115
|
|
|
107
|
-
On
|
|
108
|
-
injected faults; the full validator admitted **0/160**.
|
|
116
|
+
On both real-data benchmarks below, schema-only checks admitted **160/160**
|
|
117
|
+
injected faults; the full validator admitted **0/160**. On ClinVar, 1★ and 2★ variants
|
|
118
|
+
the validator holds back were **about 4–5× more likely** to be reclassified or put in
|
|
119
|
+
conflict three years later.
|
|
109
120
|
|
|
110
121
|
## Use it
|
|
111
122
|
|
|
@@ -178,7 +189,7 @@ jobs:
|
|
|
178
189
|
runs-on: ubuntu-latest
|
|
179
190
|
steps:
|
|
180
191
|
- uses: actions/checkout@v4
|
|
181
|
-
- uses: NingyuSUN/bioai-evidence-validator@v0.
|
|
192
|
+
- uses: NingyuSUN/bioai-evidence-validator@v0.7.0
|
|
182
193
|
with:
|
|
183
194
|
files: records/**/*.yaml # whitespace-separated globs
|
|
184
195
|
format: draft # or: record (default)
|
|
@@ -250,7 +261,36 @@ Use the [annotation templates](https://github.com/NingyuSUN/bioai-evidence-valid
|
|
|
250
261
|
source version and intended use. Document reviewer roles and whether labels are
|
|
251
262
|
single-reviewed or independently reviewed by multiple people.
|
|
252
263
|
|
|
253
|
-
## Real-data
|
|
264
|
+
## Real-data cases
|
|
265
|
+
|
|
266
|
+
### ClinVar germline classifications, three years later
|
|
267
|
+
|
|
268
|
+
The [ClinVar case](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/examples/clinvar_germline/README.md)
|
|
269
|
+
turns every 2023-09 lab submission into evidence, validates 5,026 sampled variants with a
|
|
270
|
+
ClinVar-style profile, and checks what happened to them by 2026-09.
|
|
271
|
+
|
|
272
|
+

|
|
273
|
+
|
|
274
|
+
- **Policy reproduction:** the profile matches NCBI's own 2023-09 review status on 98–99% of
|
|
275
|
+
decisions; every disagreement is listed with its cause.
|
|
276
|
+
- **Where the validator is stricter, classifications were less stable.** It sends any P/LP
|
|
277
|
+
variant with a dissenting submission to review, even when ClinVar's aggregate does not.
|
|
278
|
+
Across all 218,920 germline P/LP variants, the share later reclassified or put in conflict:
|
|
279
|
+
|
|
280
|
+
| 2023-09 ClinVar review status | No dissenting submission | With one (validator: review) |
|
|
281
|
+
|---|---:|---:|
|
|
282
|
+
| 1★ single submitter | 2.63% (2.54–2.71) | **13.77%** (11.46–16.47) |
|
|
283
|
+
| 2★ multiple submitters | 2.19% (2.05–2.33) | **8.32%** (6.51–10.59) |
|
|
284
|
+
| 3★ expert panel | 0.07% (0.03–0.16) | **1.30%** (0.63–2.66) |
|
|
285
|
+
|
|
286
|
+
Wilson 95% intervals. Stability is not correctness, and this is observational; see the
|
|
287
|
+
case's interpretation limits. Not for clinical use.
|
|
288
|
+
|
|
289
|
+
```bash
|
|
290
|
+
uv run python examples/clinvar_germline/run.py --output artifacts/clinvar
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
### VBO canine name mapping
|
|
254
294
|
|
|
255
295
|
[VBO canine name mapping](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/examples/vbo_canine/README.md) uses a frozen public ontology:
|
|
256
296
|
72 real-name cases, 160 controlled errors, and 16 separately reported trust-boundary
|
|
@@ -261,7 +301,7 @@ per-required-evidence-type validation. Source-derived labels are not expert anno
|
|
|
261
301
|
uv run python examples/vbo_canine/run.py --output artifacts/vbo-canine
|
|
262
302
|
```
|
|
263
303
|
|
|
264
|
-
|
|
304
|
+
#### Benchmark results (v0.4.1)
|
|
265
305
|
|
|
266
306
|

|
|
267
307
|
|
|
@@ -271,20 +311,31 @@ On 72 real-source name mappings, the full validator admitted all 48 unambiguous
|
|
|
271
311
|
|
|
272
312
|
## Scope
|
|
273
313
|
|
|
274
|
-
The VBO
|
|
314
|
+
The VBO and ClinVar cases use attributed public data; other fixtures are synthetic.
|
|
275
315
|
Admission means **the supplied record meets the selected
|
|
276
316
|
profile**, not that a biological claim is true. The toolkit does not retrieve papers,
|
|
277
|
-
verify reviewer identities, train models, or measure prediction accuracy. The generic core compares supplied hashes; the VBO
|
|
278
|
-
|
|
317
|
+
verify reviewer identities, train models, or measure prediction accuracy. The generic core compares supplied hashes; the VBO and ClinVar importers also hash their local source
|
|
318
|
+
projections. External source truth and cohort independence require upstream verification.
|
|
319
|
+
Neither benchmark has independent expert annotation yet; a blinded [expert-review kit](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/evaluation/clinvar_review) for the ClinVar case is ready for reviewers.
|
|
279
320
|
|
|
280
321
|
## Versions and branches
|
|
281
322
|
|
|
282
|
-
`main` is the domain-neutral framework (0.
|
|
283
|
-
and SQLite adapter from 0.3
|
|
284
|
-
[`canine-
|
|
323
|
+
`main` is the domain-neutral framework (0.7.0). The complete canine implementation
|
|
324
|
+
and SQLite adapter from 0.3 are preserved at the
|
|
325
|
+
[`canine-0.3` tag](https://github.com/NingyuSUN/bioai-evidence-validator/tree/canine-0.3)
|
|
326
|
+
(also the `canine-breed` branch);
|
|
285
327
|
see the [0.4 migration guide](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/MIGRATION-0.4.md) and
|
|
286
328
|
[changelog](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CHANGELOG.md).
|
|
287
329
|
|
|
330
|
+
## Contributing
|
|
331
|
+
|
|
332
|
+
Bug reports, domain profiles and new benchmarks are welcome. See
|
|
333
|
+
[CONTRIBUTING.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/CONTRIBUTING.md),
|
|
334
|
+
the [community profiles](https://github.com/NingyuSUN/bioai-evidence-validator/tree/main/community/profiles)
|
|
335
|
+
and issues labelled [`good first issue`](https://github.com/NingyuSUN/bioai-evidence-validator/labels/good%20first%20issue).
|
|
336
|
+
Report security problems privately as described in
|
|
337
|
+
[SECURITY.md](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/SECURITY.md).
|
|
338
|
+
|
|
288
339
|
## Citing
|
|
289
340
|
|
|
290
341
|
If you use this toolkit in research, please cite it using the metadata in
|
|
@@ -293,6 +344,7 @@ If you use this toolkit in research, please cite it using the metadata in
|
|
|
293
344
|
|
|
294
345
|
[Create a profile](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/PROFILES.md) ·
|
|
295
346
|
[Draft format](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/DRAFTS.md) ·
|
|
347
|
+
[Standards alignment](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/STANDARDS.md) ·
|
|
296
348
|
[Engineering contract](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/ENGINEERING.md) ·
|
|
297
349
|
[Design case study](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/CASE_STUDY.md) ·
|
|
298
350
|
[Architecture decision](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/ADR-002-domain-neutral-main.md) ·
|