OperonDBS 0.6.2__tar.gz → 0.7.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {operondbs-0.6.2 → operondbs-0.7.0}/.github/workflows/publish.yml +3 -3
- {operondbs-0.6.2 → operondbs-0.7.0}/AGENTS.md +22 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/PKG-INFO +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/SOURCES.txt +3 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/PKG-INFO +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/conf.py +43 -5
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/extensibility.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/external-analysis.md +2 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/metadata-and-data-model.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/overview.md +8 -6
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/qc-and-rules.md +1 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/application-release.md +2 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/development-testing.md +15 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/pypi-release.md +5 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/first-project.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/installation.md +3 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/backup-migration.md +6 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/curation-lifecycle.md +11 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/external-analysis.md +4 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/metadata-import.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/ncbi-datasets.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/qc-profiles.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/remote-storage.md +4 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/troubleshooting.md +4 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/index.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/database-compatibility.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/overview.md +2 -0
- operondbs-0.7.0/docs/en/reference/behaviors-and-limitations.md +192 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-decisions-reports.md +7 -3
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-files-qc.md +4 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-tui.md +4 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/index.md +2 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/recipe-fields.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/recipe-overview.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/recipe-parsers-examples.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/extensibility.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/external-analysis.md +2 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/metadata-and-data-model.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/overview.md +8 -6
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/qc-and-rules.md +1 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/application-release.md +2 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/development-testing.md +17 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/pypi-release.md +4 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/first-project.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/installation.md +3 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/backup-migration.md +10 -4
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/curation-lifecycle.md +15 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/external-analysis.md +9 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/metadata-import.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/ncbi-datasets.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/qc-profiles.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/remote-storage.md +7 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/troubleshooting.md +6 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/index.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/database-compatibility.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/overview.md +2 -0
- operondbs-0.7.0/docs/zh/reference/behaviors-and-limitations.md +190 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-decisions-reports.md +7 -4
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-files-qc.md +4 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-tui.md +2 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/index.md +2 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/recipe-fields.md +2 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/recipe-overview.md +2 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/recipe-parsers-examples.md +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/backup.py +50 -8
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/cli.py +76 -4
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/database.py +77 -7
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/export.py +68 -8
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/files.py +85 -45
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/import_wizard.py +1 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/lineage.py +73 -14
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/qc_module/__init__.py +97 -12
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/release.py +149 -19
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/rules.py +67 -8
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/table_import.py +11 -7
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/actions.py +33 -31
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/utils.py +0 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/pyproject.toml +2 -2
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_cli_edges.py +29 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_database_edges_more.py +23 -1
- operondbs-0.7.0/tests/unit/test_docs_versions.py +69 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_export.py +26 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_files_edges.py +52 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_import_wizard_edges.py +9 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_lineage.py +80 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_qc_and_rules.py +41 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_rules_schema_edges.py +31 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_support_edges.py +49 -1
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_table_import_edges.py +13 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tui_config.py +38 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_views_release_reports_edges.py +57 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/.github/workflows/deploy.yml +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/.gitignore +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/.readthedocs.yaml +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/LICENSE +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/MANIFEST.in +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/dependency_links.txt +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/entry_points.txt +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/requires.txt +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/top_level.txt +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/README.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/README_ZH.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/benchmarks/qc_representative_entities.tsv +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/_static/language-switcher.js +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/_static/operon.css +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/_templates/layout.html +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/files-and-storage.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/release-lifecycle.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/taxonomy-coverage.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/documentation-deployment.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/repository-guide.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/daily-workflow.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/quickstart.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/file-archiving.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/remote-execution.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/taxonomy-coverage.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/ncbi-recovery-migration.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/qc-performance.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-analysis.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-project-metadata.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-remote.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-taxonomy-lifecycle-admin.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-workflow.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/data-model.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/requirements.txt +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/files-and-storage.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/release-lifecycle.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/taxonomy-coverage.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/documentation-deployment.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/repository-guide.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/daily-workflow.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/quickstart.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/file-archiving.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/remote-execution.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/taxonomy-coverage.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/index.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/ncbi-recovery-migration.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/qc-performance.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-analysis.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-project-metadata.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-remote.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-taxonomy-lifecycle-admin.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-workflow.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/data-model.md +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/__main__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/adapters/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/adapters/ncbi_datasets.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/config.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/coverage.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/demo.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/entity_view.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/environment.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/errors.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/execution.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/lifecycle.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/metadata_files.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/ncbi_reconcile.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/profiles.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/qc_module/_parsers.pyx +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/qc_module/parsers.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/remotes.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/reports.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/schema.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/shutdown.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/taxonomy.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tools.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/app.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/app.tcss +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/data.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/common.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/config.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/decisions.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/entities.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/files.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/files_ops.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/home.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/runs.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/operon/workflow.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/setup.cfg +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/setup.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/compatibility/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/compatibility/test_python_support.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/helpers.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_analysis_resume.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_analysis_shutdown.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_analysis_tools.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_application_build.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_execution_backends.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_lineage_cascade.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_ncbi_datasets_adapter.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_pipeline_and_release.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_taxonomy_coverage.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/test_correctness.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/test_cython_parser_parity.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/test_parser_semantics.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/__init__.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_config_workflow_edges.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_coverage_edges.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_environment.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_execution.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_execution_edges.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_import_backup_show.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_lifecycle.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_ncbi_edge_cases.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_ncbi_reconcile_edges.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_parser_edge_paths.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_qc_edges_more.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_recipe_history.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_remotes.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_remotes_edges.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_schema_2_9.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_schema_and_metadata.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_shutdown.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_taxonomy_edges.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tools_edges.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tui.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tui_writes.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_workflow_cli.py +0 -0
- {operondbs-0.6.2 → operondbs-0.7.0}/tools/build.py +0 -0
|
@@ -28,7 +28,7 @@ jobs:
|
|
|
28
28
|
- run: python -m pip install --upgrade build twine
|
|
29
29
|
- run: python -m build --sdist
|
|
30
30
|
- run: python -m twine check dist/*
|
|
31
|
-
- uses: actions/upload-artifact@
|
|
31
|
+
- uses: actions/upload-artifact@v7
|
|
32
32
|
with:
|
|
33
33
|
name: python-package-sdist
|
|
34
34
|
path: dist/*.tar.gz
|
|
@@ -65,7 +65,7 @@ jobs:
|
|
|
65
65
|
CIBW_TEST_COMMAND: >-
|
|
66
66
|
python -c "import operon; import operon.qc_module._parsers;
|
|
67
67
|
print(operon.__version__)" && operon --help
|
|
68
|
-
- uses: actions/upload-artifact@
|
|
68
|
+
- uses: actions/upload-artifact@v7
|
|
69
69
|
with:
|
|
70
70
|
name: python-package-wheel-${{ matrix.artifact }}
|
|
71
71
|
path: wheelhouse/*.whl
|
|
@@ -80,7 +80,7 @@ jobs:
|
|
|
80
80
|
permissions:
|
|
81
81
|
id-token: write
|
|
82
82
|
steps:
|
|
83
|
-
- uses: actions/download-artifact@
|
|
83
|
+
- uses: actions/download-artifact@v8
|
|
84
84
|
with:
|
|
85
85
|
pattern: python-package-*
|
|
86
86
|
path: dist/
|
|
@@ -187,4 +187,25 @@ change:
|
|
|
187
187
|
`docs/*/index.md`
|
|
188
188
|
|
|
189
189
|
Version markers in docs (`operon` 0.6.2, database schema 2.9, metadata
|
|
190
|
-
schema 1.4) must match `pyproject.toml` and the code.
|
|
190
|
+
schema 1.4) must match `pyproject.toml` and the code. Do not write the
|
|
191
|
+
current values literally in Markdown sources: use the `myst_substitutions`
|
|
192
|
+
references `{{ operon_version }}`, `{{ db_schema }}`, and
|
|
193
|
+
`{{ metadata_schema }}`, which `docs/conf.py` resolves from the single
|
|
194
|
+
sources above at build time. Substitutions expand in paragraph text only,
|
|
195
|
+
never inside code spans or fenced code blocks — examples there use
|
|
196
|
+
`<version>` placeholders instead. Intentional historical pins stay literal — either on the
|
|
197
|
+
allowlisted era-pinned pages (`docs/*/operations/database-compatibility.md`,
|
|
198
|
+
`docs/*/operations/ncbi-recovery-migration.md`) or on a line carrying an
|
|
199
|
+
inline `<!-- version-pin -->` marker — and
|
|
200
|
+
`tests/unit/test_docs_versions.py` fails on any other hardcoded current
|
|
201
|
+
version. Only this `AGENTS.md` keeps literal current markers (it is not
|
|
202
|
+
Sphinx-rendered); update the list above when bumping.
|
|
203
|
+
|
|
204
|
+
## Special Note For Codex/ChatGPT
|
|
205
|
+
|
|
206
|
+
To ensure security, code testing in Codex/ChatGPT runs in a sandbox by default.
|
|
207
|
+
However, this causes the `test_tui.py` section to experience Textual/asyncio
|
|
208
|
+
cleanup blocking during testing, resulting in a "FAIL" report due to a timeout.
|
|
209
|
+
|
|
210
|
+
To run the full test suite, first exclude `test_tui.py`, then run it separately
|
|
211
|
+
outside the sandbox.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: OperonDBS
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.7.0
|
|
4
4
|
Summary: Operon the Database System: Archive, Quality-Control, Organize, Analyze and Release Your Bio-Data
|
|
5
5
|
Author-email: hyli360 <lihuanyu2003@gmail.com>
|
|
6
6
|
License-Expression: AGPL-3.0-or-later
|
|
@@ -60,6 +60,7 @@ docs/en/operations/database-compatibility.md
|
|
|
60
60
|
docs/en/operations/index.md
|
|
61
61
|
docs/en/operations/ncbi-recovery-migration.md
|
|
62
62
|
docs/en/operations/qc-performance.md
|
|
63
|
+
docs/en/reference/behaviors-and-limitations.md
|
|
63
64
|
docs/en/reference/cli-analysis.md
|
|
64
65
|
docs/en/reference/cli-decisions-reports.md
|
|
65
66
|
docs/en/reference/cli-files-qc.md
|
|
@@ -111,6 +112,7 @@ docs/zh/operations/database-compatibility.md
|
|
|
111
112
|
docs/zh/operations/index.md
|
|
112
113
|
docs/zh/operations/ncbi-recovery-migration.md
|
|
113
114
|
docs/zh/operations/qc-performance.md
|
|
115
|
+
docs/zh/reference/behaviors-and-limitations.md
|
|
114
116
|
docs/zh/reference/cli-analysis.md
|
|
115
117
|
docs/zh/reference/cli-decisions-reports.md
|
|
116
118
|
docs/zh/reference/cli-files-qc.md
|
|
@@ -197,6 +199,7 @@ tests/unit/test_cli_edges.py
|
|
|
197
199
|
tests/unit/test_config_workflow_edges.py
|
|
198
200
|
tests/unit/test_coverage_edges.py
|
|
199
201
|
tests/unit/test_database_edges_more.py
|
|
202
|
+
tests/unit/test_docs_versions.py
|
|
200
203
|
tests/unit/test_environment.py
|
|
201
204
|
tests/unit/test_execution.py
|
|
202
205
|
tests/unit/test_execution_edges.py
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: OperonDBS
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.7.0
|
|
4
4
|
Summary: Operon the Database System: Archive, Quality-Control, Organize, Analyze and Release Your Bio-Data
|
|
5
5
|
Author-email: hyli360 <lihuanyu2003@gmail.com>
|
|
6
6
|
License-Expression: AGPL-3.0-or-later
|
|
@@ -2,22 +2,60 @@
|
|
|
2
2
|
|
|
3
3
|
from __future__ import annotations
|
|
4
4
|
|
|
5
|
+
import re
|
|
5
6
|
from importlib.metadata import PackageNotFoundError, version
|
|
6
7
|
from pathlib import Path
|
|
7
8
|
|
|
9
|
+
try:
|
|
10
|
+
import tomllib
|
|
11
|
+
except ModuleNotFoundError: # Python 3.10
|
|
12
|
+
import tomli as tomllib
|
|
13
|
+
|
|
8
14
|
|
|
9
15
|
DOCS_DIR = Path(__file__).resolve().parent
|
|
16
|
+
REPO_ROOT = DOCS_DIR.parent
|
|
10
17
|
|
|
11
18
|
project = "Operon"
|
|
12
19
|
author = "Operon contributors"
|
|
13
20
|
copyright = "2026, Operon contributors"
|
|
14
21
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
22
|
+
|
|
23
|
+
def _package_version() -> str:
|
|
24
|
+
try:
|
|
25
|
+
return version("OperonDBS")
|
|
26
|
+
except PackageNotFoundError:
|
|
27
|
+
with (REPO_ROOT / "pyproject.toml").open("rb") as handle:
|
|
28
|
+
return tomllib.load(handle)["project"]["version"]
|
|
29
|
+
|
|
30
|
+
|
|
31
|
+
def _source_constant(module: str, name: str) -> str:
|
|
32
|
+
"""Read a module-level string constant, falling back to the source file."""
|
|
33
|
+
|
|
34
|
+
try:
|
|
35
|
+
imported = __import__(f"operon.{module}", fromlist=[name])
|
|
36
|
+
return str(getattr(imported, name))
|
|
37
|
+
except Exception:
|
|
38
|
+
source = (REPO_ROOT / "operon" / f"{module}.py").read_text(encoding="utf-8")
|
|
39
|
+
match = re.search(rf'^{name} = "([^"]+)"', source, re.MULTILINE)
|
|
40
|
+
if match is None:
|
|
41
|
+
raise RuntimeError(f"cannot resolve {name} from operon/{module}.py")
|
|
42
|
+
return match.group(1)
|
|
43
|
+
|
|
44
|
+
|
|
45
|
+
release = _package_version()
|
|
19
46
|
version = release
|
|
20
47
|
|
|
48
|
+
# Markdown sources reference these as {{ operon_version }} / {{ db_schema }} /
|
|
49
|
+
# {{ metadata_schema }} in paragraph text. Substitutions do not expand inside
|
|
50
|
+
# code spans or fenced code blocks, so examples there use `<version>`
|
|
51
|
+
# placeholders instead. Historical version mentions stay literal and are
|
|
52
|
+
# guarded by tests/unit/test_docs_versions.py.
|
|
53
|
+
myst_substitutions = {
|
|
54
|
+
"operon_version": release,
|
|
55
|
+
"db_schema": _source_constant("database", "SCHEMA_VERSION"),
|
|
56
|
+
"metadata_schema": _source_constant("schema", "METADATA_SCHEMA_VERSION"),
|
|
57
|
+
}
|
|
58
|
+
|
|
21
59
|
extensions = ["myst_parser"]
|
|
22
60
|
source_suffix = {".md": "markdown"}
|
|
23
61
|
root_doc = "index"
|
|
@@ -27,7 +65,7 @@ templates_path = ["_templates"]
|
|
|
27
65
|
# resolves their relative Markdown links as Sphinx cross-references, while the
|
|
28
66
|
# toctrees provide one coherent navigation hierarchy for both languages.
|
|
29
67
|
myst_heading_anchors = 4
|
|
30
|
-
myst_enable_extensions = ["colon_fence", "deflist", "fieldlist"]
|
|
68
|
+
myst_enable_extensions = ["colon_fence", "deflist", "fieldlist", "substitution"]
|
|
31
69
|
|
|
32
70
|
exclude_patterns = ["_build", "Thumbs.db", ".DS_Store"]
|
|
33
71
|
nitpicky = True
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## Current boundaries
|
|
4
4
|
|
|
5
|
-
The built-in source adapter currently covers NCBI Datasets; sources such as ENA remain part of the future extension boundary. Taxonomy coverage currently supports only NCBI Taxonomy; GTDB and the NCBI↔GTDB crosswalk are not yet implemented. Built-in QC covers file level, reads basics, assembly structure, and annotation structure. BUSCO is natively integrated through directory output and a JSON summary parser; tools without a parser yet — QUAST, Merqury, Kraken2, CheckM2, and similar — can still be integrated through `run-external` + `import-qc`. Downstream comparative-genomics analysis is done by external workflows in `analysis/`; `operon` is responsible for data admission, provenance, and publication.
|
|
5
|
+
The built-in source adapter currently covers NCBI Datasets; sources such as ENA remain part of the future extension boundary. Taxonomy coverage currently supports only NCBI Taxonomy; GTDB and the NCBI↔GTDB crosswalk are not yet implemented. Built-in QC covers file level, generic sequence basics for non-genome FASTA (CDS/protein), reads basics, assembly structure, and annotation structure. BUSCO is natively integrated through directory output and a JSON summary parser; tools without a parser yet — QUAST, Merqury, Kraken2, CheckM2, and similar — can still be integrated through `run-external` + `import-qc`. Downstream comparative-genomics analysis is done by external workflows in `analysis/`; `operon` is responsible for data admission, provenance, and publication.
|
|
6
6
|
|
|
7
7
|
The contract between downstream workflows and the database is a closed loop formed by `operon export` and `operon adopt`: export materializes the selected entities by file identity into a `data/<entity_type>/<entity_id>/<filename>` layout, accompanied by `manifest.tsv` (with SHA-256 recomputed over the materialized bytes), a `qc.tsv` QC long-table snapshot, `checksums.sha256`, and `provenance.json` (the input-side manifest); after consuming this artifact set, the external workflow uses adopt to re-register derived artifacts as first-class manifest members — materialized under `analysis/adopted/<entity_id>/`, inheriting the ingest idempotency/conflict invariants, with file-to-file lineage edges recorded in `file_lineage` (the output-side manifest). Adopted products can be QC'd, evaluated, exported, released, and selected by `analyze` as inputs of downstream recipes (cascading analysis). Downstream workflows should read and write through this contract instead of reading the database directly. Orchestration of cascading workflows (dependency graphs, parallelism, retries) belongs to workflow managers such as snakemake/nextflow; `operon` is responsible for data admission, lineage, and publication, while `run-pipeline` only covers simple single-file chaining. The batch adopt manifest format is described in the [external analysis guide](../guides/external-analysis.md). Export is semantically complementary to release: release targets publication (QC-gated, immutable snapshot), while export targets analysis inputs (arbitrary selection criteria, materialized on demand).
|
|
8
8
|
|
|
@@ -30,9 +30,9 @@ Execution environment capture (schema 2.8, `environment.py`): all three backends
|
|
|
30
30
|
|
|
31
31
|
A failed probe leaves the run's `environment_id` NULL without raising an error or affecting the run; rows from before 2.8 are likewise NULL.
|
|
32
32
|
|
|
33
|
-
Recipe versioning and snapshots (schema 2.9): a recipe gains an optional `version:` field (a positive integer, default 1; invalid values are rejected at configuration validation). As `analyze` processes each candidate file it records the current recipe together with the verbatim spec of its referenced tool into the `recipe_snapshots` table: the snapshot document is `{"recipe": <the recipe's raw mapping>, "tool": <the referenced tool spec's raw mapping>}`, content-addressed by the SHA-256 of its canonicalized JSON and deduplicated by `UNIQUE(recipe_name, recipe_version, recipe_sha256)` — so edits to the tool definition also produce a new snapshot, and cache hits record a snapshot of the current configuration as well. `analysis_jobs.recipe_snapshot_id` points back to the exact configuration that produced the job; jobs adopted during resume inherit the original job's snapshot id (they were produced by that configuration, not today's). Inspect them with `operon recipes list / history / show`; the QC-profile counterparts recorded in `qc_profiles` are inspected with `operon profiles history / show`. Restoration via the CLI is print-only in both cases: a human copies the output back into the configuration YAML, and the CLI never rewrites files in place. The audited alternative is the TUI Config screen, whose structured editors save every change as the next version with a new snapshot and can restore any recorded snapshot into the editor (saving it creates the next version); TUI recipe saves normalize `tools.yaml` formatting and drop hand-written comments.
|
|
33
|
+
Recipe versioning and snapshots (schema 2.9): a recipe gains an optional `version:` field (a positive integer, default 1; invalid values are rejected at configuration validation). As `analyze` processes each candidate file it records the current recipe together with the verbatim spec of its referenced tool into the `recipe_snapshots` table: the snapshot document is `{"recipe": <the recipe's raw mapping>, "tool": <the referenced tool spec's raw mapping>}`, content-addressed by the SHA-256 of its canonicalized JSON and deduplicated by `UNIQUE(recipe_name, recipe_version, recipe_sha256)` — so edits to the tool definition also produce a new snapshot, and cache hits record a snapshot of the current configuration as well. `analysis_jobs.recipe_snapshot_id` points back to the exact configuration that produced the job; jobs adopted during resume inherit the original job's snapshot id (they were produced by that configuration, not today's). Inspect them with `operon recipes list / history / show`; the QC-profile counterparts recorded in `qc_profiles` are inspected with `operon profiles history / show`. Restoration via the CLI is print-only in both cases: a human copies the output back into the configuration YAML, and the CLI never rewrites files in place. The audited alternative is the TUI Config screen, whose structured editors save every change as the next version with a new snapshot and can restore any recorded snapshot into the editor (saving it creates the next version); TUI recipe saves normalize `tools.yaml` formatting and drop hand-written comments. <!-- version-pin -->
|
|
34
34
|
|
|
35
|
-
Run resource-usage recording (schema 2.9): `workflow_runs` gains `duration_seconds` (wall clock; previously only present in the JSONL), `avg_rss_mb` (average RSS), and `cpu_seconds` (core-seconds), and the pre-existing `max_rss_mb` column is now actually populated. Collection is per backend:
|
|
35
|
+
Run resource-usage recording (schema 2.9): `workflow_runs` gains `duration_seconds` (wall clock; previously only present in the JSONL), `avg_rss_mb` (average RSS), and `cpu_seconds` (core-seconds), and the pre-existing `max_rss_mb` column is now actually populated. Collection is per backend: <!-- version-pin -->
|
|
36
36
|
|
|
37
37
|
- `local`: a sampling thread polls `VmRSS` in `/proc/<pid>/status` when procfs is available and otherwise uses the POSIX `ps` RSS field (including on macOS) for peak and average; core-seconds come from the `getrusage(RUSAGE_CHILDREN)` delta across the run;
|
|
38
38
|
- `slurm`: after the job finishes, `sacct` is queried with extended fields (`MaxRSS`/`AveRSS`/`Elapsed`/`TotalCPU`); remote Slurm (`ssh` with `scheduler: slurm`) follows the same path;
|
|
@@ -50,7 +50,7 @@ Identity and relationship policy:
|
|
|
50
50
|
- Records without a BioSample use an assembly-specific sample;
|
|
51
51
|
- Annotation identity includes source accession, provider, version, and release date, with files automatically assigned to the corresponding `ANN_`; pre-2.6 rows are continued with strictly identical metadata, avoiding duplicate assignment when the provider is not `NCBI *`.
|
|
52
52
|
|
|
53
|
-
Before writing metadata, the adapter computes SHA-256 for the files to be archived and checks both in-package conflicts for the same entity/role and existing manifest conflicts. Alternate genomes/reports from paired sources use controlled roles with `_genbank`/`_refseq` suffixes, so bytes from different sources can coexist without relaxing the no-overwrite constraint on the same entity and role. Original reports/ZIPs are stored by SHA-256 under `raw/metadata/ncbi_datasets/`; import summaries are written to `changes` and the workflow provenance. On a formal import into an old project, adapter-owned fields and source-file roles are merged in, and the metadata schema is upgraded to
|
|
53
|
+
Before writing metadata, the adapter computes SHA-256 for the files to be archived and checks both in-package conflicts for the same entity/role and existing manifest conflicts. Alternate genomes/reports from paired sources use controlled roles with `_genbank`/`_refseq` suffixes, so bytes from different sources can coexist without relaxing the no-overwrite constraint on the same entity and role. Original reports/ZIPs are stored by SHA-256 under `raw/metadata/ncbi_datasets/`; import summaries are written to `changes` and the workflow provenance. On a formal import into an old project, adapter-owned fields and source-file roles are merged in, and the metadata schema is upgraded to {{ metadata_schema }}; custom fields are preserved, and dry runs use only the in-memory upgraded schema.
|
|
54
54
|
|
|
55
55
|
An adapter run writes a `running` workflow before processing begins; each accession's state is kept in `adapter_run_items`. Failed or interrupted runs keep their state, and a resumed run uses a new run ID with `resumes_run_id`; a request whose SHA-256 does not match is refused. Field-level before/after values of metadata upserts are linked to the concrete run through `changes.workflow_run_id`. Anomalies from the old adapter are handled by an explicit `ncbi-reconcile` that generates and applies a compensation plan, preserving all old rows and files through `entity_supersessions`.
|
|
56
56
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Architecture overview
|
|
2
2
|
|
|
3
|
-
This document corresponds to `operon`
|
|
3
|
+
This document corresponds to `operon` {{ operon_version }}, internal database schema {{ db_schema }}, and metadata schema {{ metadata_schema }}.
|
|
4
4
|
|
|
5
5
|
## Design goals
|
|
6
6
|
|
|
@@ -22,7 +22,7 @@ How the principles map to implementations:
|
|
|
22
22
|
| Raw immutable, standardized derived | Atomic ingest + `ConflictError` + independent copies by default |
|
|
23
23
|
| Filenames contain only stable ID/role/format/compression | `canonical_filename()` |
|
|
24
24
|
| Paths are not file identity | `files.file_id + sha256 + size_bytes` |
|
|
25
|
-
| Layered QC | `file_integrity/reads_basic/assembly_basic/annotation_basic` |
|
|
25
|
+
| Layered QC | `file_integrity/reads_basic/sequence_basic/assembly_basic/annotation_basic` |
|
|
26
26
|
| Measurement separated from decision | `qc_results` long table + YAML profile rule engine |
|
|
27
27
|
| Taxonomy coverage does not drift with upstream upgrades | NCBI taxonomy snapshot + compiled reference-set TSV + SHA-256 |
|
|
28
28
|
| Automated state machine, explicit failures, idempotent resume | `entity_state` + strict transitions + atomic operations |
|
|
@@ -90,7 +90,7 @@ How the principles map to implementations:
|
|
|
90
90
|
| `operon/cli.py` | argparse command parsing, dispatch, human-readable output |
|
|
91
91
|
| `operon/config.py` | Reads `project.yaml`, locates the project root, generates the directory structure |
|
|
92
92
|
| `operon/schema.py` | Built-in metadata field definitions, type validation and normalization, derived TSV output |
|
|
93
|
-
| `operon/database.py` | SQLite DDL, WAL/foreign keys/indexes, development-time compatibility migrations and incremental schema 2.2–
|
|
93
|
+
| `operon/database.py` | SQLite DDL, WAL/foreign keys/indexes, development-time compatibility migrations and incremental schema 2.2–{{ db_schema }} migrations, transactions, read-only queries |
|
|
94
94
|
| `operon/files.py` | File format/compression detection, atomic archiving, idempotent ingest, checksum verification, standardized views |
|
|
95
95
|
| `operon/lifecycle.py` | Retire/restore plans, append-only lifecycle events, hierarchical propagation, and the current retired list |
|
|
96
96
|
| `operon/import_wizard.py` | English questionary import wizard, draft summary review, non-linear section editing, preflight and commit |
|
|
@@ -115,12 +115,12 @@ How the principles map to implementations:
|
|
|
115
115
|
|
|
116
116
|
## Project directory structure
|
|
117
117
|
|
|
118
|
-
`operon init` creates the following directories and files. The SQLite database is
|
|
118
|
+
`operon init` creates the following directories and files. The SQLite database is created eagerly by `operon init`; two entries appear only lazily when first needed (`logs/workflow.jsonl` on the first workflow run, `.operon/placeholders/` on the first remote-evict/pull pointer write).
|
|
119
119
|
|
|
120
120
|
```text
|
|
121
121
|
project/
|
|
122
122
|
├── project.yaml # project config: paths, default QC profile, resource parameters
|
|
123
|
-
├── operon.sqlite # file-based database (created
|
|
123
|
+
├── operon.sqlite # file-based database (created by operon init)
|
|
124
124
|
├── config/
|
|
125
125
|
│ ├── schemas.yaml # metadata field contract (types/required/allowed values/regex)
|
|
126
126
|
│ ├── tools.yaml # external analysis programs (BLAST/HMMER/BUSCO, artifact types)
|
|
@@ -128,6 +128,7 @@ project/
|
|
|
128
128
|
│ ├── file_integrity_v1.yaml
|
|
129
129
|
│ ├── assembly_production_v1.yaml
|
|
130
130
|
│ ├── annotation_release_v1.yaml
|
|
131
|
+
│ ├── annotation_busco_viridiplantae_odb12_v1.yaml
|
|
131
132
|
│ ├── reads_qc_v1.yaml
|
|
132
133
|
│ └── coverage_viridiplantae_v1.yaml
|
|
133
134
|
├── metadata/ # legacy layout compatibility note; no longer a read/write data source
|
|
@@ -137,8 +138,9 @@ project/
|
|
|
137
138
|
├── analysis/ # analysis workspace (external tool output, downstream analysis)
|
|
138
139
|
├── reports/ # decisions, summary exports, and coverage reports
|
|
139
140
|
├── taxonomy/reference_sets/ # compiled immutable family/genus denominators and provenance
|
|
140
|
-
├── logs/workflow.jsonl # machine-readable workflow log
|
|
141
|
+
├── logs/workflow.jsonl # machine-readable workflow log (created on the first workflow run)
|
|
141
142
|
├── .operon/placeholders/ # small, non-authoritative pointers for REMOTE_ONLY files
|
|
143
|
+
│ # (created on the first remote evict/pull)
|
|
142
144
|
└── releases/ # immutable dataset release snapshots
|
|
143
145
|
```
|
|
144
146
|
|
|
@@ -8,6 +8,7 @@ Built-in QC loads the Cython streaming parsers by default and requires them; the
|
|
|
8
8
|
|---|---|---|
|
|
9
9
|
| `file_integrity` | All files | `file_exists`, `size_bytes`, `sha256_match`, `parseable` |
|
|
10
10
|
| `assembly_basic` | genome FASTA | `total_length`, `contig_n50/n90`, `contig_l50/l90`, `gc_percent`, `n_percent` (strictly N only), `gap_count`/`gap_percent` (runs and fraction of alignment gap characters `-`), `ambiguous_base_percent`, duplicate seqids/complete headers, circular/empty sequences |
|
|
11
|
+
| `sequence_basic` | other FASTA (e.g. standalone CDS or protein FASTA) | `sequence_count`, `total_length`, `empty_sequence_count`, `duplicate_sequence_id_count` |
|
|
11
12
|
| `reads_basic` | FASTQ | `read_count`, `total_bases`, `q20_percent`, `q30_percent`, `gc_percent`, `duplicate_percent`, sampling count/strategy, `overrepresented_sequence_count`, read length N50, R1/R2 pairing |
|
|
12
13
|
| `annotation_basic` | GFF3 (+ assembly FASTA/protein FASTA) | gene/mRNA/CDS counts, CDS triplet ratio, ID/Parent integrity, coordinate errors, seqid matching, protein duplicate IDs, X ratio, internal stop codons |
|
|
13
14
|
|
|
@@ -22,13 +22,13 @@ Release content lands in a versioned directory:
|
|
|
22
22
|
A Linux build machine additionally needs the system command `patchelf`; it is a build-time tool for cx_Freeze's ELF dependency handling, not a Python runtime dependency of `operon`. When it is missing, cx_Freeze stops right at the `build_exe` stage.
|
|
23
23
|
|
|
24
24
|
```text
|
|
25
|
-
build/release/
|
|
25
|
+
build/release/v<version>/
|
|
26
26
|
├── operon # command-line executable; operon.exe on Windows
|
|
27
27
|
├── lib/ # Python runtime, the operon package, and third-party dependencies
|
|
28
28
|
├── LICENSE # Operon's own license (AGPL-3.0-or-later)
|
|
29
29
|
├── licenses/ # THIRD_PARTY_NOTICES.md and full license texts of third-party dependencies
|
|
30
30
|
├── source/
|
|
31
|
-
│ └── operondbs
|
|
31
|
+
│ └── operondbs-<version>.tar.gz # complete project source sdist corresponding to this binary
|
|
32
32
|
├── frozen_application_license.txt # license of the frozen bootstrap code automatically included by cx_Freeze
|
|
33
33
|
└── share/doc/operon/
|
|
34
34
|
├── README.md # English project overview
|
|
@@ -19,6 +19,14 @@ sphinx-build -W --keep-going -b html docs docs/_build/html
|
|
|
19
19
|
The pytest suite is organized into four categories — `unit`, `integration`, `regression`, `compatibility` — covering: Python 3.10 syntax and runtime gates, schema validation and controlled vocabularies, metadata round-trips and transaction rollback, stable IDs, default copy isolation, query read-only constraints, file-aware QC identity, profile/decision history, gzip FASTA recognition, assembly/annotation QC, rule decisions, idempotent ingest and conflict protection, checksum-tamper detection, the demo end-to-end pipeline and release verification, the NCBI Datasets adapter, wrapped BLAST/HMMER/BUSCO execution, directory artifacts, JSON summaries, conda run prefix parsing, cache hits/forced re-runs, result write-back, and input-tamper rejection.
|
|
20
20
|
The taxonomy coverage integration tests additionally cover taxonomy source-package identity conflicts, profile type/content conflicts, exclusion rules, secondary TaxIDs, denominator/report idempotence, and that active metadata modifications do not affect the release-frozen scope.
|
|
21
21
|
|
|
22
|
+
## Special Note For Codex/ChatGPT
|
|
23
|
+
|
|
24
|
+
To ensure security, code testing in Codex/ChatGPT runs in a sandbox by default. However, this causes the `test_tui.py` section to experience Textual/asyncio cleanup blocking during testing, resulting in a "FAIL" report due to a timeout.
|
|
25
|
+
|
|
26
|
+
To run the full test suite, first exclude `test_tui.py`, then run it separately outside the sandbox.
|
|
27
|
+
|
|
28
|
+
This information has also been updated in AGENTS.md.
|
|
29
|
+
|
|
22
30
|
## Documentation synchronization
|
|
23
31
|
|
|
24
32
|
When changing the CLI, configuration fields, behavior, or storage layout, update the Chinese and English documentation in the same change:
|
|
@@ -31,4 +39,10 @@ When changing the CLI, configuration fields, behavior, or storage layout, update
|
|
|
31
39
|
| `tools.yaml` recipes, placeholders, or parsers | `docs/*/reference/recipe-*.md` |
|
|
32
40
|
| Migrations, performance diagnostics, or compatibility boundaries | `docs/*/operations/` |
|
|
33
41
|
|
|
34
|
-
Software versions, database schema versions, and metadata schema versions stated in the documentation must stay consistent with `pyproject.toml` and the code.
|
|
42
|
+
Software versions, database schema versions, and metadata schema versions stated in the documentation must stay consistent with `pyproject.toml` and the code. Write current version markers in the Markdown sources as substitutions:
|
|
43
|
+
|
|
44
|
+
```text
|
|
45
|
+
{{ operon_version }} {{ db_schema }} {{ metadata_schema }}
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
`docs/conf.py` resolves them from `pyproject.toml` and the code constants at build time. Substitutions expand in paragraph text only, not inside code spans or fenced code blocks; examples there use `<version>` placeholders instead. Intentional historical pins stay literal: they either live on the allowlisted era-pinned pages under `docs/*/operations/` or carry an inline `<!-- version-pin -->` marker. `tests/unit/test_docs_versions.py` enforces the rule.
|
|
@@ -31,7 +31,11 @@ manual approval before the final upload.
|
|
|
31
31
|
|
|
32
32
|
## Release procedure
|
|
33
33
|
|
|
34
|
-
1. Update `[project].version
|
|
34
|
+
1. Update `[project].version`, plus `SCHEMA_VERSION` / `METADATA_SCHEMA_VERSION`
|
|
35
|
+
in the code when they change. Documentation version markers render from
|
|
36
|
+
these single sources through `myst_substitutions` in `docs/conf.py`, so no
|
|
37
|
+
manual sweep is needed; `tests/unit/test_docs_versions.py` rejects
|
|
38
|
+
hardcoded current versions in the Markdown sources.
|
|
35
39
|
2. Run the full pytest suite and strict documentation build.
|
|
36
40
|
3. Commit the release state and create tag `v<project.version>` on that exact
|
|
37
41
|
commit. Never reuse a tag that points to older package metadata.
|
|
@@ -18,7 +18,7 @@ metadata/ Legacy-layout compatibility note
|
|
|
18
18
|
raw/ standardized/ qc/ analysis/ reports/ logs/ releases/ taxonomy/
|
|
19
19
|
```
|
|
20
20
|
|
|
21
|
-
`operon
|
|
21
|
+
`operon init` also creates an empty `operon.sqlite` immediately, together with the directory tree.
|
|
22
22
|
|
|
23
23
|
> The global `--project` option must appear before the subcommand. It can be omitted inside the project root. Outside the project, use `operon --project /path/to/my-genome-project <subcommand>`.
|
|
24
24
|
|
|
@@ -51,11 +51,13 @@ Verify the installation:
|
|
|
51
51
|
|
|
52
52
|
```bash
|
|
53
53
|
operon --version
|
|
54
|
-
# Expected: operon 0.6.2
|
|
55
54
|
|
|
56
55
|
operon --help
|
|
57
56
|
```
|
|
58
57
|
|
|
58
|
+
The first command prints `operon` followed by the installed version —
|
|
59
|
+
{{ operon_version }} for the release this documentation matches.
|
|
60
|
+
|
|
59
61
|
To build a standalone cx_Freeze application, install the build extra and use the unified release entry point:
|
|
60
62
|
|
|
61
63
|
```bash
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## Back up and migrate a project
|
|
4
4
|
|
|
5
|
-
Use `backup` to create a consistent SQLite snapshot. Do not copy database files directly while the database may be active:
|
|
5
|
+
Use `backup` to create a consistent SQLite snapshot. Do not copy database files directly while the database may be active. The `--output` directory must be outside the project root and must not exist yet; `backup create` refuses otherwise:
|
|
6
6
|
|
|
7
7
|
```bash
|
|
8
8
|
# Configuration, SQLite, audit records, and workflow logs
|
|
@@ -17,8 +17,12 @@ operon backup create --output /backups/my-project-full --scope full
|
|
|
17
17
|
operon backup verify --input /backups/my-project-full
|
|
18
18
|
```
|
|
19
19
|
|
|
20
|
+
Note the scope boundaries: `results` excludes `raw/` and `standardized/` (the bytes you usually cannot regenerate), so it is not a restorable substitute for `full`; only `full` can restore data files.
|
|
21
|
+
|
|
20
22
|
`backup verify` validates an exact snapshot. In addition to checking size and SHA-256 for files listed in the manifest, it rejects extra files in the backup directory. Keep notes, temporary files, and recovery records outside the backup directory.
|
|
21
23
|
|
|
24
|
+
New backups use manifest format 2; format 1 backups remain verifiable. Symbolic links are recorded and verified by their target text, including broken links and directory links, without following their targets. Absolute links inside the standardized view that point into the project are rebased to relative paths within the backup, so a full backup can be moved independently. Links inside archived directory artifacts retain their exact text to preserve artifact identity. External link targets are not backed up; restoring such a link does not restore its external referent.
|
|
25
|
+
|
|
22
26
|
With `REMOTE_ONLY` files, a local backup must include the SQLite database containing `file_locations`. Back up the remote mirror root independently, including `operon-manifest.json` and actual objects. Placeholder files are not recovery evidence. Safe hydration requires both local `files` identity and the remote manifest/bytes.
|
|
23
27
|
|
|
24
28
|
`report metadata` is not a backup. It exports metadata/manifest TSV files for browsing and exchange, but does not include complete QC, decisions, changes, workflows, remote locations, or migration state.
|
|
@@ -55,6 +59,6 @@ Core steps are idempotent:
|
|
|
55
59
|
- `release` rejects an existing version directory rather than overwriting it.
|
|
56
60
|
- `taxonomy compile` reuses identical profile/taxonomy/TSV input and rejects different content under the same identity.
|
|
57
61
|
- `report coverage` validates and reuses an old report when input membership, profile, and reference-set identity match.
|
|
58
|
-
- After Ctrl+C/SIGTERM, `analyze` marks the current job `interrupted` and removes partial outputs. On rerun, completed files use the cache; an old result with unchanged input and verified output is adopted (`adopted`); only unfinished work is recomputed.
|
|
62
|
+
- After Ctrl+C/SIGTERM, `analyze` marks the current job `interrupted` and removes partial outputs (`--keep-partial` preserves them for debugging). On rerun, completed files use the cache; an old result with unchanged input and verified output is adopted (`adopted`); only unfinished work is recomputed.
|
|
59
63
|
|
|
60
64
|
Rerun the same command to continue from the interruption. Use `status` to inspect each entity's current state.
|
|
@@ -104,4 +104,15 @@ Restoration is the strict inverse operation: it appends `RESTORE` and points bac
|
|
|
104
104
|
|
|
105
105
|
For databases older than schema 2.7, run `operon migrate` first.
|
|
106
106
|
|
|
107
|
+
## Force a state transition manually
|
|
108
|
+
|
|
109
|
+
`operon set-state` performs an audited manual state change when a transition is needed that the automatic workflow does not produce (for example, recovering a stuck entity):
|
|
110
|
+
|
|
111
|
+
```bash
|
|
112
|
+
operon set-state --entity-type assembly --entity-id ASM_000001 --state QC_COMPLETE \
|
|
113
|
+
--message "Manual review confirmed metrics are complete" --force
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
The normal path enforces the legal transition table and appends a `changes` audit row with the message; `--force` bypasses the transition check for manual recovery while the audit row keeps the action traceable. Note that `RELEASED` is a terminal state: entities published in a release cannot leave it without `--force`, and setting a state that equals the current state is a silent no-op. Prefer `curate` (for decisions) or rerunning the affected step (for provenance) whenever possible.
|
|
117
|
+
|
|
107
118
|
There is currently no `purge` command. Do not use manual SQL, `rm`, or remote-object deletion as a substitute. Physical deletion requires separately designed retention periods, release/remote reference protection, a restoration window, and irreversible confirmation. Until then, auditable retirement and restoration are the supported safe-disposal path.
|
|
@@ -92,6 +92,8 @@ Preview selection and cache status:
|
|
|
92
92
|
operon analyze --analysis blastn_nt --dry-run
|
|
93
93
|
```
|
|
94
94
|
|
|
95
|
+
Useful batch controls: `--limit N` processes only the first N matching files (in `file_id` order), and `--threads` overrides the recipe default.
|
|
96
|
+
|
|
95
97
|
Inspect synchronized results:
|
|
96
98
|
|
|
97
99
|
```bash
|
|
@@ -210,6 +212,7 @@ operon run-external \
|
|
|
210
212
|
```
|
|
211
213
|
|
|
212
214
|
- `--command` is parsed with shell-style quoting but is not run through a shell.
|
|
215
|
+
- `--tool` records a tool version probed from `config/tools.yaml`; `--input` declares an input file/directory that is hashed for provenance (repeatable); `--threads`, `--cwd`, `--timeout`, and `--backend` (default `local`; also `slurm` or `ssh`) control execution.
|
|
213
216
|
- stdout and stderr are saved to `logs/<WF_ID>.stdout.log` and `.stderr.log`.
|
|
214
217
|
- Run records are written to `logs/workflow.jsonl` and `workflow_runs`.
|
|
215
218
|
- The run is `completed` only when the exit code is 0 and every `--expected-output` exists and is non-empty; otherwise it is `failed` and the command exits non-zero.
|
|
@@ -255,5 +258,5 @@ operon adopt --from-manifest adopt_manifest.json
|
|
|
255
258
|
```
|
|
256
259
|
|
|
257
260
|
- Each item requires `path`, `entity_type`, `entity_id`, `role`, and `derived_from` (at least one already-registered file_id); relative paths resolve from the project root.
|
|
258
|
-
- Artifacts are materialized under `analysis/adopted/<entity_id>/`; same entity and role with identical bytes is reused idempotently, different bytes raise `ConflictError
|
|
261
|
+
- Artifacts are materialized under `analysis/adopted/<entity_id>/`; same entity and role with identical bytes is reused idempotently, different bytes raise `ConflictError`. The whole batch is preflighted, then registered in one transaction. A failure before commit rolls back metadata, lineage, state and workflow rows and removes newly created artifacts; existing files are preserved. Resolve conflicting occupied targets explicitly before retrying. Completed JSONL records are written only after commit.
|
|
259
262
|
- Roles are freely named by the workflow; lineage edges are written to the `file_lineage` table and can be audited with `operon query`.
|
|
@@ -25,7 +25,7 @@ Import behavior:
|
|
|
25
25
|
- Schema, controlled-vocabulary, and foreign-key validation run before the row-by-row `insert/update/unchanged` preview.
|
|
26
26
|
- `--on-conflict error` rejects existing rows, `skip` skips them, and `update` updates fields with per-field audit records.
|
|
27
27
|
- Deletes and full-snapshot replacement are not supported. If any write fails, the transaction for the table is rolled back.
|
|
28
|
-
- For XLSX, the first
|
|
28
|
+
- For XLSX, the first worksheet in workbook order is imported whatever its name (templates generated by `--template` name it `data` and add a read-only `schema` documentation worksheet).
|
|
29
29
|
|
|
30
30
|
CSV example:
|
|
31
31
|
|
|
@@ -12,7 +12,7 @@ operon ncbi-datasets \
|
|
|
12
12
|
|
|
13
13
|
The output includes the number of organism/sample/assembly/annotation IDs to create in `new_ids` and the rows to upsert in `metadata_rows`. A dry run does not copy input, write the database, or create logs.
|
|
14
14
|
|
|
15
|
-
If the project still uses an old metadata schema, a formal import preserves custom fields, adds the fields and paired-source file roles required by the adapter, and upgrades the schema to
|
|
15
|
+
If the project still uses an old metadata schema, a formal import preserves custom fields, adds the fields and paired-source file roles required by the adapter, and upgrades the schema to {{ metadata_schema }}. A dry run does not modify the schema.
|
|
16
16
|
|
|
17
17
|
After review, remove `--dry-run`:
|
|
18
18
|
|
|
@@ -33,7 +33,7 @@ warnings:
|
|
|
33
33
|
code: HIGH_BUSCO_DUPLICATION
|
|
34
34
|
```
|
|
35
35
|
|
|
36
|
-
Supported operators are `>=`, `<=`, `>`, `<`, `==`, `!=`, `between` (requires `min` and `max`), `in`, `not_in` (requires `values`), and `exists`.
|
|
36
|
+
Supported operators are `>=`, `<=`, `>`, `<`, `==`, `!=`, `between` (requires `min` and `max`), `in`, `not_in` (requires `values`), and `exists`. Comparisons are inclusive at the boundaries, and `in`/`not_in` compare metric values as strings. Note two profile-validation gaps: a `between` rule missing `min`/`max`, or an `in`/`not_in` rule missing `values`, is not rejected when the profile loads — the former surfaces as a Python traceback at evaluation time, and the latter silently evaluates against an empty value set.
|
|
37
37
|
|
|
38
38
|
Hand-editing the YAML file is fully supported. The audited alternative is the
|
|
39
39
|
TUI Config screen (`operon tui`, key `6`, QC Profiles tab): a structured form
|
|
@@ -18,6 +18,7 @@ remotes:
|
|
|
18
18
|
# Alternatively pin an administrator-provided fingerprint:
|
|
19
19
|
# host_key_sha256: SHA256:base64...
|
|
20
20
|
insecure_accept_unknown_host: false
|
|
21
|
+
connect_timeout: 30 # seconds; also bounds the remote manifest lock wait
|
|
21
22
|
```
|
|
22
23
|
|
|
23
24
|
Paramiko is included in the standard `OperonDBS` installation.
|
|
@@ -54,7 +55,7 @@ The remote model preserves the raw-file invariants:
|
|
|
54
55
|
- Remote relative paths must remain safely under the remote root. By default, `pull` checks every record against local SQLite `file_id + relative_path + sha256 + size_bytes`; the remote manifest cannot rewrite local identity.
|
|
55
56
|
- Every transfer writes workflow provenance (`push:<name>` or `pull:<name>`), and successful locations are recorded in `file_locations`.
|
|
56
57
|
- A failed item does not stop the rest of a push/pull/evict batch. Every item receives a result, and the command exits with code 1 if any item has `error`.
|
|
57
|
-
- After `pull` restores a missing local file, `files.status` returns to `CHECKSUM_VERIFIED` and the change is audited in `changes`.
|
|
58
|
+
- After `pull` restores a missing local file, `files.status` returns to `CHECKSUM_VERIFIED` and the change is audited in `changes`. A file that was already `STANDARDIZED` before eviction keeps the `STANDARDIZED` status after restore.
|
|
58
59
|
|
|
59
60
|
## Keep the control plane local and large files remote
|
|
60
61
|
|
|
@@ -95,6 +96,8 @@ execution:
|
|
|
95
96
|
|
|
96
97
|
`evict` explicitly deletes local bytes; without `--file-id`, it processes every manifest object. It first validates local identity, remote manifest identity, and actual remote SHA-256/tree hash. The state change is written to `changes`. `standardize` and `release` require local bytes, so run `pull` first. External `analyze` can consume `REMOTE_ONLY` input directly.
|
|
97
98
|
|
|
99
|
+
Eviction writes a small placeholder pointer file under `.operon/placeholders/<file_id>.json` (deleted again when `pull` restores the bytes). The first remote-only status also extends `config/schemas.yaml` with the `REMOTE_ONLY` file status and bumps its `schema_version` to 1.2 — the file is rewritten with normalized formatting, so hand-written comments in it are dropped.
|
|
100
|
+
|
|
98
101
|
When a local object is missing, `verify` checks the remote in real time rather than treating `file_locations.status=AVAILABLE` as permanent proof. A deleted or damaged remote object returns `MISSING` and updates the cache. An unreachable SSH host returns `REMOTE_UNVERIFIED` and exit code 1 while preserving the last persistent state, so a network failure is not misclassified as data loss.
|
|
99
102
|
|
|
100
103
|
Remote files can also be archived directly from URLs:
|
|
@@ -18,3 +18,7 @@ Recommended actions:
|
|
|
18
18
|
| `CHECKSUM_FAILED` | Stop QC. Determine whether the file was modified and restore it from the original source. |
|
|
19
19
|
| `QC_FAILED` | Inspect files with `parseable=0` in `operon report qc`, then use `operon workflow list --step qc --status failed` and `operon workflow show WF_ID` for the recorded error and execution details. |
|
|
20
20
|
| Format parsing failure | Check with an external validator such as `seqkit stats` or a GFF3 validator. Archive the repaired file as a new version; do not overwrite raw data. |
|
|
21
|
+
|
|
22
|
+
## Further reading
|
|
23
|
+
|
|
24
|
+
For implicit semantics, edge cases, and known issues that are not defects in a single workflow — such as how missing metrics are treated, multi-file QC state semantics, or exit-code conventions — see [Implicit Behaviors, Edge Cases, and Known Issues](../reference/behaviors-and-limitations.md).
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Operon is a file-backed database for large-scale genomic data. It supports metadata management, immutable file archiving, quality control (QC), rule-based decisions, external analysis, remote storage and execution, and versioned dataset releases.
|
|
4
4
|
|
|
5
|
-
This documentation matches `operon`
|
|
5
|
+
This documentation matches `operon` {{ operon_version }}, database schema {{ db_schema }}, and metadata schema {{ metadata_schema }}. The Chinese and English documentation use the same directory structure.
|
|
6
6
|
|
|
7
7
|
## Reading paths
|
|
8
8
|
|
|
@@ -19,7 +19,7 @@ File: `operon/database.py`
|
|
|
19
19
|
|
|
20
20
|
`Database._migrate_remote_schema_2_2()` is also not part of the "development-era v1 compatibility layer" above. It upgrades a 2.1 database to 2.2 purely additively: adding `executor`, `scheduler_job_id`, and `execution_details` to `workflow_runs`, and creating `file_locations`. It must be kept as long as opening 2.1 projects is supported; if that support ever ends, it should be replaced through the formal database migration policy, not deleted together with `_migrate_pre_1_0_schema()`. The corresponding test is `test_schema_2_2_adds_remote_location_and_executor_provenance`.
|
|
21
21
|
|
|
22
|
-
`Database._migrate_taxonomy_schema_2_3()` is likewise a purely additive migration needed by current functionality, not part of `_migrate_pre_1_0_schema()`: for 2.2 projects it creates `taxonomy_snapshots`, `taxonomy_nodes`, `taxonomy_aliases`, `taxonomy_reference_sets`, `coverage_reports`, and `coverage_report_metrics` plus related indexes, without modifying existing business rows. It must be kept as long as opening 2.2 projects is supported. The corresponding regression test is `
|
|
22
|
+
`Database._migrate_taxonomy_schema_2_3()` is likewise a purely additive migration needed by current functionality, not part of `_migrate_pre_1_0_schema()`: for 2.2 projects it creates `taxonomy_snapshots`, `taxonomy_nodes`, `taxonomy_aliases`, `taxonomy_reference_sets`, `coverage_reports`, and `coverage_report_metrics` plus related indexes, without modifying existing business rows. It must be kept as long as opening 2.2 projects is supported. The corresponding regression test is `test_schema_2_3_adds_taxonomy_and_coverage_history`.
|
|
23
23
|
|
|
24
24
|
`Database._migrate_source_schema_2_4()` is another purely additive migration needed by current functionality: for 2.3 projects it creates `data_sources` and `source_links`, storing normalized external databases/repositories, citations, licenses, and their associated objects. Non-INSDC sources must contain both citation and License; source content is deduplicated by SHA-256 identity. It must be kept as long as opening 2.3 projects is supported. The corresponding regression test is `test_schema_2_4_adds_normalized_source_provenance`.
|
|
25
25
|
|
|
@@ -20,6 +20,8 @@ Operon manages data admission, identity verification, provenance, rule evaluatio
|
|
|
20
20
|
|
|
21
21
|
The current source adapter supports NCBI Datasets. Taxonomy coverage currently supports NCBI Taxonomy only; GTDB and NCBI↔GTDB crosswalks are extension work.
|
|
22
22
|
|
|
23
|
+
Implicit semantics, edge cases, and known issues that are not covered by the task-facing pages are catalogued in [Implicit Behaviors, Edge Cases, and Known Issues](reference/behaviors-and-limitations.md).
|
|
24
|
+
|
|
23
25
|
## Recommended workflow
|
|
24
26
|
|
|
25
27
|
```text
|