OperonDBS 0.6.2__tar.gz → 0.7.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (235) hide show
  1. {operondbs-0.6.2 → operondbs-0.7.0}/.github/workflows/publish.yml +3 -3
  2. {operondbs-0.6.2 → operondbs-0.7.0}/AGENTS.md +22 -1
  3. {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/PKG-INFO +1 -1
  4. {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/SOURCES.txt +3 -0
  5. {operondbs-0.6.2 → operondbs-0.7.0}/PKG-INFO +1 -1
  6. {operondbs-0.6.2 → operondbs-0.7.0}/docs/conf.py +43 -5
  7. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/extensibility.md +1 -1
  8. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/external-analysis.md +2 -2
  9. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/metadata-and-data-model.md +1 -1
  10. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/overview.md +8 -6
  11. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/qc-and-rules.md +1 -0
  12. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/application-release.md +2 -2
  13. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/development-testing.md +15 -1
  14. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/pypi-release.md +5 -1
  15. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/first-project.md +1 -1
  16. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/installation.md +3 -1
  17. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/backup-migration.md +6 -2
  18. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/curation-lifecycle.md +11 -0
  19. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/external-analysis.md +4 -1
  20. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/metadata-import.md +1 -1
  21. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/ncbi-datasets.md +1 -1
  22. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/qc-profiles.md +1 -1
  23. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/remote-storage.md +4 -1
  24. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/troubleshooting.md +4 -0
  25. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/index.md +1 -1
  26. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/database-compatibility.md +1 -1
  27. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/overview.md +2 -0
  28. operondbs-0.7.0/docs/en/reference/behaviors-and-limitations.md +192 -0
  29. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-decisions-reports.md +7 -3
  30. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-files-qc.md +4 -2
  31. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-tui.md +4 -1
  32. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/index.md +2 -0
  33. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/recipe-fields.md +1 -1
  34. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/recipe-overview.md +1 -1
  35. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/recipe-parsers-examples.md +1 -1
  36. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/extensibility.md +1 -1
  37. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/external-analysis.md +2 -2
  38. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/metadata-and-data-model.md +1 -1
  39. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/overview.md +8 -6
  40. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/qc-and-rules.md +1 -0
  41. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/application-release.md +2 -2
  42. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/development-testing.md +17 -1
  43. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/pypi-release.md +4 -1
  44. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/first-project.md +1 -1
  45. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/installation.md +3 -1
  46. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/backup-migration.md +10 -4
  47. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/curation-lifecycle.md +15 -0
  48. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/external-analysis.md +9 -2
  49. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/metadata-import.md +1 -1
  50. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/ncbi-datasets.md +1 -1
  51. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/qc-profiles.md +1 -1
  52. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/remote-storage.md +7 -1
  53. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/troubleshooting.md +6 -0
  54. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/index.md +1 -1
  55. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/database-compatibility.md +1 -1
  56. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/overview.md +2 -0
  57. operondbs-0.7.0/docs/zh/reference/behaviors-and-limitations.md +190 -0
  58. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-decisions-reports.md +7 -4
  59. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-files-qc.md +4 -2
  60. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-tui.md +2 -0
  61. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/index.md +2 -0
  62. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/recipe-fields.md +2 -2
  63. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/recipe-overview.md +2 -2
  64. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/recipe-parsers-examples.md +1 -1
  65. {operondbs-0.6.2 → operondbs-0.7.0}/operon/backup.py +50 -8
  66. {operondbs-0.6.2 → operondbs-0.7.0}/operon/cli.py +76 -4
  67. {operondbs-0.6.2 → operondbs-0.7.0}/operon/database.py +77 -7
  68. {operondbs-0.6.2 → operondbs-0.7.0}/operon/export.py +68 -8
  69. {operondbs-0.6.2 → operondbs-0.7.0}/operon/files.py +85 -45
  70. {operondbs-0.6.2 → operondbs-0.7.0}/operon/import_wizard.py +1 -1
  71. {operondbs-0.6.2 → operondbs-0.7.0}/operon/lineage.py +73 -14
  72. {operondbs-0.6.2 → operondbs-0.7.0}/operon/qc_module/__init__.py +97 -12
  73. {operondbs-0.6.2 → operondbs-0.7.0}/operon/release.py +149 -19
  74. {operondbs-0.6.2 → operondbs-0.7.0}/operon/rules.py +67 -8
  75. {operondbs-0.6.2 → operondbs-0.7.0}/operon/table_import.py +11 -7
  76. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/actions.py +33 -31
  77. {operondbs-0.6.2 → operondbs-0.7.0}/operon/utils.py +0 -2
  78. {operondbs-0.6.2 → operondbs-0.7.0}/pyproject.toml +2 -2
  79. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_cli_edges.py +29 -1
  80. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_database_edges_more.py +23 -1
  81. operondbs-0.7.0/tests/unit/test_docs_versions.py +69 -0
  82. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_export.py +26 -0
  83. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_files_edges.py +52 -0
  84. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_import_wizard_edges.py +9 -0
  85. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_lineage.py +80 -0
  86. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_qc_and_rules.py +41 -0
  87. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_rules_schema_edges.py +31 -0
  88. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_support_edges.py +49 -1
  89. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_table_import_edges.py +13 -0
  90. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tui_config.py +38 -0
  91. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_views_release_reports_edges.py +57 -0
  92. {operondbs-0.6.2 → operondbs-0.7.0}/.github/workflows/deploy.yml +0 -0
  93. {operondbs-0.6.2 → operondbs-0.7.0}/.gitignore +0 -0
  94. {operondbs-0.6.2 → operondbs-0.7.0}/.readthedocs.yaml +0 -0
  95. {operondbs-0.6.2 → operondbs-0.7.0}/LICENSE +0 -0
  96. {operondbs-0.6.2 → operondbs-0.7.0}/MANIFEST.in +0 -0
  97. {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/dependency_links.txt +0 -0
  98. {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/entry_points.txt +0 -0
  99. {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/requires.txt +0 -0
  100. {operondbs-0.6.2 → operondbs-0.7.0}/OperonDBS.egg-info/top_level.txt +0 -0
  101. {operondbs-0.6.2 → operondbs-0.7.0}/README.md +0 -0
  102. {operondbs-0.6.2 → operondbs-0.7.0}/README_ZH.md +0 -0
  103. {operondbs-0.6.2 → operondbs-0.7.0}/benchmarks/qc_representative_entities.tsv +0 -0
  104. {operondbs-0.6.2 → operondbs-0.7.0}/docs/_static/language-switcher.js +0 -0
  105. {operondbs-0.6.2 → operondbs-0.7.0}/docs/_static/operon.css +0 -0
  106. {operondbs-0.6.2 → operondbs-0.7.0}/docs/_templates/layout.html +0 -0
  107. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/files-and-storage.md +0 -0
  108. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/index.md +0 -0
  109. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/release-lifecycle.md +0 -0
  110. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/architecture/taxonomy-coverage.md +0 -0
  111. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/documentation-deployment.md +0 -0
  112. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/index.md +0 -0
  113. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/contributor/repository-guide.md +0 -0
  114. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/daily-workflow.md +0 -0
  115. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/index.md +0 -0
  116. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/getting-started/quickstart.md +0 -0
  117. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/file-archiving.md +0 -0
  118. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/index.md +0 -0
  119. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/remote-execution.md +0 -0
  120. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/guides/taxonomy-coverage.md +0 -0
  121. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/index.md +0 -0
  122. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/ncbi-recovery-migration.md +0 -0
  123. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/operations/qc-performance.md +0 -0
  124. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-analysis.md +0 -0
  125. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-project-metadata.md +0 -0
  126. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-remote.md +0 -0
  127. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-taxonomy-lifecycle-admin.md +0 -0
  128. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/cli-workflow.md +0 -0
  129. {operondbs-0.6.2 → operondbs-0.7.0}/docs/en/reference/data-model.md +0 -0
  130. {operondbs-0.6.2 → operondbs-0.7.0}/docs/index.md +0 -0
  131. {operondbs-0.6.2 → operondbs-0.7.0}/docs/requirements.txt +0 -0
  132. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/files-and-storage.md +0 -0
  133. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/index.md +0 -0
  134. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/release-lifecycle.md +0 -0
  135. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/architecture/taxonomy-coverage.md +0 -0
  136. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/documentation-deployment.md +0 -0
  137. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/index.md +0 -0
  138. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/contributor/repository-guide.md +0 -0
  139. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/daily-workflow.md +0 -0
  140. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/index.md +0 -0
  141. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/getting-started/quickstart.md +0 -0
  142. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/file-archiving.md +0 -0
  143. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/index.md +0 -0
  144. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/remote-execution.md +0 -0
  145. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/guides/taxonomy-coverage.md +0 -0
  146. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/index.md +0 -0
  147. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/ncbi-recovery-migration.md +0 -0
  148. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/operations/qc-performance.md +0 -0
  149. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-analysis.md +0 -0
  150. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-project-metadata.md +0 -0
  151. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-remote.md +0 -0
  152. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-taxonomy-lifecycle-admin.md +0 -0
  153. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/cli-workflow.md +0 -0
  154. {operondbs-0.6.2 → operondbs-0.7.0}/docs/zh/reference/data-model.md +0 -0
  155. {operondbs-0.6.2 → operondbs-0.7.0}/operon/__init__.py +0 -0
  156. {operondbs-0.6.2 → operondbs-0.7.0}/operon/__main__.py +0 -0
  157. {operondbs-0.6.2 → operondbs-0.7.0}/operon/adapters/__init__.py +0 -0
  158. {operondbs-0.6.2 → operondbs-0.7.0}/operon/adapters/ncbi_datasets.py +0 -0
  159. {operondbs-0.6.2 → operondbs-0.7.0}/operon/config.py +0 -0
  160. {operondbs-0.6.2 → operondbs-0.7.0}/operon/coverage.py +0 -0
  161. {operondbs-0.6.2 → operondbs-0.7.0}/operon/demo.py +0 -0
  162. {operondbs-0.6.2 → operondbs-0.7.0}/operon/entity_view.py +0 -0
  163. {operondbs-0.6.2 → operondbs-0.7.0}/operon/environment.py +0 -0
  164. {operondbs-0.6.2 → operondbs-0.7.0}/operon/errors.py +0 -0
  165. {operondbs-0.6.2 → operondbs-0.7.0}/operon/execution.py +0 -0
  166. {operondbs-0.6.2 → operondbs-0.7.0}/operon/lifecycle.py +0 -0
  167. {operondbs-0.6.2 → operondbs-0.7.0}/operon/metadata_files.py +0 -0
  168. {operondbs-0.6.2 → operondbs-0.7.0}/operon/ncbi_reconcile.py +0 -0
  169. {operondbs-0.6.2 → operondbs-0.7.0}/operon/profiles.py +0 -0
  170. {operondbs-0.6.2 → operondbs-0.7.0}/operon/qc_module/_parsers.pyx +0 -0
  171. {operondbs-0.6.2 → operondbs-0.7.0}/operon/qc_module/parsers.py +0 -0
  172. {operondbs-0.6.2 → operondbs-0.7.0}/operon/remotes.py +0 -0
  173. {operondbs-0.6.2 → operondbs-0.7.0}/operon/reports.py +0 -0
  174. {operondbs-0.6.2 → operondbs-0.7.0}/operon/schema.py +0 -0
  175. {operondbs-0.6.2 → operondbs-0.7.0}/operon/shutdown.py +0 -0
  176. {operondbs-0.6.2 → operondbs-0.7.0}/operon/taxonomy.py +0 -0
  177. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tools.py +0 -0
  178. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/__init__.py +0 -0
  179. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/app.py +0 -0
  180. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/app.tcss +0 -0
  181. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/data.py +0 -0
  182. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/__init__.py +0 -0
  183. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/common.py +0 -0
  184. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/config.py +0 -0
  185. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/decisions.py +0 -0
  186. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/entities.py +0 -0
  187. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/files.py +0 -0
  188. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/files_ops.py +0 -0
  189. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/home.py +0 -0
  190. {operondbs-0.6.2 → operondbs-0.7.0}/operon/tui/screens/runs.py +0 -0
  191. {operondbs-0.6.2 → operondbs-0.7.0}/operon/workflow.py +0 -0
  192. {operondbs-0.6.2 → operondbs-0.7.0}/setup.cfg +0 -0
  193. {operondbs-0.6.2 → operondbs-0.7.0}/setup.py +0 -0
  194. {operondbs-0.6.2 → operondbs-0.7.0}/tests/__init__.py +0 -0
  195. {operondbs-0.6.2 → operondbs-0.7.0}/tests/compatibility/__init__.py +0 -0
  196. {operondbs-0.6.2 → operondbs-0.7.0}/tests/compatibility/test_python_support.py +0 -0
  197. {operondbs-0.6.2 → operondbs-0.7.0}/tests/helpers.py +0 -0
  198. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/__init__.py +0 -0
  199. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_analysis_resume.py +0 -0
  200. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_analysis_shutdown.py +0 -0
  201. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_analysis_tools.py +0 -0
  202. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_application_build.py +0 -0
  203. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_execution_backends.py +0 -0
  204. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_lineage_cascade.py +0 -0
  205. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_ncbi_datasets_adapter.py +0 -0
  206. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_pipeline_and_release.py +0 -0
  207. {operondbs-0.6.2 → operondbs-0.7.0}/tests/integration/test_taxonomy_coverage.py +0 -0
  208. {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/__init__.py +0 -0
  209. {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/test_correctness.py +0 -0
  210. {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/test_cython_parser_parity.py +0 -0
  211. {operondbs-0.6.2 → operondbs-0.7.0}/tests/regression/test_parser_semantics.py +0 -0
  212. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/__init__.py +0 -0
  213. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_config_workflow_edges.py +0 -0
  214. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_coverage_edges.py +0 -0
  215. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_environment.py +0 -0
  216. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_execution.py +0 -0
  217. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_execution_edges.py +0 -0
  218. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_import_backup_show.py +0 -0
  219. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_lifecycle.py +0 -0
  220. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_ncbi_edge_cases.py +0 -0
  221. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_ncbi_reconcile_edges.py +0 -0
  222. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_parser_edge_paths.py +0 -0
  223. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_qc_edges_more.py +0 -0
  224. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_recipe_history.py +0 -0
  225. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_remotes.py +0 -0
  226. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_remotes_edges.py +0 -0
  227. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_schema_2_9.py +0 -0
  228. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_schema_and_metadata.py +0 -0
  229. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_shutdown.py +0 -0
  230. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_taxonomy_edges.py +0 -0
  231. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tools_edges.py +0 -0
  232. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tui.py +0 -0
  233. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_tui_writes.py +0 -0
  234. {operondbs-0.6.2 → operondbs-0.7.0}/tests/unit/test_workflow_cli.py +0 -0
  235. {operondbs-0.6.2 → operondbs-0.7.0}/tools/build.py +0 -0
@@ -28,7 +28,7 @@ jobs:
28
28
  - run: python -m pip install --upgrade build twine
29
29
  - run: python -m build --sdist
30
30
  - run: python -m twine check dist/*
31
- - uses: actions/upload-artifact@v5
31
+ - uses: actions/upload-artifact@v7
32
32
  with:
33
33
  name: python-package-sdist
34
34
  path: dist/*.tar.gz
@@ -65,7 +65,7 @@ jobs:
65
65
  CIBW_TEST_COMMAND: >-
66
66
  python -c "import operon; import operon.qc_module._parsers;
67
67
  print(operon.__version__)" && operon --help
68
- - uses: actions/upload-artifact@v5
68
+ - uses: actions/upload-artifact@v7
69
69
  with:
70
70
  name: python-package-wheel-${{ matrix.artifact }}
71
71
  path: wheelhouse/*.whl
@@ -80,7 +80,7 @@ jobs:
80
80
  permissions:
81
81
  id-token: write
82
82
  steps:
83
- - uses: actions/download-artifact@v6
83
+ - uses: actions/download-artifact@v8
84
84
  with:
85
85
  pattern: python-package-*
86
86
  path: dist/
@@ -187,4 +187,25 @@ change:
187
187
  `docs/*/index.md`
188
188
 
189
189
  Version markers in docs (`operon` 0.6.2, database schema 2.9, metadata
190
- schema 1.4) must match `pyproject.toml` and the code.
190
+ schema 1.4) must match `pyproject.toml` and the code. Do not write the
191
+ current values literally in Markdown sources: use the `myst_substitutions`
192
+ references `{{ operon_version }}`, `{{ db_schema }}`, and
193
+ `{{ metadata_schema }}`, which `docs/conf.py` resolves from the single
194
+ sources above at build time. Substitutions expand in paragraph text only,
195
+ never inside code spans or fenced code blocks — examples there use
196
+ `<version>` placeholders instead. Intentional historical pins stay literal — either on the
197
+ allowlisted era-pinned pages (`docs/*/operations/database-compatibility.md`,
198
+ `docs/*/operations/ncbi-recovery-migration.md`) or on a line carrying an
199
+ inline `<!-- version-pin -->` marker — and
200
+ `tests/unit/test_docs_versions.py` fails on any other hardcoded current
201
+ version. Only this `AGENTS.md` keeps literal current markers (it is not
202
+ Sphinx-rendered); update the list above when bumping.
203
+
204
+ ## Special Note For Codex/ChatGPT
205
+
206
+ To ensure security, code testing in Codex/ChatGPT runs in a sandbox by default.
207
+ However, this causes the `test_tui.py` section to experience Textual/asyncio
208
+ cleanup blocking during testing, resulting in a "FAIL" report due to a timeout.
209
+
210
+ To run the full test suite, first exclude `test_tui.py`, then run it separately
211
+ outside the sandbox.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: OperonDBS
3
- Version: 0.6.2
3
+ Version: 0.7.0
4
4
  Summary: Operon the Database System: Archive, Quality-Control, Organize, Analyze and Release Your Bio-Data
5
5
  Author-email: hyli360 <lihuanyu2003@gmail.com>
6
6
  License-Expression: AGPL-3.0-or-later
@@ -60,6 +60,7 @@ docs/en/operations/database-compatibility.md
60
60
  docs/en/operations/index.md
61
61
  docs/en/operations/ncbi-recovery-migration.md
62
62
  docs/en/operations/qc-performance.md
63
+ docs/en/reference/behaviors-and-limitations.md
63
64
  docs/en/reference/cli-analysis.md
64
65
  docs/en/reference/cli-decisions-reports.md
65
66
  docs/en/reference/cli-files-qc.md
@@ -111,6 +112,7 @@ docs/zh/operations/database-compatibility.md
111
112
  docs/zh/operations/index.md
112
113
  docs/zh/operations/ncbi-recovery-migration.md
113
114
  docs/zh/operations/qc-performance.md
115
+ docs/zh/reference/behaviors-and-limitations.md
114
116
  docs/zh/reference/cli-analysis.md
115
117
  docs/zh/reference/cli-decisions-reports.md
116
118
  docs/zh/reference/cli-files-qc.md
@@ -197,6 +199,7 @@ tests/unit/test_cli_edges.py
197
199
  tests/unit/test_config_workflow_edges.py
198
200
  tests/unit/test_coverage_edges.py
199
201
  tests/unit/test_database_edges_more.py
202
+ tests/unit/test_docs_versions.py
200
203
  tests/unit/test_environment.py
201
204
  tests/unit/test_execution.py
202
205
  tests/unit/test_execution_edges.py
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: OperonDBS
3
- Version: 0.6.2
3
+ Version: 0.7.0
4
4
  Summary: Operon the Database System: Archive, Quality-Control, Organize, Analyze and Release Your Bio-Data
5
5
  Author-email: hyli360 <lihuanyu2003@gmail.com>
6
6
  License-Expression: AGPL-3.0-or-later
@@ -2,22 +2,60 @@
2
2
 
3
3
  from __future__ import annotations
4
4
 
5
+ import re
5
6
  from importlib.metadata import PackageNotFoundError, version
6
7
  from pathlib import Path
7
8
 
9
+ try:
10
+ import tomllib
11
+ except ModuleNotFoundError: # Python 3.10
12
+ import tomli as tomllib
13
+
8
14
 
9
15
  DOCS_DIR = Path(__file__).resolve().parent
16
+ REPO_ROOT = DOCS_DIR.parent
10
17
 
11
18
  project = "Operon"
12
19
  author = "Operon contributors"
13
20
  copyright = "2026, Operon contributors"
14
21
 
15
- try:
16
- release = version("OperonDBS")
17
- except PackageNotFoundError:
18
- release = "0.6.2"
22
+
23
+ def _package_version() -> str:
24
+ try:
25
+ return version("OperonDBS")
26
+ except PackageNotFoundError:
27
+ with (REPO_ROOT / "pyproject.toml").open("rb") as handle:
28
+ return tomllib.load(handle)["project"]["version"]
29
+
30
+
31
+ def _source_constant(module: str, name: str) -> str:
32
+ """Read a module-level string constant, falling back to the source file."""
33
+
34
+ try:
35
+ imported = __import__(f"operon.{module}", fromlist=[name])
36
+ return str(getattr(imported, name))
37
+ except Exception:
38
+ source = (REPO_ROOT / "operon" / f"{module}.py").read_text(encoding="utf-8")
39
+ match = re.search(rf'^{name} = "([^"]+)"', source, re.MULTILINE)
40
+ if match is None:
41
+ raise RuntimeError(f"cannot resolve {name} from operon/{module}.py")
42
+ return match.group(1)
43
+
44
+
45
+ release = _package_version()
19
46
  version = release
20
47
 
48
+ # Markdown sources reference these as {{ operon_version }} / {{ db_schema }} /
49
+ # {{ metadata_schema }} in paragraph text. Substitutions do not expand inside
50
+ # code spans or fenced code blocks, so examples there use `<version>`
51
+ # placeholders instead. Historical version mentions stay literal and are
52
+ # guarded by tests/unit/test_docs_versions.py.
53
+ myst_substitutions = {
54
+ "operon_version": release,
55
+ "db_schema": _source_constant("database", "SCHEMA_VERSION"),
56
+ "metadata_schema": _source_constant("schema", "METADATA_SCHEMA_VERSION"),
57
+ }
58
+
21
59
  extensions = ["myst_parser"]
22
60
  source_suffix = {".md": "markdown"}
23
61
  root_doc = "index"
@@ -27,7 +65,7 @@ templates_path = ["_templates"]
27
65
  # resolves their relative Markdown links as Sphinx cross-references, while the
28
66
  # toctrees provide one coherent navigation hierarchy for both languages.
29
67
  myst_heading_anchors = 4
30
- myst_enable_extensions = ["colon_fence", "deflist", "fieldlist"]
68
+ myst_enable_extensions = ["colon_fence", "deflist", "fieldlist", "substitution"]
31
69
 
32
70
  exclude_patterns = ["_build", "Thumbs.db", ".DS_Store"]
33
71
  nitpicky = True
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## Current boundaries
4
4
 
5
- The built-in source adapter currently covers NCBI Datasets; sources such as ENA remain part of the future extension boundary. Taxonomy coverage currently supports only NCBI Taxonomy; GTDB and the NCBI↔GTDB crosswalk are not yet implemented. Built-in QC covers file level, reads basics, assembly structure, and annotation structure. BUSCO is natively integrated through directory output and a JSON summary parser; tools without a parser yet — QUAST, Merqury, Kraken2, CheckM2, and similar — can still be integrated through `run-external` + `import-qc`. Downstream comparative-genomics analysis is done by external workflows in `analysis/`; `operon` is responsible for data admission, provenance, and publication.
5
+ The built-in source adapter currently covers NCBI Datasets; sources such as ENA remain part of the future extension boundary. Taxonomy coverage currently supports only NCBI Taxonomy; GTDB and the NCBI↔GTDB crosswalk are not yet implemented. Built-in QC covers file level, generic sequence basics for non-genome FASTA (CDS/protein), reads basics, assembly structure, and annotation structure. BUSCO is natively integrated through directory output and a JSON summary parser; tools without a parser yet — QUAST, Merqury, Kraken2, CheckM2, and similar — can still be integrated through `run-external` + `import-qc`. Downstream comparative-genomics analysis is done by external workflows in `analysis/`; `operon` is responsible for data admission, provenance, and publication.
6
6
 
7
7
  The contract between downstream workflows and the database is a closed loop formed by `operon export` and `operon adopt`: export materializes the selected entities by file identity into a `data/<entity_type>/<entity_id>/<filename>` layout, accompanied by `manifest.tsv` (with SHA-256 recomputed over the materialized bytes), a `qc.tsv` QC long-table snapshot, `checksums.sha256`, and `provenance.json` (the input-side manifest); after consuming this artifact set, the external workflow uses adopt to re-register derived artifacts as first-class manifest members — materialized under `analysis/adopted/<entity_id>/`, inheriting the ingest idempotency/conflict invariants, with file-to-file lineage edges recorded in `file_lineage` (the output-side manifest). Adopted products can be QC'd, evaluated, exported, released, and selected by `analyze` as inputs of downstream recipes (cascading analysis). Downstream workflows should read and write through this contract instead of reading the database directly. Orchestration of cascading workflows (dependency graphs, parallelism, retries) belongs to workflow managers such as snakemake/nextflow; `operon` is responsible for data admission, lineage, and publication, while `run-pipeline` only covers simple single-file chaining. The batch adopt manifest format is described in the [external analysis guide](../guides/external-analysis.md). Export is semantically complementary to release: release targets publication (QC-gated, immutable snapshot), while export targets analysis inputs (arbitrary selection criteria, materialized on demand).
8
8
 
@@ -30,9 +30,9 @@ Execution environment capture (schema 2.8, `environment.py`): all three backends
30
30
 
31
31
  A failed probe leaves the run's `environment_id` NULL without raising an error or affecting the run; rows from before 2.8 are likewise NULL.
32
32
 
33
- Recipe versioning and snapshots (schema 2.9): a recipe gains an optional `version:` field (a positive integer, default 1; invalid values are rejected at configuration validation). As `analyze` processes each candidate file it records the current recipe together with the verbatim spec of its referenced tool into the `recipe_snapshots` table: the snapshot document is `{"recipe": <the recipe's raw mapping>, "tool": <the referenced tool spec's raw mapping>}`, content-addressed by the SHA-256 of its canonicalized JSON and deduplicated by `UNIQUE(recipe_name, recipe_version, recipe_sha256)` — so edits to the tool definition also produce a new snapshot, and cache hits record a snapshot of the current configuration as well. `analysis_jobs.recipe_snapshot_id` points back to the exact configuration that produced the job; jobs adopted during resume inherit the original job's snapshot id (they were produced by that configuration, not today's). Inspect them with `operon recipes list / history / show`; the QC-profile counterparts recorded in `qc_profiles` are inspected with `operon profiles history / show`. Restoration via the CLI is print-only in both cases: a human copies the output back into the configuration YAML, and the CLI never rewrites files in place. The audited alternative is the TUI Config screen, whose structured editors save every change as the next version with a new snapshot and can restore any recorded snapshot into the editor (saving it creates the next version); TUI recipe saves normalize `tools.yaml` formatting and drop hand-written comments.
33
+ Recipe versioning and snapshots (schema 2.9): a recipe gains an optional `version:` field (a positive integer, default 1; invalid values are rejected at configuration validation). As `analyze` processes each candidate file it records the current recipe together with the verbatim spec of its referenced tool into the `recipe_snapshots` table: the snapshot document is `{"recipe": <the recipe's raw mapping>, "tool": <the referenced tool spec's raw mapping>}`, content-addressed by the SHA-256 of its canonicalized JSON and deduplicated by `UNIQUE(recipe_name, recipe_version, recipe_sha256)` — so edits to the tool definition also produce a new snapshot, and cache hits record a snapshot of the current configuration as well. `analysis_jobs.recipe_snapshot_id` points back to the exact configuration that produced the job; jobs adopted during resume inherit the original job's snapshot id (they were produced by that configuration, not today's). Inspect them with `operon recipes list / history / show`; the QC-profile counterparts recorded in `qc_profiles` are inspected with `operon profiles history / show`. Restoration via the CLI is print-only in both cases: a human copies the output back into the configuration YAML, and the CLI never rewrites files in place. The audited alternative is the TUI Config screen, whose structured editors save every change as the next version with a new snapshot and can restore any recorded snapshot into the editor (saving it creates the next version); TUI recipe saves normalize `tools.yaml` formatting and drop hand-written comments. <!-- version-pin -->
34
34
 
35
- Run resource-usage recording (schema 2.9): `workflow_runs` gains `duration_seconds` (wall clock; previously only present in the JSONL), `avg_rss_mb` (average RSS), and `cpu_seconds` (core-seconds), and the pre-existing `max_rss_mb` column is now actually populated. Collection is per backend:
35
+ Run resource-usage recording (schema 2.9): `workflow_runs` gains `duration_seconds` (wall clock; previously only present in the JSONL), `avg_rss_mb` (average RSS), and `cpu_seconds` (core-seconds), and the pre-existing `max_rss_mb` column is now actually populated. Collection is per backend: <!-- version-pin -->
36
36
 
37
37
  - `local`: a sampling thread polls `VmRSS` in `/proc/<pid>/status` when procfs is available and otherwise uses the POSIX `ps` RSS field (including on macOS) for peak and average; core-seconds come from the `getrusage(RUSAGE_CHILDREN)` delta across the run;
38
38
  - `slurm`: after the job finishes, `sacct` is queried with extended fields (`MaxRSS`/`AveRSS`/`Elapsed`/`TotalCPU`); remote Slurm (`ssh` with `scheduler: slurm`) follows the same path;
@@ -50,7 +50,7 @@ Identity and relationship policy:
50
50
  - Records without a BioSample use an assembly-specific sample;
51
51
  - Annotation identity includes source accession, provider, version, and release date, with files automatically assigned to the corresponding `ANN_`; pre-2.6 rows are continued with strictly identical metadata, avoiding duplicate assignment when the provider is not `NCBI *`.
52
52
 
53
- Before writing metadata, the adapter computes SHA-256 for the files to be archived and checks both in-package conflicts for the same entity/role and existing manifest conflicts. Alternate genomes/reports from paired sources use controlled roles with `_genbank`/`_refseq` suffixes, so bytes from different sources can coexist without relaxing the no-overwrite constraint on the same entity and role. Original reports/ZIPs are stored by SHA-256 under `raw/metadata/ncbi_datasets/`; import summaries are written to `changes` and the workflow provenance. On a formal import into an old project, adapter-owned fields and source-file roles are merged in, and the metadata schema is upgraded to 1.4; custom fields are preserved, and dry runs use only the in-memory upgraded schema.
53
+ Before writing metadata, the adapter computes SHA-256 for the files to be archived and checks both in-package conflicts for the same entity/role and existing manifest conflicts. Alternate genomes/reports from paired sources use controlled roles with `_genbank`/`_refseq` suffixes, so bytes from different sources can coexist without relaxing the no-overwrite constraint on the same entity and role. Original reports/ZIPs are stored by SHA-256 under `raw/metadata/ncbi_datasets/`; import summaries are written to `changes` and the workflow provenance. On a formal import into an old project, adapter-owned fields and source-file roles are merged in, and the metadata schema is upgraded to {{ metadata_schema }}; custom fields are preserved, and dry runs use only the in-memory upgraded schema.
54
54
 
55
55
  An adapter run writes a `running` workflow before processing begins; each accession's state is kept in `adapter_run_items`. Failed or interrupted runs keep their state, and a resumed run uses a new run ID with `resumes_run_id`; a request whose SHA-256 does not match is refused. Field-level before/after values of metadata upserts are linked to the concrete run through `changes.workflow_run_id`. Anomalies from the old adapter are handled by an explicit `ncbi-reconcile` that generates and applies a compensation plan, preserving all old rows and files through `entity_supersessions`.
56
56
 
@@ -1,6 +1,6 @@
1
1
  # Architecture overview
2
2
 
3
- This document corresponds to `operon` 0.6.2, internal database schema 2.9, and metadata schema 1.4.
3
+ This document corresponds to `operon` {{ operon_version }}, internal database schema {{ db_schema }}, and metadata schema {{ metadata_schema }}.
4
4
 
5
5
  ## Design goals
6
6
 
@@ -22,7 +22,7 @@ How the principles map to implementations:
22
22
  | Raw immutable, standardized derived | Atomic ingest + `ConflictError` + independent copies by default |
23
23
  | Filenames contain only stable ID/role/format/compression | `canonical_filename()` |
24
24
  | Paths are not file identity | `files.file_id + sha256 + size_bytes` |
25
- | Layered QC | `file_integrity/reads_basic/assembly_basic/annotation_basic` |
25
+ | Layered QC | `file_integrity/reads_basic/sequence_basic/assembly_basic/annotation_basic` |
26
26
  | Measurement separated from decision | `qc_results` long table + YAML profile rule engine |
27
27
  | Taxonomy coverage does not drift with upstream upgrades | NCBI taxonomy snapshot + compiled reference-set TSV + SHA-256 |
28
28
  | Automated state machine, explicit failures, idempotent resume | `entity_state` + strict transitions + atomic operations |
@@ -90,7 +90,7 @@ How the principles map to implementations:
90
90
  | `operon/cli.py` | argparse command parsing, dispatch, human-readable output |
91
91
  | `operon/config.py` | Reads `project.yaml`, locates the project root, generates the directory structure |
92
92
  | `operon/schema.py` | Built-in metadata field definitions, type validation and normalization, derived TSV output |
93
- | `operon/database.py` | SQLite DDL, WAL/foreign keys/indexes, development-time compatibility migrations and incremental schema 2.2–2.9 migrations, transactions, read-only queries |
93
+ | `operon/database.py` | SQLite DDL, WAL/foreign keys/indexes, development-time compatibility migrations and incremental schema 2.2–{{ db_schema }} migrations, transactions, read-only queries |
94
94
  | `operon/files.py` | File format/compression detection, atomic archiving, idempotent ingest, checksum verification, standardized views |
95
95
  | `operon/lifecycle.py` | Retire/restore plans, append-only lifecycle events, hierarchical propagation, and the current retired list |
96
96
  | `operon/import_wizard.py` | English questionary import wizard, draft summary review, non-linear section editing, preflight and commit |
@@ -115,12 +115,12 @@ How the principles map to implementations:
115
115
 
116
116
  ## Project directory structure
117
117
 
118
- `operon init` creates the following directories and files. The SQLite database is not created at init time, but on the first command that needs it.
118
+ `operon init` creates the following directories and files. The SQLite database is created eagerly by `operon init`; two entries appear only lazily when first needed (`logs/workflow.jsonl` on the first workflow run, `.operon/placeholders/` on the first remote-evict/pull pointer write).
119
119
 
120
120
  ```text
121
121
  project/
122
122
  ├── project.yaml # project config: paths, default QC profile, resource parameters
123
- ├── operon.sqlite # file-based database (created on first command use)
123
+ ├── operon.sqlite # file-based database (created by operon init)
124
124
  ├── config/
125
125
  │ ├── schemas.yaml # metadata field contract (types/required/allowed values/regex)
126
126
  │ ├── tools.yaml # external analysis programs (BLAST/HMMER/BUSCO, artifact types)
@@ -128,6 +128,7 @@ project/
128
128
  │ ├── file_integrity_v1.yaml
129
129
  │ ├── assembly_production_v1.yaml
130
130
  │ ├── annotation_release_v1.yaml
131
+ │ ├── annotation_busco_viridiplantae_odb12_v1.yaml
131
132
  │ ├── reads_qc_v1.yaml
132
133
  │ └── coverage_viridiplantae_v1.yaml
133
134
  ├── metadata/ # legacy layout compatibility note; no longer a read/write data source
@@ -137,8 +138,9 @@ project/
137
138
  ├── analysis/ # analysis workspace (external tool output, downstream analysis)
138
139
  ├── reports/ # decisions, summary exports, and coverage reports
139
140
  ├── taxonomy/reference_sets/ # compiled immutable family/genus denominators and provenance
140
- ├── logs/workflow.jsonl # machine-readable workflow log
141
+ ├── logs/workflow.jsonl # machine-readable workflow log (created on the first workflow run)
141
142
  ├── .operon/placeholders/ # small, non-authoritative pointers for REMOTE_ONLY files
143
+ │ # (created on the first remote evict/pull)
142
144
  └── releases/ # immutable dataset release snapshots
143
145
  ```
144
146
 
@@ -8,6 +8,7 @@ Built-in QC loads the Cython streaming parsers by default and requires them; the
8
8
  |---|---|---|
9
9
  | `file_integrity` | All files | `file_exists`, `size_bytes`, `sha256_match`, `parseable` |
10
10
  | `assembly_basic` | genome FASTA | `total_length`, `contig_n50/n90`, `contig_l50/l90`, `gc_percent`, `n_percent` (strictly N only), `gap_count`/`gap_percent` (runs and fraction of alignment gap characters `-`), `ambiguous_base_percent`, duplicate seqids/complete headers, circular/empty sequences |
11
+ | `sequence_basic` | other FASTA (e.g. standalone CDS or protein FASTA) | `sequence_count`, `total_length`, `empty_sequence_count`, `duplicate_sequence_id_count` |
11
12
  | `reads_basic` | FASTQ | `read_count`, `total_bases`, `q20_percent`, `q30_percent`, `gc_percent`, `duplicate_percent`, sampling count/strategy, `overrepresented_sequence_count`, read length N50, R1/R2 pairing |
12
13
  | `annotation_basic` | GFF3 (+ assembly FASTA/protein FASTA) | gene/mRNA/CDS counts, CDS triplet ratio, ID/Parent integrity, coordinate errors, seqid matching, protein duplicate IDs, X ratio, internal stop codons |
13
14
 
@@ -22,13 +22,13 @@ Release content lands in a versioned directory:
22
22
  A Linux build machine additionally needs the system command `patchelf`; it is a build-time tool for cx_Freeze's ELF dependency handling, not a Python runtime dependency of `operon`. When it is missing, cx_Freeze stops right at the `build_exe` stage.
23
23
 
24
24
  ```text
25
- build/release/v0.6.2/
25
+ build/release/v<version>/
26
26
  ├── operon # command-line executable; operon.exe on Windows
27
27
  ├── lib/ # Python runtime, the operon package, and third-party dependencies
28
28
  ├── LICENSE # Operon's own license (AGPL-3.0-or-later)
29
29
  ├── licenses/ # THIRD_PARTY_NOTICES.md and full license texts of third-party dependencies
30
30
  ├── source/
31
- │ └── operondbs-0.6.2.tar.gz # complete project source sdist corresponding to this binary
31
+ │ └── operondbs-<version>.tar.gz # complete project source sdist corresponding to this binary
32
32
  ├── frozen_application_license.txt # license of the frozen bootstrap code automatically included by cx_Freeze
33
33
  └── share/doc/operon/
34
34
  ├── README.md # English project overview
@@ -19,6 +19,14 @@ sphinx-build -W --keep-going -b html docs docs/_build/html
19
19
  The pytest suite is organized into four categories — `unit`, `integration`, `regression`, `compatibility` — covering: Python 3.10 syntax and runtime gates, schema validation and controlled vocabularies, metadata round-trips and transaction rollback, stable IDs, default copy isolation, query read-only constraints, file-aware QC identity, profile/decision history, gzip FASTA recognition, assembly/annotation QC, rule decisions, idempotent ingest and conflict protection, checksum-tamper detection, the demo end-to-end pipeline and release verification, the NCBI Datasets adapter, wrapped BLAST/HMMER/BUSCO execution, directory artifacts, JSON summaries, conda run prefix parsing, cache hits/forced re-runs, result write-back, and input-tamper rejection.
20
20
  The taxonomy coverage integration tests additionally cover taxonomy source-package identity conflicts, profile type/content conflicts, exclusion rules, secondary TaxIDs, denominator/report idempotence, and that active metadata modifications do not affect the release-frozen scope.
21
21
 
22
+ ## Special Note For Codex/ChatGPT
23
+
24
+ To ensure security, code testing in Codex/ChatGPT runs in a sandbox by default. However, this causes the `test_tui.py` section to experience Textual/asyncio cleanup blocking during testing, resulting in a "FAIL" report due to a timeout.
25
+
26
+ To run the full test suite, first exclude `test_tui.py`, then run it separately outside the sandbox.
27
+
28
+ This information has also been updated in AGENTS.md.
29
+
22
30
  ## Documentation synchronization
23
31
 
24
32
  When changing the CLI, configuration fields, behavior, or storage layout, update the Chinese and English documentation in the same change:
@@ -31,4 +39,10 @@ When changing the CLI, configuration fields, behavior, or storage layout, update
31
39
  | `tools.yaml` recipes, placeholders, or parsers | `docs/*/reference/recipe-*.md` |
32
40
  | Migrations, performance diagnostics, or compatibility boundaries | `docs/*/operations/` |
33
41
 
34
- Software versions, database schema versions, and metadata schema versions stated in the documentation must stay consistent with `pyproject.toml` and the code.
42
+ Software versions, database schema versions, and metadata schema versions stated in the documentation must stay consistent with `pyproject.toml` and the code. Write current version markers in the Markdown sources as substitutions:
43
+
44
+ ```text
45
+ {{ operon_version }} {{ db_schema }} {{ metadata_schema }}
46
+ ```
47
+
48
+ `docs/conf.py` resolves them from `pyproject.toml` and the code constants at build time. Substitutions expand in paragraph text only, not inside code spans or fenced code blocks; examples there use `<version>` placeholders instead. Intentional historical pins stay literal: they either live on the allowlisted era-pinned pages under `docs/*/operations/` or carry an inline `<!-- version-pin -->` marker. `tests/unit/test_docs_versions.py` enforces the rule.
@@ -31,7 +31,11 @@ manual approval before the final upload.
31
31
 
32
32
  ## Release procedure
33
33
 
34
- 1. Update `[project].version` and every documented version marker together.
34
+ 1. Update `[project].version`, plus `SCHEMA_VERSION` / `METADATA_SCHEMA_VERSION`
35
+ in the code when they change. Documentation version markers render from
36
+ these single sources through `myst_substitutions` in `docs/conf.py`, so no
37
+ manual sweep is needed; `tests/unit/test_docs_versions.py` rejects
38
+ hardcoded current versions in the Markdown sources.
35
39
  2. Run the full pytest suite and strict documentation build.
36
40
  3. Commit the release state and create tag `v<project.version>` on that exact
37
41
  commit. Never reuse a tag that points to older package metadata.
@@ -18,7 +18,7 @@ metadata/ Legacy-layout compatibility note
18
18
  raw/ standardized/ qc/ analysis/ reports/ logs/ releases/ taxonomy/
19
19
  ```
20
20
 
21
- `operon.sqlite` is created the first time a command needs the database.
21
+ `operon init` also creates an empty `operon.sqlite` immediately, together with the directory tree.
22
22
 
23
23
  > The global `--project` option must appear before the subcommand. It can be omitted inside the project root. Outside the project, use `operon --project /path/to/my-genome-project <subcommand>`.
24
24
 
@@ -51,11 +51,13 @@ Verify the installation:
51
51
 
52
52
  ```bash
53
53
  operon --version
54
- # Expected: operon 0.6.2
55
54
 
56
55
  operon --help
57
56
  ```
58
57
 
58
+ The first command prints `operon` followed by the installed version —
59
+ {{ operon_version }} for the release this documentation matches.
60
+
59
61
  To build a standalone cx_Freeze application, install the build extra and use the unified release entry point:
60
62
 
61
63
  ```bash
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## Back up and migrate a project
4
4
 
5
- Use `backup` to create a consistent SQLite snapshot. Do not copy database files directly while the database may be active:
5
+ Use `backup` to create a consistent SQLite snapshot. Do not copy database files directly while the database may be active. The `--output` directory must be outside the project root and must not exist yet; `backup create` refuses otherwise:
6
6
 
7
7
  ```bash
8
8
  # Configuration, SQLite, audit records, and workflow logs
@@ -17,8 +17,12 @@ operon backup create --output /backups/my-project-full --scope full
17
17
  operon backup verify --input /backups/my-project-full
18
18
  ```
19
19
 
20
+ Note the scope boundaries: `results` excludes `raw/` and `standardized/` (the bytes you usually cannot regenerate), so it is not a restorable substitute for `full`; only `full` can restore data files.
21
+
20
22
  `backup verify` validates an exact snapshot. In addition to checking size and SHA-256 for files listed in the manifest, it rejects extra files in the backup directory. Keep notes, temporary files, and recovery records outside the backup directory.
21
23
 
24
+ New backups use manifest format 2; format 1 backups remain verifiable. Symbolic links are recorded and verified by their target text, including broken links and directory links, without following their targets. Absolute links inside the standardized view that point into the project are rebased to relative paths within the backup, so a full backup can be moved independently. Links inside archived directory artifacts retain their exact text to preserve artifact identity. External link targets are not backed up; restoring such a link does not restore its external referent.
25
+
22
26
  With `REMOTE_ONLY` files, a local backup must include the SQLite database containing `file_locations`. Back up the remote mirror root independently, including `operon-manifest.json` and actual objects. Placeholder files are not recovery evidence. Safe hydration requires both local `files` identity and the remote manifest/bytes.
23
27
 
24
28
  `report metadata` is not a backup. It exports metadata/manifest TSV files for browsing and exchange, but does not include complete QC, decisions, changes, workflows, remote locations, or migration state.
@@ -55,6 +59,6 @@ Core steps are idempotent:
55
59
  - `release` rejects an existing version directory rather than overwriting it.
56
60
  - `taxonomy compile` reuses identical profile/taxonomy/TSV input and rejects different content under the same identity.
57
61
  - `report coverage` validates and reuses an old report when input membership, profile, and reference-set identity match.
58
- - After Ctrl+C/SIGTERM, `analyze` marks the current job `interrupted` and removes partial outputs. On rerun, completed files use the cache; an old result with unchanged input and verified output is adopted (`adopted`); only unfinished work is recomputed.
62
+ - After Ctrl+C/SIGTERM, `analyze` marks the current job `interrupted` and removes partial outputs (`--keep-partial` preserves them for debugging). On rerun, completed files use the cache; an old result with unchanged input and verified output is adopted (`adopted`); only unfinished work is recomputed.
59
63
 
60
64
  Rerun the same command to continue from the interruption. Use `status` to inspect each entity's current state.
@@ -104,4 +104,15 @@ Restoration is the strict inverse operation: it appends `RESTORE` and points bac
104
104
 
105
105
  For databases older than schema 2.7, run `operon migrate` first.
106
106
 
107
+ ## Force a state transition manually
108
+
109
+ `operon set-state` performs an audited manual state change when a transition is needed that the automatic workflow does not produce (for example, recovering a stuck entity):
110
+
111
+ ```bash
112
+ operon set-state --entity-type assembly --entity-id ASM_000001 --state QC_COMPLETE \
113
+ --message "Manual review confirmed metrics are complete" --force
114
+ ```
115
+
116
+ The normal path enforces the legal transition table and appends a `changes` audit row with the message; `--force` bypasses the transition check for manual recovery while the audit row keeps the action traceable. Note that `RELEASED` is a terminal state: entities published in a release cannot leave it without `--force`, and setting a state that equals the current state is a silent no-op. Prefer `curate` (for decisions) or rerunning the affected step (for provenance) whenever possible.
117
+
107
118
  There is currently no `purge` command. Do not use manual SQL, `rm`, or remote-object deletion as a substitute. Physical deletion requires separately designed retention periods, release/remote reference protection, a restoration window, and irreversible confirmation. Until then, auditable retirement and restoration are the supported safe-disposal path.
@@ -92,6 +92,8 @@ Preview selection and cache status:
92
92
  operon analyze --analysis blastn_nt --dry-run
93
93
  ```
94
94
 
95
+ Useful batch controls: `--limit N` processes only the first N matching files (in `file_id` order), and `--threads` overrides the recipe default.
96
+
95
97
  Inspect synchronized results:
96
98
 
97
99
  ```bash
@@ -210,6 +212,7 @@ operon run-external \
210
212
  ```
211
213
 
212
214
  - `--command` is parsed with shell-style quoting but is not run through a shell.
215
+ - `--tool` records a tool version probed from `config/tools.yaml`; `--input` declares an input file/directory that is hashed for provenance (repeatable); `--threads`, `--cwd`, `--timeout`, and `--backend` (default `local`; also `slurm` or `ssh`) control execution.
213
216
  - stdout and stderr are saved to `logs/<WF_ID>.stdout.log` and `.stderr.log`.
214
217
  - Run records are written to `logs/workflow.jsonl` and `workflow_runs`.
215
218
  - The run is `completed` only when the exit code is 0 and every `--expected-output` exists and is non-empty; otherwise it is `failed` and the command exits non-zero.
@@ -255,5 +258,5 @@ operon adopt --from-manifest adopt_manifest.json
255
258
  ```
256
259
 
257
260
  - Each item requires `path`, `entity_type`, `entity_id`, `role`, and `derived_from` (at least one already-registered file_id); relative paths resolve from the project root.
258
- - Artifacts are materialized under `analysis/adopted/<entity_id>/`; same entity and role with identical bytes is reused idempotently, different bytes raise `ConflictError`; if any item is invalid (e.g. an unregistered `derived_from` or a retired entity), nothing in the batch is registered.
261
+ - Artifacts are materialized under `analysis/adopted/<entity_id>/`; same entity and role with identical bytes is reused idempotently, different bytes raise `ConflictError`. The whole batch is preflighted, then registered in one transaction. A failure before commit rolls back metadata, lineage, state and workflow rows and removes newly created artifacts; existing files are preserved. Resolve conflicting occupied targets explicitly before retrying. Completed JSONL records are written only after commit.
259
262
  - Roles are freely named by the workflow; lineage edges are written to the `file_lineage` table and can be audited with `operon query`.
@@ -25,7 +25,7 @@ Import behavior:
25
25
  - Schema, controlled-vocabulary, and foreign-key validation run before the row-by-row `insert/update/unchanged` preview.
26
26
  - `--on-conflict error` rejects existing rows, `skip` skips them, and `update` updates fields with per-field audit records.
27
27
  - Deletes and full-snapshot replacement are not supported. If any write fails, the transaction for the table is rolled back.
28
- - For XLSX, the first `data` worksheet is imported; the template's second `schema` worksheet is read-only documentation.
28
+ - For XLSX, the first worksheet in workbook order is imported whatever its name (templates generated by `--template` name it `data` and add a read-only `schema` documentation worksheet).
29
29
 
30
30
  CSV example:
31
31
 
@@ -12,7 +12,7 @@ operon ncbi-datasets \
12
12
 
13
13
  The output includes the number of organism/sample/assembly/annotation IDs to create in `new_ids` and the rows to upsert in `metadata_rows`. A dry run does not copy input, write the database, or create logs.
14
14
 
15
- If the project still uses an old metadata schema, a formal import preserves custom fields, adds the fields and paired-source file roles required by the adapter, and upgrades the schema to 1.4. A dry run does not modify the schema.
15
+ If the project still uses an old metadata schema, a formal import preserves custom fields, adds the fields and paired-source file roles required by the adapter, and upgrades the schema to {{ metadata_schema }}. A dry run does not modify the schema.
16
16
 
17
17
  After review, remove `--dry-run`:
18
18
 
@@ -33,7 +33,7 @@ warnings:
33
33
  code: HIGH_BUSCO_DUPLICATION
34
34
  ```
35
35
 
36
- Supported operators are `>=`, `<=`, `>`, `<`, `==`, `!=`, `between` (requires `min` and `max`), `in`, `not_in` (requires `values`), and `exists`.
36
+ Supported operators are `>=`, `<=`, `>`, `<`, `==`, `!=`, `between` (requires `min` and `max`), `in`, `not_in` (requires `values`), and `exists`. Comparisons are inclusive at the boundaries, and `in`/`not_in` compare metric values as strings. Note two profile-validation gaps: a `between` rule missing `min`/`max`, or an `in`/`not_in` rule missing `values`, is not rejected when the profile loads — the former surfaces as a Python traceback at evaluation time, and the latter silently evaluates against an empty value set.
37
37
 
38
38
  Hand-editing the YAML file is fully supported. The audited alternative is the
39
39
  TUI Config screen (`operon tui`, key `6`, QC Profiles tab): a structured form
@@ -18,6 +18,7 @@ remotes:
18
18
  # Alternatively pin an administrator-provided fingerprint:
19
19
  # host_key_sha256: SHA256:base64...
20
20
  insecure_accept_unknown_host: false
21
+ connect_timeout: 30 # seconds; also bounds the remote manifest lock wait
21
22
  ```
22
23
 
23
24
  Paramiko is included in the standard `OperonDBS` installation.
@@ -54,7 +55,7 @@ The remote model preserves the raw-file invariants:
54
55
  - Remote relative paths must remain safely under the remote root. By default, `pull` checks every record against local SQLite `file_id + relative_path + sha256 + size_bytes`; the remote manifest cannot rewrite local identity.
55
56
  - Every transfer writes workflow provenance (`push:<name>` or `pull:<name>`), and successful locations are recorded in `file_locations`.
56
57
  - A failed item does not stop the rest of a push/pull/evict batch. Every item receives a result, and the command exits with code 1 if any item has `error`.
57
- - After `pull` restores a missing local file, `files.status` returns to `CHECKSUM_VERIFIED` and the change is audited in `changes`.
58
+ - After `pull` restores a missing local file, `files.status` returns to `CHECKSUM_VERIFIED` and the change is audited in `changes`. A file that was already `STANDARDIZED` before eviction keeps the `STANDARDIZED` status after restore.
58
59
 
59
60
  ## Keep the control plane local and large files remote
60
61
 
@@ -95,6 +96,8 @@ execution:
95
96
 
96
97
  `evict` explicitly deletes local bytes; without `--file-id`, it processes every manifest object. It first validates local identity, remote manifest identity, and actual remote SHA-256/tree hash. The state change is written to `changes`. `standardize` and `release` require local bytes, so run `pull` first. External `analyze` can consume `REMOTE_ONLY` input directly.
97
98
 
99
+ Eviction writes a small placeholder pointer file under `.operon/placeholders/<file_id>.json` (deleted again when `pull` restores the bytes). The first remote-only status also extends `config/schemas.yaml` with the `REMOTE_ONLY` file status and bumps its `schema_version` to 1.2 — the file is rewritten with normalized formatting, so hand-written comments in it are dropped.
100
+
98
101
  When a local object is missing, `verify` checks the remote in real time rather than treating `file_locations.status=AVAILABLE` as permanent proof. A deleted or damaged remote object returns `MISSING` and updates the cache. An unreachable SSH host returns `REMOTE_UNVERIFIED` and exit code 1 while preserving the last persistent state, so a network failure is not misclassified as data loss.
99
102
 
100
103
  Remote files can also be archived directly from URLs:
@@ -18,3 +18,7 @@ Recommended actions:
18
18
  | `CHECKSUM_FAILED` | Stop QC. Determine whether the file was modified and restore it from the original source. |
19
19
  | `QC_FAILED` | Inspect files with `parseable=0` in `operon report qc`, then use `operon workflow list --step qc --status failed` and `operon workflow show WF_ID` for the recorded error and execution details. |
20
20
  | Format parsing failure | Check with an external validator such as `seqkit stats` or a GFF3 validator. Archive the repaired file as a new version; do not overwrite raw data. |
21
+
22
+ ## Further reading
23
+
24
+ For implicit semantics, edge cases, and known issues that are not defects in a single workflow — such as how missing metrics are treated, multi-file QC state semantics, or exit-code conventions — see [Implicit Behaviors, Edge Cases, and Known Issues](../reference/behaviors-and-limitations.md).
@@ -2,7 +2,7 @@
2
2
 
3
3
  Operon is a file-backed database for large-scale genomic data. It supports metadata management, immutable file archiving, quality control (QC), rule-based decisions, external analysis, remote storage and execution, and versioned dataset releases.
4
4
 
5
- This documentation matches `operon` 0.6.2, database schema 2.9, and metadata schema 1.4. The Chinese and English documentation use the same directory structure.
5
+ This documentation matches `operon` {{ operon_version }}, database schema {{ db_schema }}, and metadata schema {{ metadata_schema }}. The Chinese and English documentation use the same directory structure.
6
6
 
7
7
  ## Reading paths
8
8
 
@@ -19,7 +19,7 @@ File: `operon/database.py`
19
19
 
20
20
  `Database._migrate_remote_schema_2_2()` is also not part of the "development-era v1 compatibility layer" above. It upgrades a 2.1 database to 2.2 purely additively: adding `executor`, `scheduler_job_id`, and `execution_details` to `workflow_runs`, and creating `file_locations`. It must be kept as long as opening 2.1 projects is supported; if that support ever ends, it should be replaced through the formal database migration policy, not deleted together with `_migrate_pre_1_0_schema()`. The corresponding test is `test_schema_2_2_adds_remote_location_and_executor_provenance`.
21
21
 
22
- `Database._migrate_taxonomy_schema_2_3()` is likewise a purely additive migration needed by current functionality, not part of `_migrate_pre_1_0_schema()`: for 2.2 projects it creates `taxonomy_snapshots`, `taxonomy_nodes`, `taxonomy_aliases`, `taxonomy_reference_sets`, `coverage_reports`, and `coverage_report_metrics` plus related indexes, without modifying existing business rows. It must be kept as long as opening 2.2 projects is supported. The corresponding regression test is `test_schema_2_3_adds_taxonomy_coverage_history`.
22
+ `Database._migrate_taxonomy_schema_2_3()` is likewise a purely additive migration needed by current functionality, not part of `_migrate_pre_1_0_schema()`: for 2.2 projects it creates `taxonomy_snapshots`, `taxonomy_nodes`, `taxonomy_aliases`, `taxonomy_reference_sets`, `coverage_reports`, and `coverage_report_metrics` plus related indexes, without modifying existing business rows. It must be kept as long as opening 2.2 projects is supported. The corresponding regression test is `test_schema_2_3_adds_taxonomy_and_coverage_history`.
23
23
 
24
24
  `Database._migrate_source_schema_2_4()` is another purely additive migration needed by current functionality: for 2.3 projects it creates `data_sources` and `source_links`, storing normalized external databases/repositories, citations, licenses, and their associated objects. Non-INSDC sources must contain both citation and License; source content is deduplicated by SHA-256 identity. It must be kept as long as opening 2.3 projects is supported. The corresponding regression test is `test_schema_2_4_adds_normalized_source_provenance`.
25
25
 
@@ -20,6 +20,8 @@ Operon manages data admission, identity verification, provenance, rule evaluatio
20
20
 
21
21
  The current source adapter supports NCBI Datasets. Taxonomy coverage currently supports NCBI Taxonomy only; GTDB and NCBI↔GTDB crosswalks are extension work.
22
22
 
23
+ Implicit semantics, edge cases, and known issues that are not covered by the task-facing pages are catalogued in [Implicit Behaviors, Edge Cases, and Known Issues](reference/behaviors-and-limitations.md).
24
+
23
25
  ## Recommended workflow
24
26
 
25
27
  ```text