docspan 0.3.0__tar.gz → 0.5.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {docspan-0.3.0 → docspan-0.5.0}/.github/workflows/publish.yml +6 -7
- docspan-0.5.0/.github/workflows/release-please.yml +39 -0
- docspan-0.5.0/.release-please-manifest.json +3 -0
- {docspan-0.3.0 → docspan-0.5.0}/CHANGELOG.md +44 -1
- {docspan-0.3.0 → docspan-0.5.0}/PKG-INFO +6 -3
- {docspan-0.3.0 → docspan-0.5.0}/README.md +4 -2
- {docspan-0.3.0 → docspan-0.5.0}/docs/backends/confluence.md +1 -0
- docspan-0.5.0/docs/backends/google-docs.md +67 -0
- {docspan-0.3.0 → docspan-0.5.0}/docs/index.md +3 -1
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/decisions/ADR-001-manifest-yaml-sidecar-keyed-by-heading-id.md +24 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/decisions/ADR-002-reorder-as-in-place-move.md +25 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/decisions/ADR-003-sectioned-pull-always-structural-path.md +19 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/implementation/adversarial-review.md +28 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/implementation/architecture-review.md +31 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/implementation/plan.md +296 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/implementation/pre-mortem.md +16 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/implementation/validation.md +62 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/requirements.md +80 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/research/architecture.md +216 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/research/build-vs-buy.md +64 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/research/features.md +257 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/research/pitfalls.md +278 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/research/stack.md +54 -0
- docspan-0.5.0/project_plans/gdocs-sectioned-sync/research/ux.md +146 -0
- docspan-0.5.0/project_plans/wedding-planning-workflow/decisions/ADR-003-no-comment-anchor-migration.md +128 -0
- {docspan-0.3.0 → docspan-0.5.0}/pyproject.toml +1 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/base.py +46 -2
- docspan-0.5.0/src/docspan/backends/confluence/anchors.py +111 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/backend.py +32 -2
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/config/models.py +5 -1
- docspan-0.5.0/src/docspan/backends/google_docs/backend.py +1744 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/client.py +72 -1
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/comments.py +83 -2
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/converter.py +6 -3
- docspan-0.5.0/src/docspan/backends/google_docs/cross_doc_links.py +286 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/docs_request_builder.py +857 -265
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/docs_structure_parser.py +95 -3
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/heading_anchors.py +26 -11
- docspan-0.5.0/src/docspan/backends/google_docs/image_source.py +260 -0
- docspan-0.5.0/src/docspan/backends/google_docs/manifest.py +193 -0
- docspan-0.5.0/src/docspan/backends/google_docs/markdown_to_paragraph_parser.py +603 -0
- docspan-0.5.0/src/docspan/backends/google_docs/mermaid_renderer.py +100 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/nodes_to_markdown.py +236 -65
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/projection.py +2 -1
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/push_preview.py +138 -28
- docspan-0.5.0/src/docspan/backends/google_docs/registry.py +67 -0
- docspan-0.5.0/src/docspan/backends/google_docs/section_splitter.py +194 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/tabs.py +46 -1
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/cli/main.py +203 -6
- docspan-0.5.0/src/docspan/config.py +281 -0
- docspan-0.5.0/src/docspan/core/atomic_dir.py +82 -0
- docspan-0.5.0/src/docspan/core/orchestrator.py +867 -0
- docspan-0.5.0/tests/__init__.py +0 -0
- docspan-0.5.0/tests/test_atomic_dir.py +108 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_cli.py +296 -3
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_code_block_granularity.py +255 -22
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_config.py +164 -2
- docspan-0.5.0/tests/test_confluence_anchors.py +175 -0
- docspan-0.5.0/tests/test_confluence_backend.py +47 -0
- docspan-0.5.0/tests/test_confluence_mermaid_push_pipeline.py +101 -0
- docspan-0.5.0/tests/test_confluence_push_dead_anchors.py +90 -0
- docspan-0.5.0/tests/test_content_key_pooling_performance.py +115 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_converter.py +18 -0
- docspan-0.5.0/tests/test_cross_doc_link_issues.py +135 -0
- docspan-0.5.0/tests/test_cross_doc_links.py +310 -0
- docspan-0.5.0/tests/test_cross_doc_links_backend.py +481 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_docs_request_builder.py +588 -5
- docspan-0.5.0/tests/test_gdocs_images.py +427 -0
- docspan-0.5.0/tests/test_gdocs_mermaid.py +119 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_gdocs_push_pipeline.py +260 -1
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_gdocs_tables_and_styles.py +98 -3
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_google_comments.py +107 -1
- docspan-0.5.0/tests/test_google_docs_backend.py +2661 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_google_onboarding.py +8 -2
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_heading_anchors.py +166 -0
- docspan-0.5.0/tests/test_heading_identity.py +982 -0
- docspan-0.5.0/tests/test_manifest.py +115 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_markdown_to_paragraph_parser.py +8 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_nodes_to_markdown.py +112 -0
- docspan-0.5.0/tests/test_orchestrator.py +944 -0
- docspan-0.5.0/tests/test_push_preview.py +689 -0
- docspan-0.5.0/tests/test_registry.py +93 -0
- docspan-0.5.0/tests/test_restyle_destruction_rate.py +149 -0
- docspan-0.5.0/tests/test_section_splitter.py +185 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_state.py +27 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_table_cell_spans.py +153 -22
- docspan-0.5.0/tests/test_tabs.py +373 -0
- {docspan-0.3.0 → docspan-0.5.0}/uv.lock +85 -74
- docspan-0.3.0/.github/workflows/release-please.yml +0 -18
- docspan-0.3.0/.release-please-manifest.json +0 -3
- docspan-0.3.0/docs/backends/google-docs.md +0 -67
- docspan-0.3.0/src/docspan/backends/google_docs/backend.py +0 -939
- docspan-0.3.0/src/docspan/backends/google_docs/markdown_to_paragraph_parser.py +0 -357
- docspan-0.3.0/src/docspan/config.py +0 -133
- docspan-0.3.0/src/docspan/core/orchestrator.py +0 -338
- docspan-0.3.0/tests/test_google_docs_backend.py +0 -1147
- docspan-0.3.0/tests/test_heading_identity.py +0 -456
- docspan-0.3.0/tests/test_orchestrator.py +0 -333
- docspan-0.3.0/tests/test_push_preview.py +0 -338
- docspan-0.3.0/tests/test_tabs.py +0 -120
- {docspan-0.3.0 → docspan-0.5.0}/.github/workflows/ci.yml +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/.gitignore +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/.mypy-error-baseline +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/CONTRIBUTING.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/Procfile +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/RAILWAY_SETUP.md +0 -0
- /docspan-0.3.0/src/docspan/__main__.py → /docspan-0.5.0/doc.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/docs/commands.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/docs/configuration.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/docs/contributing.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/docs/install.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/docspan.yaml.example +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/markgate.yaml.example +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/mkdocs.yml +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/modules/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/modules/auth.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/modules/conflict_handler.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/modules/converter.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/modules/gdrive_client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/modules/sync_engine.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/bidirectional-comments/plan.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/implementation/adversarial-review.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/implementation/plan.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/implementation/release-checklist.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/implementation/validation.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/requirements.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/research/architecture.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/research/features.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/research/google-docs-push.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/research/pitfalls.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/docspan-release/research/stack.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/gdocs-tables-inline-styles/plan.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/decisions/ADR-001-merge3-dependency.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/decisions/ADR-002-base-content-sidecar-store.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/implementation/adversarial-review.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/implementation/plan.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/implementation/validation.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/requirements.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/research/architecture.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/research/features.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/research/pitfalls.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/markgate-sync/research/stack.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/decisions/ADR-001-checklist-state-as-literal-text.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/decisions/ADR-002-comment-risk-flagging-not-anchor-preservation.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/feature-gap-report.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/implementation/adversarial-review.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/implementation/architecture-review.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/implementation/plan.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/implementation/pre-mortem.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/implementation/validation.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/requirements.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/research/architecture.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/research/build-vs-buy.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/research/features.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/research/pitfalls.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/research/stack.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/research/ux.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/project_plans/wedding-planning-workflow/workflow-runbook.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/release-please-config.json +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/requirements.txt +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/runtime.txt +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/__init__.py +0 -0
- /docspan-0.3.0/src/docspan/backends/confluence/__init__.py → /docspan-0.5.0/src/docspan/__main__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/__init__.py +0 -0
- {docspan-0.3.0/src/docspan/backends/confluence/services → docspan-0.5.0/src/docspan/backends/confluence}/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/comparator.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/converter.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/converters.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/interfaces.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/nodes.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/parser.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/validators.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/adf/visitors.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/config/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/config/loader.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/config/validation.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/ast.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/extensions/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/extensions/frontmatter.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/extensions/mermaid.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/extensions/wikilinks.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/inline_parser.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/markdown/parser.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/models/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/models/markdown_file.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/models/page.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/models/path_utils.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/models/results.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/models/sync_status.py +0 -0
- {docspan-0.3.0/src/docspan/backends/google_docs → docspan-0.5.0/src/docspan/backends/confluence/services}/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/attachment_client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/base_client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/comment_client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/crawler.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/label_client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/page_client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/space_client.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/confluence/services/confluence/url_parser.py +0 -0
- {docspan-0.3.0/src/docspan/cli → docspan-0.5.0/src/docspan/backends/google_docs}/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/auth.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/checkbox_state.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/backends/google_docs/onboarding.py +0 -0
- {docspan-0.3.0/tests → docspan-0.5.0/src/docspan/cli}/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/core/__init__.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/core/merge.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/core/paths.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/core/state.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/src/docspan/core/xdg.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/sync.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/terraform/gcp/README.md +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/terraform/gcp/main.tf +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/terraform/gcp/outputs.tf +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/terraform/gcp/variables.tf +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/terraform/main.tf +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/terraform/variables.tf +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/conftest.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/fixtures/github_slugger_vectors.json +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_checkbox_state.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_conflict_resolution.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_docs_structure_parser.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_google_oauth.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_merge.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_span_trailing_newline.py +0 -0
- {docspan-0.3.0 → docspan-0.5.0}/tests/test_xdg_central_config.py +0 -0
|
@@ -3,6 +3,11 @@ name: Publish to PyPI
|
|
|
3
3
|
on:
|
|
4
4
|
release:
|
|
5
5
|
types: [published]
|
|
6
|
+
workflow_dispatch:
|
|
7
|
+
inputs:
|
|
8
|
+
ref:
|
|
9
|
+
description: "Git tag to build and publish (e.g. docspan-v0.4.0)"
|
|
10
|
+
required: true
|
|
6
11
|
|
|
7
12
|
jobs:
|
|
8
13
|
build:
|
|
@@ -10,19 +15,16 @@ jobs:
|
|
|
10
15
|
steps:
|
|
11
16
|
- uses: actions/checkout@v4
|
|
12
17
|
with:
|
|
18
|
+
ref: ${{ inputs.ref || github.ref }}
|
|
13
19
|
fetch-depth: 0 # needed for hatch-vcs version from git tags
|
|
14
|
-
|
|
15
20
|
- name: Install uv
|
|
16
21
|
uses: astral-sh/setup-uv@v4
|
|
17
|
-
|
|
18
22
|
- name: Build package
|
|
19
23
|
run: uv build
|
|
20
|
-
|
|
21
24
|
- uses: actions/upload-artifact@v4.6.2
|
|
22
25
|
with:
|
|
23
26
|
name: dist
|
|
24
27
|
path: dist/
|
|
25
|
-
|
|
26
28
|
publish-testpypi:
|
|
27
29
|
needs: build
|
|
28
30
|
runs-on: ubuntu-latest
|
|
@@ -38,11 +40,9 @@ jobs:
|
|
|
38
40
|
with:
|
|
39
41
|
name: dist
|
|
40
42
|
path: dist/
|
|
41
|
-
|
|
42
43
|
- uses: pypa/gh-action-pypi-publish@release/v1
|
|
43
44
|
with:
|
|
44
45
|
repository-url: https://test.pypi.org/legacy/
|
|
45
|
-
|
|
46
46
|
publish-pypi:
|
|
47
47
|
needs: publish-testpypi
|
|
48
48
|
runs-on: ubuntu-latest
|
|
@@ -58,5 +58,4 @@ jobs:
|
|
|
58
58
|
with:
|
|
59
59
|
name: dist
|
|
60
60
|
path: dist/
|
|
61
|
-
|
|
62
61
|
- uses: pypa/gh-action-pypi-publish@release/v1
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
name: Release Please
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
branches: ["main"]
|
|
6
|
+
|
|
7
|
+
permissions:
|
|
8
|
+
contents: write
|
|
9
|
+
pull-requests: write
|
|
10
|
+
|
|
11
|
+
jobs:
|
|
12
|
+
release-please:
|
|
13
|
+
runs-on: ubuntu-latest
|
|
14
|
+
permissions:
|
|
15
|
+
contents: write
|
|
16
|
+
pull-requests: write
|
|
17
|
+
actions: write # to dispatch publish.yml below
|
|
18
|
+
steps:
|
|
19
|
+
- uses: googleapis/release-please-action@v4
|
|
20
|
+
id: release
|
|
21
|
+
with:
|
|
22
|
+
config-file: release-please-config.json
|
|
23
|
+
manifest-file: .release-please-manifest.json
|
|
24
|
+
|
|
25
|
+
# release-please-action creates the GitHub Release using the default
|
|
26
|
+
# GITHUB_TOKEN. GitHub doesn't let events produced by that token
|
|
27
|
+
# trigger other workflows (loop-prevention), so publish.yml's
|
|
28
|
+
# `on: release: published` trigger never fires for these releases —
|
|
29
|
+
# dispatch it explicitly instead. workflow_dispatch is exempted from
|
|
30
|
+
# that restriction even when invoked with GITHUB_TOKEN.
|
|
31
|
+
- name: Trigger PyPI publish
|
|
32
|
+
if: ${{ steps.release.outputs.release_created == 'true' }}
|
|
33
|
+
env:
|
|
34
|
+
GH_TOKEN: ${{ github.token }}
|
|
35
|
+
run: |
|
|
36
|
+
gh workflow run publish.yml \
|
|
37
|
+
--repo "${{ github.repository }}" \
|
|
38
|
+
--ref main \
|
|
39
|
+
-f ref="${{ steps.release.outputs.tag_name }}"
|
|
@@ -5,6 +5,47 @@ All notable changes to this project will be documented in this file.
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
|
|
6
6
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
7
|
|
|
8
|
+
## [0.5.0](https://github.com/tstapler/docspan/compare/docspan-v0.4.0...docspan-v0.5.0) (2026-08-14)
|
|
9
|
+
|
|
10
|
+
|
|
11
|
+
### Features
|
|
12
|
+
|
|
13
|
+
* **google-docs:** sectioned sync for large document mappings ([#106](https://github.com/tstapler/docspan/issues/106)) ([fe12122](https://github.com/tstapler/docspan/commit/fe121229c8b6b957254020fd2c06241f7506aa80))
|
|
14
|
+
|
|
15
|
+
|
|
16
|
+
### Bug Fixes
|
|
17
|
+
|
|
18
|
+
* **ci:** dispatch PyPI publish from release-please via workflow_dispatch ([e76ffb8](https://github.com/tstapler/docspan/commit/e76ffb8547bfd7d2d2b453a64a5fb3899445fc5f))
|
|
19
|
+
* **confluence:** report internal anchors instead of writing a link to nowhere ([#105](https://github.com/tstapler/docspan/issues/105)) ([cd38f24](https://github.com/tstapler/docspan/commit/cd38f2415e37814e071a180182454d761cfcbfa9))
|
|
20
|
+
* **google-docs:** reset table cell paragraph style to NORMAL_TEXT on fill ([e6db797](https://github.com/tstapler/docspan/commit/e6db797d51746c738586c3cab171367e23bd5be0))
|
|
21
|
+
|
|
22
|
+
## [0.4.0](https://github.com/tstapler/docspan/compare/docspan-v0.3.0...docspan-v0.4.0) (2026-08-13)
|
|
23
|
+
|
|
24
|
+
|
|
25
|
+
### Features
|
|
26
|
+
|
|
27
|
+
* **docspan:** add docspan map command and push auto-create for unmapped files ([#80](https://github.com/tstapler/docspan/issues/80)) ([aa4258b](https://github.com/tstapler/docspan/commit/aa4258b3ff381c3e6afc389ad55c426abb5850ba))
|
|
28
|
+
* **google-docs:** add inline image push/pull support ([#101](https://github.com/tstapler/docspan/issues/101)) ([e5256af](https://github.com/tstapler/docspan/commit/e5256afa18076051b891f44511786e4f2750a489))
|
|
29
|
+
* **google-docs:** render mermaid diagrams as inline PNGs on push ([9298d1b](https://github.com/tstapler/docspan/commit/9298d1b0c645231056571a2c5afaa296e14d751b))
|
|
30
|
+
|
|
31
|
+
|
|
32
|
+
### Bug Fixes
|
|
33
|
+
|
|
34
|
+
* **config:** round-trip YAML comments on save_config writes ([780ddcc](https://github.com/tstapler/docspan/commit/780ddcc18b760f184627eb9c9a2b870249c6ed51))
|
|
35
|
+
* **confluence:** confirm mermaid render pipeline is a stub, not a rasterizer ([#94](https://github.com/tstapler/docspan/issues/94)) ([fb5ca80](https://github.com/tstapler/docspan/commit/fb5ca8002e91e050447766d7e498ff30eadae9dc))
|
|
36
|
+
* **google-docs:** bound difflib's cubic-ish blowup on duplicate-heavy documents ([#84](https://github.com/tstapler/docspan/issues/84)) ([834df71](https://github.com/tstapler/docspan/commit/834df71d2737b09ee68b14dd1f509bc3928249da))
|
|
37
|
+
* **google-docs:** distinguish delete-and-reinsert churn from real removal in push preview ([#86](https://github.com/tstapler/docspan/issues/86)) ([052b64d](https://github.com/tstapler/docspan/commit/052b64d591bc88ab6962e5173c2494b2fedab1c3))
|
|
38
|
+
* **google-docs:** escape backticks in monospace spans on both pull paths ([#103](https://github.com/tstapler/docspan/issues/103)) ([74d007d](https://github.com/tstapler/docspan/commit/74d007d162d511641650d2afb8b9f8b482c3e48e))
|
|
39
|
+
* **google-docs:** land PR [#70](https://github.com/tstapler/docspan/issues/70) restyle repair, verify AC0-8 (issue [#52](https://github.com/tstapler/docspan/issues/52)) ([#99](https://github.com/tstapler/docspan/issues/99)) ([8203c93](https://github.com/tstapler/docspan/commit/8203c93567a3d962bb3bb5754338298260f610aa))
|
|
40
|
+
* **google-docs:** order same-anchor insert groups after restyle/delete groups ([#83](https://github.com/tstapler/docspan/issues/83)) ([b8eace0](https://github.com/tstapler/docspan/commit/b8eace0505f7196eeb5bfe659f2d284a4b8c07ad))
|
|
41
|
+
* **google-docs:** render multi-paragraph table cells as HTML tables ([#79](https://github.com/tstapler/docspan/issues/79)) ([45e072d](https://github.com/tstapler/docspan/commit/45e072de2ae845443d6b5862a6c70234a55fb501))
|
|
42
|
+
* **google-docs:** resolve cross-tab heading anchors on push ([#102](https://github.com/tstapler/docspan/issues/102)) ([c1be541](https://github.com/tstapler/docspan/commit/c1be541d7e8d227369367ac8d99700405ac31195))
|
|
43
|
+
* **google-docs:** resolve relative cross-document markdown links to target Google Doc URLs ([#98](https://github.com/tstapler/docspan/issues/98)) ([27e7f15](https://github.com/tstapler/docspan/commit/27e7f1551549d7bff9d64b5e43abfd09f4f19e36))
|
|
44
|
+
* **google-docs:** richer at-risk-comment warning, no anchor migration ([#92](https://github.com/tstapler/docspan/issues/92)) ([#95](https://github.com/tstapler/docspan/issues/95)) ([be854db](https://github.com/tstapler/docspan/commit/be854db8cc4dd5a8252685847455e51d444decee))
|
|
45
|
+
* **google-docs:** split/preserve fenced code blocks in list items and blockquotes ([#87](https://github.com/tstapler/docspan/issues/87)) ([0a01f9f](https://github.com/tstapler/docspan/commit/0a01f9f6c927489c9234049aed56b4fe8c895b58))
|
|
46
|
+
* **google-docs:** stop force-push from corrupting tab-scoped checkbox docs ([#97](https://github.com/tstapler/docspan/issues/97)) ([63ab43b](https://github.com/tstapler/docspan/commit/63ab43bc755dad67ce22d516b6ea4f1eaeba9c3e))
|
|
47
|
+
* **google-docs:** stop replace branch from duplicating the doc-end-clamped newline ([#85](https://github.com/tstapler/docspan/issues/85)) ([68f0de9](https://github.com/tstapler/docspan/commit/68f0de985be5b13df0d4e1b3c4b0177a2e0eced8))
|
|
48
|
+
|
|
8
49
|
## [0.3.0](https://github.com/tstapler/docspan/compare/docspan-v0.2.0...docspan-v0.3.0) (2026-08-11)
|
|
9
50
|
|
|
10
51
|
|
|
@@ -104,7 +145,9 @@ Each of these is tracked as a follow-up rather than half-addressed here.
|
|
|
104
145
|
- A pull cannot express a `bookmark`/`bookmarkId` link, a link to a tab, or any link inside
|
|
105
146
|
a table cell, so those are dropped from the pulled file without a report.
|
|
106
147
|
- Confluence writes an internal anchor as a literal `#fragment` href, which it does not
|
|
107
|
-
resolve.
|
|
148
|
+
resolve. `push` now reports this as a warning naming the anchor(s) instead of shipping it
|
|
149
|
+
silently; the href itself is unchanged, since no live instance was available to establish
|
|
150
|
+
what Confluence actually generates for a heading.
|
|
108
151
|
- An anchor that resolves to nothing is written as plain text, so a later pull replaces the
|
|
109
152
|
author's `[text](#anchor)` with `text`. The push reports it; nothing does afterwards.
|
|
110
153
|
- Such a push exits non-zero on every run, with no flag to suppress it.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: docspan
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.5.0
|
|
4
4
|
Summary: Push and pull markdown to Google Docs and Confluence from a single CLI
|
|
5
5
|
Project-URL: Homepage, https://github.com/tstapler/docspan
|
|
6
6
|
Project-URL: Repository, https://github.com/tstapler/docspan
|
|
@@ -36,6 +36,7 @@ Requires-Dist: python-dateutil>=2.8.2
|
|
|
36
36
|
Requires-Dist: pyyaml>=6.0
|
|
37
37
|
Requires-Dist: requests>=2.25.0
|
|
38
38
|
Requires-Dist: rich>=13.0.0
|
|
39
|
+
Requires-Dist: ruamel-yaml>=0.18.0
|
|
39
40
|
Requires-Dist: typer>=0.9.0
|
|
40
41
|
Provides-Extra: dev
|
|
41
42
|
Requires-Dist: mypy>=1.0.0; extra == 'dev'
|
|
@@ -306,8 +307,10 @@ docspan generates these files in your project directory after first sync:
|
|
|
306
307
|
> [!NOTE]
|
|
307
308
|
> **Known limitations in v0.1.0**
|
|
308
309
|
>
|
|
309
|
-
> - Google Docs: comments on edited paragraphs are lost on push (paragraph-level structural diff; comments on unchanged paragraphs are preserved). `docspan push --dry-run` and a default fail-closed `--force`-gated block now warn before this happens — it is still not prevented.
|
|
310
|
-
> -
|
|
310
|
+
> - Google Docs: comments on edited paragraphs are lost on push (paragraph-level structural diff; comments on unchanged paragraphs are preserved). `docspan push --dry-run` and a default fail-closed `--force`-gated block now warn before this happens, naming every at-risk comment (id, author, quoted snippet) per flagged paragraph — it is still not prevented. Re-anchoring or recreating the comment was investigated and rejected: neither `comments().update` nor `comments().create` can produce a comment Google Docs' own UI renders as anchored to arbitrary text (see [ADR-003](project_plans/wedding-planning-workflow/decisions/ADR-003-no-comment-anchor-migration.md)), so no migration ships.
|
|
311
|
+
> - Google Docs: images push and pull (`` uploads to Drive; `https://` URLs are referenced directly). SVG, missing, and oversized (>50MB) images are reported as warnings rather than blocking the push. An image mixed into a paragraph alongside running text is not supported — only a standalone `` on its own line
|
|
312
|
+
> - Google Docs: a pulled image's markdown link is Google's `contentUri` for that embedded object, which Google's API docs say may change over time even when the image itself hasn't. docspan's push diff keys image identity on `alt`/size rather than this URI, so a rotated URI alone will not cause a paragraph to be needlessly deleted and reinserted (and its comments lost) on the next push — but the URI written into your markdown file can itself go stale between pulls, and a stale-but-unchanged URI line can still show up as a one-sided edit in `docspan conflicts resolve`'s three-way diff even though nothing meaningful changed
|
|
313
|
+
> - Push: no image support for Confluence — local images cannot be pushed
|
|
311
314
|
> - Push: no table support — markdown tables are not rendered in Google Docs
|
|
312
315
|
> - Confluence: requires an Atlassian API token; no OAuth flow
|
|
313
316
|
> - Confluence: the comment sidecar (`{file}.comments.md`) is informational only; comments cannot be pushed back
|
|
@@ -255,8 +255,10 @@ docspan generates these files in your project directory after first sync:
|
|
|
255
255
|
> [!NOTE]
|
|
256
256
|
> **Known limitations in v0.1.0**
|
|
257
257
|
>
|
|
258
|
-
> - Google Docs: comments on edited paragraphs are lost on push (paragraph-level structural diff; comments on unchanged paragraphs are preserved). `docspan push --dry-run` and a default fail-closed `--force`-gated block now warn before this happens — it is still not prevented.
|
|
259
|
-
> -
|
|
258
|
+
> - Google Docs: comments on edited paragraphs are lost on push (paragraph-level structural diff; comments on unchanged paragraphs are preserved). `docspan push --dry-run` and a default fail-closed `--force`-gated block now warn before this happens, naming every at-risk comment (id, author, quoted snippet) per flagged paragraph — it is still not prevented. Re-anchoring or recreating the comment was investigated and rejected: neither `comments().update` nor `comments().create` can produce a comment Google Docs' own UI renders as anchored to arbitrary text (see [ADR-003](project_plans/wedding-planning-workflow/decisions/ADR-003-no-comment-anchor-migration.md)), so no migration ships.
|
|
259
|
+
> - Google Docs: images push and pull (`` uploads to Drive; `https://` URLs are referenced directly). SVG, missing, and oversized (>50MB) images are reported as warnings rather than blocking the push. An image mixed into a paragraph alongside running text is not supported — only a standalone `` on its own line
|
|
260
|
+
> - Google Docs: a pulled image's markdown link is Google's `contentUri` for that embedded object, which Google's API docs say may change over time even when the image itself hasn't. docspan's push diff keys image identity on `alt`/size rather than this URI, so a rotated URI alone will not cause a paragraph to be needlessly deleted and reinserted (and its comments lost) on the next push — but the URI written into your markdown file can itself go stale between pulls, and a stale-but-unchanged URI line can still show up as a one-sided edit in `docspan conflicts resolve`'s three-way diff even though nothing meaningful changed
|
|
261
|
+
> - Push: no image support for Confluence — local images cannot be pushed
|
|
260
262
|
> - Push: no table support — markdown tables are not rendered in Google Docs
|
|
261
263
|
> - Confluence: requires an Atlassian API token; no OAuth flow
|
|
262
264
|
> - Confluence: the comment sidecar (`{file}.comments.md`) is informational only; comments cannot be pushed back
|
|
@@ -65,3 +65,4 @@ export CONFLUENCE_API_TOKEN=your-token
|
|
|
65
65
|
- **Comment sidecar is informational only**: Comments pulled from Confluence are written to `{file}.comments.md` but cannot be pushed back via docspan.
|
|
66
66
|
- **Push replaces page content**: The full page body is replaced on every push. Inline comment positions in Confluence may shift after a push.
|
|
67
67
|
- **Complex macros not preserved faithfully**: Confluence macros (status, panels, expand, etc.) are converted to approximate markdown equivalents on pull and may not round-trip cleanly on push.
|
|
68
|
+
- **Mermaid diagrams are not rendered**: a ` ```mermaid ` fence is pushed as a plain ADF code block (language `mermaid`), i.e. the raw diagram source as visible text — not an image, and not Confluence's native mermaid macro. `render_mermaid_diagrams` in `PublishConfig` currently has no effect on this: the code path that would rasterize a diagram (`MermaidParser.render_diagram()` in `markdown/extensions/mermaid.py`) is a stub that never runs — it's dead code, unreachable from the real parse pipeline (`markdown/parser.py` builds `MermaidNode` directly). Fixing this requires a real renderer (e.g. `mermaid-cli` or a hosted rendering service) wired to upload the result as a Confluence attachment before ADF conversion; tracked as follow-up work, not yet implemented.
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# Google Docs Backend
|
|
2
|
+
|
|
3
|
+
## How it works
|
|
4
|
+
|
|
5
|
+
The Google Docs backend authenticates either via a Google service account JSON key or via per-user OAuth (an `InstalledAppFlow` that acts as you, similar to `gws`) — whichever `markgate.yaml` configures (`credentials_path` for the service account, `oauth_client_secret_path` for OAuth). Push uses a paragraph-level structural diff that computes the minimal set of `batchUpdate` requests needed to transform the current document into the target content. This approach preserves comments attached to paragraphs that have not changed. Pull exports the Google Doc as HTML and converts it to markdown.
|
|
6
|
+
|
|
7
|
+
## Auth Setup
|
|
8
|
+
|
|
9
|
+
Run `docspan auth setup google_docs` to see setup instructions.
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
Google Docs Auth Setup
|
|
13
|
+
========================================
|
|
14
|
+
Run this in an interactive terminal for a guided setup, or configure manually:
|
|
15
|
+
|
|
16
|
+
Per-user OAuth (recommended — acts as you, like gws):
|
|
17
|
+
1. Create an OAuth client (Desktop app); download client_secret.json
|
|
18
|
+
2. docspan auth setup google_docs --oauth --client-secret /path/to/client_secret.json
|
|
19
|
+
(or set backends.google_docs.oauth_client_secret_path in markgate.yaml)
|
|
20
|
+
|
|
21
|
+
Service account (automation):
|
|
22
|
+
1. Create a service account + JSON key; enable the Docs & Drive APIs
|
|
23
|
+
2. Share your docs with the service-account email
|
|
24
|
+
3. Set credentials_path in markgate.yaml (or ACCOUNT_A_CREDENTIALS_PATH env)
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Service account credentials can also be provided inline via `ACCOUNT_A_CREDENTIALS` (the JSON itself, not a path) instead of `ACCOUNT_A_CREDENTIALS_PATH`.
|
|
28
|
+
|
|
29
|
+
## Required Scopes
|
|
30
|
+
|
|
31
|
+
Every credential path (`GoogleAuthenticator`, `OAuthAuthenticator`) requests the same read/write scopes (`PUSH_SCOPES`, aliased as `SCOPES`/`DEFAULT_SCOPES`), whether the operation is push or pull:
|
|
32
|
+
|
|
33
|
+
- `https://www.googleapis.com/auth/documents` — read and write Google Docs
|
|
34
|
+
- `https://www.googleapis.com/auth/drive` — read/write Drive (comment reads/writes, file metadata; not just export)
|
|
35
|
+
- `https://www.googleapis.com/auth/spreadsheets.readonly` — read Sheets embedded/linked in a Doc
|
|
36
|
+
|
|
37
|
+
`auth.py` also defines a narrower read-only `PULL_SCOPES`, but nothing in the codebase wires it up today — pull requests the same full grant as push, not a readonly subset. Comment reads/writes reuse this same grant too; no separate scope is added for them.
|
|
38
|
+
|
|
39
|
+
## `markgate.yaml` Example
|
|
40
|
+
|
|
41
|
+
```yaml
|
|
42
|
+
backends:
|
|
43
|
+
google_docs:
|
|
44
|
+
credentials_path: /path/to/service-account.json
|
|
45
|
+
# token_path: .markgate/google_token.json # default, rarely changed
|
|
46
|
+
|
|
47
|
+
mappings:
|
|
48
|
+
- local: docs/design-doc.md
|
|
49
|
+
backend: google_docs
|
|
50
|
+
remote_id: 1BxiMVs0XRA5nFMdKvBdBZjgmUUqptlbs74O
|
|
51
|
+
direction: both
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
## Limitations
|
|
55
|
+
|
|
56
|
+
!!! warning
|
|
57
|
+
- **Comments destroyed on push for edited paragraphs**: The structural diff preserves comments on unchanged paragraphs, but any paragraph that is deleted and reinserted loses its comments. This is a known v0.1.0 limitation.
|
|
58
|
+
- **Images push and pull**: `` uploads the local file to Drive and references it by URI; an `https://` URL is referenced directly, bypassing upload. Only a standalone image on its own line is supported — one mixed into a paragraph alongside running text is left as plain text. Missing files, files over 50MB, and unsupported formats (SVG) are reported as push warnings rather than blocking the write or crashing.
|
|
59
|
+
- **Mermaid diagrams push as rendered PNGs**: a fenced ` ```mermaid ` block is rendered to a raster PNG (via `mermaid-cli`/`mmdc`, shelled out to — install with `npm install -g @mermaid-js/mermaid-cli`, or it's fetched on demand through `npx`) and pushed as an inline image, since `insertInlineImage` has no native mermaid or SVG support. A render failure (missing Node.js/mermaid-cli, invalid diagram syntax, timeout) is reported as a push warning, not a crash. There is no pull-side reconstruction — a mermaid diagram round-trips back to markdown as a plain image reference, not a ` ```mermaid ` fence.
|
|
60
|
+
- **Pulled image URIs can go stale**: a pulled `` link is Google's `contentUri` for that embedded object, which Google's API docs say may change over time even when the image itself is unchanged. The push structural diff keys image identity on `alt`/width/height, not this URI, so a rotated `contentUri` alone will not cause the paragraph to be deleted and reinserted (which would destroy any comment anchored to it) — but the stale URI persisted in your markdown file can still surface as a one-sided edit in `docspan conflicts resolve`'s three-way diff.
|
|
61
|
+
- **Table cells hold one paragraph**: a markdown table cell is pushed as a single
|
|
62
|
+
paragraph, and inline formatting inside it (bold, monospace, links, internal
|
|
63
|
+
`#anchor` references) is applied on the second pass. Two limits follow: a cell
|
|
64
|
+
whose content spans more than one paragraph in the Doc cannot be styled, and a
|
|
65
|
+
table created by the current push gets its cell styling on the *next* push —
|
|
66
|
+
docspan reports both rather than failing silently.
|
|
67
|
+
- **Rate limiting**: The Google Docs API allows 300 requests per minute per project. Large documents with many changed paragraphs may trigger rate limit errors.
|
|
@@ -57,7 +57,9 @@ See the [Install](install.md) page for full auth setup instructions and the [Com
|
|
|
57
57
|
|
|
58
58
|
!!! warning "Known limitations in v0.1.0"
|
|
59
59
|
- Google Docs: comments on edited paragraphs are lost on push (paragraph-level structural diff; comments on unchanged paragraphs are preserved)
|
|
60
|
-
-
|
|
60
|
+
- Google Docs: images push and pull (local files upload to Drive; `https://` URLs are referenced directly)
|
|
61
|
+
- Google Docs: a pulled image's markdown link is Google's `contentUri`, which can change over time even when the image hasn't — push doesn't misdetect this as a real change, but the stale URI in your local file can still show up as a one-sided edit during conflict resolution
|
|
62
|
+
- Push: no image support for Confluence — local images cannot be pushed
|
|
61
63
|
- Push: no table support — markdown tables are not rendered in Google Docs
|
|
62
64
|
- Confluence: requires an Atlassian API token; no OAuth flow
|
|
63
65
|
- Confluence: the comment sidecar (`{file}.comments.md`) is informational only; comments cannot be pushed back
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
# ADR-001: Manifest is a `_manifest.yaml` sidecar keyed by `heading_id`, never renumbered
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
Accepted
|
|
5
|
+
|
|
6
|
+
## Context
|
|
7
|
+
A sectioned mapping needs a durable record of (a) which on-disk section files correspond to which sections of the Google Doc, and (b) their canonical order, so that renames, reorders, and add/delete can be detected reliably instead of guessed from filenames or content.
|
|
8
|
+
|
|
9
|
+
Three candidate identity/order mechanisms were evaluated (`project_plans/gdocs-sectioned-sync/research/build-vs-buy.md`):
|
|
10
|
+
|
|
11
|
+
1. **Google Docs `NamedRange`/`namedRanges`** — visible to all collaborators, but has no Docs UI (a human can't see or fix it directly), names aren't required to be unique (so lookup-by-name isn't authoritative on its own — a `namedRangeId` still has to be tracked externally, meaning it doesn't eliminate the need for a manifest), and a single named range can silently split into multiple discontiguous ranges when a collaborator edits across its boundary — which complicates both splitting and reorder detection. It would also require new `CreateNamedRangeRequest`/`DeleteNamedRangeRequest` batchUpdate request types that `docs_request_builder.py` doesn't currently emit.
|
|
12
|
+
2. **Filesystem/filename order** (bare slug or numeric prefix as sole order authority) — `ls`/`git diff --stat` order is convenient for humans but is not a stable identity: renumbering files on every insert/delete makes every unrelated file show as changed in git history (research/ux.md's "sharpest UX risk").
|
|
13
|
+
3. **A flat, git-tracked YAML sidecar (`_manifest.yaml`) keyed by Google Docs' own persistent `heading_id`**, with `slug` and `filename` as companion, non-authoritative fields.
|
|
14
|
+
|
|
15
|
+
## Decision
|
|
16
|
+
The manifest is `_manifest.yaml`, a plain YAML file (not embedded front matter, not `.md`) living alongside section files in the sectioned mapping's directory. It is the single source of truth for section order and identity, keyed by `heading_id` (`DocsParagraphNode.heading_id`, `docs_structure_parser.py:188`). Filenames (`NN-slug.md`) are a human-readable cache of that order, not the authority — on push, manifest order wins over `ls` order or file-content order. `heading_id`s are assigned once (by Google Docs, not by docspan) and are never renumbered on reorder.
|
|
17
|
+
|
|
18
|
+
Written/read atomically via temp-file-then-`os.replace`, mirroring the existing pattern in `config.py:126-167`'s `save_config`.
|
|
19
|
+
|
|
20
|
+
## Consequences
|
|
21
|
+
- Manifest and section files must be kept in sync at all times; any code path that mutates one without the other risks the desync failure mode called out in pitfalls research. Atomic directory-level writes (Epic 6 in `implementation/plan.md`) exist specifically to bound this risk.
|
|
22
|
+
- `heading_id` does not survive delete+reinsert in Google Docs (confirmed via `heading_anchors.py` and `tabs.py:38-55`), so reorder must be implemented as an in-place move, not delete+reinsert — see ADR-002. This ADR and ADR-002 are coupled: choosing `heading_id` as identity is only safe because ADR-002 commits to preserving it across reorders.
|
|
23
|
+
- Manifest is git-trackable plain text, giving full diff visibility — but that also means a manifest is one more file a human could hand-edit incorrectly; a lint/warning for manifest/directory-contents mismatch is recommended future work (noted in `research/ux.md`'s "discoverability" unstated need) but not required for v1.
|
|
24
|
+
- `NamedRange` remains available as a *future, redundant* corroborating identity mechanism if manifest drift proves to be a real-world problem — explicitly not needed for v1, so this decision doesn't foreclose it.
|
docspan-0.5.0/project_plans/gdocs-sectioned-sync/decisions/ADR-002-reorder-as-in-place-move.md
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# ADR-002: Section reorder is implemented as an in-place move, not delete+reinsert
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
Accepted — Task 3.2.2's spike concluded (see Consequences): the Docs API batchUpdate surface has no move-equivalent primitive, so rung 3 of the fallback ladder shipped (`_classify_section_reorder` / `push_sectioned` in `src/docspan/backends/google_docs/backend.py`, ~lines 940-986): `heading_id` churn is accepted as a documented limitation and surfaced as a push warning when a reordered section also carries a content edit. No move primitive or copy-then-delete fallback was implemented.
|
|
5
|
+
|
|
6
|
+
## Context
|
|
7
|
+
When a sectioned mapping's local directory reflects two sections having swapped order (or a section relocated elsewhere in the document) relative to the stored manifest, `docspan push` must translate that into Google Docs batchUpdate requests.
|
|
8
|
+
|
|
9
|
+
The most naive implementation — delete the moved section's content range and reinsert it (as fresh markdown) at its new position — is the same mechanism the existing single-file push already uses for ordinary content edits (`_build_push_plan`'s diff-and-emit pipeline, `docs_request_builder.py`). But `heading_id` is confirmed (via `heading_anchors.py`'s docstring and `tabs.py:38-55`) to **not survive delete+reinsert**: Google Docs assigns a fresh `heading_id` to any newly-inserted heading paragraph, even if its text is byte-identical to what was deleted.
|
|
10
|
+
|
|
11
|
+
This matters because:
|
|
12
|
+
- ADR-001 makes `heading_id` the manifest's identity key. Regenerating it on every reorder would mean the manifest's identity mapping goes stale on the very operation (reorder) it exists to detect.
|
|
13
|
+
- Any cross-section anchor link pointing at that heading (`heading_anchors.py`'s anchor resolution) would silently break — the link would point at a `heading_id` that no longer exists — with no error surfaced at push time, only a broken link discovered later.
|
|
14
|
+
- This is exactly the kind of failure mode research (`research/pitfalls.md`) flags as needing a design-time decision, since it changes what "detecting a reorder" has to compile down to in `docs_request_builder.py`, not something safe to discover mid-implementation.
|
|
15
|
+
|
|
16
|
+
## Decision
|
|
17
|
+
Reorder of a section (a move that changes position without changing content) is implemented as an in-place move of the existing content range, preserving its `heading_id`, rather than as delete-old-content + insert-new-content-from-markdown. Add/delete/reorder classification (Task 3.2.1 in `implementation/plan.md`) uses `difflib.SequenceMatcher` over `heading_id` sequences (stored-manifest-order vs. current-local-directory-order) specifically so that "moved" is a distinguishable classification from "deleted + inserted," and the move path is only taken for entries the classifier marks as moved-not-changed.
|
|
18
|
+
|
|
19
|
+
The concrete Docs API batchUpdate request shape for the move itself is **not yet finalized** — `docs_request_builder.py` today only expresses insert/delete/style-update requests, no "move a range" primitive. Task 3.2.2 in the implementation plan is an explicit design spike to determine whether the Docs API exposes a true move-equivalent request, or whether the safest available approximation is copy-content-to-new-position followed by delete-of-the-old-range *only after* the insert succeeds (still avoiding a naive delete-then-reinsert-from-markdown, which is what would regenerate the `heading_id`).
|
|
20
|
+
|
|
21
|
+
## Consequences
|
|
22
|
+
- Task 3.2.2's spike (recorded here rather than left as an open Unresolved Question in `implementation/plan.md`) found no Docs API batchUpdate request that expresses "move a range" — `docs_request_builder.py`'s `build()` only ever emits equal/delete/insert/replace opcodes. Rung 1 (a true move primitive) is therefore unavailable, and rung 2 (copy-then-delete-after-insert-succeeds) was not implemented — it would need new request-shape code beyond anything demonstrated necessary once rung 3 covered the observed cases.
|
|
23
|
+
- Rung 3 shipped instead: for a *pure* reorder (manifest order changed, section content itself untouched), `docs_request_builder.py`'s `_repair` step folds the swapped run's content pairs back to `equal` opcodes, so `heading_id` is preserved for free and zero batch_update requests are emitted (a true no-op push, `status="skipped"`). `heading_id` churn only actually occurs when the reordered section *also* carries a genuine content edit the differ cannot fold back to `equal` — that case is written as ordinary delete+insert, the `heading_id` does not survive, and `push_sectioned` attaches a `status="warning"` message naming the reordered sections. The next `pull_sectioned` re-derives the manifest fresh, so the churn is self-healing rather than a lasting inconsistency; cross-section anchors pointing at a churned heading are not automatically re-resolved and can go stale until that next pull.
|
|
24
|
+
- Reorder-only pushes remain simpler than the move primitive this ADR originally anticipated, since the reused diff tail (`docs_request_builder.py:1133-1195`'s write-backwards, highest-anchor-first ordering) needed no new request type — `_classify_section_reorder` only changes what warning is attached, never what gets emitted.
|
|
25
|
+
- Edit-and-rename-in-the-same-cycle (a heading both moved and heavily content-edited at once) remains an accepted, named limitation — the classifier cannot always disambiguate "moved-and-edited" from "deleted-and-a-different-section-inserted-at-that-position" from content alone. This is the same case that now hits the rung-3 heading_id-churn path above; this ADR does not attempt further disambiguation.
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# ADR-003: Sectioned pull always uses the structural path, never Drive HTML export
|
|
2
|
+
|
|
3
|
+
## Status
|
|
4
|
+
Accepted
|
|
5
|
+
|
|
6
|
+
## Context
|
|
7
|
+
`GoogleDocsBackend.pull()` today has two distinct code paths (`backend.py`):
|
|
8
|
+
- **Default path** (`tab_id is None`, lines ~910-950): fetches the doc via Drive's HTML export and converts it with `DocumentConverter().html_to_markdown()`. `DocsStructureParser` is used here only throwaway, for anchor upgrading — its parsed node list is discarded, not what gets written to disk.
|
|
9
|
+
- **Structural path** (`tab_id is not None`, lines ~874-908): parses the doc into a flat `List[DocsParagraphNode]` via `DocsStructureParser`, projects it through `project()`, and writes the *parsed node list itself* (via `render_nodes_to_markdown()`) to disk.
|
|
10
|
+
|
|
11
|
+
Sectioned pull needs to partition the document at a configured heading level (`split_level`) into N separate markdown files. Drive's HTML export is an opaque, whole-document conversion — architecture research (`research/architecture.md`) confirmed it cannot be scoped to a heading range or otherwise handed a "start here, stop there" instruction. Only the structural path's flat node list is something `section_splitter.py` (a new module) can walk and cut at heading boundaries, because the split happens against the same node representation `render_nodes_to_markdown()` consumes.
|
|
12
|
+
|
|
13
|
+
## Decision
|
|
14
|
+
`backend.pull_sectioned()` always parses via `DocsStructureParser` + `project()` (the structural path), regardless of whether the mapping targets a tab or the default document — never via Drive HTML export, even though non-sectioned pull on the default (non-tab) path still uses HTML export today. Splitting happens after `project()` runs, not before, so each section's node list is already reduced to "what markdown can represent" form before being cut into groups — avoiding disagreement about residue handling at section boundaries (e.g. an empty paragraph immediately before a heading).
|
|
15
|
+
|
|
16
|
+
## Consequences
|
|
17
|
+
- Sectioned pull's output may have subtly different formatting-fidelity characteristics than a non-sectioned pull of the same document would have had via the default HTML-export path, since it's a genuinely different conversion pipeline (structural vs. HTML-export). This is judged acceptable because sectioned mode is opt-in and net-new — there is no existing sectioned-mode behavior to regress relative to.
|
|
18
|
+
- This makes the structural path (already used for tab-scoped docs) load-bearing for a second, independent reason (sectioning). Any future bug fix to the structural path's fidelity now has two justifications to preserve behavior for, and tests should cover both tab-scoped and sectioned-but-non-tab-scoped documents through it — two structural pull paths existing simultaneously (tab-scoping and sectioning) was explicitly flagged as a "Rabbit Hole" in requirements.md and needs a test matrix covering their interaction (a sectioned mapping targeting a specific tab).
|
|
19
|
+
- `render_nodes_to_markdown()`'s only stateful pass (`_group_code_runs`) was confirmed to be a local pass over the given list — safe to invoke once per section rather than once per document — so no shared cross-section state needs to be threaded through the splitter.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Adversarial Review: gdocs-sectioned-sync (iteration 2)
|
|
2
|
+
**Date**: 2026-08-13
|
|
3
|
+
**Verdict**: CONCERNS
|
|
4
|
+
|
|
5
|
+
## Blockers
|
|
6
|
+
(none)
|
|
7
|
+
|
|
8
|
+
## Concerns
|
|
9
|
+
- [x] Task 3.1.3 (`implementation/plan.md:202`) — **Resolved.** Plan.md now specifies only "pass a section file's own path (never the bare sectioned directory)" as `markdown_path`, dropping the previously-offered directory option that resolved one level too high against `build_source`'s `.parent` behavior. Also added a fixture requirement (Story 7.1/7.2) for colliding same-named images across two sections, surfaced as a `warning`/`error` per the Observability Plan.
|
|
10
|
+
- [x] Citation in Task 5.2.1 — **Resolved.** plan.md now cites `MappingState` at `core/state.py:12` and `SyncState.get`/`update` at `core/state.py:40-43`, not `orchestrator.py:81-143`.
|
|
11
|
+
- [ ] **Accepted limitation, not fixed.** No dedicated task for a manifest-vs-on-disk consistency pre-diff check (e.g., manifest lists a filename absent from disk, or a stray `.md` file not in the manifest) before push-time diffing begins. Coverage is implicit via Task 2.2.1 (pull-side matching) and Story 3.2 (add/delete/reorder detection derived from directory listing vs. manifest); this is judged sufficient for v1 given the appetite, and is not named as a separate pre-diff validation step.
|
|
12
|
+
- [ ] **Accepted limitation, not fixed.** Concurrent pull/push race handling has no explicit task or acceptance criterion (e.g., two processes racing on the same sectioned directory, or a push racing a manifest write). Single-file docspan mappings have the same unaddressed race today; sectioned mode does not regress this, and dedicated locking is out of scope for v1.
|
|
13
|
+
- [ ] Partial-API-failure coverage for the reorder fallback (copy-to-new-position + delete-old-range-after-insert) remains thin. Task 6.1.2's "keep the full reassembled request list within a single `batchUpdate` call wherever possible" implicitly gives the fallback's two operations atomicity (Docs batchUpdate applies all sub-requests as one transaction), but the plan never states this connection explicitly for Story 3.2/ADR-002's fallback path — a reader has to infer it. Recommend a one-line cross-reference in Task 3.2.3 or ADR-002 Consequences.
|
|
14
|
+
|
|
15
|
+
## Minors
|
|
16
|
+
- Preamble sentinel-key format (Task 2.1.2, `plan.md:187`) is still only "e.g. a fixed sentinel key" — no concrete value chosen. Left as-is.
|
|
17
|
+
- YAML library choice for `manifest.py` (Task 1.2.1) is still unspecified; `config.py` uses `import yaml` (PyYAML) while `pyproject.toml` also lists `ruamel.yaml>=0.18.0` as a dependency — the plan doesn't say which one `ManifestStore` should use. Left as-is.
|
|
18
|
+
- Comment-bucketing misassignment edge case (near-duplicate quoted text appearing in two sections) is not more deeply treated than iteration 1; Task 4.1.1's "first match wins in manifest order" is a defined-but-not-obviously-correct tiebreak for that case. Left as-is.
|
|
19
|
+
|
|
20
|
+
## Prior-blocker resolution status
|
|
21
|
+
|
|
22
|
+
1. **Resolved.** Step 0.5 (`plan.md:21`) and Task 3.1.3 (`plan.md:202`) now correctly describe `_build_push_plan`'s real signature/behavior, verified independently against `backend.py:172-299`: `_build_push_plan(local_path, doc_id, tab_id=None)` reads `pathlib.Path(local_path).read_text()` at line 195 and calls `resolve_document_images(image_nodes, local_path, ...)` at line 208 — matching the plan's citations. Task 3.1.3 adds a concrete `content: Optional[str] = None` parameter to skip the read when pre-assembled content is supplied, leaving single-file `push()`/`preview_push()` unaffected (`content=None`). The "unchanged" claim is now correctly scoped to only the diff/request-emission tail (`_build_push_plan`'s post-front-half code, confirmed unchanged in the read source at `backend.py:232-299`), not the whole function. One residual inaccuracy in the same task's image-path guidance is flagged above as a concern (not a blocker — the correct alternative is also given in the same sentence).
|
|
23
|
+
|
|
24
|
+
2. **Resolved.** Task 1.1.2 (`plan.md:172`) now specifies a concrete `@model_validator(mode="after")` on `Mapping` enforcing `sectioned == (split_level is not None)`, raising `ValueError` for both the `sectioned=True, split_level=None` and `sectioned=False, split_level=<set>` cases, and Story 1.1's Given/When/Then acceptance criteria (`plan.md:167-170`) reflect both invalid combinations plus the two valid ones. `config.py`'s current `Mapping` class (confirmed at `src/docspan/config.py:78-88`) has no existing validator, matching the plan's note that this introduces the pattern.
|
|
25
|
+
|
|
26
|
+
3. **Resolved.** New Story 7.5 (`plan.md:268-272`) and Task 7.5.1 deliver a `sectioned` × `tab_id` test matrix, explicitly citing ADR-003. It covers the three combinations that matter: both set (correct tab + same split as unscoped), `sectioned` true with `tab_id` unset (regression vs. Story 7.1), and `tab_id` set with `sectioned` false (regression check that sectioned code doesn't affect the existing tab-scoped path). The fourth combination (neither set) is already exercised by pre-existing tests, so its omission here is reasonable.
|
|
27
|
+
|
|
28
|
+
4. **Resolved.** A "Go/no-go gate" section (`plan.md:210-214`) now sits explicitly between Task 3.2.2 and Task 3.2.3, naming three tiers: a real move primitive, ADR-002's documented copy-then-delete-after-insert fallback, and (if even that is infeasible) accepting `heading_id` churn as a documented limitation — matching what ADR-002's Decision/Consequences sections actually say (verified by reading `decisions/ADR-002-reorder-as-in-place-move.md`). Story 7.3's acceptance criteria (`plan.md:261`) are rewritten to hold under either outcome ("go" preserves both `heading_id`s; "no-go" asserts the documented churn explicitly), so they don't need rewriting once the spike resolves.
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# Architecture Review: gdocs-sectioned-sync
|
|
2
|
+
**Date**: 2026-08-13
|
|
3
|
+
**Verdict**: CONCERNS
|
|
4
|
+
|
|
5
|
+
## Constitution Violations
|
|
6
|
+
- N/A — no `docs/adr/ADR-000-architecture-constitution.md` found in the repo (checked `docs/adr/` and repo-wide for `ADR-000*`).
|
|
7
|
+
|
|
8
|
+
## Blockers
|
|
9
|
+
|
|
10
|
+
None. All plan citations against actual source were verified accurate (`Mapping` at `src/docspan/config.py:78`, the three `m.local == file` sites at `src/docspan/cli/main.py:466,578,789`, `DiffTooExpensive`/`_bounded_opcodes` at `docs_request_builder.py:53-113`, write-backwards ordering at `build()` starting `docs_request_builder.py:1133`, `heading_id` capture at `docs_structure_parser.py:188/578`, and the claim that the request builder has no "move" primitive today — confirmed only insert/delete/style-update requests exist). No story bakes in a correctness or safety defect severe enough to block; the issues below are all fixable by restructuring specific tasks before Epic 1/2 implementation starts.
|
|
11
|
+
|
|
12
|
+
## Concerns
|
|
13
|
+
|
|
14
|
+
- [ ] **Task 1.1.1/1.1.2 (illegal state: `sectioned`/`split_level` pair)** — `Mapping.sectioned: bool = False` and `Mapping.split_level: Optional[str] = None` are two independently-optional fields, so `sectioned=True, split_level=None` and `sectioned=False, split_level="HEADING_1"` are both representable, and Task 1.1.2 only validates the *value* of `split_level` (one of HEADING_1..6), not the *conjunction* with `sectioned`. Remediation: add a pydantic `model_validator(mode="after")` on `Mapping` that raises when `sectioned and split_level is None`, and either clears or rejects a set `split_level` when `sectioned` is False. This is a config-parse-boundary fix (Lens 2 point 7), consistent with the existing "parse at the boundary" pattern `load_config` already follows for Confluence env-var defaults.
|
|
15
|
+
|
|
16
|
+
- [ ] **Task 1.2.1 (`heading_id` sentinel conflation for preamble)** — the manifest's `heading_id` field is documented to hold either a real Docs `headingId` or a synthetic sentinel for the preamble section (Domain Glossary: "Preamble / section 0"). Loading a real Docs ID and a manufactured sentinel into the same `str` field is exactly the kind of illegal-state-representable case Lens 2 point 6 flags — a future reader (or a diff routine) can't tell "real heading_id" from "sentinel" without out-of-band knowledge, and a pathological doc with a heading whose ID collides with the sentinel string is a real (if rare) correctness bug with no defense. Remediation: make `SectionManifestEntry.heading_id: Optional[str]` (`None` = preamble, matching how `DocsParagraphNode.heading_id` already models "no heading" as `None` at `docs_structure_parser.py:188`), and special-case the preamble entry positionally (always index 0, matched by position not key) in the Epic 3 SequenceMatcher pass rather than round-tripping a magic string through the identity system.
|
|
17
|
+
|
|
18
|
+
- [ ] **Task 4.1.1 (comment bucketing: silent mis-assignment, not just non-assignment)** — bucketing is "first section (in manifest order) whose rendered text contains the quoted substring wins." The plan only names a residue case for *zero* matches ("unassigned"); it has no case for the quoted text substring-matching *more than one* section (a repeated phrase, a shared boilerplate line, or a short/generic quote). Today's `first-match-wins` will silently attach the comment to the wrong section with no warning — worse than "unassigned," since a human won't know to look for it. Remediation: track match count per comment across sections; when count > 1, emit an "ambiguous — assigned to section N by manifest-order tiebreak" warning-level residue (same status shape as the existing "unmatched" residue) so at least it's visible, rather than a silent success.
|
|
19
|
+
|
|
20
|
+
- [ ] **Story 3.2 / DDD aggregate boundary (no `SectionedDocument` aggregate)** — the plan spreads "sectioned document" invariants (manifest order matches on-disk files, every manifest entry has a corresponding file, section count is consistent) across `ManifestStore`, `section_splitter.py`, and procedural code inside `backend.py`'s `pull_sectioned`/`push_sectioned` bodies, with no single object whose constructor enforces them. This is a missing aggregate root/value-object (Lens 1 point 3): nothing stops `push_sectioned` from reading a manifest with 5 entries against a directory with 4 files except an ad hoc check wherever someone remembers to add one. Remediation: introduce a `SectionedDocument` value object with a single factory (`SectionedDocument.load(directory) -> SectionedDocument`) that validates manifest/file consistency once at load time and raises/reports a clear error, and have both `pull_sectioned` and `push_sectioned` construct/return this type rather than juggling loose `(manifest, list[Path])` tuples inline.
|
|
21
|
+
|
|
22
|
+
- [ ] **Task 5.1.1 / cli/main.py path-resolution (SRP creep risk, not the pattern itself)** — replacing the three `m.local == file` exact-match sites with a "Specification pattern" predicate is the right direction, but the plan doesn't specify where this predicate lives. If it's added as a private helper duplicated near each of the three call sites (natural path of least resistance during implementation), the three sites drift independently over time the same way the current exact-match duplication already has (three separate `next((m for m in config.mappings if ...), None)` idioms doing the same lookup). Remediation: land it as one function, e.g. `resolve_mapping_for_path(config, path) -> Optional[Mapping]`, used by all three CLI call sites plus `orchestrator.py`'s dispatch — not three independent predicates.
|
|
23
|
+
|
|
24
|
+
- [ ] **Epic 5, Story 5.2 (state-keying migration is additive, but undocumented for existing sectioned re-runs)** — `SyncState.mappings` is a flat `dict[str, MappingState]` keyed by arbitrary path strings (`src/docspan/core/state.py:22,40-44`), so keying one entry per section-file path instead of one per `mapping.local` is mechanically trivial, confirming the plan's claim that `merge.py`/the content-hash store need no change. But the plan doesn't say what happens to the *existing* single `mapping.local`-keyed entry if a mapping is later switched from non-sectioned to sectioned (Migration Plan only covers the reverse: sectioned-absent → unchanged). Since that conversion is out of migration scope per the requirements ("No auto-migration of existing single-file mappings"), this is likely fine as long as `.markgate-state.json` isn't left with a dangling stale entry under the old path — worth one line in Epic 5 to state that a converted mapping's stale `mapping.local` entry is simply orphaned (harmless, ignorable) rather than silently reused.
|
|
25
|
+
|
|
26
|
+
## Nitpicks
|
|
27
|
+
|
|
28
|
+
- **"Chain-of-Responsibility" label on comment bucketing (Task 4.1.1)** and **"Specification pattern" label on CLI path resolution (Task 5.1.1)** are both GoF-pattern names applied to what implementation will likely be a single linear-scan function and a single predicate function, respectively (Lens 3 point 9). Naming the pattern in the plan is fine as design vocabulary, but implementers should not build handler-class hierarchies or a `Specification` interface/composite for either — a plain function is the right amount of structure here, and the plan should say so explicitly to head off over-engineering during Epic 4/5 implementation.
|
|
29
|
+
- **`_bounded_opcodes` is a private (underscore-free but module-private-by-convention) helper** in `docs_request_builder.py` (`docs_request_builder.py:80-113`) that Task 3.2.1 plans to reuse for section-level `heading_id` sequence alignment. Its signature (`List[Tuple]` in, generic) is already reusable as-is, but it's imported today only from within its own module. Worth a one-line task to confirm/export it as an intentional shared utility (or move it to a small `diff_utils.py`) before Epic 3 imports it cross-module, rather than reaching into another module's underscore-adjacent internals ad hoc.
|
|
30
|
+
- **`pull_sectioned`/`push_sectioned` as new `GoogleDocsBackend` methods (winning Architecture A)** is the right call given `push()`/`pull()` are already 130-380 lines each (confirmed: `backend.py` is 1228 lines total, `pull` starts at line 852, `push` at line 377 running to `pull`'s start — a ~475-line span for `push` alone including its helpers). Recommend Epic 2/3 acceptance criteria explicitly cap `pull_sectioned`/`push_sectioned` themselves to thin dispatch (manifest load → split/concat → delegate), so the size problem the plan correctly avoided in Option B doesn't quietly reappear inside the new methods instead of the old ones.
|
|
31
|
+
- **Primitive-obsession lens (heading_id/slug/filename as bare `str`)** is technically present (Lens 2 point 5) but is consistent with the rest of the codebase's convention — `DocsParagraphNode.heading_id: Optional[str]`, `style: str`, etc. are already unwrapped strings throughout `docs_structure_parser.py`. Not worth introducing `NewType` wrappers against the grain of the existing style; the sentinel-conflation concern above is the sharper, actionable version of this issue and supersedes it.
|