mdfetch 0.5.2__tar.gz → 0.8.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- mdfetch-0.8.0/.specify/feature.json +3 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/memory/changelog.md +46 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/memory/plan.md +25 -5
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/memory/spec.md +117 -1
- {mdfetch-0.5.2 → mdfetch-0.8.0}/CLAUDE.md +21 -2
- {mdfetch-0.5.2 → mdfetch-0.8.0}/PKG-INFO +8 -1
- {mdfetch-0.5.2 → mdfetch-0.8.0}/README.md +7 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/pyproject.toml +1 -1
- mdfetch-0.8.0/specs/010-boomi-blog-provider/checklists/requirements.md +37 -0
- mdfetch-0.8.0/specs/010-boomi-blog-provider/contracts/extractor-contract.md +57 -0
- mdfetch-0.8.0/specs/010-boomi-blog-provider/data-model.md +45 -0
- mdfetch-0.8.0/specs/010-boomi-blog-provider/plan.md +115 -0
- mdfetch-0.8.0/specs/010-boomi-blog-provider/quickstart.md +47 -0
- mdfetch-0.8.0/specs/010-boomi-blog-provider/research.md +76 -0
- mdfetch-0.8.0/specs/010-boomi-blog-provider/spec.md +99 -0
- mdfetch-0.8.0/specs/010-boomi-blog-provider/tasks.md +177 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/checklists/requirements.md +35 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/contracts/extractor-contract.md +60 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/data-model.md +66 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/plan.md +136 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/quickstart.md +60 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/research.md +127 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/spec.md +105 -0
- mdfetch-0.8.0/specs/011-konghq-blog-provider/tasks.md +198 -0
- mdfetch-0.8.0/specs/012-list-platforms/checklists/requirements.md +35 -0
- mdfetch-0.8.0/specs/012-list-platforms/contracts/cli.md +66 -0
- mdfetch-0.8.0/specs/012-list-platforms/data-model.md +40 -0
- mdfetch-0.8.0/specs/012-list-platforms/plan.md +113 -0
- mdfetch-0.8.0/specs/012-list-platforms/quickstart.md +47 -0
- mdfetch-0.8.0/specs/012-list-platforms/research.md +91 -0
- mdfetch-0.8.0/specs/012-list-platforms/spec.md +106 -0
- mdfetch-0.8.0/specs/012-list-platforms/tasks.md +116 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/cli.py +20 -3
- mdfetch-0.8.0/src/mdfetch/providers/boomi.py +43 -0
- mdfetch-0.8.0/src/mdfetch/providers/kong.py +104 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/router.py +11 -0
- mdfetch-0.8.0/tests/integration/snapshots/boomi-data-consistency-saas-on-prem.md +29 -0
- mdfetch-0.8.0/tests/integration/snapshots/boomi-gartner-magic-quadrant-ipaas-2026.md +29 -0
- mdfetch-0.8.0/tests/integration/snapshots/boomi-real-time-vs-batch-data-integration.md +29 -0
- mdfetch-0.8.0/tests/integration/snapshots/kong-ai-gateway-vs-litellm.md +29 -0
- mdfetch-0.8.0/tests/integration/snapshots/kong-gateway-3-14.md +30 -0
- mdfetch-0.8.0/tests/integration/snapshots/kong-insomnia-12-6.md +30 -0
- mdfetch-0.8.0/tests/integration/test_boomi_integration.py +56 -0
- mdfetch-0.8.0/tests/integration/test_kong_integration.py +62 -0
- mdfetch-0.8.0/tests/unit/test_boomi_extractor.py +138 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_cli.py +46 -0
- mdfetch-0.8.0/tests/unit/test_kong_extractor.py +241 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_router.py +27 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/uv.lock +1 -1
- mdfetch-0.5.2/.specify/feature.json +0 -3
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-analyze/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-archive-run/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-checklist/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-clarify/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-constitution/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-git-commit/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-git-feature/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-git-initialize/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-git-remote/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-git-validate/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-implement/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-plan/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-reconcile-run/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-specify/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-tasks/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.claude/skills/speckit-taskstoissues/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.analyze.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.archive.run.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.checklist.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.clarify.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.constitution.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.implement.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.plan.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.reconcile.run.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.specify.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.tasks.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gemini/commands/speckit.taskstoissues.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gitattributes +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.github/copilot-instructions.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.github/workflows/ci.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.github/workflows/integration.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.github/workflows/publish.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.gitignore +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.python-version +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/.registry +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/archive/LICENSE +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/archive/README.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/archive/commands/archive.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/archive/extension.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/README.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/commands/speckit.git.commit.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/commands/speckit.git.feature.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/commands/speckit.git.initialize.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/commands/speckit.git.remote.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/commands/speckit.git.validate.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/config-template.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/extension.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/git-config.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/bash/auto-commit.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/bash/create-new-feature.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/bash/git-common.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/bash/initialize-repo.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/powershell/auto-commit.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/powershell/create-new-feature.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/powershell/git-common.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/git/scripts/powershell/initialize-repo.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/reconcile/LICENSE +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/reconcile/README.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/reconcile/commands/reconcile.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions/reconcile/extension.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/extensions.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/init-options.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/integration.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/integrations/claude.manifest.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/integrations/gemini.manifest.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/integrations/speckit.manifest.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/memory/constitution.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/scripts/bash/check-prerequisites.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/scripts/bash/common.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/scripts/bash/create-new-feature.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/scripts/bash/setup-plan.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/scripts/bash/setup-tasks.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/templates/checklist-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/templates/constitution-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/templates/plan-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/templates/spec-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/templates/tasks-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/workflows/speckit/workflow.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.specify/workflows/workflow-registry.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/.vscode/settings.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/GEMINI.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/LICENSE +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/Makefile +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/contracts/api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/001-mdfetch-medium-extractor/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/contracts/public-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/002-devto-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/003-medium-freedium-fallback/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/003-medium-freedium-fallback/contracts/extract-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/003-medium-freedium-fallback/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/003-medium-freedium-fallback/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/003-medium-freedium-fallback/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/003-medium-freedium-fallback/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/004-remove-backoff/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/004-remove-backoff/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/004-remove-backoff/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/004-remove-backoff/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/004-remove-backoff/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/contracts/extractor-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/005-substack-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/contracts/public-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/006-thenewstack-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/contracts/public-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/007-dzone-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/contracts/cli.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/008-mdfetch-cli/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/contracts/formula.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/contracts/tap-update-job.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/specs/009-homebrew-tap-formula/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/base.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/exceptions.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/providers/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/providers/devto.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/providers/dzone.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/providers/medium.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/providers/substack.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/src/mdfetch/providers/thenewstack.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/conftest.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/conftest.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/architecting-the-asynchronous-agent.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/devto-integration-digest-december-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/devto-integration-digest-july-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/devto-integration-digest-march-2026.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/dzone-image-classification-pipeline-camel-djl.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/dzone-integration-patterns-fail-production.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/dzone-kiro-feature-to-requirements-design-tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/from-drift-to-parity.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/integration-digest-december-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/substack-api-trends-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/substack-kafka-topic-types.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/thenewstack-api-mcp-agent.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/thenewstack-async-apis.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/thenewstack-developer-portal-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/thenewstack-json-schema-ai.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/snapshots/thenewstack-mcp-api-governance.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/test_cli_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/test_devto_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/test_dzone_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/test_medium_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/test_substack_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/integration/test_thenewstack_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_devto_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_dzone_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_fetch_errors.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_medium_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_silent.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_substack_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.8.0}/tests/unit/test_thenewstack_extractor.py +0 -0
|
@@ -1,3 +1,49 @@
|
|
|
1
|
+
### mdfetch — CLI: List Supported Platforms — 2026-06-02
|
|
2
|
+
|
|
3
|
+
**Branch**: `012-list-platforms`
|
|
4
|
+
**Spec**: specs/012-list-platforms
|
|
5
|
+
|
|
6
|
+
**What was added**:
|
|
7
|
+
- `md-fetch --list-platforms` flag that prints every supported platform domain (alphabetical) to stdout and exits 0 — no URL, no network request.
|
|
8
|
+
- Multi-tenant platforms annotated with their wildcard (e.g. `medium.com (and *.medium.com)`); exact-match platforms shown plain.
|
|
9
|
+
- `URL` argument made optional: required only when `--list-platforms` is absent, preserving the existing `md-fetch <URL>` contract (no breaking change). Missing URL without the flag → Click usage error, exit 2.
|
|
10
|
+
- New subdomain-aware router accessor `supported_platforms() -> list[tuple[str, bool]]` (sorted, registry-derived); existing `supported_domains()` unchanged.
|
|
11
|
+
- `README.md` usage example + `pyproject.toml` version bump `0.7.1` → `0.8.0`.
|
|
12
|
+
|
|
13
|
+
**New Components**:
|
|
14
|
+
- `src/mdfetch/router.py` — `supported_platforms()` accessor (modified, no new file)
|
|
15
|
+
- `src/mdfetch/cli.py` — `--list-platforms` flag, optional URL (modified, no new file)
|
|
16
|
+
- `tests/unit/test_router.py` — 3 new tests; `tests/unit/test_cli.py` — 5 new tests
|
|
17
|
+
- No integration test (operation is offline)
|
|
18
|
+
|
|
19
|
+
**Tasks Completed**: 10/10
|
|
20
|
+
|
|
21
|
+
**Note**: Archived pre-merge at the user's request (deviates from the usual post-merge archival).
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
### mdfetch — Kong Blog Provider — 2026-06-02
|
|
26
|
+
|
|
27
|
+
**Branch**: `011-konghq-blog-provider`
|
|
28
|
+
**Spec**: specs/011-konghq-blog-provider
|
|
29
|
+
|
|
30
|
+
**What was added**:
|
|
31
|
+
- `KongExtractor` for `konghq.com/blog/<category>/<slug>` articles. Kong is a Next.js site with hashed CSS-module class names, so selection pins to stable companion classes only.
|
|
32
|
+
- Article discriminator: `<main class="type-article">` (absent on the blog index and category listings). Body = the `<main>` `<section>` richest in `.rich-text-block`.
|
|
33
|
+
- In-body chrome stripped: `.toc-wrap` (TOC sidebar — NOT `[class*=TableOfContents]`, which wraps the whole body), `.component.video`, `.component.more-on-this`, `.order-top`, and trailing non-`intro` `.section-header-block`.
|
|
34
|
+
- `.agent` "agent mode" spans stripped (they inject literal Markdown duplicating styled content; also fixed a `# #` double-heading on the title).
|
|
35
|
+
- Title `<h1>` + publication date prepended (date matched by a month-name regex); author bylines and read time dropped per spec clarification.
|
|
36
|
+
- `README.md` (Supported platforms row + usage example), `pyproject.toml` version bump `0.6.0` → `0.7.0`, and a `CLAUDE.md` gotcha documenting the discriminator and CSS-module fragility.
|
|
37
|
+
|
|
38
|
+
**New Components**:
|
|
39
|
+
- `src/mdfetch/providers/kong.py` — `KongExtractor`
|
|
40
|
+
- `tests/unit/test_kong_extractor.py` — 12 unit tests
|
|
41
|
+
- `tests/integration/test_kong_integration.py` + 3 snapshots — real-URL tests (3 articles + index/category rejection)
|
|
42
|
+
|
|
43
|
+
**Tasks Completed**: 26/26
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
1
47
|
### mdfetch — Homebrew Tap Formula — 2026-05-16
|
|
2
48
|
|
|
3
49
|
**Branch**: `009-homebrew-tap-formula`
|
|
@@ -90,6 +90,16 @@ DZoneExtractor(BaseExtractor) — src/mdfetch/providers/dzone.py
|
|
|
90
90
|
├── DOMAINS = frozenset({"dzone.com"})
|
|
91
91
|
├── _markdownify_kwargs() → overrides to add code_language_callback
|
|
92
92
|
└── clean_html() → locates div.content-html; raises UnsupportedContentTypeError if absent;
|
|
93
|
+
|
|
94
|
+
KongExtractor(BaseExtractor) — src/mdfetch/providers/kong.py [011-konghq-blog-provider]
|
|
95
|
+
├── DOMAINS = frozenset({"konghq.com"})
|
|
96
|
+
└── clean_html() → requires <main class="type-article"> (else UnsupportedContentTypeError);
|
|
97
|
+
body = the <main> <section> richest in .rich-text-block;
|
|
98
|
+
strips in-body chrome (.toc-wrap, .component.video, .component.more-on-this,
|
|
99
|
+
.order-top, trailing non-intro .section-header-block) and all .agent spans;
|
|
100
|
+
prepends copy(h1) + publication date (class-less hero <div>, month-name regex);
|
|
101
|
+
Note: Next.js CSS-module hashes — pin to stable classes only; TOC is .toc-wrap,
|
|
102
|
+
NOT [class*=TableOfContents] (that wraps the whole body)
|
|
93
103
|
```
|
|
94
104
|
|
|
95
105
|
### Router / Auto-Discovery
|
|
@@ -122,16 +132,17 @@ src/
|
|
|
122
132
|
└── mdfetch/
|
|
123
133
|
├── __init__.py # Public surface: exposes extract(), exception re-exports
|
|
124
134
|
├── exceptions.py # MdfetchError hierarchy (6 exception classes)
|
|
125
|
-
├── router.py # @register, _autodiscover_providers(), route()
|
|
135
|
+
├── router.py # @register, _autodiscover_providers(), route(), supported_domains(), supported_platforms() [012-list-platforms]
|
|
126
136
|
├── base.py # BaseExtractor ABC + fetch_html() + convert_to_markdown() + extract() template
|
|
127
|
-
├── cli.py # click-
|
|
137
|
+
├── cli.py # click CLI; `md-fetch <URL>` + `--list-platforms` (URL optional with flag) [012-list-platforms]
|
|
128
138
|
└── providers/
|
|
129
139
|
├── __init__.py # Empty — auto-discovery handles registration
|
|
130
140
|
├── medium.py # MediumExtractor
|
|
131
141
|
├── devto.py # DevToExtractor [002-devto-provider]
|
|
132
142
|
├── substack.py # SubstackExtractor [005-substack-provider]
|
|
133
143
|
├── thenewstack.py # TheNewStackExtractor [006-thenewstack-provider]
|
|
134
|
-
|
|
144
|
+
├── dzone.py # DZoneExtractor
|
|
145
|
+
└── kong.py # KongExtractor [011-konghq-blog-provider]
|
|
135
146
|
|
|
136
147
|
tests/
|
|
137
148
|
├── unit/
|
|
@@ -222,13 +233,15 @@ update-homebrew-tap job (NEW, needs: publish)
|
|
|
222
233
|
## Testing Strategy
|
|
223
234
|
|
|
224
235
|
**Unit tests** (101 tests, offline):
|
|
225
|
-
- Router: domain routing, subdomain suffix matching, duplicate registration, invalid URLs, unsupported platforms
|
|
236
|
+
- Router: domain routing, subdomain suffix matching, duplicate registration, invalid URLs, unsupported platforms; `supported_platforms()` accessor (sorted `(domain, bool)`, covers every `supported_domains()` entry, subdomain flags) [012-list-platforms]
|
|
226
237
|
- MediumExtractor: clean_html, convert_to_markdown, empty content, non-article pages, _parse_freedium (heading remap, missing main-content), fallback on 403/429 (URL construction, exc.url contract, no-sleep on 429), no-fallback on 200, UnsupportedContentTypeError.url on Freedium path [003-medium-freedium-fallback]
|
|
227
238
|
- DevToExtractor: clean_html (title/cover/heading/image preservation, iframe/ltag embed→link, anchor stripping, non-article error), convert_to_markdown (headings/code/lists/images, no raw HTML, empty content error) [002-devto-provider]
|
|
228
239
|
- SubstackExtractor: routing (subdomain + root domain + _no_retry_status_codes assertion), clean_html (body.markup tag return, subscription-widget strip, title prepend, subtitle prepend, prose preservation, iframe→anchor), convert_to_markdown (title heading, no triple blank lines, image syntax, link preservation), paywalled post (non-empty, Subscribe text absent, free preview present), error cases (UnsupportedContentTypeError on no body, EmptyContentError on whitespace body) [005-substack-provider]
|
|
229
240
|
- TheNewStackExtractor: routing (thenewstack.io domain), clean_html (body div return, title prepend, deck-as-paragraph prepend, sponsor note strip, all 3 disclosure variant strips, iframe→anchor, no deck when absent), convert_to_markdown (title heading, deck after title, no triple blank lines, image syntax, link preservation), error cases (UnsupportedContentTypeError on no body, EmptyContentError on whitespace body) [006-thenewstack-provider]
|
|
241
|
+
- KongExtractor: routing (konghq.com domain), clean_html (title+date prepend order, chrome strip — .toc-wrap/.component.video/.component.more-on-this/.order-top/non-intro .section-header-block, .agent affordance strip, body-block preservation, raises without type-article, raises without .rich-text-block), convert_to_markdown (title then date, structure + inline code preserved, chrome/authors excluded, no triple blank lines, EmptyContentError on whitespace body) [011-konghq-blog-provider]
|
|
230
242
|
- Fetch errors: HTTP 404, 503, timeout, connection error, size limit exceeded; `_no_retry_status_codes` immediate-raise + `_no_retry_codes` override [003-medium-freedium-fallback]
|
|
231
243
|
- Silent: no stdout/stderr output, no logging during extraction
|
|
244
|
+
- CLI: `--list-platforms` lists all registered domains (subdomain-annotated), exits 0 with no network and no `extract` call, list takes precedence over a supplied URL, and missing-URL-without-flag exits 2 [012-list-platforms]
|
|
232
245
|
|
|
233
246
|
**Integration tests** (15 tests, network required):
|
|
234
247
|
- Parametrized over 3 real stn1slv.medium.com articles (including a known paywalled URL that exercises the Freedium fallback when medium.com returns 403) [003-medium-freedium-fallback]
|
|
@@ -299,4 +312,11 @@ update-homebrew-tap job (NEW, needs: publish)
|
|
|
299
312
|
|
|
300
313
|
---
|
|
301
314
|
|
|
302
|
-
|
|
315
|
+
### Revision: Archival 2026-06-02
|
|
316
|
+
- Archived **011-konghq-blog-provider**: added the `KongExtractor` architecture block, `kong.py` to the project-structure tree, and its unit/integration test coverage. No new runtime dependencies. [Source: specs/011-konghq-blog-provider]
|
|
317
|
+
- Note: features **007-dzone-provider** (partially present) and **010-boomi-blog-provider** (absent) are not fully archived in this memory plan (gap pre-dating this run).
|
|
318
|
+
|
|
319
|
+
### Revision: Archival 2026-06-02 (012-list-platforms)
|
|
320
|
+
- Archived **012-list-platforms**: annotated `cli.py` (`--list-platforms`, URL optional) and `router.py` (`supported_platforms()`) in the project-structure tree; added CLI + router-accessor unit-test coverage notes. No new runtime dependencies; no integration test (operation is offline). Reconciled pre-merge at the user's request. [Source: specs/012-list-platforms]
|
|
321
|
+
|
|
322
|
+
*Last Updated: 2026-06-02 | Sources appended: [specs/004-remove-backoff/plan.md], [specs/005-substack-provider/plan.md], [specs/006-thenewstack-provider/plan.md], [specs/009-homebrew-tap-formula/plan.md], [specs/011-konghq-blog-provider/plan.md], [specs/012-list-platforms/plan.md]*
|
|
@@ -223,6 +223,51 @@ A developer runs the integration test suite and all dev.to integration tests pas
|
|
|
223
223
|
|
|
224
224
|
---
|
|
225
225
|
|
|
226
|
+
### US-019 — Extract a Kong Blog Article to Markdown (P1)
|
|
227
|
+
[Source: specs/011-konghq-blog-provider]
|
|
228
|
+
|
|
229
|
+
A user passes a public Kong (`konghq.com`) blog article URL and receives the article's title, publication date, and full body as clean Markdown, with all site chrome (navigation, breadcrumbs, table-of-contents, demo CTAs, recommended-posts carousel, author bios, newsletter, footer) removed.
|
|
230
|
+
|
|
231
|
+
**Acceptance Scenarios**:
|
|
232
|
+
1. Given a public Kong blog article URL, when `extract()` is called, then the result is Markdown whose first line is `# <article title>`, followed by the publication date, then the body content.
|
|
233
|
+
2. Given a Kong article with H2/H3 sections, lists, and inline code, when extracted, then those are preserved as Markdown in document order, and no author byline or read-time text is present.
|
|
234
|
+
|
|
235
|
+
---
|
|
236
|
+
|
|
237
|
+
### US-020 — Reject Non-Article Kong URLs (P2)
|
|
238
|
+
[Source: specs/011-konghq-blog-provider]
|
|
239
|
+
|
|
240
|
+
A user passes a Kong URL that is not a readable article (the blog index or a category listing). The library raises a typed error rather than returning garbage.
|
|
241
|
+
|
|
242
|
+
**Acceptance Scenarios**:
|
|
243
|
+
1. Given the Kong blog index or a category listing URL, when `extract()` is called, then `UnsupportedContentTypeError` is raised.
|
|
244
|
+
2. Given a Kong article whose body yields no text after stripping, when `extract()` is called, then `EmptyContentError` is raised.
|
|
245
|
+
|
|
246
|
+
---
|
|
247
|
+
|
|
248
|
+
### US-021 — List Supported Platforms from the CLI (P1)
|
|
249
|
+
[Source: specs/012-list-platforms]
|
|
250
|
+
|
|
251
|
+
A user runs the `md-fetch` CLI with a dedicated flag and receives a readable list of every supported platform domain, printed to standard output with a success exit code, without supplying a URL or making any network request.
|
|
252
|
+
|
|
253
|
+
**Acceptance Scenarios**:
|
|
254
|
+
1. Given the CLI is installed, when the user runs `md-fetch --list-platforms`, then every supported platform domain is printed to standard output and the process exits 0, with no network request made.
|
|
255
|
+
2. Given a new provider has been registered, when the user runs `md-fetch --list-platforms`, then the newly supported domain appears with no change to the list operation itself.
|
|
256
|
+
3. Given neither a URL nor `--list-platforms` is supplied, when `md-fetch` is run, then a usage error is printed and the process exits with code 2 (the existing `md-fetch <URL>` behaviour is unchanged).
|
|
257
|
+
|
|
258
|
+
---
|
|
259
|
+
|
|
260
|
+
### US-022 — Distinguish Multi-Tenant Platforms in the Listing (P2)
|
|
261
|
+
[Source: specs/012-list-platforms]
|
|
262
|
+
|
|
263
|
+
A user listing platforms can tell which platforms also accept per-author subdomains (e.g. Medium, Substack) from those that match an exact domain only.
|
|
264
|
+
|
|
265
|
+
**Acceptance Scenarios**:
|
|
266
|
+
1. Given a provider that matches subdomains, when the user lists platforms, then the output indicates subdomains of that platform are supported (e.g. `medium.com (and *.medium.com)`).
|
|
267
|
+
2. Given a provider that matches an exact domain only, when the user lists platforms, then the output shows the bare domain with no subdomain indicator.
|
|
268
|
+
|
|
269
|
+
---
|
|
270
|
+
|
|
226
271
|
## Functional Requirements
|
|
227
272
|
|
|
228
273
|
### Extraction
|
|
@@ -302,6 +347,28 @@ A developer runs the integration test suite and all dev.to integration tests pas
|
|
|
302
347
|
- **FR-018**: The dev.to provider MUST raise `UnsupportedContentTypeError` when a `dev.to` URL is provided but the page is not an article (e.g., an author profile, a tag listing, or an organisation page). [Source: specs/002-devto-provider]
|
|
303
348
|
- **FR-019**: The dev.to provider MUST replace embedded third-party content (GitHub Gists, CodePen demos, YouTube videos, liquid-tag embeds) with a plain Markdown link to the embedded resource URL. Embeds must not be silently dropped. [Source: specs/002-devto-provider]
|
|
304
349
|
|
|
350
|
+
### Kong Platform
|
|
351
|
+
- **FR-060**: The library MUST route `konghq.com` blog URLs to the Kong provider using the existing domain-registration mechanism (`@register` decorator + `DOMAINS` frozenset). [Source: specs/011-konghq-blog-provider]
|
|
352
|
+
- **FR-061**: The library MUST extract the main article body from a Kong blog page and return it as clean Markdown. The body is the `<main>` `<section>` richest in `.rich-text-block` blocks. [Source: specs/011-konghq-blog-provider]
|
|
353
|
+
- **FR-062**: The library MUST strip all non-content elements before conversion: navigation, breadcrumbs, topic tags, the table-of-contents sidebar (`.toc-wrap`), demo CTAs, recommended-posts carousel, author byline/bio blocks, the read-time estimate, newsletter signup, and footer. [Source: specs/011-konghq-blog-provider]
|
|
354
|
+
- **FR-063**: The library MUST prepend the article title as a top-level Markdown heading (`# Title`), followed by the publication date when present. Author bylines and read time MUST NOT be included. [Source: specs/011-konghq-blog-provider]
|
|
355
|
+
- **FR-064**: The library MUST preserve structural body content: H2/H3 headings, paragraphs, ordered/unordered lists, inline and block code, blockquotes, hyperlinks, and images inside the body. Images outside the body (hero/banner) MUST be excluded. [Source: specs/011-konghq-blog-provider]
|
|
356
|
+
- **FR-065**: The library MUST raise `UnsupportedContentTypeError` when the page is not a recognisable article — i.e. `<main>` lacks the `type-article` class (blog index, category listing, non-article page). [Source: specs/011-konghq-blog-provider]
|
|
357
|
+
- **FR-066**: The library MUST raise `EmptyContentError` when the Kong article body yields no extractable text after stripping. [Source: specs/011-konghq-blog-provider]
|
|
358
|
+
- **FR-067**: The library MUST collapse runs of three or more consecutive blank lines to a single blank line in the Kong output Markdown. [Source: specs/011-konghq-blog-provider]
|
|
359
|
+
- **FR-068**: The Kong provider MUST depend only on stable, human-authored CSS classes (never Next.js hashed CSS-module suffixes), and MUST strip `.agent` "agent mode" spans that inject literal Markdown duplicating the styled content. [Source: specs/011-konghq-blog-provider]
|
|
360
|
+
|
|
361
|
+
### CLI — List Platforms
|
|
362
|
+
- **FR-069**: The CLI MUST provide a `--list-platforms` flag on the existing `md-fetch` command that lists all currently supported platform domains. [Source: specs/012-list-platforms]
|
|
363
|
+
- **FR-070**: When `--list-platforms` is supplied, the `URL` argument MUST be optional; without the flag, the existing behaviour (URL required) MUST be unchanged. [Source: specs/012-list-platforms]
|
|
364
|
+
- **FR-071**: The list output MUST be derived from the live provider registry so adding or removing a provider changes the output with no separately maintained list. [Source: specs/012-list-platforms]
|
|
365
|
+
- **FR-072**: The list operation MUST print to standard output and exit with code 0. [Source: specs/012-list-platforms]
|
|
366
|
+
- **FR-073**: The list operation MUST NOT require a URL and MUST NOT perform any network request. [Source: specs/012-list-platforms]
|
|
367
|
+
- **FR-074**: The list output ordering MUST be deterministic (alphabetical by domain). [Source: specs/012-list-platforms]
|
|
368
|
+
- **FR-075**: The list output MUST indicate, for multi-tenant platforms, that subdomains are also supported (`<domain> (and *.<domain>)`), while showing exact-match-only platforms as the bare domain. [Source: specs/012-list-platforms]
|
|
369
|
+
- **FR-076**: The `--list-platforms` flag MUST be documented in the CLI's own `--help` text. [Source: specs/012-list-platforms]
|
|
370
|
+
- **FR-077**: When `--list-platforms` is requested, the CLI MUST NOT attempt extraction in the same invocation, even if a URL is also supplied (the list takes precedence). [Source: specs/012-list-platforms]
|
|
371
|
+
|
|
305
372
|
---
|
|
306
373
|
|
|
307
374
|
## Key Entities
|
|
@@ -369,6 +436,18 @@ A developer runs the integration test suite and all dev.to integration tests pas
|
|
|
369
436
|
|
|
370
437
|
A GitHub fine-grained Personal Access Token with `Contents: read+write` on `stn1slv/homebrew-tap` only. Stored as a repository secret in `stn1slv/md-fetch` and consumed exclusively by the `update-homebrew-tap` CI job.
|
|
371
438
|
|
|
439
|
+
### Kong Blog Post
|
|
440
|
+
[Source: specs/011-konghq-blog-provider]
|
|
441
|
+
| Attribute | Type | Description |
|
|
442
|
+
|-----------|------|-------------|
|
|
443
|
+
| `is-article` | discriminator | `<main>` carries the stable `type-article` class; absence ⇒ non-article |
|
|
444
|
+
| `title` | `str` | Single `<h1>` in the hero section; prepended as `# Title` (with `.agent` `#` affordance stripped) |
|
|
445
|
+
| `date` | `str \| None` | Publication date — a class-less hero `<div>` matched by `^[A-Z][a-z]+ \d{1,2}, \d{4}$`; rendered under the title |
|
|
446
|
+
| `body` | `Tag` | The `<main>` `<section>` richest in `.rich-text-block`; in-body chrome stripped (`.toc-wrap`, `.component.video`, `.component.more-on-this`, `.order-top`, trailing non-`intro` `.section-header-block`) |
|
|
447
|
+
| authors / read time | chrome | Dropped (clarification: keep date only) |
|
|
448
|
+
|
|
449
|
+
**Validation**: `main.type-article` must exist and a content section must contain ≥1 `.rich-text-block` → else `UnsupportedContentTypeError`. Body must yield non-empty text → else `EmptyContentError`. No paywall — all Kong blog articles are publicly accessible.
|
|
450
|
+
|
|
372
451
|
### ExtractionResult (Output)
|
|
373
452
|
| Attribute | Type | Description |
|
|
374
453
|
|-----------|------|-------------|
|
|
@@ -389,6 +468,18 @@ MdfetchError (base)
|
|
|
389
468
|
|
|
390
469
|
All exceptions carry `message: str` and `url: str | None`. `HTTPStatusError` additionally carries `status_code: int`.
|
|
391
470
|
|
|
471
|
+
### Supported Platform
|
|
472
|
+
[Source: specs/012-list-platforms]
|
|
473
|
+
|
|
474
|
+
A read-only projection of one entry in the provider registry, exposed by `router.supported_platforms() -> list[tuple[str, bool]]` (sorted ascending by domain).
|
|
475
|
+
|
|
476
|
+
| Attribute | Type | Description |
|
|
477
|
+
|-----------|------|-------------|
|
|
478
|
+
| `domain` | `str` | Registered domain (e.g. `medium.com`, `dev.to`) |
|
|
479
|
+
| `matches_subdomains` | `bool` | Whether subdomains of `domain` also route to this provider (the provider's `MATCH_SUBDOMAINS` flag) |
|
|
480
|
+
|
|
481
|
+
**Validation**: One tuple per registered domain — no missing, no extras; domains unique. This is the subdomain-aware superset of `supported_domains() -> frozenset[str]`, which is retained unchanged.
|
|
482
|
+
|
|
392
483
|
---
|
|
393
484
|
|
|
394
485
|
## Call Lifecycle / State Transitions
|
|
@@ -450,6 +541,11 @@ caller provides URL string
|
|
|
450
541
|
- **Concurrent releases**: Addressed by `concurrency: group=homebrew-tap-update, cancel-in-progress=false` — runs are serialized so the second release waits for the first tap-update to complete. [Source: specs/009-homebrew-tap-formula]
|
|
451
542
|
- **brew test fails after install**: `brew test md-fetch` exits non-zero when a transitive dependency is missing or `md-fetch --version` fails — installation validation fails. [Source: specs/009-homebrew-tap-formula]
|
|
452
543
|
- **Formula structure changed (sed no-op)**: If the formula is manually edited and the url/sha256 line format changes, `sed` produces no diff and `git commit` finds nothing staged → exits non-zero → CI job fails. [Source: specs/009-homebrew-tap-formula]
|
|
544
|
+
- **Kong non-article pages**: The blog index (`/blog`) and category listings (`/blog/product-releases`) render `<main>` without the `type-article` class (and zero `.rich-text-block`) → `UnsupportedContentTypeError`. [Source: specs/011-konghq-blog-provider]
|
|
545
|
+
- **Kong Next.js CSS-module hashes**: Per-component class names (e.g. `Section_section__Grz_Y`) change between builds; selection pins to stable companion classes only. The TOC sidebar is `.toc-wrap` — NOT `[class*=TableOfContents]`, which wraps the entire body and would delete the article. [Source: specs/011-konghq-blog-provider]
|
|
546
|
+
- **Kong "agent mode" duplicates**: `<span class="agent">` elements inject literal Markdown (`**`, `- `, `# `) duplicating styled content (and cause a `# #` double-heading on the title); all `.agent` spans are decomposed before conversion. [Source: specs/011-konghq-blog-provider]
|
|
547
|
+
- **CLI list + output/fetch options**: When `--list-platforms` is combined with `-o/--output`, `--retries`, or `--force`, the list is printed to stdout and those options are ignored — the list operation returns before any extraction. [Source: specs/012-list-platforms]
|
|
548
|
+
- **CLI missing URL**: Running `md-fetch` with neither a URL nor `--list-platforms` raises a Click usage error and exits with code 2. [Source: specs/012-list-platforms]
|
|
453
549
|
|
|
454
550
|
---
|
|
455
551
|
|
|
@@ -486,6 +582,19 @@ caller provides URL string
|
|
|
486
582
|
- **SC-034**: Formula update failures produce a visible CI job failure on every failed attempt, with zero silent failures. [Source: specs/009-homebrew-tap-formula]
|
|
487
583
|
- **SC-035**: The README install section includes the Homebrew installation command, enabling users to discover and use it without prior knowledge of the project's Python packaging. [Source: specs/009-homebrew-tap-formula]
|
|
488
584
|
|
|
585
|
+
- **SC-036**: Each of the three reference Kong articles returns Markdown containing the full title and all body section headings and paragraph text, with zero non-content elements. [Source: specs/011-konghq-blog-provider]
|
|
586
|
+
- **SC-037**: Extraction of a Kong article completes within the base class 30-second fetch timeout on a stable connection. [Source: specs/011-konghq-blog-provider]
|
|
587
|
+
- **SC-038**: A non-article Kong URL (the blog index or a category listing) raises `UnsupportedContentTypeError` within the normal fetch timeout. [Source: specs/011-konghq-blog-provider]
|
|
588
|
+
- **SC-039**: The extracted Markdown for any Kong article contains no consecutive blank-line runs of three or more lines. [Source: specs/011-konghq-blog-provider]
|
|
589
|
+
- **SC-040**: The Kong provider is exercised by integration tests using real network requests against the reference URLs, matching the pattern established by existing providers. [Source: specs/011-konghq-blog-provider]
|
|
590
|
+
- **SC-041**: For a Kong article with a visible publication date, the extracted Markdown contains the date directly beneath the title and contains no author byline or read-time text. [Source: specs/011-konghq-blog-provider]
|
|
591
|
+
|
|
592
|
+
- **SC-042**: A user can list every supported platform with a single `md-fetch --list-platforms` invocation and no URL, completing in well under one second with no network access. [Source: specs/012-list-platforms]
|
|
593
|
+
- **SC-043**: The listed domains exactly match the set of domains registered to providers — no missing entries and no extras. [Source: specs/012-list-platforms]
|
|
594
|
+
- **SC-044**: After a new provider is added, its domain appears in the list output with no edit to the list operation's code. [Source: specs/012-list-platforms]
|
|
595
|
+
- **SC-045**: Multi-tenant platforms (those accepting subdomains) are visually distinguishable from exact-match platforms in the output. [Source: specs/012-list-platforms]
|
|
596
|
+
- **SC-046**: The list operation is exercised by at least one automated unit test that asserts the output contains the known supported domains and exits successfully. [Source: specs/012-list-platforms]
|
|
597
|
+
|
|
489
598
|
---
|
|
490
599
|
|
|
491
600
|
## Assumptions
|
|
@@ -513,4 +622,11 @@ caller provides URL string
|
|
|
513
622
|
|
|
514
623
|
---
|
|
515
624
|
|
|
516
|
-
|
|
625
|
+
### Revision: Archival 2026-06-02
|
|
626
|
+
- Archived **011-konghq-blog-provider**: added US-019/US-020, FR-060–FR-068 (Kong Platform), the Kong Blog Post entity, Kong edge cases, and SC-036–SC-041. [Source: specs/011-konghq-blog-provider]
|
|
627
|
+
- Note: features **007-dzone-provider** and **010-boomi-blog-provider** are not present in this memory spec (un-archived gap pre-dating this run).
|
|
628
|
+
|
|
629
|
+
### Revision: Archival 2026-06-02 (012-list-platforms)
|
|
630
|
+
- Archived **012-list-platforms**: added US-021/US-022, FR-069–FR-077 (CLI — List Platforms), the Supported Platform entity, two CLI edge cases, and SC-042–SC-046. Reconciled pre-merge at the user's request (deviates from the usual post-merge archival). [Source: specs/012-list-platforms]
|
|
631
|
+
|
|
632
|
+
*Last Updated: 2026-06-02 | Sources appended: [specs/004-remove-backoff/spec.md], [specs/005-substack-provider/spec.md], [specs/006-thenewstack-provider/spec.md], [specs/009-homebrew-tap-formula/spec.md], [specs/011-konghq-blog-provider/spec.md]*
|
|
@@ -17,11 +17,13 @@ src/mdfetch/
|
|
|
17
17
|
├── devto.py # DevToExtractor (dev.to)
|
|
18
18
|
├── substack.py # SubstackExtractor (substack.com + *.substack.com)
|
|
19
19
|
├── thenewstack.py # TheNewStackExtractor (thenewstack.io)
|
|
20
|
-
|
|
20
|
+
├── dzone.py # DZoneExtractor (dzone.com)
|
|
21
|
+
├── boomi.py # BoomiExtractor (boomi.com/blog)
|
|
22
|
+
└── kong.py # KongExtractor (konghq.com/blog)
|
|
21
23
|
|
|
22
24
|
tests/
|
|
23
25
|
├── unit/ # pytest unit tests (no network)
|
|
24
|
-
└── integration/ # real network tests (Medium + dev.to + Substack + TheNewStack URLs + snapshots)
|
|
26
|
+
└── integration/ # real network tests (Medium + dev.to + Substack + TheNewStack + DZone + Boomi + Kong URLs + snapshots)
|
|
25
27
|
|
|
26
28
|
.github/workflows/
|
|
27
29
|
├── ci.yml # lint + unit tests on push/PR (Python 3.12–3.14)
|
|
@@ -40,6 +42,8 @@ Makefile # setup / test / lint / format / build / upgrade-deps / cl
|
|
|
40
42
|
- **No logging**: all failures communicated via typed exceptions only (FR-013)
|
|
41
43
|
- **Type hints**: strict (`mypy src/` must pass with zero errors)
|
|
42
44
|
- **Linter/formatter**: `ruff` (`make lint` / `make format`)
|
|
45
|
+
- **Version bump**: every user-facing change MUST bump `version` in `pyproject.toml` (SemVer — new provider/feature = minor, bug fix = patch). Adding a provider follows the dev.to/Substack/Boomi minor-bump precedent. Don't forget this — it gates the PyPI release (`publish.yml`) and the Homebrew tap auto-update.
|
|
46
|
+
- **README sync**: any change to supported platforms or public usage MUST update `README.md` (add the provider to the Supported platforms table + a usage example). Keep the `## Project structure` tree above and the integration-test description in this file in sync too.
|
|
43
47
|
|
|
44
48
|
## Common commands
|
|
45
49
|
|
|
@@ -76,4 +80,19 @@ make typecheck # type check
|
|
|
76
80
|
**Issue:** `brew audit --strict --new Formula/md-fetch.rb` fails with missing system library declarations when `lxml` is a resource block.
|
|
77
81
|
**Root Cause:** `lxml` requires `libxml2` and `libxslt`, which are macOS system libraries. Homebrew requires these to be declared explicitly via `uses_from_macos`.
|
|
78
82
|
**Prevention Rule:** Any Homebrew formula that includes `lxml` as a resource MUST declare `uses_from_macos "libxml2"` and `uses_from_macos "libxslt"` after the `depends_on` lines. Discovered via `brew audit --strict --new` during implementation.
|
|
83
|
+
|
|
84
|
+
### ⚠️ boomi.com article detection relies on `div.post-content`, NOT `section.wysiwyg-section`
|
|
85
|
+
**Issue:** The Boomi blog index (`/blog/`) renders a `section.wysiwyg-section` intro but NO `div.post-content`. Selecting `wysiwyg-section` as the body container would fail to raise `UnsupportedContentTypeError` for the index and other non-article pages.
|
|
86
|
+
**Root Cause:** Confirmed via live DOM inspection: only article pages render `div.post-content` (containing `section.wysiwyg-section` + a `div.blog-nav` prev/next block); the index renders `wysiwyg-section` only.
|
|
87
|
+
**Prevention Rule:** `BoomiExtractor.clean_html()` MUST select `div.post-content` as the body container (its presence is the article discriminator) and strip the inner `div.blog-nav`. Do not switch to `wysiwyg-section`. The title `<h1>` lives in the page hero outside the body and is prepended.
|
|
88
|
+
|
|
89
|
+
### ⚠️ konghq.com is a Next.js site — pin to stable classes, NOT hashed CSS-module suffixes
|
|
90
|
+
**Issue:** Kong's per-component class names are Next.js CSS-module build hashes (e.g. `Section_section__Grz_Y`, `Article_toc__LOyCI`) that change between site builds. Selecting on a `__xxxxx` suffix would silently break on the next Kong deploy.
|
|
91
|
+
**Root Cause:** Confirmed via live DOM inspection: every block carries a hashed module class plus a stable, human-authored companion class. Only the companion classes are durable.
|
|
92
|
+
**Prevention Rule:** `KongExtractor.clean_html()` MUST use only stable classes: `main.type-article` is the **article discriminator** (absent on the blog index and category listings, which also have 0 `.rich-text-block`); the body is the `<section>` richest in `.rich-text-block`; strip in-body chrome via `.component.video`, `.component.more-on-this`, `.toc-wrap`, `.order-top`, and `.section-header-block:not(.intro)`. NOTE: `.toc-wrap` (not `[class*=TableOfContents]`) is the TOC sidebar — the `TableOfContents` component WRAPS the whole body, so matching it would delete the article. Also strip `.agent` spans everywhere: they are "agent mode" affordances that inject literal Markdown (`**`, `- `, `# `) duplicating the styled content. Title `<h1>` + the publication date (a class-less hero `<div>` matched by a month-name regex) are prepended; author bylines and read time are dropped.
|
|
93
|
+
|
|
94
|
+
### ⚠️ CLI `url` argument is `required=False` to support `--list-platforms`
|
|
95
|
+
**Issue:** `md-fetch --list-platforms` takes no URL, so the Click `url` argument is declared `required=False`. A future refactor that restores `required=True` (or drops the manual `url is None` check) would break `--list-platforms` and change the missing-URL exit code.
|
|
96
|
+
**Root Cause:** Click cannot express "required unless another flag is set"; requiredness is enforced manually in `main()` (`raise click.UsageError(...)` → exit 2) after the early `--list-platforms` return.
|
|
97
|
+
**Prevention Rule:** Keep `@click.argument("url", required=False)` and the manual `if url is None: raise click.UsageError(...)` guard. The list branch MUST `return` before any extraction so `--list-platforms` never makes a network call (FR-077). Use `router.supported_platforms()` (subdomain-aware) for the listing — NOT `supported_domains()`. [012-list-platforms]
|
|
79
98
|
<!-- SPECKIT END -->
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: mdfetch
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.8.0
|
|
4
4
|
Summary: Extract article content from web platforms and return it as clean Markdown.
|
|
5
5
|
Project-URL: Homepage, https://github.com/stn1slv/md-fetch
|
|
6
6
|
Project-URL: Source, https://github.com/stn1slv/md-fetch
|
|
@@ -61,6 +61,9 @@ md-fetch https://medium.com/example/article
|
|
|
61
61
|
|
|
62
62
|
# Fetch and save Markdown to a file
|
|
63
63
|
md-fetch https://dev.to/example/article --output article.md
|
|
64
|
+
|
|
65
|
+
# List all supported platforms (no URL or network needed)
|
|
66
|
+
md-fetch --list-platforms
|
|
64
67
|
```
|
|
65
68
|
|
|
66
69
|
## Python Usage
|
|
@@ -74,6 +77,8 @@ markdown = extract("https://dev.to/username/article-slug")
|
|
|
74
77
|
markdown = extract("https://example.substack.com/p/article-slug")
|
|
75
78
|
markdown = extract("https://thenewstack.io/article-slug")
|
|
76
79
|
markdown = extract("https://dzone.com/articles/article-slug")
|
|
80
|
+
markdown = extract("https://boomi.com/blog/article-slug")
|
|
81
|
+
markdown = extract("https://konghq.com/blog/category/article-slug")
|
|
77
82
|
print(markdown)
|
|
78
83
|
```
|
|
79
84
|
|
|
@@ -117,6 +122,8 @@ except EmptyContentError as e:
|
|
|
117
122
|
| Substack | `substack.com`, `*.substack.com` |
|
|
118
123
|
| The New Stack | `thenewstack.io` |
|
|
119
124
|
| DZone | `dzone.com` |
|
|
125
|
+
| Boomi | `boomi.com` |
|
|
126
|
+
| Kong | `konghq.com` |
|
|
120
127
|
|
|
121
128
|
## Development
|
|
122
129
|
|
|
@@ -28,6 +28,9 @@ md-fetch https://medium.com/example/article
|
|
|
28
28
|
|
|
29
29
|
# Fetch and save Markdown to a file
|
|
30
30
|
md-fetch https://dev.to/example/article --output article.md
|
|
31
|
+
|
|
32
|
+
# List all supported platforms (no URL or network needed)
|
|
33
|
+
md-fetch --list-platforms
|
|
31
34
|
```
|
|
32
35
|
|
|
33
36
|
## Python Usage
|
|
@@ -41,6 +44,8 @@ markdown = extract("https://dev.to/username/article-slug")
|
|
|
41
44
|
markdown = extract("https://example.substack.com/p/article-slug")
|
|
42
45
|
markdown = extract("https://thenewstack.io/article-slug")
|
|
43
46
|
markdown = extract("https://dzone.com/articles/article-slug")
|
|
47
|
+
markdown = extract("https://boomi.com/blog/article-slug")
|
|
48
|
+
markdown = extract("https://konghq.com/blog/category/article-slug")
|
|
44
49
|
print(markdown)
|
|
45
50
|
```
|
|
46
51
|
|
|
@@ -84,6 +89,8 @@ except EmptyContentError as e:
|
|
|
84
89
|
| Substack | `substack.com`, `*.substack.com` |
|
|
85
90
|
| The New Stack | `thenewstack.io` |
|
|
86
91
|
| DZone | `dzone.com` |
|
|
92
|
+
| Boomi | `boomi.com` |
|
|
93
|
+
| Kong | `konghq.com` |
|
|
87
94
|
|
|
88
95
|
## Development
|
|
89
96
|
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Specification Quality Checklist: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
|
4
|
+
**Created**: 2026-06-02
|
|
5
|
+
**Feature**: [spec.md](../spec.md)
|
|
6
|
+
|
|
7
|
+
## Content Quality
|
|
8
|
+
|
|
9
|
+
- [x] No implementation details (languages, frameworks, APIs)
|
|
10
|
+
- [x] Focused on user value and business needs
|
|
11
|
+
- [x] Written for non-technical stakeholders
|
|
12
|
+
- [x] All mandatory sections completed
|
|
13
|
+
|
|
14
|
+
## Requirement Completeness
|
|
15
|
+
|
|
16
|
+
- [x] No [NEEDS CLARIFICATION] markers remain
|
|
17
|
+
- [x] Requirements are testable and unambiguous
|
|
18
|
+
- [x] Success criteria are measurable
|
|
19
|
+
- [x] Success criteria are technology-agnostic (no implementation details)
|
|
20
|
+
- [x] All acceptance scenarios are defined
|
|
21
|
+
- [x] Edge cases are identified
|
|
22
|
+
- [x] Scope is clearly bounded
|
|
23
|
+
- [x] Dependencies and assumptions identified
|
|
24
|
+
|
|
25
|
+
## Feature Readiness
|
|
26
|
+
|
|
27
|
+
- [x] All functional requirements have clear acceptance criteria
|
|
28
|
+
- [x] User scenarios cover primary flows
|
|
29
|
+
- [x] Feature meets measurable outcomes defined in Success Criteria
|
|
30
|
+
- [x] No implementation details leak into specification
|
|
31
|
+
|
|
32
|
+
## Notes
|
|
33
|
+
|
|
34
|
+
- Items marked incomplete require spec updates before `/speckit-clarify` or `/speckit-plan`
|
|
35
|
+
- Note: FR-001 and the Key Entities section reference the existing provider/`@register` mechanism. This is an
|
|
36
|
+
intentional, established convention in this repository's spec template (see prior provider specs 002–007),
|
|
37
|
+
describing the routing contract rather than implementation internals. Retained for consistency.
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# Provider Contract: BoomiExtractor
|
|
2
|
+
|
|
3
|
+
This is a library; its external contracts are (1) the public `extract()` API and
|
|
4
|
+
(2) the `BaseExtractor` subclass contract the new provider must satisfy.
|
|
5
|
+
|
|
6
|
+
## 1. Public API (unchanged)
|
|
7
|
+
|
|
8
|
+
```python
|
|
9
|
+
from mdfetch import extract
|
|
10
|
+
|
|
11
|
+
markdown: str = extract("https://boomi.com/blog/<slug>/", retries=3, retry_delay=2.0)
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
- **Input**: a `boomi.com` blog article URL (`str`).
|
|
15
|
+
- **Output**: clean Markdown (`str`), title-first.
|
|
16
|
+
- **Raises**: `UnsupportedPlatformError` (domain not registered — N/A once Boomi is registered),
|
|
17
|
+
`UnsupportedContentTypeError`, `EmptyContentError`, `HTTPStatusError`, `FetchError`
|
|
18
|
+
(all from `mdfetch.exceptions`). No new exception types.
|
|
19
|
+
|
|
20
|
+
## 2. Subclass contract
|
|
21
|
+
|
|
22
|
+
```python
|
|
23
|
+
@register
|
|
24
|
+
class BoomiExtractor(BaseExtractor):
|
|
25
|
+
DOMAINS: frozenset[str] = frozenset({"boomi.com"})
|
|
26
|
+
# MATCH_SUBDOMAINS stays False (default)
|
|
27
|
+
|
|
28
|
+
def clean_html(self, soup: BeautifulSoup) -> Tag: ...
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
### `clean_html(soup)` obligations
|
|
32
|
+
|
|
33
|
+
| # | Obligation |
|
|
34
|
+
|---|------------|
|
|
35
|
+
| C1 | Return the `div.post-content` `Tag` with chrome stripped and the title prepended. |
|
|
36
|
+
| C2 | Raise `UnsupportedContentTypeError` when `div.post-content` is absent. |
|
|
37
|
+
| C3 | Decompose every `div.blog-nav` descendant before returning. |
|
|
38
|
+
| C4 | Prepend a copy of the page `<h1>` as the first child of the returned body. |
|
|
39
|
+
| C5 | MUST NOT modify the base class or shared utilities. |
|
|
40
|
+
| C6 | MUST NOT add image-specific logic (body images flow through default conversion). |
|
|
41
|
+
|
|
42
|
+
`extract()` and `convert_to_markdown()` are inherited unchanged; the latter enforces the
|
|
43
|
+
ATX-heading, blank-line-collapsing, and `EmptyContentError` behavior.
|
|
44
|
+
|
|
45
|
+
## 3. Acceptance contract (maps to spec SC-00x)
|
|
46
|
+
|
|
47
|
+
| Check | Maps to |
|
|
48
|
+
|-------|---------|
|
|
49
|
+
| Each reference article → Markdown begins with `# <title>` and contains body section headings + paragraph text; no chrome | SC-001, FR-002/003/004/005 |
|
|
50
|
+
| `extract("https://boomi.com/blog/")` (index) raises `UnsupportedContentTypeError` | SC-003, FR-006, FR-009 |
|
|
51
|
+
| Output has no 3+ consecutive blank lines | SC-004, FR-008 |
|
|
52
|
+
| ≥1 integration test hits a real reference URL with snapshot containment | SC-005 |
|
|
53
|
+
|
|
54
|
+
## 4. Routing contract
|
|
55
|
+
|
|
56
|
+
- After `@register`, `route("https://boomi.com/blog/...")` resolves to `BoomiExtractor`.
|
|
57
|
+
- `test_router.py` unsupported-domain fixture (`wordpress.com`) remains valid — no change needed.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# Phase 1 Data Model: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
This library is stateless; there are no persisted entities. The "data model" describes
|
|
4
|
+
the DOM structures the extractor reads and the in-memory shapes it produces.
|
|
5
|
+
|
|
6
|
+
## Entity: Boomi Blog Post (input DOM)
|
|
7
|
+
|
|
8
|
+
A single article page at `https://boomi.com/blog/<slug>/`.
|
|
9
|
+
|
|
10
|
+
| Field | Source selector | Notes |
|
|
11
|
+
|-------|-----------------|-------|
|
|
12
|
+
| title | `h1` (within `section.post-detail-hero`) | Exactly one per page; outside the body container |
|
|
13
|
+
| body | `div.post-content` | Body container; presence = "is an article" |
|
|
14
|
+
| content | `section.wysiwyg-section.bullet-styled` | Real article content (child of body) |
|
|
15
|
+
| nav (chrome) | `div.blog-nav` | Prev/next post links (child of body) — stripped |
|
|
16
|
+
| images | `img` inside `div.post-content` | Preserved (body-only, per clarification) |
|
|
17
|
+
| blockquotes | `blockquote` inside content | Preserved natively |
|
|
18
|
+
| subtitle/deck | — | Not present on any reference article |
|
|
19
|
+
|
|
20
|
+
**Validation rules**:
|
|
21
|
+
- `div.post-content` MUST exist → else `UnsupportedContentTypeError`.
|
|
22
|
+
- After stripping and conversion, Markdown MUST be non-empty → else `EmptyContentError`.
|
|
23
|
+
|
|
24
|
+
## Entity: Extraction Result (output)
|
|
25
|
+
|
|
26
|
+
A single Markdown `str` (the public `extract()` return value):
|
|
27
|
+
- Line 1: `# <title>` (top-level ATX heading).
|
|
28
|
+
- Followed by the converted body: headings, paragraphs, lists, blockquotes, links, and
|
|
29
|
+
any in-body images, in document order.
|
|
30
|
+
- No runs of 3+ consecutive blank lines (collapsed by base `convert_to_markdown`).
|
|
31
|
+
- Contains no site chrome (nav, language selector, CTAs, TOC, share buttons, blog-nav,
|
|
32
|
+
sidebar promos, footer).
|
|
33
|
+
|
|
34
|
+
## DOM Element Taxonomy (selector → action)
|
|
35
|
+
|
|
36
|
+
| Selector | Action | Reason |
|
|
37
|
+
|----------|--------|--------|
|
|
38
|
+
| `div.post-content` | **keep** (body root) | Article body; absence ⇒ non-article |
|
|
39
|
+
| `div.blog-nav` (inside body) | **strip** (`decompose`) | Prev/next post chrome |
|
|
40
|
+
| `h1` (hero) | **prepend** (copy into body) | Article title heading |
|
|
41
|
+
| everything outside `div.post-content` | **excluded** implicitly | Nav, TOC, share, sidebar promos, footer, hero image |
|
|
42
|
+
|
|
43
|
+
## State Transitions
|
|
44
|
+
|
|
45
|
+
None. Each `extract(url)` call is independent: `fetch → parse → clean → convert → return | raise`.
|