mdfetch 0.5.2__tar.gz → 0.6.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- mdfetch-0.6.0/.specify/feature.json +3 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/CLAUDE.md +8 -2
- {mdfetch-0.5.2 → mdfetch-0.6.0}/PKG-INFO +3 -1
- {mdfetch-0.5.2 → mdfetch-0.6.0}/README.md +2 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/pyproject.toml +1 -1
- mdfetch-0.6.0/specs/010-boomi-blog-provider/checklists/requirements.md +37 -0
- mdfetch-0.6.0/specs/010-boomi-blog-provider/contracts/extractor-contract.md +57 -0
- mdfetch-0.6.0/specs/010-boomi-blog-provider/data-model.md +45 -0
- mdfetch-0.6.0/specs/010-boomi-blog-provider/plan.md +115 -0
- mdfetch-0.6.0/specs/010-boomi-blog-provider/quickstart.md +47 -0
- mdfetch-0.6.0/specs/010-boomi-blog-provider/research.md +76 -0
- mdfetch-0.6.0/specs/010-boomi-blog-provider/spec.md +99 -0
- mdfetch-0.6.0/specs/010-boomi-blog-provider/tasks.md +177 -0
- mdfetch-0.6.0/src/mdfetch/providers/boomi.py +43 -0
- mdfetch-0.6.0/tests/integration/snapshots/boomi-data-consistency-saas-on-prem.md +29 -0
- mdfetch-0.6.0/tests/integration/snapshots/boomi-gartner-magic-quadrant-ipaas-2026.md +29 -0
- mdfetch-0.6.0/tests/integration/snapshots/boomi-real-time-vs-batch-data-integration.md +29 -0
- mdfetch-0.6.0/tests/integration/test_boomi_integration.py +56 -0
- mdfetch-0.6.0/tests/unit/test_boomi_extractor.py +138 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/uv.lock +1 -1
- mdfetch-0.5.2/.specify/feature.json +0 -3
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-analyze/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-archive-run/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-checklist/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-clarify/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-constitution/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-git-commit/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-git-feature/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-git-initialize/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-git-remote/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-git-validate/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-implement/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-plan/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-reconcile-run/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-specify/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-tasks/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.claude/skills/speckit-taskstoissues/SKILL.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.analyze.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.archive.run.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.checklist.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.clarify.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.constitution.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.implement.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.plan.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.reconcile.run.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.specify.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.tasks.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gemini/commands/speckit.taskstoissues.toml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gitattributes +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.github/copilot-instructions.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.github/workflows/ci.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.github/workflows/integration.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.github/workflows/publish.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.gitignore +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.python-version +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/.registry +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/archive/LICENSE +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/archive/README.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/archive/commands/archive.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/archive/extension.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/README.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/commands/speckit.git.commit.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/commands/speckit.git.feature.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/commands/speckit.git.initialize.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/commands/speckit.git.remote.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/commands/speckit.git.validate.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/config-template.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/extension.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/git-config.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/bash/auto-commit.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/bash/create-new-feature.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/bash/git-common.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/bash/initialize-repo.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/powershell/auto-commit.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/powershell/create-new-feature.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/powershell/git-common.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/git/scripts/powershell/initialize-repo.ps1 +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/reconcile/LICENSE +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/reconcile/README.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/reconcile/commands/reconcile.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions/reconcile/extension.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/extensions.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/init-options.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/integration.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/integrations/claude.manifest.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/integrations/gemini.manifest.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/integrations/speckit.manifest.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/memory/changelog.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/memory/constitution.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/memory/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/memory/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/scripts/bash/check-prerequisites.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/scripts/bash/common.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/scripts/bash/create-new-feature.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/scripts/bash/setup-plan.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/scripts/bash/setup-tasks.sh +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/templates/checklist-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/templates/constitution-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/templates/plan-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/templates/spec-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/templates/tasks-template.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/workflows/speckit/workflow.yml +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.specify/workflows/workflow-registry.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/.vscode/settings.json +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/GEMINI.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/LICENSE +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/Makefile +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/contracts/api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/001-mdfetch-medium-extractor/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/contracts/public-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/002-devto-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/003-medium-freedium-fallback/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/003-medium-freedium-fallback/contracts/extract-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/003-medium-freedium-fallback/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/003-medium-freedium-fallback/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/003-medium-freedium-fallback/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/003-medium-freedium-fallback/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/004-remove-backoff/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/004-remove-backoff/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/004-remove-backoff/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/004-remove-backoff/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/004-remove-backoff/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/contracts/extractor-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/005-substack-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/contracts/public-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/006-thenewstack-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/contracts/public-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/007-dzone-provider/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/contracts/cli.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/008-mdfetch-cli/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/checklists/requirements.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/contracts/formula.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/contracts/tap-update-job.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/data-model.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/plan.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/quickstart.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/research.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/spec.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/specs/009-homebrew-tap-formula/tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/base.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/cli.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/exceptions.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/providers/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/providers/devto.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/providers/dzone.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/providers/medium.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/providers/substack.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/providers/thenewstack.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/src/mdfetch/router.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/conftest.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/conftest.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/architecting-the-asynchronous-agent.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/devto-integration-digest-december-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/devto-integration-digest-july-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/devto-integration-digest-march-2026.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/dzone-image-classification-pipeline-camel-djl.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/dzone-integration-patterns-fail-production.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/dzone-kiro-feature-to-requirements-design-tasks.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/from-drift-to-parity.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/integration-digest-december-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/substack-api-trends-2025.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/substack-kafka-topic-types.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/thenewstack-api-mcp-agent.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/thenewstack-async-apis.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/thenewstack-developer-portal-api.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/thenewstack-json-schema-ai.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/snapshots/thenewstack-mcp-api-governance.md +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/test_cli_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/test_devto_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/test_dzone_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/test_medium_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/test_substack_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/integration/test_thenewstack_integration.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/__init__.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_cli.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_devto_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_dzone_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_fetch_errors.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_medium_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_router.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_silent.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_substack_extractor.py +0 -0
- {mdfetch-0.5.2 → mdfetch-0.6.0}/tests/unit/test_thenewstack_extractor.py +0 -0
|
@@ -17,11 +17,12 @@ src/mdfetch/
|
|
|
17
17
|
├── devto.py # DevToExtractor (dev.to)
|
|
18
18
|
├── substack.py # SubstackExtractor (substack.com + *.substack.com)
|
|
19
19
|
├── thenewstack.py # TheNewStackExtractor (thenewstack.io)
|
|
20
|
-
|
|
20
|
+
├── dzone.py # DZoneExtractor (dzone.com)
|
|
21
|
+
└── boomi.py # BoomiExtractor (boomi.com/blog)
|
|
21
22
|
|
|
22
23
|
tests/
|
|
23
24
|
├── unit/ # pytest unit tests (no network)
|
|
24
|
-
└── integration/ # real network tests (Medium + dev.to + Substack + TheNewStack URLs + snapshots)
|
|
25
|
+
└── integration/ # real network tests (Medium + dev.to + Substack + TheNewStack + DZone + Boomi URLs + snapshots)
|
|
25
26
|
|
|
26
27
|
.github/workflows/
|
|
27
28
|
├── ci.yml # lint + unit tests on push/PR (Python 3.12–3.14)
|
|
@@ -76,4 +77,9 @@ make typecheck # type check
|
|
|
76
77
|
**Issue:** `brew audit --strict --new Formula/md-fetch.rb` fails with missing system library declarations when `lxml` is a resource block.
|
|
77
78
|
**Root Cause:** `lxml` requires `libxml2` and `libxslt`, which are macOS system libraries. Homebrew requires these to be declared explicitly via `uses_from_macos`.
|
|
78
79
|
**Prevention Rule:** Any Homebrew formula that includes `lxml` as a resource MUST declare `uses_from_macos "libxml2"` and `uses_from_macos "libxslt"` after the `depends_on` lines. Discovered via `brew audit --strict --new` during implementation.
|
|
80
|
+
|
|
81
|
+
### ⚠️ boomi.com article detection relies on `div.post-content`, NOT `section.wysiwyg-section`
|
|
82
|
+
**Issue:** The Boomi blog index (`/blog/`) renders a `section.wysiwyg-section` intro but NO `div.post-content`. Selecting `wysiwyg-section` as the body container would fail to raise `UnsupportedContentTypeError` for the index and other non-article pages.
|
|
83
|
+
**Root Cause:** Confirmed via live DOM inspection: only article pages render `div.post-content` (containing `section.wysiwyg-section` + a `div.blog-nav` prev/next block); the index renders `wysiwyg-section` only.
|
|
84
|
+
**Prevention Rule:** `BoomiExtractor.clean_html()` MUST select `div.post-content` as the body container (its presence is the article discriminator) and strip the inner `div.blog-nav`. Do not switch to `wysiwyg-section`. The title `<h1>` lives in the page hero outside the body and is prepended.
|
|
79
85
|
<!-- SPECKIT END -->
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: mdfetch
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.6.0
|
|
4
4
|
Summary: Extract article content from web platforms and return it as clean Markdown.
|
|
5
5
|
Project-URL: Homepage, https://github.com/stn1slv/md-fetch
|
|
6
6
|
Project-URL: Source, https://github.com/stn1slv/md-fetch
|
|
@@ -74,6 +74,7 @@ markdown = extract("https://dev.to/username/article-slug")
|
|
|
74
74
|
markdown = extract("https://example.substack.com/p/article-slug")
|
|
75
75
|
markdown = extract("https://thenewstack.io/article-slug")
|
|
76
76
|
markdown = extract("https://dzone.com/articles/article-slug")
|
|
77
|
+
markdown = extract("https://boomi.com/blog/article-slug")
|
|
77
78
|
print(markdown)
|
|
78
79
|
```
|
|
79
80
|
|
|
@@ -117,6 +118,7 @@ except EmptyContentError as e:
|
|
|
117
118
|
| Substack | `substack.com`, `*.substack.com` |
|
|
118
119
|
| The New Stack | `thenewstack.io` |
|
|
119
120
|
| DZone | `dzone.com` |
|
|
121
|
+
| Boomi | `boomi.com` |
|
|
120
122
|
|
|
121
123
|
## Development
|
|
122
124
|
|
|
@@ -41,6 +41,7 @@ markdown = extract("https://dev.to/username/article-slug")
|
|
|
41
41
|
markdown = extract("https://example.substack.com/p/article-slug")
|
|
42
42
|
markdown = extract("https://thenewstack.io/article-slug")
|
|
43
43
|
markdown = extract("https://dzone.com/articles/article-slug")
|
|
44
|
+
markdown = extract("https://boomi.com/blog/article-slug")
|
|
44
45
|
print(markdown)
|
|
45
46
|
```
|
|
46
47
|
|
|
@@ -84,6 +85,7 @@ except EmptyContentError as e:
|
|
|
84
85
|
| Substack | `substack.com`, `*.substack.com` |
|
|
85
86
|
| The New Stack | `thenewstack.io` |
|
|
86
87
|
| DZone | `dzone.com` |
|
|
88
|
+
| Boomi | `boomi.com` |
|
|
87
89
|
|
|
88
90
|
## Development
|
|
89
91
|
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Specification Quality Checklist: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
|
4
|
+
**Created**: 2026-06-02
|
|
5
|
+
**Feature**: [spec.md](../spec.md)
|
|
6
|
+
|
|
7
|
+
## Content Quality
|
|
8
|
+
|
|
9
|
+
- [x] No implementation details (languages, frameworks, APIs)
|
|
10
|
+
- [x] Focused on user value and business needs
|
|
11
|
+
- [x] Written for non-technical stakeholders
|
|
12
|
+
- [x] All mandatory sections completed
|
|
13
|
+
|
|
14
|
+
## Requirement Completeness
|
|
15
|
+
|
|
16
|
+
- [x] No [NEEDS CLARIFICATION] markers remain
|
|
17
|
+
- [x] Requirements are testable and unambiguous
|
|
18
|
+
- [x] Success criteria are measurable
|
|
19
|
+
- [x] Success criteria are technology-agnostic (no implementation details)
|
|
20
|
+
- [x] All acceptance scenarios are defined
|
|
21
|
+
- [x] Edge cases are identified
|
|
22
|
+
- [x] Scope is clearly bounded
|
|
23
|
+
- [x] Dependencies and assumptions identified
|
|
24
|
+
|
|
25
|
+
## Feature Readiness
|
|
26
|
+
|
|
27
|
+
- [x] All functional requirements have clear acceptance criteria
|
|
28
|
+
- [x] User scenarios cover primary flows
|
|
29
|
+
- [x] Feature meets measurable outcomes defined in Success Criteria
|
|
30
|
+
- [x] No implementation details leak into specification
|
|
31
|
+
|
|
32
|
+
## Notes
|
|
33
|
+
|
|
34
|
+
- Items marked incomplete require spec updates before `/speckit-clarify` or `/speckit-plan`
|
|
35
|
+
- Note: FR-001 and the Key Entities section reference the existing provider/`@register` mechanism. This is an
|
|
36
|
+
intentional, established convention in this repository's spec template (see prior provider specs 002–007),
|
|
37
|
+
describing the routing contract rather than implementation internals. Retained for consistency.
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# Provider Contract: BoomiExtractor
|
|
2
|
+
|
|
3
|
+
This is a library; its external contracts are (1) the public `extract()` API and
|
|
4
|
+
(2) the `BaseExtractor` subclass contract the new provider must satisfy.
|
|
5
|
+
|
|
6
|
+
## 1. Public API (unchanged)
|
|
7
|
+
|
|
8
|
+
```python
|
|
9
|
+
from mdfetch import extract
|
|
10
|
+
|
|
11
|
+
markdown: str = extract("https://boomi.com/blog/<slug>/", retries=3, retry_delay=2.0)
|
|
12
|
+
```
|
|
13
|
+
|
|
14
|
+
- **Input**: a `boomi.com` blog article URL (`str`).
|
|
15
|
+
- **Output**: clean Markdown (`str`), title-first.
|
|
16
|
+
- **Raises**: `UnsupportedPlatformError` (domain not registered — N/A once Boomi is registered),
|
|
17
|
+
`UnsupportedContentTypeError`, `EmptyContentError`, `HTTPStatusError`, `FetchError`
|
|
18
|
+
(all from `mdfetch.exceptions`). No new exception types.
|
|
19
|
+
|
|
20
|
+
## 2. Subclass contract
|
|
21
|
+
|
|
22
|
+
```python
|
|
23
|
+
@register
|
|
24
|
+
class BoomiExtractor(BaseExtractor):
|
|
25
|
+
DOMAINS: frozenset[str] = frozenset({"boomi.com"})
|
|
26
|
+
# MATCH_SUBDOMAINS stays False (default)
|
|
27
|
+
|
|
28
|
+
def clean_html(self, soup: BeautifulSoup) -> Tag: ...
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
### `clean_html(soup)` obligations
|
|
32
|
+
|
|
33
|
+
| # | Obligation |
|
|
34
|
+
|---|------------|
|
|
35
|
+
| C1 | Return the `div.post-content` `Tag` with chrome stripped and the title prepended. |
|
|
36
|
+
| C2 | Raise `UnsupportedContentTypeError` when `div.post-content` is absent. |
|
|
37
|
+
| C3 | Decompose every `div.blog-nav` descendant before returning. |
|
|
38
|
+
| C4 | Prepend a copy of the page `<h1>` as the first child of the returned body. |
|
|
39
|
+
| C5 | MUST NOT modify the base class or shared utilities. |
|
|
40
|
+
| C6 | MUST NOT add image-specific logic (body images flow through default conversion). |
|
|
41
|
+
|
|
42
|
+
`extract()` and `convert_to_markdown()` are inherited unchanged; the latter enforces the
|
|
43
|
+
ATX-heading, blank-line-collapsing, and `EmptyContentError` behavior.
|
|
44
|
+
|
|
45
|
+
## 3. Acceptance contract (maps to spec SC-00x)
|
|
46
|
+
|
|
47
|
+
| Check | Maps to |
|
|
48
|
+
|-------|---------|
|
|
49
|
+
| Each reference article → Markdown begins with `# <title>` and contains body section headings + paragraph text; no chrome | SC-001, FR-002/003/004/005 |
|
|
50
|
+
| `extract("https://boomi.com/blog/")` (index) raises `UnsupportedContentTypeError` | SC-003, FR-006, FR-009 |
|
|
51
|
+
| Output has no 3+ consecutive blank lines | SC-004, FR-008 |
|
|
52
|
+
| ≥1 integration test hits a real reference URL with snapshot containment | SC-005 |
|
|
53
|
+
|
|
54
|
+
## 4. Routing contract
|
|
55
|
+
|
|
56
|
+
- After `@register`, `route("https://boomi.com/blog/...")` resolves to `BoomiExtractor`.
|
|
57
|
+
- `test_router.py` unsupported-domain fixture (`wordpress.com`) remains valid — no change needed.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# Phase 1 Data Model: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
This library is stateless; there are no persisted entities. The "data model" describes
|
|
4
|
+
the DOM structures the extractor reads and the in-memory shapes it produces.
|
|
5
|
+
|
|
6
|
+
## Entity: Boomi Blog Post (input DOM)
|
|
7
|
+
|
|
8
|
+
A single article page at `https://boomi.com/blog/<slug>/`.
|
|
9
|
+
|
|
10
|
+
| Field | Source selector | Notes |
|
|
11
|
+
|-------|-----------------|-------|
|
|
12
|
+
| title | `h1` (within `section.post-detail-hero`) | Exactly one per page; outside the body container |
|
|
13
|
+
| body | `div.post-content` | Body container; presence = "is an article" |
|
|
14
|
+
| content | `section.wysiwyg-section.bullet-styled` | Real article content (child of body) |
|
|
15
|
+
| nav (chrome) | `div.blog-nav` | Prev/next post links (child of body) — stripped |
|
|
16
|
+
| images | `img` inside `div.post-content` | Preserved (body-only, per clarification) |
|
|
17
|
+
| blockquotes | `blockquote` inside content | Preserved natively |
|
|
18
|
+
| subtitle/deck | — | Not present on any reference article |
|
|
19
|
+
|
|
20
|
+
**Validation rules**:
|
|
21
|
+
- `div.post-content` MUST exist → else `UnsupportedContentTypeError`.
|
|
22
|
+
- After stripping and conversion, Markdown MUST be non-empty → else `EmptyContentError`.
|
|
23
|
+
|
|
24
|
+
## Entity: Extraction Result (output)
|
|
25
|
+
|
|
26
|
+
A single Markdown `str` (the public `extract()` return value):
|
|
27
|
+
- Line 1: `# <title>` (top-level ATX heading).
|
|
28
|
+
- Followed by the converted body: headings, paragraphs, lists, blockquotes, links, and
|
|
29
|
+
any in-body images, in document order.
|
|
30
|
+
- No runs of 3+ consecutive blank lines (collapsed by base `convert_to_markdown`).
|
|
31
|
+
- Contains no site chrome (nav, language selector, CTAs, TOC, share buttons, blog-nav,
|
|
32
|
+
sidebar promos, footer).
|
|
33
|
+
|
|
34
|
+
## DOM Element Taxonomy (selector → action)
|
|
35
|
+
|
|
36
|
+
| Selector | Action | Reason |
|
|
37
|
+
|----------|--------|--------|
|
|
38
|
+
| `div.post-content` | **keep** (body root) | Article body; absence ⇒ non-article |
|
|
39
|
+
| `div.blog-nav` (inside body) | **strip** (`decompose`) | Prev/next post chrome |
|
|
40
|
+
| `h1` (hero) | **prepend** (copy into body) | Article title heading |
|
|
41
|
+
| everything outside `div.post-content` | **excluded** implicitly | Nav, TOC, share, sidebar promos, footer, hero image |
|
|
42
|
+
|
|
43
|
+
## State Transitions
|
|
44
|
+
|
|
45
|
+
None. Each `extract(url)` call is independent: `fetch → parse → clean → convert → return | raise`.
|
|
@@ -0,0 +1,115 @@
|
|
|
1
|
+
# Implementation Plan: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
**Branch**: `010-boomi-blog-provider` | **Date**: 2026-06-02 | **Spec**: [spec.md](./spec.md)
|
|
4
|
+
|
|
5
|
+
**Input**: Feature specification from `/specs/010-boomi-blog-provider/spec.md`
|
|
6
|
+
|
|
7
|
+
## Summary
|
|
8
|
+
|
|
9
|
+
Add a new provider that extracts public Boomi blog articles (`boomi.com/blog/<slug>/`) and returns them as clean Markdown. The site runs WordPress; all three reference articles share an identical, stable structure: the article body is `div.post-content`, which contains the real content (`section.wysiwyg-section`) plus a previous/next post navigation block (`div.blog-nav`) to strip. The article title is an `<h1>` rendered in the page hero (`section.post-detail-hero`), outside the body. The blog index, non-article pages, and other `boomi.com` pages do **not** contain `div.post-content`, so its absence cleanly signals "not an article" → `UnsupportedContentTypeError`. The implementation is a single new provider file subclassing `BaseExtractor`, reusing all shared fetch/convert logic.
|
|
10
|
+
|
|
11
|
+
## Technical Context
|
|
12
|
+
|
|
13
|
+
**Language/Version**: Python 3.12+ (matches CI matrix: 3.12–3.14)
|
|
14
|
+
|
|
15
|
+
**Primary Dependencies**: `httpx` (HTTP fetch), `BeautifulSoup` / `lxml` (HTML parsing), `markdownify` (Markdown conversion), `pytest` (testing)
|
|
16
|
+
|
|
17
|
+
**Storage**: N/A — stateless extraction library
|
|
18
|
+
|
|
19
|
+
**Testing**: `pytest` via `uv run pytest` — unit tests (no network) + integration tests (`-m integration`, real URLs + snapshots)
|
|
20
|
+
|
|
21
|
+
**Target Platform**: PyPI library (cross-platform)
|
|
22
|
+
|
|
23
|
+
**Project Type**: Library
|
|
24
|
+
|
|
25
|
+
**Performance Goals**: Inherits base class 30-second fetch timeout; no additional targets
|
|
26
|
+
|
|
27
|
+
**Constraints**: One new provider file; no changes to shared infrastructure
|
|
28
|
+
|
|
29
|
+
**Scale/Scope**: Single-article extraction per call
|
|
30
|
+
|
|
31
|
+
## Constitution Check
|
|
32
|
+
|
|
33
|
+
*GATE: Must pass before implementation. Re-check after design phase.*
|
|
34
|
+
|
|
35
|
+
- [x] Validates Provider Pattern Architecture (new `BoomiExtractor(BaseExtractor)`; no base-class changes; no code duplication — reuses `fetch_html`/`convert_to_markdown`)
|
|
36
|
+
- [x] Confirms Technology Stack (`httpx`, `BeautifulSoup`, `Markdownify`, `pytest` — all inherited)
|
|
37
|
+
- [x] Adheres to Coding Standards (PEP 8, strict type hints, clear vocabulary)
|
|
38
|
+
- [x] Incorporates Integration Testing (real reference URLs + snapshot containment, matching existing providers)
|
|
39
|
+
- [x] Respects Packaging and Distribution standards (`pyproject.toml`, `src/` layout, `uv` for all dev commands)
|
|
40
|
+
|
|
41
|
+
**Result**: PASS — no violations. Complexity Tracking not required.
|
|
42
|
+
|
|
43
|
+
## Project Structure
|
|
44
|
+
|
|
45
|
+
### Documentation (this feature)
|
|
46
|
+
|
|
47
|
+
```text
|
|
48
|
+
specs/010-boomi-blog-provider/
|
|
49
|
+
├── plan.md # This file
|
|
50
|
+
├── research.md # Phase 0 output
|
|
51
|
+
├── data-model.md # Phase 1 output
|
|
52
|
+
├── quickstart.md # Phase 1 output
|
|
53
|
+
├── contracts/
|
|
54
|
+
│ └── extractor-contract.md # Phase 1 output (provider contract)
|
|
55
|
+
└── tasks.md # Phase 2 output (/speckit-tasks — NOT created here)
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
### Source Code
|
|
59
|
+
|
|
60
|
+
```text
|
|
61
|
+
src/mdfetch/providers/
|
|
62
|
+
└── boomi.py # NEW: BoomiExtractor
|
|
63
|
+
|
|
64
|
+
tests/unit/
|
|
65
|
+
└── test_boomi_extractor.py # NEW: unit tests (no network)
|
|
66
|
+
|
|
67
|
+
tests/integration/
|
|
68
|
+
├── snapshots/
|
|
69
|
+
│ ├── boomi-gartner-magic-quadrant-ipaas-2026.md # NEW
|
|
70
|
+
│ ├── boomi-real-time-vs-batch-data-integration.md # NEW
|
|
71
|
+
│ └── boomi-data-consistency-saas-on-prem.md # NEW
|
|
72
|
+
└── test_boomi_integration.py # NEW: integration tests (real URLs)
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
**Non-runtime changes**: `README.md` (add Boomi row to Supported platforms table + usage example); `pyproject.toml` (version bump `0.5.2` → `0.6.0`, new provider = minor per dev.to/Substack precedent); `CLAUDE.md` (provider tree + integration-test description + Boomi discriminator gotcha). No `test_router.py` fixture change required — the unsupported-domain fixture uses `wordpress.com`, which Boomi does not register.
|
|
76
|
+
|
|
77
|
+
### Revision: Implementation Sync 2026-06-02
|
|
78
|
+
- Reason: Reconciled release/docs drift found after implementation — version was not bumped and `CLAUDE.md` (project structure, integration-test description, gotchas) did not reflect the new Boomi provider. No behavioral drift; extraction output verified clean against all 3 reference articles.
|
|
79
|
+
|
|
80
|
+
## Extraction Algorithm
|
|
81
|
+
|
|
82
|
+
```
|
|
83
|
+
extract(url): # inherited from BaseExtractor
|
|
84
|
+
html ← fetch_html(url) # 30s timeout, 3 retries on transient errors
|
|
85
|
+
soup ← BeautifulSoup(html, "lxml")
|
|
86
|
+
body_tag ← clean_html(soup)
|
|
87
|
+
return convert_to_markdown(body_tag) # ATX headings, collapse 3+ blank lines, EmptyContentError if blank
|
|
88
|
+
|
|
89
|
+
clean_html(soup): # BoomiExtractor implementation
|
|
90
|
+
1. body ← soup.find("div", class_="post-content")
|
|
91
|
+
→ if not a Tag: raise UnsupportedContentTypeError # index / non-article / non-blog page
|
|
92
|
+
2. for nav in body.find_all("div", class_="blog-nav"): nav.decompose() # strip prev/next chrome
|
|
93
|
+
3. title ← soup.find("h1") # hero title, outside body
|
|
94
|
+
→ if Tag: body.insert(0, copy.copy(title)) # prepend as top-level heading
|
|
95
|
+
4. return body
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
**Design notes**:
|
|
99
|
+
- `div.post-content` is the *discriminating* selector: present on every article, absent on the blog index and non-article pages (verified live across 3 articles + the `/blog/` index). This is why no URL-path filtering is needed (FR-009 satisfied via body-container presence, consistent with the existing route-by-domain pattern).
|
|
100
|
+
- Body images live inside `post-content` and are preserved by default markdownify (Option A clarification). The hero/banner image sits in `section.post-detail-hero` — outside the body — and is therefore excluded automatically. No image-specific code is needed.
|
|
101
|
+
- Blockquotes (e.g., the Gartner pull-quote) are inside `wysiwyg-section` and convert natively. No embed/iframe handling observed in references; if present, base `_replace_iframes_with_links` is available but not wired in unless a future article requires it.
|
|
102
|
+
- `MATCH_SUBDOMAINS` stays `False` (default): only `boomi.com` is in scope.
|
|
103
|
+
|
|
104
|
+
## Error Mapping
|
|
105
|
+
|
|
106
|
+
| Condition | Exception |
|
|
107
|
+
|-----------|-----------|
|
|
108
|
+
| `div.post-content` not found (index / non-article / non-blog page) | `UnsupportedContentTypeError` |
|
|
109
|
+
| Body found but no extractable text after conversion | `EmptyContentError` |
|
|
110
|
+
| HTTP error (any non-2xx after retries) | `HTTPStatusError` |
|
|
111
|
+
| Network / timeout failure | `FetchError` |
|
|
112
|
+
|
|
113
|
+
## Complexity Tracking
|
|
114
|
+
|
|
115
|
+
> No Constitution Check violations — section intentionally empty.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Quickstart: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
## Use it
|
|
4
|
+
|
|
5
|
+
```python
|
|
6
|
+
from mdfetch import extract
|
|
7
|
+
|
|
8
|
+
md = extract("https://boomi.com/blog/real-time-vs-batch-data-integration-choosing-the-right-approach/")
|
|
9
|
+
print(md) # "# Real-Time vs Batch Data Integration: ..." followed by the article body
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
## Develop it (TDD-friendly order)
|
|
13
|
+
|
|
14
|
+
```bash
|
|
15
|
+
make setup # uv sync --all-extras
|
|
16
|
+
|
|
17
|
+
# 1. Add src/mdfetch/providers/boomi.py (BoomiExtractor) — see plan.md "Extraction Algorithm"
|
|
18
|
+
# 2. Write tests/unit/test_boomi_extractor.py (no network) using sample HTML fixtures
|
|
19
|
+
make test # unit tests only — should pass without network
|
|
20
|
+
make typecheck # mypy src/ — zero errors
|
|
21
|
+
make lint # ruff check
|
|
22
|
+
|
|
23
|
+
# 3. Add integration test + snapshots
|
|
24
|
+
make integration # network required; hits real boomi.com URLs
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
## Reference URLs
|
|
28
|
+
|
|
29
|
+
| Slug | Use |
|
|
30
|
+
|------|-----|
|
|
31
|
+
| `gartner-magic-quadrant-ipaas-2026` | has 1 in-body image + blockquote |
|
|
32
|
+
| `real-time-vs-batch-data-integration-choosing-the-right-approach` | headings + lists |
|
|
33
|
+
| `how-to-maintain-data-consistency-across-saas-and-on-prem-systems` | headings + blockquote |
|
|
34
|
+
| `https://boomi.com/blog/` (index) | non-article → expect `UnsupportedContentTypeError` |
|
|
35
|
+
|
|
36
|
+
## Generate a snapshot (first 30 lines, blank lines preserved)
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
uv run python -c "from mdfetch import extract; c=extract('<url>'); \
|
|
40
|
+
open('tests/integration/snapshots/boomi-<slug>.md','w',encoding='utf-8').write('\n'.join(c.split('\n')[:30]).rstrip())"
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
## Done when
|
|
44
|
+
|
|
45
|
+
- `make test`, `make typecheck`, `make lint` all green.
|
|
46
|
+
- `make integration` passes (snapshot containment for the 3 articles; index raises `UnsupportedContentTypeError`).
|
|
47
|
+
- `README.md` Supported platforms table includes a `Boomi | boomi.com` row.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Phase 0 Research: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
All findings below come from live DOM inspection (2026-06-02) of the three reference
|
|
4
|
+
articles plus the `/blog/` index, using the project's browser-like User-Agent.
|
|
5
|
+
|
|
6
|
+
## Decision: Article body container = `div.post-content`
|
|
7
|
+
|
|
8
|
+
- **Decision**: Isolate the article body by selecting `soup.find("div", class_="post-content")`.
|
|
9
|
+
- **Rationale**: Present on every article page; its direct children are exactly
|
|
10
|
+
`section.wysiwyg-section.bullet-styled` (the real content) and `div.blog-nav`
|
|
11
|
+
(prev/next chrome). Crucially, the `/blog/` **index page does NOT contain
|
|
12
|
+
`div.post-content`** (it has a `wysiwyg-section` intro but no `post-content`),
|
|
13
|
+
so the selector doubles as the article-vs-non-article discriminator.
|
|
14
|
+
- **Alternatives considered**:
|
|
15
|
+
- `section.wysiwyg-section` — **rejected**: also present on the blog index, so it
|
|
16
|
+
would fail to raise `UnsupportedContentTypeError` for non-articles.
|
|
17
|
+
- `<article>` / `<main>` landmarks — **rejected**: not emitted with usable classes
|
|
18
|
+
in this WordPress theme.
|
|
19
|
+
|
|
20
|
+
## Decision: Strip `div.blog-nav`
|
|
21
|
+
|
|
22
|
+
- **Decision**: `decompose()` every `div.blog-nav` inside the body before conversion.
|
|
23
|
+
- **Rationale**: It is the only chrome element living *inside* `post-content`; it holds
|
|
24
|
+
the "Previous / Next" post links. All other chrome (top nav, language selector,
|
|
25
|
+
login/demo CTAs, "On this page" TOC, social share, sidebar report promos, footer)
|
|
26
|
+
lives **outside** `post-content` and is excluded automatically by isolating the body.
|
|
27
|
+
- **Alternatives considered**: Stripping each chrome class individually — **rejected** as
|
|
28
|
+
unnecessary; isolating `post-content` already removes everything except `blog-nav`.
|
|
29
|
+
|
|
30
|
+
## Decision: Title from hero `<h1>`, prepended to body
|
|
31
|
+
|
|
32
|
+
- **Decision**: `soup.find("h1")` and prepend a copy to the body as the top-level heading.
|
|
33
|
+
- **Rationale**: The title `<h1>` is rendered in `section.post-detail-hero` (the page hero),
|
|
34
|
+
structurally outside `post-content`, so an unconditional prepend includes it exactly once
|
|
35
|
+
with no body-scan deduplication (same approach as `SubstackExtractor`). There is exactly
|
|
36
|
+
one `<h1>` per page. No subtitle/deck element is present on any reference article.
|
|
37
|
+
- **Alternatives considered**: `og:title` meta — **rejected**: prefer the on-page `<h1>` for
|
|
38
|
+
consistency with other providers and to avoid meta/visible-title drift.
|
|
39
|
+
|
|
40
|
+
## Decision: Image handling — body-only (Option A)
|
|
41
|
+
|
|
42
|
+
- **Decision**: Convert only images inside `post-content`; do not hoist the hero image.
|
|
43
|
+
- **Rationale**: Matches the spec clarification (2026-06-02). The hero/banner image lives in
|
|
44
|
+
`section.post-detail-hero` (outside the body) and is excluded automatically; genuine inline
|
|
45
|
+
content images (e.g., 1 image inside reference #1's body) are preserved by default markdownify.
|
|
46
|
+
No image-specific code required.
|
|
47
|
+
- **Evidence**: img count inside `post-content` — ref#1 (gartner): 1; ref#2 (real-time-vs-batch): 0;
|
|
48
|
+
ref#3 (data-consistency): 0. All have 1 blockquote (Gartner / pull-quotes), converted natively.
|
|
49
|
+
|
|
50
|
+
## Decision: No URL path filtering; route by domain only
|
|
51
|
+
|
|
52
|
+
- **Decision**: Register `DOMAINS = frozenset({"boomi.com"})`, `MATCH_SUBDOMAINS = False`.
|
|
53
|
+
Rely on `post-content` presence to reject non-`/blog/` pages.
|
|
54
|
+
- **Rationale**: Consistent with the codebase's route-by-domain + validate-body pattern. Non-blog
|
|
55
|
+
`boomi.com` pages and the blog index lack `post-content` → `UnsupportedContentTypeError`,
|
|
56
|
+
satisfying FR-009 without bespoke path logic.
|
|
57
|
+
|
|
58
|
+
## Decision: No new infrastructure / markdownify overrides
|
|
59
|
+
|
|
60
|
+
- **Decision**: Inherit `fetch_html`, `convert_to_markdown`, and default markdownify kwargs.
|
|
61
|
+
- **Rationale**: No code blocks observed in references (so no `code_language_callback` like DZone),
|
|
62
|
+
no embeds requiring link conversion. Keeps the provider minimal per the Provider Pattern principle.
|
|
63
|
+
|
|
64
|
+
## Platform facts summary
|
|
65
|
+
|
|
66
|
+
| Fact | Value |
|
|
67
|
+
|------|-------|
|
|
68
|
+
| CMS | WordPress 6.9.4 |
|
|
69
|
+
| Domain | `boomi.com` (articles under `/blog/<slug>/`) |
|
|
70
|
+
| Body container | `div.post-content` |
|
|
71
|
+
| Real content | `section.wysiwyg-section.bullet-styled` |
|
|
72
|
+
| In-body chrome | `div.blog-nav` (prev/next) |
|
|
73
|
+
| Title | single `<h1>` in `section.post-detail-hero` (outside body) |
|
|
74
|
+
| Subtitle/deck | none |
|
|
75
|
+
| Paywall | none (freely readable) |
|
|
76
|
+
| Index page has `post-content`? | No (→ clean non-article detection) |
|
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
# Feature Specification: Boomi Blog Provider
|
|
2
|
+
|
|
3
|
+
**Feature Branch**: `010-boomi-blog-provider`
|
|
4
|
+
|
|
5
|
+
**Created**: 2026-06-02
|
|
6
|
+
|
|
7
|
+
**Status**: Draft
|
|
8
|
+
|
|
9
|
+
**Input**: User description: "I would like to add support of the articles from https://boomi.com/blog/. Please use the following ones as references: 1. https://boomi.com/blog/gartner-magic-quadrant-ipaas-2026/ 2. https://boomi.com/blog/real-time-vs-batch-data-integration-choosing-the-right-approach/ 3. https://boomi.com/blog/how-to-maintain-data-consistency-across-saas-and-on-prem-systems/"
|
|
10
|
+
|
|
11
|
+
## User Scenarios & Testing *(mandatory)*
|
|
12
|
+
|
|
13
|
+
### User Story 1 - Extract a Boomi blog article as clean Markdown (Priority: P1)
|
|
14
|
+
|
|
15
|
+
A user passes the URL of a public Boomi blog article to the library and receives the article's title and full body rendered as clean Markdown, with all site chrome (navigation, CTAs, share buttons, related-post links, footer) removed.
|
|
16
|
+
|
|
17
|
+
**Why this priority**: This is the core value of the feature — without it, Boomi blog articles cannot be consumed at all. It is the minimum viable slice and delivers the complete user benefit on its own.
|
|
18
|
+
|
|
19
|
+
**Independent Test**: Can be fully tested by calling `extract()` with one of the three reference URLs and verifying the returned Markdown begins with the article title as a top-level heading and contains the body's section headings and paragraph text, with no navigation, subscription/demo CTAs, or footer text.
|
|
20
|
+
|
|
21
|
+
**Acceptance Scenarios**:
|
|
22
|
+
|
|
23
|
+
1. **Given** a public Boomi blog article URL, **When** `extract()` is called, **Then** the result is Markdown whose first line is `# <article title>` followed by the body content.
|
|
24
|
+
2. **Given** a Boomi blog article containing H2 sections and bullet lists, **When** extracted, **Then** the headings and lists are preserved as Markdown headings and lists in document order.
|
|
25
|
+
3. **Given** a Boomi blog article, **When** extracted, **Then** the output contains none of: top navigation, language selector, login/demo CTAs, "on this page" jump links, social share buttons, related/previous/next post links, sidebar report promotions, or footer links.
|
|
26
|
+
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
### User Story 2 - Reject non-article Boomi URLs cleanly (Priority: P2)
|
|
30
|
+
|
|
31
|
+
A user passes a Boomi URL that is not a readable blog article (the blog index, a category listing, or a marketing page on `boomi.com`). The library raises a typed error rather than returning garbage or partial chrome.
|
|
32
|
+
|
|
33
|
+
**Why this priority**: Protects users from silently receiving meaningless output and keeps behavior consistent with existing providers. Valuable but secondary to the happy path.
|
|
34
|
+
|
|
35
|
+
**Independent Test**: Can be tested by calling `extract()` with a non-article Boomi URL (e.g., the blog index `https://boomi.com/blog/`) and asserting that a typed extraction error is raised.
|
|
36
|
+
|
|
37
|
+
**Acceptance Scenarios**:
|
|
38
|
+
|
|
39
|
+
1. **Given** the Boomi blog index URL, **When** `extract()` is called, **Then** an `UnsupportedContentTypeError` is raised because no single article body is present.
|
|
40
|
+
2. **Given** a Boomi article whose body element is present but yields no text after stripping, **When** `extract()` is called, **Then** an `EmptyContentError` is raised.
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
### Edge Cases
|
|
45
|
+
|
|
46
|
+
- **Non-article pages**: The Boomi blog index (`/blog/`), category/tag listings, and non-blog marketing pages on `boomi.com` do not contain a single recognizable article body → `UnsupportedContentTypeError`.
|
|
47
|
+
- **Hero/inline images**: Reference articles include a featured hero image and may include inline images. Only images inside the article body container are preserved as Markdown images; the featured hero/banner image (rendered outside the body) and other chrome images are stripped with their containers.
|
|
48
|
+
- **Blockquotes**: Some articles embed pull quotes (e.g., the Gartner quote in reference #1). These are preserved as Markdown blockquotes.
|
|
49
|
+
- **Legal disclaimers**: Trailing legal/attribution text that is part of the article body (e.g., the Gartner disclaimer) is preserved as body content; site-wide footer legal links are not.
|
|
50
|
+
- **HTTP errors**: A 404/410 for a removed article, or a 429/503 transient error, is surfaced via the base class's existing fetch/error behavior — this feature introduces no new network handling.
|
|
51
|
+
- **No embedded media in references**: The reference articles contain no code blocks, videos, tweets, or iframes; should such embeds appear in other Boomi articles, they are handled by the inherited base conversion behavior, not by feature-specific logic.
|
|
52
|
+
|
|
53
|
+
## Requirements *(mandatory)*
|
|
54
|
+
|
|
55
|
+
### Functional Requirements
|
|
56
|
+
|
|
57
|
+
- **FR-001**: The system MUST route `boomi.com` blog URLs to the new provider using the existing domain-registration mechanism (`@register` decorator + `DOMAINS` frozenset).
|
|
58
|
+
- **FR-002**: The system MUST extract the main article body from a Boomi blog post page and return it as clean Markdown.
|
|
59
|
+
- **FR-003**: The system MUST strip all non-content elements before conversion: top navigation, language selector, login/demo CTAs, "on this page" jump links, social share buttons, related/previous/next post navigation, sidebar report/promo callouts, and the site footer.
|
|
60
|
+
- **FR-004**: The system MUST prepend the article title as a top-level Markdown heading (`# Title`).
|
|
61
|
+
- **FR-005**: The system MUST preserve structural body content: section headings, paragraphs, ordered and unordered lists, blockquotes, hyperlinks, and images located inside the article body container. Images outside the body container (e.g., a featured hero/banner image rendered in a page header) MUST be treated as page chrome and excluded.
|
|
62
|
+
- **FR-006**: The system MUST raise `UnsupportedContentTypeError` when the page does not contain a recognizable single article body element (e.g., the blog index or a non-article page).
|
|
63
|
+
- **FR-007**: The system MUST raise `EmptyContentError` when the article body yields no extractable text after stripping.
|
|
64
|
+
- **FR-008**: The system MUST collapse runs of three or more consecutive blank lines to a single blank line.
|
|
65
|
+
- **FR-009**: The system MUST yield article output only for pages that contain a recognizable article body container. Pages on `boomi.com` without one — the blog index, category/tag listings, and non-blog pages — MUST NOT yield article output and MUST raise `UnsupportedContentTypeError`. (In practice all article pages live under `boomi.com/blog/<slug>/`; scoping is enforced by body-container presence rather than URL path matching.)
|
|
66
|
+
|
|
67
|
+
### Key Entities *(include if feature involves data)*
|
|
68
|
+
|
|
69
|
+
- **Boomi Blog Post**: A single article page under `https://boomi.com/blog/<slug>/`. Key attributes: title, body content, author (e.g., "Ed Macosky", "Boomi"), publication date, category, optional hero image. Subtitle/deck is generally absent.
|
|
70
|
+
- **DOM Element Taxonomy**: Mapping of CSS selectors for the Boomi blog layout to actions (keep the article body container; strip navigation, language selector, CTAs, jump links, share buttons, related-post navigation, sidebar promos, and footer).
|
|
71
|
+
|
|
72
|
+
## Success Criteria *(mandatory)*
|
|
73
|
+
|
|
74
|
+
### Measurable Outcomes
|
|
75
|
+
|
|
76
|
+
- **SC-001**: Each of the three reference articles returns Markdown containing the full title and all body section headings and paragraph text, with zero non-content elements (navigation, CTAs, share buttons, related-post links, footer).
|
|
77
|
+
- **SC-002**: Extraction completes within the base class 30-second fetch timeout on a stable connection.
|
|
78
|
+
- **SC-003**: A non-article Boomi URL (the blog index) raises `UnsupportedContentTypeError` within the normal fetch timeout.
|
|
79
|
+
- **SC-004**: The extracted Markdown contains no consecutive blank-line runs of three or more lines.
|
|
80
|
+
- **SC-005**: The new provider is exercised by at least one integration test using a real network request against a reference URL, matching the pattern established by existing providers.
|
|
81
|
+
|
|
82
|
+
## Clarifications
|
|
83
|
+
|
|
84
|
+
### Session 2026-06-02
|
|
85
|
+
|
|
86
|
+
- Q: How should the featured/hero image be handled? → A: Body-only images — include only images inside the article body container; the featured hero/banner image (outside the body) is treated as page chrome and excluded.
|
|
87
|
+
|
|
88
|
+
## Assumptions
|
|
89
|
+
|
|
90
|
+
- The Boomi blog's public HTML structure is stable enough for CSS-selector-based extraction; a major redesign would require an extractor update.
|
|
91
|
+
- All target content is freely readable; no authentication or paywall handling is required (confirmed across the three reference articles).
|
|
92
|
+
- Only the primary `boomi.com` domain is in scope. Extraction is limited to article pages under the `/blog/` path; the blog index and other `boomi.com` pages are treated as unsupported content rather than routed to a different provider.
|
|
93
|
+
- Subtitle/deck handling is best-effort: when no deck is present (as in the references), only the title heading is prepended.
|
|
94
|
+
- The implementation follows the existing provider pattern: one new file under `src/mdfetch/providers/` subclassing `BaseExtractor` and reusing its shared conversion methods, with no changes to shared infrastructure.
|
|
95
|
+
- The base class retry/timeout behavior (existing defaults) is inherited without modification.
|
|
96
|
+
- The `test_router.py` unsupported-domain fixture remains valid (Boomi registers `boomi.com`, which is not currently used as the unsupported-domain example — `wordpress.com` is); no fixture update is required unless that changes.
|
|
97
|
+
|
|
98
|
+
### Revision: Implementation Sync 2026-06-02
|
|
99
|
+
- Reason: Post-implementation audit confirmed the shipped extraction matches all functional requirements with no behavioral drift. Reconciled release/docs drift only (version bump + `CLAUDE.md` provider listing); no spec requirements changed.
|