mdfetch 0.3.0__tar.gz → 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (190) hide show
  1. mdfetch-0.4.0/.github/copilot-instructions.md +1 -0
  2. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/git-config.yml +17 -17
  3. mdfetch-0.4.0/.specify/feature.json +3 -0
  4. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/memory/changelog.md +29 -0
  5. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/memory/plan.md +41 -14
  6. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/memory/spec.md +62 -5
  7. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/templates/checklist-template.md +1 -0
  8. mdfetch-0.4.0/.specify/templates/constitution-template.md +44 -0
  9. mdfetch-0.4.0/.specify/templates/plan-template.md +128 -0
  10. mdfetch-0.4.0/.specify/templates/spec-template.md +151 -0
  11. mdfetch-0.4.0/.specify/templates/tasks-template.md +220 -0
  12. {mdfetch-0.3.0 → mdfetch-0.4.0}/CLAUDE.md +17 -2
  13. {mdfetch-0.3.0 → mdfetch-0.4.0}/PKG-INFO +5 -1
  14. {mdfetch-0.3.0 → mdfetch-0.4.0}/README.md +4 -0
  15. {mdfetch-0.3.0 → mdfetch-0.4.0}/pyproject.toml +1 -1
  16. mdfetch-0.4.0/specs/006-thenewstack-provider/checklists/requirements.md +37 -0
  17. mdfetch-0.4.0/specs/006-thenewstack-provider/contracts/public-api.md +54 -0
  18. mdfetch-0.4.0/specs/006-thenewstack-provider/data-model.md +60 -0
  19. mdfetch-0.4.0/specs/006-thenewstack-provider/plan.md +139 -0
  20. mdfetch-0.4.0/specs/006-thenewstack-provider/quickstart.md +81 -0
  21. mdfetch-0.4.0/specs/006-thenewstack-provider/research.md +79 -0
  22. mdfetch-0.4.0/specs/006-thenewstack-provider/spec.md +101 -0
  23. mdfetch-0.4.0/specs/006-thenewstack-provider/tasks.md +224 -0
  24. mdfetch-0.4.0/specs/007-dzone-provider/checklists/requirements.md +37 -0
  25. mdfetch-0.4.0/specs/007-dzone-provider/contracts/public-api.md +53 -0
  26. mdfetch-0.4.0/specs/007-dzone-provider/data-model.md +70 -0
  27. mdfetch-0.4.0/specs/007-dzone-provider/plan.md +136 -0
  28. mdfetch-0.4.0/specs/007-dzone-provider/quickstart.md +107 -0
  29. mdfetch-0.4.0/specs/007-dzone-provider/research.md +85 -0
  30. mdfetch-0.4.0/specs/007-dzone-provider/spec.md +118 -0
  31. mdfetch-0.4.0/specs/007-dzone-provider/tasks.md +272 -0
  32. mdfetch-0.4.0/src/mdfetch/providers/dzone.py +88 -0
  33. mdfetch-0.4.0/src/mdfetch/providers/thenewstack.py +76 -0
  34. mdfetch-0.4.0/tests/integration/conftest.py +9 -0
  35. mdfetch-0.4.0/tests/integration/snapshots/dzone-image-classification-pipeline-camel-djl.md +29 -0
  36. mdfetch-0.4.0/tests/integration/snapshots/dzone-integration-patterns-fail-production.md +29 -0
  37. mdfetch-0.4.0/tests/integration/snapshots/dzone-kiro-feature-to-requirements-design-tasks.md +29 -0
  38. mdfetch-0.4.0/tests/integration/snapshots/thenewstack-api-mcp-agent.md +29 -0
  39. mdfetch-0.4.0/tests/integration/snapshots/thenewstack-async-apis.md +29 -0
  40. mdfetch-0.4.0/tests/integration/snapshots/thenewstack-developer-portal-api.md +29 -0
  41. mdfetch-0.4.0/tests/integration/snapshots/thenewstack-json-schema-ai.md +29 -0
  42. mdfetch-0.4.0/tests/integration/snapshots/thenewstack-mcp-api-governance.md +29 -0
  43. mdfetch-0.4.0/tests/integration/test_dzone_integration.py +56 -0
  44. mdfetch-0.4.0/tests/integration/test_thenewstack_integration.py +64 -0
  45. mdfetch-0.4.0/tests/unit/test_dzone_extractor.py +162 -0
  46. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/unit/test_router.py +12 -0
  47. mdfetch-0.4.0/tests/unit/test_thenewstack_extractor.py +207 -0
  48. {mdfetch-0.3.0 → mdfetch-0.4.0}/uv.lock +52 -52
  49. mdfetch-0.3.0/.specify/feature.json +0 -3
  50. mdfetch-0.3.0/.specify/templates/constitution-template.md +0 -50
  51. mdfetch-0.3.0/.specify/templates/plan-template.md +0 -117
  52. mdfetch-0.3.0/.specify/templates/spec-template.md +0 -129
  53. mdfetch-0.3.0/.specify/templates/tasks-template.md +0 -252
  54. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-analyze/SKILL.md +0 -0
  55. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-archive-run/SKILL.md +0 -0
  56. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-checklist/SKILL.md +0 -0
  57. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-clarify/SKILL.md +0 -0
  58. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-constitution/SKILL.md +0 -0
  59. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-git-commit/SKILL.md +0 -0
  60. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-git-feature/SKILL.md +0 -0
  61. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-git-initialize/SKILL.md +0 -0
  62. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-git-remote/SKILL.md +0 -0
  63. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-git-validate/SKILL.md +0 -0
  64. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-implement/SKILL.md +0 -0
  65. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-plan/SKILL.md +0 -0
  66. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-reconcile-run/SKILL.md +0 -0
  67. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-specify/SKILL.md +0 -0
  68. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-tasks/SKILL.md +0 -0
  69. {mdfetch-0.3.0 → mdfetch-0.4.0}/.claude/skills/speckit-taskstoissues/SKILL.md +0 -0
  70. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.analyze.toml +0 -0
  71. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.archive.run.toml +0 -0
  72. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.checklist.toml +0 -0
  73. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.clarify.toml +0 -0
  74. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.constitution.toml +0 -0
  75. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.implement.toml +0 -0
  76. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.plan.toml +0 -0
  77. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.reconcile.run.toml +0 -0
  78. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.specify.toml +0 -0
  79. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.tasks.toml +0 -0
  80. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gemini/commands/speckit.taskstoissues.toml +0 -0
  81. {mdfetch-0.3.0 → mdfetch-0.4.0}/.github/workflows/ci.yml +0 -0
  82. {mdfetch-0.3.0 → mdfetch-0.4.0}/.github/workflows/integration.yml +0 -0
  83. {mdfetch-0.3.0 → mdfetch-0.4.0}/.github/workflows/publish.yml +0 -0
  84. {mdfetch-0.3.0 → mdfetch-0.4.0}/.gitignore +0 -0
  85. {mdfetch-0.3.0 → mdfetch-0.4.0}/.python-version +0 -0
  86. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/.registry +0 -0
  87. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/archive/LICENSE +0 -0
  88. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/archive/README.md +0 -0
  89. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/archive/commands/archive.md +0 -0
  90. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/archive/extension.yml +0 -0
  91. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/README.md +0 -0
  92. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/commands/speckit.git.commit.md +0 -0
  93. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/commands/speckit.git.feature.md +0 -0
  94. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/commands/speckit.git.initialize.md +0 -0
  95. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/commands/speckit.git.remote.md +0 -0
  96. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/commands/speckit.git.validate.md +0 -0
  97. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/config-template.yml +0 -0
  98. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/extension.yml +0 -0
  99. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/bash/auto-commit.sh +0 -0
  100. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/bash/create-new-feature.sh +0 -0
  101. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/bash/git-common.sh +0 -0
  102. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/bash/initialize-repo.sh +0 -0
  103. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/powershell/auto-commit.ps1 +0 -0
  104. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/powershell/create-new-feature.ps1 +0 -0
  105. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/powershell/git-common.ps1 +0 -0
  106. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/git/scripts/powershell/initialize-repo.ps1 +0 -0
  107. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/reconcile/LICENSE +0 -0
  108. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/reconcile/README.md +0 -0
  109. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/reconcile/commands/reconcile.md +0 -0
  110. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions/reconcile/extension.yml +0 -0
  111. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/extensions.yml +0 -0
  112. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/init-options.json +0 -0
  113. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/integration.json +0 -0
  114. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/integrations/claude.manifest.json +0 -0
  115. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/integrations/gemini.manifest.json +0 -0
  116. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/integrations/speckit.manifest.json +0 -0
  117. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/memory/constitution.md +0 -0
  118. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/scripts/bash/check-prerequisites.sh +0 -0
  119. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/scripts/bash/common.sh +0 -0
  120. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/scripts/bash/create-new-feature.sh +0 -0
  121. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/scripts/bash/setup-plan.sh +0 -0
  122. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/scripts/bash/setup-tasks.sh +0 -0
  123. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/workflows/speckit/workflow.yml +0 -0
  124. {mdfetch-0.3.0 → mdfetch-0.4.0}/.specify/workflows/workflow-registry.json +0 -0
  125. {mdfetch-0.3.0 → mdfetch-0.4.0}/GEMINI.md +0 -0
  126. {mdfetch-0.3.0 → mdfetch-0.4.0}/LICENSE +0 -0
  127. {mdfetch-0.3.0 → mdfetch-0.4.0}/Makefile +0 -0
  128. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/checklists/requirements.md +0 -0
  129. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/contracts/api.md +0 -0
  130. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/data-model.md +0 -0
  131. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/plan.md +0 -0
  132. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/quickstart.md +0 -0
  133. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/research.md +0 -0
  134. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/spec.md +0 -0
  135. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/001-mdfetch-medium-extractor/tasks.md +0 -0
  136. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/checklists/requirements.md +0 -0
  137. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/contracts/public-api.md +0 -0
  138. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/data-model.md +0 -0
  139. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/plan.md +0 -0
  140. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/quickstart.md +0 -0
  141. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/research.md +0 -0
  142. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/spec.md +0 -0
  143. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/002-devto-provider/tasks.md +0 -0
  144. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/003-medium-freedium-fallback/checklists/requirements.md +0 -0
  145. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/003-medium-freedium-fallback/contracts/extract-api.md +0 -0
  146. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/003-medium-freedium-fallback/plan.md +0 -0
  147. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/003-medium-freedium-fallback/research.md +0 -0
  148. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/003-medium-freedium-fallback/spec.md +0 -0
  149. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/003-medium-freedium-fallback/tasks.md +0 -0
  150. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/004-remove-backoff/checklists/requirements.md +0 -0
  151. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/004-remove-backoff/plan.md +0 -0
  152. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/004-remove-backoff/research.md +0 -0
  153. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/004-remove-backoff/spec.md +0 -0
  154. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/004-remove-backoff/tasks.md +0 -0
  155. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/checklists/requirements.md +0 -0
  156. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/contracts/extractor-api.md +0 -0
  157. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/data-model.md +0 -0
  158. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/plan.md +0 -0
  159. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/quickstart.md +0 -0
  160. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/research.md +0 -0
  161. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/spec.md +0 -0
  162. {mdfetch-0.3.0 → mdfetch-0.4.0}/specs/005-substack-provider/tasks.md +0 -0
  163. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/__init__.py +0 -0
  164. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/base.py +0 -0
  165. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/exceptions.py +0 -0
  166. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/providers/__init__.py +0 -0
  167. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/providers/devto.py +0 -0
  168. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/providers/medium.py +0 -0
  169. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/providers/substack.py +0 -0
  170. {mdfetch-0.3.0 → mdfetch-0.4.0}/src/mdfetch/router.py +0 -0
  171. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/__init__.py +0 -0
  172. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/conftest.py +0 -0
  173. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/__init__.py +0 -0
  174. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/architecting-the-asynchronous-agent.md +0 -0
  175. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/devto-integration-digest-december-2025.md +0 -0
  176. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/devto-integration-digest-july-2025.md +0 -0
  177. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/devto-integration-digest-march-2026.md +0 -0
  178. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/from-drift-to-parity.md +0 -0
  179. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/integration-digest-december-2025.md +0 -0
  180. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/substack-api-trends-2025.md +0 -0
  181. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/snapshots/substack-kafka-topic-types.md +0 -0
  182. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/test_devto_integration.py +0 -0
  183. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/test_medium_integration.py +0 -0
  184. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/integration/test_substack_integration.py +0 -0
  185. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/unit/__init__.py +0 -0
  186. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/unit/test_devto_extractor.py +0 -0
  187. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/unit/test_fetch_errors.py +0 -0
  188. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/unit/test_medium_extractor.py +0 -0
  189. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/unit/test_silent.py +0 -0
  190. {mdfetch-0.3.0 → mdfetch-0.4.0}/tests/unit/test_substack_extractor.py +0 -0
@@ -0,0 +1 @@
1
+ ../CLAUDE.md
@@ -11,52 +11,52 @@ init_commit_message: "[Spec Kit] Initial commit"
11
11
  # Set "default" to enable for all commands, then override per-command.
12
12
  # Each key can be true/false. Message is customizable per-command.
13
13
  auto_commit:
14
- default: false
14
+ default: true
15
15
  before_clarify:
16
- enabled: false
16
+ enabled: true
17
17
  message: "[Spec Kit] Save progress before clarification"
18
18
  before_plan:
19
- enabled: false
19
+ enabled: true
20
20
  message: "[Spec Kit] Save progress before planning"
21
21
  before_tasks:
22
- enabled: false
22
+ enabled: true
23
23
  message: "[Spec Kit] Save progress before task generation"
24
24
  before_implement:
25
- enabled: false
25
+ enabled: true
26
26
  message: "[Spec Kit] Save progress before implementation"
27
27
  before_checklist:
28
- enabled: false
28
+ enabled: true
29
29
  message: "[Spec Kit] Save progress before checklist"
30
30
  before_analyze:
31
- enabled: false
31
+ enabled: true
32
32
  message: "[Spec Kit] Save progress before analysis"
33
33
  before_taskstoissues:
34
- enabled: false
34
+ enabled: true
35
35
  message: "[Spec Kit] Save progress before issue sync"
36
36
  after_constitution:
37
- enabled: false
37
+ enabled: true
38
38
  message: "[Spec Kit] Add project constitution"
39
39
  after_specify:
40
- enabled: false
40
+ enabled: true
41
41
  message: "[Spec Kit] Add specification"
42
42
  after_clarify:
43
- enabled: false
43
+ enabled: true
44
44
  message: "[Spec Kit] Clarify specification"
45
45
  after_plan:
46
- enabled: false
46
+ enabled: true
47
47
  message: "[Spec Kit] Add implementation plan"
48
48
  after_tasks:
49
- enabled: false
49
+ enabled: true
50
50
  message: "[Spec Kit] Add tasks"
51
51
  after_implement:
52
- enabled: false
52
+ enabled: true
53
53
  message: "[Spec Kit] Implementation progress"
54
54
  after_checklist:
55
- enabled: false
55
+ enabled: true
56
56
  message: "[Spec Kit] Add checklist"
57
57
  after_analyze:
58
- enabled: false
58
+ enabled: true
59
59
  message: "[Spec Kit] Add analysis report"
60
60
  after_taskstoissues:
61
- enabled: false
61
+ enabled: true
62
62
  message: "[Spec Kit] Sync tasks to issues"
@@ -0,0 +1,3 @@
1
+ {
2
+ "feature_directory": "specs/007-dzone-provider"
3
+ }
@@ -2,6 +2,35 @@
2
2
 
3
3
  ---
4
4
 
5
+ ### mdfetch — The New Stack Provider — 2026-05-16
6
+
7
+ **Branch**: `006-thenewstack-provider`
8
+ **Spec**: specs/006-thenewstack-provider
9
+
10
+ **What was added**:
11
+ - `TheNewStackExtractor` provider for `thenewstack.io` articles, auto-discovered via `@register` decorator
12
+ - Article body isolation from `div#tns-post-body-content`; title prepended from `h1.title` in `div#tns-post-headline`; optional deck/subtitle prepended as plain `<p>` tag from `div.post-excerpt`
13
+ - 4 sponsored-content selectors decomposed from body: `div.sponsored-post-disclosure`, `div.tns-sponsored-post-disclosure`, `div.sponsor-disclosure`, `div.tns-sponsor-note`
14
+ - iframes converted to plain anchor links (defensive; not observed in reference articles)
15
+ - `UnsupportedContentTypeError` raised when `div#tns-post-body-content` is absent; `EmptyContentError` raised when body yields no extractable text
16
+ - 14 unit tests in `tests/unit/test_thenewstack_extractor.py`; 6 integration tests (5 article snapshots + 1 homepage error) in `tests/integration/test_thenewstack_integration.py`
17
+ - VoxPop polls (`div.tns-voxpop-screen`) confirmed as page-level modals outside body — no stripping needed
18
+ - Snapshots use verbatim first-30-line prefix format (preserves blank lines for containment assertion)
19
+
20
+ **New Components**:
21
+ - `src/mdfetch/providers/thenewstack.py` — TheNewStackExtractor
22
+ - `tests/unit/test_thenewstack_extractor.py` — 14 unit tests
23
+ - `tests/integration/test_thenewstack_integration.py` — 6 integration tests
24
+ - `tests/integration/snapshots/thenewstack-developer-portal-api.md`
25
+ - `tests/integration/snapshots/thenewstack-async-apis.md`
26
+ - `tests/integration/snapshots/thenewstack-json-schema-ai.md`
27
+ - `tests/integration/snapshots/thenewstack-mcp-api-governance.md`
28
+ - `tests/integration/snapshots/thenewstack-api-mcp-agent.md`
29
+
30
+ **Tasks Completed**: 20/20
31
+
32
+ ---
33
+
5
34
  ### mdfetch — Substack Provider — 2026-05-15
6
35
 
7
36
  **Branch**: `005-substack-provider`
@@ -1,7 +1,7 @@
1
1
  # mdfetch — Main Implementation Plan
2
2
 
3
- **Last Updated**: 2026-05-15
4
- **Sources**: [specs/001-mdfetch-medium-extractor/plan.md], [specs/002-devto-provider/plan.md], [specs/003-medium-freedium-fallback/plan.md], [specs/004-remove-backoff/plan.md], [specs/005-substack-provider/plan.md]
3
+ **Last Updated**: 2026-05-16
4
+ **Sources**: [specs/001-mdfetch-medium-extractor/plan.md], [specs/002-devto-provider/plan.md], [specs/003-medium-freedium-fallback/plan.md], [specs/004-remove-backoff/plan.md], [specs/005-substack-provider/plan.md], [specs/006-thenewstack-provider/plan.md]
5
5
 
6
6
  ---
7
7
 
@@ -74,6 +74,18 @@ SubstackExtractor(BaseExtractor) — src/mdfetch/providers/substack.py [005-su
74
74
  │ prepends h3.subtitle from div.post-header (if present);
75
75
  │ prepends h1.post-title from div.post-header (unconditional — structurally outside body)
76
76
  └── convert_to_markdown() → markdownify with ATX headings; collapses 3+ newlines to 2; raises EmptyContentError if empty
77
+
78
+ TheNewStackExtractor(BaseExtractor) — src/mdfetch/providers/thenewstack.py [006-thenewstack-provider]
79
+ ├── DOMAINS = frozenset({"thenewstack.io"}) — no subdomain routing needed (single-site WordPress)
80
+ ├── _no_retry_status_codes = frozenset() — inherits base class default (no overrides needed)
81
+ ├── clean_html() → locates div#tns-post-body-content; raises UnsupportedContentTypeError if absent;
82
+ │ decomposes 4 sponsored-content selectors: div.sponsored-post-disclosure,
83
+ │ div.tns-sponsored-post-disclosure, div.sponsor-disclosure, div.tns-sponsor-note;
84
+ │ replaces iframes with anchor links using src/data-src (defensive; not observed in reference articles);
85
+ │ prepends div.post-excerpt text as new <p> tag (deck) from div#tns-post-headline (if present);
86
+ │ prepends copy.copy(h1.title) from div#tns-post-headline (if present)
87
+ │ Note: VoxPop polls (div.tns-voxpop-screen) are page-level modals outside body — no stripping needed
88
+ └── convert_to_markdown() → markdownify with ATX headings; collapses 3+ newlines to 2; raises EmptyContentError if empty
77
89
  ```
78
90
 
79
91
  ### Router / Auto-Discovery
@@ -111,7 +123,8 @@ src/
111
123
  ├── __init__.py # Empty — auto-discovery handles registration
112
124
  ├── medium.py # MediumExtractor
113
125
  ├── devto.py # DevToExtractor [002-devto-provider]
114
- └── substack.py # SubstackExtractor [005-substack-provider]
126
+ ├── substack.py # SubstackExtractor [005-substack-provider]
127
+ └── thenewstack.py # TheNewStackExtractor [006-thenewstack-provider]
115
128
 
116
129
  tests/
117
130
  ├── unit/
@@ -120,20 +133,27 @@ tests/
120
133
  │ ├── test_fetch_errors.py
121
134
  │ ├── test_silent.py
122
135
  │ ├── test_devto_extractor.py # [002-devto-provider]
123
- │ └── test_substack_extractor.py # [005-substack-provider]
136
+ │ ├── test_substack_extractor.py # [005-substack-provider]
137
+ │ └── test_thenewstack_extractor.py # [006-thenewstack-provider]
124
138
  └── integration/
125
139
  ├── snapshots/ # Golden Markdown files (article body snapshots)
126
140
  │ ├── from-drift-to-parity.md
127
141
  │ ├── architecting-the-asynchronous-agent.md
128
142
  │ ├── integration-digest-december-2025.md
129
- │ ├── devto-integration-digest-december-2025.md # [002-devto-provider]
130
- │ ├── devto-integration-digest-july-2025.md # [002-devto-provider]
131
- │ ├── devto-integration-digest-march-2026.md # [002-devto-provider]
132
- │ ├── substack-kafka-topic-types.md # [005-substack-provider]
133
- │ └── substack-api-trends-2025.md # [005-substack-provider]
143
+ │ ├── devto-integration-digest-december-2025.md # [002-devto-provider]
144
+ │ ├── devto-integration-digest-july-2025.md # [002-devto-provider]
145
+ │ ├── devto-integration-digest-march-2026.md # [002-devto-provider]
146
+ │ ├── substack-kafka-topic-types.md # [005-substack-provider]
147
+ │ ├── substack-api-trends-2025.md # [005-substack-provider]
148
+ │ ├── thenewstack-developer-portal-api.md # [006-thenewstack-provider]
149
+ │ ├── thenewstack-async-apis.md # [006-thenewstack-provider]
150
+ │ ├── thenewstack-json-schema-ai.md # [006-thenewstack-provider]
151
+ │ ├── thenewstack-mcp-api-governance.md # [006-thenewstack-provider]
152
+ │ └── thenewstack-api-mcp-agent.md # [006-thenewstack-provider]
134
153
  ├── test_medium_integration.py
135
- ├── test_devto_integration.py # [002-devto-provider]
136
- └── test_substack_integration.py # [005-substack-provider]
154
+ ├── test_devto_integration.py # [002-devto-provider]
155
+ ├── test_substack_integration.py # [005-substack-provider]
156
+ └── test_thenewstack_integration.py # [006-thenewstack-provider]
137
157
 
138
158
  specs/ # Speckit feature specifications
139
159
  pyproject.toml # hatchling build backend, uv package manager
@@ -156,18 +176,20 @@ Makefile # setup / test / integration / lint / typecheck / f
156
176
 
157
177
  ## Testing Strategy
158
178
 
159
- **Unit tests** (84 tests, offline):
179
+ **Unit tests** (101 tests, offline):
160
180
  - Router: domain routing, subdomain suffix matching, duplicate registration, invalid URLs, unsupported platforms
161
181
  - MediumExtractor: clean_html, convert_to_markdown, empty content, non-article pages, _parse_freedium (heading remap, missing main-content), fallback on 403/429 (URL construction, exc.url contract, no-sleep on 429), no-fallback on 200, UnsupportedContentTypeError.url on Freedium path [003-medium-freedium-fallback]
162
182
  - DevToExtractor: clean_html (title/cover/heading/image preservation, iframe/ltag embed→link, anchor stripping, non-article error), convert_to_markdown (headings/code/lists/images, no raw HTML, empty content error) [002-devto-provider]
163
183
  - SubstackExtractor: routing (subdomain + root domain + _no_retry_status_codes assertion), clean_html (body.markup tag return, subscription-widget strip, title prepend, subtitle prepend, prose preservation, iframe→anchor), convert_to_markdown (title heading, no triple blank lines, image syntax, link preservation), paywalled post (non-empty, Subscribe text absent, free preview present), error cases (UnsupportedContentTypeError on no body, EmptyContentError on whitespace body) [005-substack-provider]
184
+ - TheNewStackExtractor: routing (thenewstack.io domain), clean_html (body div return, title prepend, deck-as-paragraph prepend, sponsor note strip, all 3 disclosure variant strips, iframe→anchor, no deck when absent), convert_to_markdown (title heading, deck after title, no triple blank lines, image syntax, link preservation), error cases (UnsupportedContentTypeError on no body, EmptyContentError on whitespace body) [006-thenewstack-provider]
164
185
  - Fetch errors: HTTP 404, 503, timeout, connection error, size limit exceeded; `_no_retry_status_codes` immediate-raise + `_no_retry_codes` override [003-medium-freedium-fallback]
165
186
  - Silent: no stdout/stderr output, no logging during extraction
166
187
 
167
- **Integration tests** (9 tests, network required):
188
+ **Integration tests** (15 tests, network required):
168
189
  - Parametrized over 3 real stn1slv.medium.com articles (including a known paywalled URL that exercises the Freedium fallback when medium.com returns 403) [003-medium-freedium-fallback]
169
190
  - Parametrized over 3 real dev.to/stn1slv articles [002-devto-provider]
170
191
  - Parametrized over 2 real Substack articles + 1 homepage error test (`UnsupportedContentTypeError`) [005-substack-provider]
192
+ - Parametrized over 5 real thenewstack.io articles + 1 homepage error test (`UnsupportedContentTypeError`); snapshots are verbatim first-30-line prefixes of the full extraction output [006-thenewstack-provider]
171
193
  - Snapshot-based containment check: `expected_body in extracted_result` — tests pass regardless of which source served the content
172
194
  - 3 retries with 2-second **fixed** delay on `FetchError` (hardcoded; not env-var configurable) [004-remove-backoff]
173
195
  - Run with: `make integration` or `uv run pytest tests/integration/ --override-ini=addopts=`
@@ -210,6 +232,11 @@ Makefile # setup / test / integration / lint / typecheck / f
210
232
  | Substack title prepend | Unconditional prepend of `h1.post-title` from `div.post-header` | Structurally guaranteed outside `div.body.markup`; section headings use distinct class `header-anchor-post` — no deduplication needed | [005-substack-provider]
211
233
  | Substack subtitle | Prepend `h3.subtitle` after title (inserted at index 0 first, then title at index 0 displaces it to index 1) | Author intent preserved; subtitle rendered as `###` heading | [005-substack-provider]
212
234
  | Substack HTTP 429 | No `_no_retry_status_codes` override — base class `frozenset()` applies | Unlike Medium, Substack has no Freedium-style mirror; retry is the correct fallback | [005-substack-provider]
235
+ | thenewstack.io article body | `div#tns-post-body-content` | Innermost element containing only prose (29 direct `<p>` children in reference articles); parent chain includes several wrapper divs that add no content | [006-thenewstack-provider]
236
+ | thenewstack.io deck element | Create new `<p>` tag with deck text rather than copying `div.post-excerpt` directly | `div.post-excerpt` is a `<div>`, not a semantic subtitle; wrapping text in `<p>` ensures proper paragraph rendering in Markdown | [006-thenewstack-provider]
237
+ | thenewstack.io VoxPop polls | No explicit stripping required | `div.tns-voxpop-screen` confirmed absent from `div#tns-post-body-content` across all 5 reference articles — page-level overlay modal, not inline content | [006-thenewstack-provider]
238
+ | thenewstack.io router test | No update to `test_router.py` unsupported-domain fixture | `wordpress.com` was already the fixture after the Substack provider; `thenewstack.io` registration requires no change | [006-thenewstack-provider]
239
+ | thenewstack.io snapshot format | Verbatim first-30-line prefix (not stripped/compacted) | Blank lines must be preserved for `snapshot in result` containment assertion to pass | [006-thenewstack-provider]
213
240
  | Medium 403/429 fallback | Override `extract()` in `MediumExtractor`; `_no_retry_status_codes=frozenset({403,429})` on class | Immediate fallback with no medium.com retries; `BaseExtractor` extended with `_no_retry_codes` param for thread safety | [003-medium-freedium-fallback]
214
241
  | Freedium HTML parsing | Dedicated `_parse_freedium()` method; `div.main-content`; h4→h3 remap | Freedium HTML is structurally incompatible with `clean_html()` (no `<article>`); heading remap ensures snapshot tests pass for both paths | [003-medium-freedium-fallback]
215
242
  | Freedium exc.url contract | `inner_exc.url = url` unconditionally; error message is source-agnostic ("Fallback page…") | Preserves transparent-fallback contract (FR-028); `exc.url` is the authoritative field; message content is internal | [003-medium-freedium-fallback]
@@ -227,4 +254,4 @@ Makefile # setup / test / integration / lint / typecheck / f
227
254
 
228
255
  ---
229
256
 
230
- *Last Updated: 2026-05-15 | Sources appended: [specs/004-remove-backoff/plan.md], [specs/005-substack-provider/plan.md]*
257
+ *Last Updated: 2026-05-16 | Sources appended: [specs/004-remove-backoff/plan.md], [specs/005-substack-provider/plan.md], [specs/006-thenewstack-provider/plan.md]*
@@ -1,13 +1,13 @@
1
1
  # mdfetch — Main Specification
2
2
 
3
- **Last Updated**: 2026-05-15
4
- **Sources**: [specs/001-mdfetch-medium-extractor/spec.md], [specs/002-devto-provider/spec.md], [specs/003-medium-freedium-fallback/spec.md], [specs/004-remove-backoff/spec.md], [specs/005-substack-provider/spec.md]
3
+ **Last Updated**: 2026-05-16
4
+ **Sources**: [specs/001-mdfetch-medium-extractor/spec.md], [specs/002-devto-provider/spec.md], [specs/003-medium-freedium-fallback/spec.md], [specs/004-remove-backoff/spec.md], [specs/005-substack-provider/spec.md], [specs/006-thenewstack-provider/spec.md]
5
5
 
6
6
  ---
7
7
 
8
8
  ## Overview
9
9
 
10
- `mdfetch` is a Python library that extracts article content from web platforms and returns it as clean, well-structured Markdown. The library enforces a provider pattern — an abstract base defines the extraction contract, and each supported platform is implemented as a separate, independent provider. Supported platforms: `medium.com` (and subdomains), `dev.to`, `substack.com` (and `*.substack.com` subdomains).
10
+ `mdfetch` is a Python library that extracts article content from web platforms and returns it as clean, well-structured Markdown. The library enforces a provider pattern — an abstract base defines the extraction contract, and each supported platform is implemented as a separate, independent provider. Supported platforms: `medium.com` (and subdomains), `dev.to`, `substack.com` (and `*.substack.com` subdomains), `thenewstack.io`.
11
11
 
12
12
  ---
13
13
 
@@ -151,6 +151,30 @@ A developer accidentally passes a Substack URL that does not point to an article
151
151
 
152
152
  ---
153
153
 
154
+ ### US-014 — Extract a Public thenewstack.io Article to Markdown (P1)
155
+ [Source: specs/006-thenewstack-provider]
156
+
157
+ A developer calls `extract()` with a thenewstack.io article URL. The function fetches the article and returns its content as clean Markdown, with no navigation menus, subscription banners, poll widgets, author bios, social share buttons, related articles sections, or any other page chrome.
158
+
159
+ **Acceptance Scenarios**:
160
+ 1. Given a valid URL pointing to a public thenewstack.io article, when `extract()` is called, then it returns a non-empty Markdown string containing the article title as a top-level heading followed by the body content.
161
+ 2. Given a thenewstack.io article with multiple headings, paragraphs, lists, and inline links, when `extract()` is called, then the returned Markdown preserves all headings, paragraphs, lists, code blocks, and hyperlinks while stripping navigation, subscription CTAs, poll widgets, and social share buttons.
162
+ 3. Given a thenewstack.io article containing images, when `extract()` is called, then images appear in the output as Markdown image syntax (`![alt](url)`).
163
+
164
+ ---
165
+
166
+ ### US-015 — Reject Non-Article thenewstack.io Pages (P2)
167
+ [Source: specs/006-thenewstack-provider]
168
+
169
+ A developer passes a thenewstack.io URL that does not point to an article (e.g., the homepage, a category listing page, or a tag archive). The function raises a typed exception rather than returning empty or garbage Markdown.
170
+
171
+ **Acceptance Scenarios**:
172
+ 1. Given the thenewstack.io homepage URL, when `extract()` is called, then `UnsupportedContentTypeError` is raised.
173
+ 2. Given a thenewstack.io category/tag listing page URL, when `extract()` is called, then `UnsupportedContentTypeError` is raised.
174
+ 3. Given an article page whose extractable body text is empty after stripping all chrome, when `extract()` is called, then `EmptyContentError` is raised.
175
+
176
+ ---
177
+
154
178
  ### US-006 — Integration Tests Pass Against Real dev.to Article URLs (P3)
155
179
  [Source: specs/002-devto-provider]
156
180
 
@@ -211,6 +235,18 @@ A developer runs the integration test suite and all dev.to integration tests pas
211
235
  - **FR-039**: The library MUST NOT treat HTTP 429 responses from Substack as a non-retryable condition; 429 MUST be retried up to the configured retry count with the standard fixed delay. [Source: specs/005-substack-provider]
212
236
  - **FR-040**: The library MUST convert embedded third-party content in Substack posts (e.g., tweet embeds, YouTube video iframes, and similar rich-media widgets) to plain anchor links using the embed's source URL, matching the pattern used by the dev.to provider. [Source: specs/005-substack-provider]
213
237
 
238
+ ### The New Stack Platform
239
+ - **FR-041**: The library MUST route all `thenewstack.io` URLs to the TheNewStack provider using the existing domain-registration mechanism (`@register` decorator + `DOMAINS` frozenset). [Source: specs/006-thenewstack-provider]
240
+ - **FR-042**: The library MUST extract the main article body from a thenewstack.io page (`div#tns-post-body-content`) and return it as clean Markdown. [Source: specs/006-thenewstack-provider]
241
+ - **FR-043**: The library MUST strip all non-content elements from within the article body before Markdown conversion, including: sponsored content disclosures (`div.sponsored-post-disclosure`, `div.tns-sponsored-post-disclosure`, `div.sponsor-disclosure`) and injected sponsor notes (`div.tns-sponsor-note`). Navigation, VoxPop polls, social buttons, related posts, sidebar, and footer are outside the body container and require no explicit stripping. [Source: specs/006-thenewstack-provider]
242
+ - **FR-044**: The library MUST prepend the article title as a top-level Markdown heading (`# Title`) from `h1.title` in `div#tns-post-headline`, followed immediately by the article deck/subtitle as a plain paragraph (from `div.post-excerpt` in `div#tns-post-headline`) when one is present. [Source: specs/006-thenewstack-provider]
243
+ - **FR-045**: The library MUST preserve the thenewstack.io article's structural content: headings (all levels), paragraphs, ordered and unordered lists, inline code, fenced code blocks, blockquotes, hyperlinks, and images. [Source: specs/006-thenewstack-provider]
244
+ - **FR-046**: The library MUST raise `UnsupportedContentTypeError` when the fetched thenewstack.io page does not contain a recognisable article body element (`div#tns-post-body-content`). [Source: specs/006-thenewstack-provider]
245
+ - **FR-047**: The library MUST raise `EmptyContentError` when the thenewstack.io article body is present but yields no extractable text after stripping. [Source: specs/006-thenewstack-provider]
246
+ - **FR-048**: The library MUST collapse runs of three or more consecutive blank lines to a single blank line in the thenewstack.io output Markdown. [Source: specs/006-thenewstack-provider]
247
+ - **FR-049**: The library MUST convert embedded third-party content (e.g., YouTube video iframes) in thenewstack.io articles to plain anchor links using the embed's source URL, discarding the embed wrapper — matching the pattern used by existing providers. [Source: specs/006-thenewstack-provider]
248
+ - **FR-050**: The library MUST treat sponsored and native-advertising thenewstack.io article pages identically to editorial articles — extracting content as-is with no special detection, marking, or rejection. [Source: specs/006-thenewstack-provider]
249
+
214
250
  ### dev.to Platform
215
251
  - **FR-015**: The library MUST add `dev.to` to the provider router so that any URL with the `dev.to` domain is dispatched to the dev.to provider without any change to the caller's code. [Source: specs/002-devto-provider]
216
252
  - **FR-016**: The library MUST include a dev.to provider that fetches the article page, isolates the main article body from `<div id="article-body">`, removes all non-content elements (navigation, social reaction widgets, comments, author sidebar, tag links), and returns the body as Markdown. [Source: specs/002-devto-provider]
@@ -237,7 +273,7 @@ A developer runs the integration test suite and all dev.to integration tests pas
237
273
  |-----------|------|-------------|
238
274
  | `DOMAINS` | `frozenset[str]` | Domain suffixes this provider handles (e.g., `{"medium.com"}`, `{"dev.to"}`) |
239
275
 
240
- **Invariants**: Each domain suffix registered to exactly one provider. Stateless — every call is independent. Registered providers: `MediumExtractor` (medium.com), `DevToExtractor` (dev.to), `SubstackExtractor` (substack.com and all `*.substack.com` subdomains).
276
+ **Invariants**: Each domain suffix registered to exactly one provider. Stateless — every call is independent. Registered providers: `MediumExtractor` (medium.com), `DevToExtractor` (dev.to), `SubstackExtractor` (substack.com and all `*.substack.com` subdomains), `TheNewStackExtractor` (thenewstack.io).
241
277
 
242
278
  ### Substack Post
243
279
  [Source: specs/005-substack-provider]
@@ -258,6 +294,16 @@ A developer runs the integration test suite and all dev.to integration tests pas
258
294
 
259
295
  **Boundary**: In the DOM, bounded by the last `div.subscription-widget-wrap` at the truncation point. Stripping that element silently achieves truncation.
260
296
 
297
+ ### TheNewStack Article
298
+ [Source: specs/006-thenewstack-provider]
299
+ | Attribute | Type | Description |
300
+ |-----------|------|-------------|
301
+ | `title` | `str` | Article title from `h1.title` in `div#tns-post-headline`; prepended as `# Title` |
302
+ | `deck` | `str \| None` | Optional subtitle from `div.post-excerpt` in `div#tns-post-headline`; rendered as plain paragraph after title |
303
+ | `body` | `Tag` | Prose content inside `div#tns-post-body-content` |
304
+
305
+ **Validation**: Page must contain `div#tns-post-body-content`; absent → `UnsupportedContentTypeError`. Body must yield non-empty text after stripping → else `EmptyContentError`. No paywall — all thenewstack.io articles are publicly accessible.
306
+
261
307
  ### ExtractionResult (Output)
262
308
  | Attribute | Type | Description |
263
309
  |-----------|------|-------------|
@@ -328,6 +374,11 @@ caller provides URL string
328
374
  - **Substack rich embeds**: `<iframe>` elements and `div[data-component-name]` containers (excluding `SubscribeWidget` and `Image2ToDOM`) are converted to plain anchor links using the embed's source URL. [Source: specs/005-substack-provider]
329
375
  - **Substack HTTP 429**: Treated as a retryable transient error (no `_no_retry_status_codes` override) — contrasts with `MediumExtractor` which uses `frozenset({403, 429})` to trigger Freedium fallback. [Source: specs/005-substack-provider]
330
376
  - **Substack HTML structure changes**: If Substack redesigns and removes `div.body.markup`, the extractor will require an update.
377
+ - **thenewstack.io non-article pages**: Homepage, category/tag listing, and author archive pages do not render `div#tns-post-body-content` → `UnsupportedContentTypeError` is raised immediately. [Source: specs/006-thenewstack-provider]
378
+ - **thenewstack.io VoxPop polls**: `div.tns-voxpop-screen` and `div.tns-voxpop-modal` are page-level overlay modals injected outside `div#tns-post-body-content` — confirmed via live DOM inspection. No explicit stripping is required; scoping extraction to the body container naturally excludes them. [Source: specs/006-thenewstack-provider]
379
+ - **thenewstack.io sponsored content**: `div.tns-sponsor-note` (mid-article sponsor injection) and three disclosure div variants are inside `div#tns-post-body-content` and must be decomposed before conversion. Sponsored article pages are extracted identically to editorial articles (FR-050). [Source: specs/006-thenewstack-provider]
380
+ - **thenewstack.io deck element**: `div.post-excerpt` is a `<div>`, not a semantic subtitle element; the extractor creates a new `<p>` tag with the deck text rather than copying the div directly, to ensure proper paragraph rendering. [Source: specs/006-thenewstack-provider]
381
+ - **thenewstack.io HTML structure changes**: If the site redesign moves content outside `div#tns-post-body-content`, the extractor will require an update.
331
382
 
332
383
  ---
333
384
 
@@ -352,6 +403,12 @@ caller provides URL string
352
403
  - **SC-024**: The extracted Markdown for any Substack article contains no consecutive blank-line runs of three or more lines. [Source: specs/005-substack-provider]
353
404
  - **SC-025**: The Substack provider is exercised by at least one integration test using a real network request, matching the pattern established by existing providers. [Source: specs/005-substack-provider]
354
405
 
406
+ - **SC-026**: A public thenewstack.io article returns Markdown that contains the full article title and body text with zero non-content element fragments (navigation link text, subscription prompts, poll questions, author bio text). [Source: specs/006-thenewstack-provider]
407
+ - **SC-027**: Extraction of a thenewstack.io article completes within the base class 30-second fetch timeout on stable internet. [Source: specs/006-thenewstack-provider]
408
+ - **SC-028**: A thenewstack.io homepage URL raises `UnsupportedContentTypeError` within the normal fetch timeout. [Source: specs/006-thenewstack-provider]
409
+ - **SC-029**: The extracted Markdown for any thenewstack.io article contains no consecutive blank-line runs of three or more lines. [Source: specs/006-thenewstack-provider]
410
+ - **SC-030**: The TheNewStack provider is exercised by integration tests using real network requests against all five reference article URLs, matching the pattern established by existing providers. [Source: specs/006-thenewstack-provider]
411
+
355
412
  ---
356
413
 
357
414
  ## Assumptions
@@ -374,4 +431,4 @@ caller provides URL string
374
431
 
375
432
  ---
376
433
 
377
- *Last Updated: 2026-05-15 | Sources appended: [specs/004-remove-backoff/spec.md], [specs/005-substack-provider/spec.md]*
434
+ *Last Updated: 2026-05-16 | Sources appended: [specs/004-remove-backoff/spec.md], [specs/005-substack-provider/spec.md], [specs/006-thenewstack-provider/spec.md]*
@@ -21,6 +21,7 @@
21
21
  -->
22
22
 
23
23
  ## [Category 1]
24
+ <!-- Example categories: Provider Compliance, Extraction Quality, Test Coverage, Constitution Gate -->
24
25
 
25
26
  - [ ] CHK001 First checklist item with clear action
26
27
  - [ ] CHK002 Second checklist item
@@ -0,0 +1,44 @@
1
+ # [PROJECT_NAME] Constitution
2
+ <!-- Example: mdfetch Constitution -->
3
+
4
+ ## Core Principles
5
+
6
+ ### [PRINCIPLE_1_NAME]
7
+ <!-- Example: I. Provider Pattern Architecture -->
8
+ [PRINCIPLE_1_DESCRIPTION]
9
+ <!-- Example: The system MUST enforce a strict Provider Pattern. An abstract base class MUST be defined for all extractors. Adding new platforms MUST only require creating a new subclass, adhering to the Open/Closed Principle. Code duplication is PROHIBITED; shared logic MUST reside in the base class. -->
10
+
11
+ ### [PRINCIPLE_2_NAME]
12
+ <!-- Example: II. Technology Stack -->
13
+ [PRINCIPLE_2_DESCRIPTION]
14
+ <!-- Example: The project MUST exclusively use: httpx (network), BeautifulSoup (parsing), Markdownify (conversion), pytest (testing), uv (package management). Direct use of pip, venv, or pip-tools is PROHIBITED. -->
15
+
16
+ ### [PRINCIPLE_3_NAME]
17
+ <!-- Example: III. Coding Standards -->
18
+ [PRINCIPLE_3_DESCRIPTION]
19
+ <!-- Example: All functions MUST use strict Python type hinting. Codebase MUST adhere to PEP 8. Variable names, docstrings, and comments MUST use clear English vocabulary. -->
20
+
21
+ ### [PRINCIPLE_4_NAME]
22
+ <!-- Example: IV. Testing Requirements -->
23
+ [PRINCIPLE_4_DESCRIPTION]
24
+ <!-- Example: The test suite MUST include integration tests. These tests MUST verify functionality by providing real links and asserting returned Markdown matches expected output. -->
25
+
26
+ ### [PRINCIPLE_5_NAME]
27
+ <!-- Example: V. Packaging and Distribution -->
28
+ [PRINCIPLE_5_DESCRIPTION]
29
+ <!-- Example: The project MUST use pyproject.toml and src/ layout. All Makefile targets MUST invoke uv run <tool> rather than calling tools directly. -->
30
+
31
+ ## [SECTION_2_NAME]
32
+ <!-- Example: Additional Constraints, Error Handling Policy, etc. -->
33
+
34
+ [SECTION_2_CONTENT]
35
+ <!-- Example: All failures communicated via typed exceptions only (no logging). Custom exception hierarchy with MdfetchError as base. -->
36
+
37
+ ## Governance
38
+ <!-- Constitution supersedes all other practices; Amendments require documentation, approval, migration plan -->
39
+
40
+ [GOVERNANCE_RULES]
41
+ <!-- Example: All PRs must verify compliance. Complexity must be justified. Amendment procedure: increment constitution version. Semantic versioning for governance changes. -->
42
+
43
+ **Version**: [CONSTITUTION_VERSION] | **Ratified**: [RATIFICATION_DATE] | **Last Amended**: [LAST_AMENDED_DATE]
44
+ <!-- Example: Version: 1.0.0 | Ratified: 2026-05-14 | Last Amended: 2026-05-14 -->
@@ -0,0 +1,128 @@
1
+ # Implementation Plan: [FEATURE]
2
+
3
+ **Branch**: `[###-feature-name]` | **Date**: [DATE] | **Spec**: [link]
4
+
5
+ **Input**: Feature specification from `/specs/[###-feature-name]/spec.md`
6
+
7
+ **Note**: This template is filled in by the `/speckit-plan` command. See `.specify/templates/plan-template.md` for the execution workflow.
8
+
9
+ ## Summary
10
+
11
+ [Extract from feature spec: primary requirement + technical approach from research]
12
+
13
+ ## Technical Context
14
+
15
+ <!--
16
+ The values below reflect mdfetch's established stack.
17
+ Override only if this feature deviates from the norm.
18
+ -->
19
+
20
+ **Language/Version**: Python 3.12+ (matches CI matrix: 3.12–3.14)
21
+
22
+ **Primary Dependencies**: `httpx` (HTTP fetch), `BeautifulSoup` / `lxml` (HTML parsing), `markdownify` (Markdown conversion), `pytest` (testing)
23
+
24
+ **Storage**: N/A — stateless extraction library
25
+
26
+ **Testing**: `pytest` via `uv run pytest` — unit tests (no network) + integration tests (`-m integration`, real URLs + snapshots)
27
+
28
+ **Target Platform**: PyPI library (cross-platform)
29
+
30
+ **Project Type**: Library
31
+
32
+ **Performance Goals**: Inherits base class 30-second fetch timeout; no additional targets unless spec overrides
33
+
34
+ **Constraints**: [e.g., "One new provider file; no changes to shared infrastructure" or NEEDS CLARIFICATION]
35
+
36
+ **Scale/Scope**: Single-article extraction per call
37
+
38
+ ## Constitution Check
39
+
40
+ *GATE: Must pass before implementation. Re-check after design phase.*
41
+
42
+ - [ ] Validates Provider Pattern Architecture (No code duplication, adheres to Open/Closed Principle)
43
+ - [ ] Confirms Technology Stack (`httpx`, `BeautifulSoup`, `Markdownify`, `pytest`)
44
+ - [ ] Adheres to Coding Standards (PEP 8, Type Hinting, Clear Vocabulary)
45
+ - [ ] Incorporates Integration Testing (Real links matching expected Markdown)
46
+ - [ ] Respects Packaging and Distribution standards (`pyproject.toml`, `src/` layout, `uv` for all dev workflow commands)
47
+
48
+ ## Project Structure
49
+
50
+ ### Documentation (this feature)
51
+
52
+ ```text
53
+ specs/[###-feature]/
54
+ ├── plan.md # This file (/speckit-plan command output)
55
+ ├── research.md # Phase 0 output (/speckit-plan command)
56
+ ├── data-model.md # Phase 1 output (/speckit-plan command)
57
+ ├── quickstart.md # Phase 1 output (/speckit-plan command)
58
+ ├── contracts/ # Phase 1 output (/speckit-plan command)
59
+ └── tasks.md # Phase 2 output (/speckit-tasks command - NOT created by /speckit-plan)
60
+ ```
61
+
62
+ ### Source Code
63
+ <!--
64
+ ACTION REQUIRED: Replace the placeholder tree below with the concrete file
65
+ list for this feature. The layout always follows the established provider pattern.
66
+ -->
67
+
68
+ ```text
69
+ src/mdfetch/providers/
70
+ └── [platform].py # NEW: [Platform]Extractor
71
+
72
+ tests/unit/
73
+ └── test_[platform]_extractor.py # NEW: unit tests (no network)
74
+
75
+ tests/integration/
76
+ ├── snapshots/
77
+ │ └── [platform]-[article-slug].md # NEW: snapshot(s)
78
+ └── test_[platform]_integration.py # NEW: integration tests (real URLs)
79
+ ```
80
+
81
+ **Non-runtime changes** (if any): [e.g., "README.md (supported platforms table), tests/unit/test_router.py (unsupported-domain fixture update)"]
82
+
83
+ ## Extraction Algorithm
84
+
85
+ <!--
86
+ ACTION REQUIRED: Replace the pseudocode below with the actual extraction
87
+ pipeline for this platform, derived from research.md analysis.
88
+ -->
89
+
90
+ ```
91
+ extract(url):
92
+ html ← fetch_html(url) # base class; retries on all transient errors
93
+ soup ← BeautifulSoup(html, "lxml")
94
+ body_tag ← clean_html(soup) # → article body with chrome stripped
95
+ return convert_to_markdown(body_tag)
96
+
97
+ clean_html(soup):
98
+ 1. Find [article body container] → raise UnsupportedContentTypeError if absent
99
+ 2. Strip [non-content elements]
100
+ 3. Convert [embedded content] → anchor links
101
+ 4. Find [title element] → prepend to body
102
+ 5. Find [subtitle element] (optional) → prepend after title
103
+ 6. Return body tag
104
+
105
+ convert_to_markdown(tag):
106
+ md ← markdownify(str(tag), heading_style="ATX", code_language="", strip=["script","style"])
107
+ md ← strip leading/trailing whitespace
108
+ md ← collapse 3+ blank lines → single blank line
109
+ if md empty → raise EmptyContentError
110
+ return md
111
+ ```
112
+
113
+ ## Error Mapping
114
+
115
+ | Condition | Exception |
116
+ |-----------|-----------|
117
+ | Article body container not found | `UnsupportedContentTypeError` |
118
+ | Body found but no extractable text | `EmptyContentError` |
119
+ | HTTP error (any non-2xx after retries) | `HTTPStatusError` |
120
+ | Network / timeout failure | `FetchError` |
121
+
122
+ ## Complexity Tracking
123
+
124
+ > **Fill ONLY if Constitution Check has violations that must be justified**
125
+
126
+ | Violation | Why Needed | Simpler Alternative Rejected Because |
127
+ |-----------|------------|--------------------------------------|
128
+ | [e.g., modifies base class] | [current need] | [why provider-only approach insufficient] |