docpluck 2.4.72__tar.gz → 2.4.73__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (370) hide show
  1. {docpluck-2.4.72 → docpluck-2.4.73}/CHANGELOG.md +15 -0
  2. {docpluck-2.4.72 → docpluck-2.4.73}/PKG-INFO +1 -1
  3. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/__init__.py +1 -1
  4. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/detect.py +59 -2
  5. docpluck-2.4.73/docs/HANDOFF_2026-05-25_haiku-orchestration-pretest.md +164 -0
  6. docpluck-2.4.73/docs/HANDOFF_2026-05-25_pretest-followups.md +135 -0
  7. docpluck-2.4.73/docs/superpowers/handoffs/2026-05-23-bundled-residual-cycle-CLOSED.md +69 -0
  8. docpluck-2.4.73/docs/superpowers/plans/2026-05-23-haiku-orchestration-pretest.md +949 -0
  9. {docpluck-2.4.72 → docpluck-2.4.73}/pyproject.toml +1 -1
  10. docpluck-2.4.73/scripts/pretest_capture_tokens.py +121 -0
  11. docpluck-2.4.73/tests/fixtures/structured/.gitkeep +0 -0
  12. docpluck-2.4.73/tests/test_pretest_capture_tokens.py +75 -0
  13. docpluck-2.4.73/tests/test_r1_whitespace_cells_wiring_real_pdf.py +130 -0
  14. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/_project/canary.json +0 -0
  15. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/_project/lessons.md +0 -0
  16. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-cleanup/SKILL.md +0 -0
  17. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-deploy/SKILL.md +0 -0
  18. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/LEARNINGS.md +0 -0
  19. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/SKILL.md +0 -0
  20. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/ai-full-doc-verify.md +0 -0
  21. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/cycle-report-template.md +0 -0
  22. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/local-verification.md +0 -0
  23. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/rationalizations.md +0 -0
  24. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/real-library-real-pdf.md +0 -0
  25. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/release-flow.md +0 -0
  26. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/self-improvement.md +0 -0
  27. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-iterate/references/three-tier-parity.md +0 -0
  28. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-qa/SKILL.md +0 -0
  29. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-qa/references/benchmark-mode.md +0 -0
  30. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-qa/references/check-11-hard-rules.md +0 -0
  31. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-qa/references/check-13-escicheck-production.md +0 -0
  32. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-qa/references/check-5-escicheck-library.md +0 -0
  33. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-qa/references/check-6-escicheck-local-webapp.md +0 -0
  34. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-qa/references/check-7-batch-smoke.md +0 -0
  35. {docpluck-2.4.72 → docpluck-2.4.73}/.claude/skills/docpluck-review/SKILL.md +0 -0
  36. {docpluck-2.4.72 → docpluck-2.4.73}/.github/workflows/bump-app-pin.yml +0 -0
  37. {docpluck-2.4.72 → docpluck-2.4.73}/.github/workflows/publish.yml +0 -0
  38. {docpluck-2.4.72 → docpluck-2.4.73}/.github/workflows/test.yml +0 -0
  39. {docpluck-2.4.72 → docpluck-2.4.73}/.gitignore +0 -0
  40. {docpluck-2.4.72 → docpluck-2.4.73}/CLAUDE.md +0 -0
  41. {docpluck-2.4.72 → docpluck-2.4.73}/HANDOFF_SECTIONS_APP_INTEGRATION.md +0 -0
  42. {docpluck-2.4.72 → docpluck-2.4.73}/LESSONS.md +0 -0
  43. {docpluck-2.4.72 → docpluck-2.4.73}/LICENSE +0 -0
  44. {docpluck-2.4.72 → docpluck-2.4.73}/REPLY_FROM_DOCPLUCK_v1.4.5.md +0 -0
  45. {docpluck-2.4.72 → docpluck-2.4.73}/REPLY_FROM_DOCPLUCK_v1.5.0.md +0 -0
  46. {docpluck-2.4.72 → docpluck-2.4.73}/REQUEST_08_CHUNKING_ENDPOINT.md +0 -0
  47. {docpluck-2.4.72 → docpluck-2.4.73}/REQUEST_09_REFERENCE_LIST_NORMALIZATION.md +0 -0
  48. {docpluck-2.4.72 → docpluck-2.4.73}/TODO.md +0 -0
  49. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/__main__.py +0 -0
  50. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/batch.py +0 -0
  51. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/cli.py +0 -0
  52. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/extract.py +0 -0
  53. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/extract_docx.py +0 -0
  54. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/extract_html.py +0 -0
  55. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/extract_layout.py +0 -0
  56. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/extract_structured.py +0 -0
  57. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/figures/__init__.py +0 -0
  58. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/figures/detect.py +0 -0
  59. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/normalize.py +0 -0
  60. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/quality.py +0 -0
  61. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/render.py +0 -0
  62. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/__init__.py +0 -0
  63. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/annotators/__init__.py +0 -0
  64. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/annotators/docx.py +0 -0
  65. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/annotators/html.py +0 -0
  66. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/annotators/pdf.py +0 -0
  67. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/annotators/text.py +0 -0
  68. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/blocks.py +0 -0
  69. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/boundaries.py +0 -0
  70. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/core.py +0 -0
  71. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/taxonomy.py +0 -0
  72. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/sections/types.py +0 -0
  73. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/__init__.py +0 -0
  74. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/bbox_utils.py +0 -0
  75. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/camelot_extract.py +0 -0
  76. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/captions.py +0 -0
  77. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/cell_cleaning.py +0 -0
  78. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/cluster.py +0 -0
  79. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/confidence.py +0 -0
  80. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/render.py +0 -0
  81. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/tables/whitespace.py +0 -0
  82. {docpluck-2.4.72 → docpluck-2.4.73}/docpluck/version.py +0 -0
  83. {docpluck-2.4.72 → docpluck-2.4.73}/docs/BENCHMARKS.md +0 -0
  84. {docpluck-2.4.72 → docpluck-2.4.73}/docs/DESIGN.md +0 -0
  85. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-07_sections_strict_iteration.md +0 -0
  86. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-09_session_state_and_followups.md +0 -0
  87. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-09_unified_extraction_brainstorm.md +0 -0
  88. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-10_table_rendering_iteration.md +0 -0
  89. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-10_table_rendering_iteration_2.md +0 -0
  90. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-10_table_rendering_iteration_3.md +0 -0
  91. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-10_table_rendering_iteration_4.md +0 -0
  92. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-10_table_rendering_iteration_5.md +0 -0
  93. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-10_table_rendering_iteration_6.md +0 -0
  94. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-10_table_rendering_iteration_7.md +0 -0
  95. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-11_PROMOTE_SPIKE_TO_LIBRARY.md +0 -0
  96. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-11_table_rendering_iteration_8.md +0 -0
  97. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-11_visual_review_findings.md +0 -0
  98. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-12_phase2_101pdf_corpus.md +0 -0
  99. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-12_remaining_ui_and_chrome_verification.md +0 -0
  100. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-12_visual_verify_results.md +0 -0
  101. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-13_apa_50_expansion.md +0 -0
  102. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-13_apa_50_expansion_iter_1.md +0 -0
  103. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-13_apa_50_expansion_iter_2.md +0 -0
  104. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-13_iterate_skill_first_use.md +0 -0
  105. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-13_iterative_1.md +0 -0
  106. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-13_iterative_library_improvement.md +0 -0
  107. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-13_table_extraction_next_iteration.md +0 -0
  108. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-14_continue_iterations_v2_4_30_to_15n.md +0 -0
  109. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-14_full_corpus_iteration_v2_4_30.md +0 -0
  110. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-14_iterate_6_cycles_complete.md +0 -0
  111. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-14_iterate_9_cycle_run.md +0 -0
  112. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-14_iterate_resume_4_cycles.md +0 -0
  113. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-14_iterate_v2_4_31_cycle_15n.md +0 -0
  114. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-14_phase_5d_gold_audit_v2_4_29.md +0 -0
  115. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-15_autonomous_apa_first_10h.md +0 -0
  116. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-15_iterate_apa_run_1.md +0 -0
  117. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-16_ai-gold-instructions.md +0 -0
  118. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-16_iterate_apa_run_2.md +0 -0
  119. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-16_iterate_apa_run_3.md +0 -0
  120. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-16_iterate_run_4_final.md +0 -0
  121. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-16_iterate_run_4_fix_and_continue.md +0 -0
  122. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-16_iterate_run_5.md +0 -0
  123. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-16_iterate_run_6.md +0 -0
  124. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-17_iterate_run_7.md +0 -0
  125. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-17_iterate_run_8.md +0 -0
  126. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-17_iterate_run_9.md +0 -0
  127. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-18_iterate_run_9_cont.md +0 -0
  128. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-18_iterate_run_9_cont2.md +0 -0
  129. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-20_iterate_run_9_cont3.md +0 -0
  130. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-22_iterate_run_9_session4_final.md +0 -0
  131. {docpluck-2.4.72 → docpluck-2.4.73}/docs/HANDOFF_2026-05-22_iterate_run_9_session5_close.md +0 -0
  132. {docpluck-2.4.72 → docpluck-2.4.73}/docs/ITERATION_VERIFICATION_LESSONS.md +0 -0
  133. {docpluck-2.4.72 → docpluck-2.4.73}/docs/LIBRARY_APP_SYNC.md +0 -0
  134. {docpluck-2.4.72 → docpluck-2.4.73}/docs/NORMALIZATION.md +0 -0
  135. {docpluck-2.4.72 → docpluck-2.4.73}/docs/README.md +0 -0
  136. {docpluck-2.4.72 → docpluck-2.4.73}/docs/TRIAGE_2026-05-10_corpus_assessment.md +0 -0
  137. {docpluck-2.4.72 → docpluck-2.4.73}/docs/TRIAGE_2026-05-14_phase_5d_gold_audit.md +0 -0
  138. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/handoffs/2026-05-22-b1-next-iteration.md +0 -0
  139. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/handoffs/2026-05-22-b2-remaining-halluc-head.md +0 -0
  140. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/handoffs/2026-05-22-b3-b7-structural-defects.md +0 -0
  141. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/handoffs/2026-05-22-residual-after-locally-doable-pass.md +0 -0
  142. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/handoffs/2026-05-23-residual-after-iterate-spine-cycles-1-3.md +0 -0
  143. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/2026-05-06-section-identification.md +0 -0
  144. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/2026-05-06-table-extraction.md +0 -0
  145. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/2026-05-07-sections-strict-iteration-progress.md +0 -0
  146. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/2026-05-08-unified-extraction-phase-0-splice-spike.md +0 -0
  147. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/sections-deferred-items.md +0 -0
  148. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/sections-issues-backlog.md +0 -0
  149. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/2026-05-07_spot-01_apa.md +0 -0
  150. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/2026-05-07_spot-02_pattern-A-shipped.md +0 -0
  151. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/2026-05-08_spot-final_all-styles.md +0 -0
  152. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/COMPARISON.md +0 -0
  153. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-a/korbmacher_table1.md +0 -0
  154. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-a/option-a.py +0 -0
  155. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-a/ziano_table1.md +0 -0
  156. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-b/korbmacher_notes_raw.txt +0 -0
  157. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-b/korbmacher_table1.md +0 -0
  158. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-b/notes.md +0 -0
  159. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-b/option-b.py +0 -0
  160. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-b/ziano_notes_raw.txt +0 -0
  161. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-b/ziano_table1.md +0 -0
  162. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-c/korbmacher_table1.md +0 -0
  163. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-c/notes.md +0 -0
  164. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-c/option-c.py +0 -0
  165. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-c/sample-pdftotext-bbox.html +0 -0
  166. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-c/ziano_table1.md +0 -0
  167. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-d/korbmacher_table1.md +0 -0
  168. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-d/notes.md +0 -0
  169. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-d/option-d.py +0 -0
  170. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-d/ziano_table1.md +0 -0
  171. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/korbmacher_2022_kruger_bbox.html +0 -0
  172. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/korbmacher_bbox.html +0 -0
  173. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/korbmacher_table1.md +0 -0
  174. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/option-e.py +0 -0
  175. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/sample-bbox.html +0 -0
  176. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/ziano_2021_joep_bbox.html +0 -0
  177. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/ziano_bbox.html +0 -0
  178. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/experiments/option-e/ziano_table1.md +0 -0
  179. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/html-fallback-demo.md +0 -0
  180. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/chandrashekar_2023_mp.err +0 -0
  181. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/chandrashekar_2023_mp.md +0 -0
  182. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/efendic_2022_affect.err +0 -0
  183. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/efendic_2022_affect.md +0 -0
  184. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/ieee_access_2.err +0 -0
  185. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/ieee_access_2.md +0 -0
  186. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/ip_feldman_2025_pspb.err +0 -0
  187. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/ip_feldman_2025_pspb.md +0 -0
  188. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/korbmacher_2022_kruger.err +0 -0
  189. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/korbmacher_2022_kruger.md +0 -0
  190. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/nat_comms_1.err +0 -0
  191. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/nat_comms_1.md +0 -0
  192. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/ziano_2021_joep.err +0 -0
  193. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs/ziano_2021_joep.md +0 -0
  194. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/am_sociol_rev_3.err +0 -0
  195. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/am_sociol_rev_3.md +0 -0
  196. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/amc_1.err +0 -0
  197. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/amc_1.md +0 -0
  198. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/amj_1.err +0 -0
  199. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/amj_1.md +0 -0
  200. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/amle_1.err +0 -0
  201. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/amle_1.md +0 -0
  202. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ar_apa_j_jesp_2009_12_010.err +0 -0
  203. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ar_apa_j_jesp_2009_12_010.md +0 -0
  204. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ar_royal_society_rsos_140066.err +0 -0
  205. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ar_royal_society_rsos_140066.md +0 -0
  206. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ar_royal_society_rsos_140072.err +0 -0
  207. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ar_royal_society_rsos_140072.md +0 -0
  208. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/bjps_1.err +0 -0
  209. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/bjps_1.md +0 -0
  210. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/chan_feldman_2025_cogemo.err +0 -0
  211. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/chan_feldman_2025_cogemo.md +0 -0
  212. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/chen_2021_jesp.err +0 -0
  213. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/chen_2021_jesp.md +0 -0
  214. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/demography_1.err +0 -0
  215. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/demography_1.md +0 -0
  216. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ieee_access_3.err +0 -0
  217. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ieee_access_3.md +0 -0
  218. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ieee_access_4.err +0 -0
  219. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/ieee_access_4.md +0 -0
  220. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/jama_open_1.err +0 -0
  221. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/jama_open_1.md +0 -0
  222. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/jama_open_2.err +0 -0
  223. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/jama_open_2.md +0 -0
  224. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/jmf_1.err +0 -0
  225. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/jmf_1.md +0 -0
  226. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/nat_comms_2.err +0 -0
  227. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/nat_comms_2.md +0 -0
  228. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/sci_rep_1.err +0 -0
  229. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/sci_rep_1.md +0 -0
  230. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/social_forces_1.err +0 -0
  231. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/outputs-new/social_forces_1.md +0 -0
  232. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/papers.md +0 -0
  233. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/report.md +0 -0
  234. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/splice_spike.py +0 -0
  235. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/plans/spot-checks/splice-spike/test_splice_spike.py +0 -0
  236. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/specs/2026-04-27-request-09-reference-normalization-design.md +0 -0
  237. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/specs/2026-05-06-section-identification-design.md +0 -0
  238. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/specs/2026-05-06-table-extraction-design.md +0 -0
  239. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/specs/2026-05-08-unified-extraction-design.md +0 -0
  240. {docpluck-2.4.72 → docpluck-2.4.73}/docs/superpowers/specs/2026-05-23-haiku-orchestration-pretest-design.md +0 -0
  241. {docpluck-2.4.72/tests → docpluck-2.4.73/scripts}/__init__.py +0 -0
  242. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/README.md +0 -0
  243. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/VERIFIER_PROMPT.md +0 -0
  244. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/__init__.py +0 -0
  245. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/baseline_matrix.json +0 -0
  246. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/checks.py +0 -0
  247. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/corpus.py +0 -0
  248. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/corpus_manifest.json +0 -0
  249. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/extract.py +0 -0
  250. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/gold_keys.json +0 -0
  251. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/harness/inspect.py +0 -0
  252. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/lint_rendered_corpus.py +0 -0
  253. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/verify_corpus.py +0 -0
  254. {docpluck-2.4.72 → docpluck-2.4.73}/scripts/verify_corpus_full.py +0 -0
  255. {docpluck-2.4.72/tests/fixtures → docpluck-2.4.73/tests}/__init__.py +0 -0
  256. {docpluck-2.4.72 → docpluck-2.4.73}/tests/conftest.py +0 -0
  257. {docpluck-2.4.72/tests/fixtures/sections → docpluck-2.4.73/tests/fixtures}/__init__.py +0 -0
  258. /docpluck-2.4.72/tests/fixtures/structured/.gitkeep → /docpluck-2.4.73/tests/fixtures/sections/__init__.py +0 -0
  259. {docpluck-2.4.72 → docpluck-2.4.73}/tests/fixtures/sections/builders.py +0 -0
  260. {docpluck-2.4.72 → docpluck-2.4.73}/tests/fixtures/structured/MANIFEST.json +0 -0
  261. {docpluck-2.4.72 → docpluck-2.4.73}/tests/fixtures/structured/README.md +0 -0
  262. {docpluck-2.4.72 → docpluck-2.4.73}/tests/golden/sections/apa_multi_study_pdf.json +0 -0
  263. {docpluck-2.4.72 → docpluck-2.4.73}/tests/golden/sections/apa_single_study_pdf.json +0 -0
  264. {docpluck-2.4.72 → docpluck-2.4.73}/tests/golden/sections/html_real_headings.json +0 -0
  265. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/amj_lattice.txt +0 -0
  266. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/apa_chan_feldman_lineless.txt +0 -0
  267. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/apa_chen_jesp_lineless.txt +0 -0
  268. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/apa_efendic_affect.txt +0 -0
  269. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/apa_ip_feldman_pspb.txt +0 -0
  270. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/bmc_lattice.txt +0 -0
  271. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/ieee_figure_heavy.txt +0 -0
  272. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/ieee_lattice.txt +0 -0
  273. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/jama_lattice.txt +0 -0
  274. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/nat_comms_figure_only.txt +0 -0
  275. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/nature_minimal_rule.txt +0 -0
  276. {docpluck-2.4.72 → docpluck-2.4.73}/tests/snapshots/scirep_minimal_rule.txt +0 -0
  277. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_a3c_leading_zero_decimal_real_pdf.py +0 -0
  278. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_all_caps_section_promote_real_pdf.py +0 -0
  279. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_bbox_utils.py +0 -0
  280. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_benchmark_docx_html.py +0 -0
  281. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_cambridge_footer_strip_real_pdf.py +0 -0
  282. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_caption_only_table_heading_real_pdf.py +0 -0
  283. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_caption_regex.py +0 -0
  284. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_chart_data_trim_real_pdf.py +0 -0
  285. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_cid_minus_recovery_real_pdf.py +0 -0
  286. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_cli_sections.py +0 -0
  287. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_cli_structured.py +0 -0
  288. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_confidence.py +0 -0
  289. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_corpus_smoke.py +0 -0
  290. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_d5_normalization_audit.py +0 -0
  291. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_edge_cases.py +0 -0
  292. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_elsevier_footer_strip_real_pdf.py +0 -0
  293. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_equation_page_header_strip_real_pdf.py +0 -0
  294. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_extract_docx.py +0 -0
  295. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_extract_filter_sugar.py +0 -0
  296. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_extract_html.py +0 -0
  297. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_extract_layout.py +0 -0
  298. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_extract_pdf_structured.py +0 -0
  299. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_extraction.py +0 -0
  300. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_f0_table_region_aware.py +0 -0
  301. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_fffd_comparison_recovery_real_pdf.py +0 -0
  302. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_figure_caption_trim_real_pdf.py +0 -0
  303. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_figure_detect.py +0 -0
  304. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_fixtures_manifest.py +0 -0
  305. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_hallucinated_heading_continuation_guard.py +0 -0
  306. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_harness_text_loss_reflow.py +0 -0
  307. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_lattice_cluster.py +0 -0
  308. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_letterspaced_label_real_pdf.py +0 -0
  309. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_ligature_decomposition_real_pdf.py +0 -0
  310. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_lt_operator_recovery_real_pdf.py +0 -0
  311. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_mathitalic_greek_real_pdf.py +0 -0
  312. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_metaesci_followups.py +0 -0
  313. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_minus_sign_recovery_real_pdf.py +0 -0
  314. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalization.py +0 -0
  315. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalize_a3_r2_body_integer_real_pdf.py +0 -0
  316. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalize_f0_footnote_strip.py +0 -0
  317. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalize_idempotent_real_pdf.py +0 -0
  318. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalize_layout_param.py +0 -0
  319. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalize_metadata_leak_real_pdf.py +0 -0
  320. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalize_report_layout_fields.py +0 -0
  321. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_normalize_v18_strips.py +0 -0
  322. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_numbered_heading_promotion_real_pdf.py +0 -0
  323. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_numbered_section_promotion_real_pdf.py +0 -0
  324. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_orphan_multilevel_number_real_pdf.py +0 -0
  325. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_orphan_section_number_real_pdf.py +0 -0
  326. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_p0r_recurring_running_header_strip.py +0 -0
  327. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_preserve_math_glyphs_real_pdf.py +0 -0
  328. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_pua_glyph_recovery_real_pdf.py +0 -0
  329. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_quality.py +0 -0
  330. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_render.py +0 -0
  331. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_render_html.py +0 -0
  332. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_request_09_reference_normalization.py +0 -0
  333. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_residual_2026_05_23_bundled.py +0 -0
  334. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_roman_numeral_section_promote_real_pdf.py +0 -0
  335. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_section_row_label_no_merge_real_pdf.py +0 -0
  336. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_boundaries.py +0 -0
  337. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_boundary_truncation.py +0 -0
  338. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_core_partition.py +0 -0
  339. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_docx_annotator.py +0 -0
  340. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_extract_text.py +0 -0
  341. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_footnote_section.py +0 -0
  342. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_golden.py +0 -0
  343. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_html_annotator.py +0 -0
  344. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_pdf_annotator.py +0 -0
  345. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_public_api.py +0 -0
  346. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_real_corpus.py +0 -0
  347. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_taxonomy.py +0 -0
  348. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_text_annotator.py +0 -0
  349. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_types.py +0 -0
  350. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_unit_corpus.py +0 -0
  351. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_v161_coalesce.py +0 -0
  352. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_v161_subheadings.py +0 -0
  353. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_v161_taxonomy.py +0 -0
  354. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_v161_text_annotator.py +0 -0
  355. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_sections_version.py +0 -0
  356. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_smoke_fixtures.py +0 -0
  357. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_structured_result_type.py +0 -0
  358. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_structured_types.py +0 -0
  359. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_structured_version.py +0 -0
  360. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_table_caption_cell_region_real_pdf.py +0 -0
  361. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_table_detect.py +0 -0
  362. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_tables_cell_cleaning.py +0 -0
  363. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_text_mode.py +0 -0
  364. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_v23_1_fixes.py +0 -0
  365. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_v23_bug_fixes.py +0 -0
  366. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_v23_post_corpus.py +0 -0
  367. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_v23_post_corpus_v2.py +0 -0
  368. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_v2_backwards_compat.py +0 -0
  369. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_v2_top_level_exports.py +0 -0
  370. {docpluck-2.4.72 → docpluck-2.4.73}/tests/test_whitespace_cluster.py +0 -0
@@ -1,5 +1,20 @@
1
1
  # Changelog
2
2
 
3
+ ## [2.4.73] — 2026-05-25
4
+
5
+ **R1-repair — wake up the dead `whitespace_cells` wiring.** The §A R1 fix in v2.4.72 (`docpluck/extract_structured.py` `whitespace_cells` fallback for caption-detected tables Camelot couldn't recover) shipped structurally dead in production: a 2026-05-25 AI-gold sweep across the 11 B1 papers found `_region_for_caption` returned `None` in 100% of unmatched-caption cases, so `whitespace_cells` was never invoked. Root cause: `_bbox_of_caption_line` in `docpluck/tables/detect.py` matched the first-20-char `cap.line_text` prefix (e.g. `'Table 5. Reflection'`) against joined layout chars — but layout chars on a single y-row drop inter-word whitespace and keep raw PDF ligatures (e.g. `'Table5.Reflection…'`), so the prefix never matched. The silent fallback to `_isolated_table_from_caption` hid the no-op behind v2.4.71-identical output.
6
+
7
+ **Fix (`docpluck/tables/detect.py::_bbox_of_caption_line`):** three-pass matcher.
8
+ - Pass 1: exact prefix match against joined chars (legacy path, preserved for compatibility).
9
+ - Pass 2: normalized prefix match — fold ligatures (`fi`→`fi`, `fl`→`fl`, etc.), strip whitespace, lowercase on both sides. Catches the dominant B1 failure shape.
10
+ - Pass 3: label-only fallback — search for the normalized `cap.label` (e.g. `table5`) anywhere in the joined row, gated to start near the left margin and within the first ~4 chars of the row to avoid false positives on body prose ("(see Table 5)") and right-column 2-column-page rows.
11
+
12
+ **Verification:** post-fix region resolution went from 0/22 captions (jdm_.2023.16 / chan_feldman / maier) to 22/22 (100%). `whitespace_cells` now fires and yields **72 cells on chan_feldman_2025_cogemo (8 captions)** and **100 cells on maier_2023_collabra (11 captions)** — verified by the new `tests/test_r1_whitespace_cells_wiring_real_pdf.py` regression test (3 cases: 2 real-PDF + 1 unit ligature/whitespace normalization).
13
+
14
+ No `NORMALIZATION_VERSION` bump (no normalize.py change). Real-PDF regression test asserts the wiring stays live (catches future regressions in ligature handling, caption-line shape, or region bbox sizing).
15
+
16
+ Follow-ups (queued in `todo.md`, not blocking this ship): (a) thread `_layout_doc` through `extract_pdf_structured` to eliminate the second `extract_pdf_layout(pdf_bytes)` pass `render.py` already triggers (perf only — measured 2x on every render path with unmatched caps); (b) for jdm_.2023.16-shape narrow tables, regions resolve but `whitespace_cells`'s ≥3-stable-column threshold leaves cells empty — needs per-page table-region detection per the 2026-05-22 R1 decision table.
17
+
3
18
  ## [2.4.72] — 2026-05-23
4
19
 
5
20
  **Bundled cycle — 2026-05-23 residual handoff (§A R1, R3a, R3b, R4, R5; §B-new-1..5; §C P0r-F).** `NORMALIZATION_VERSION` 1.9.22 → 1.9.23. Eleven fixes landed in one cycle per user directive ("implement and fix all in one go, leave nothing behind"):
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: docpluck
3
- Version: 2.4.72
3
+ Version: 2.4.73
4
4
  Summary: PDF, DOCX, and HTML text extraction and normalization for academic papers
5
5
  Project-URL: Homepage, https://docpluck.app
6
6
  Project-URL: Documentation, https://docpluck.app/api-docs
@@ -71,7 +71,7 @@ from .figures import Figure
71
71
  from .extract_structured import TABLE_EXTRACTION_VERSION, StructuredResult, extract_pdf_structured
72
72
  from .render import render_pdf_to_markdown
73
73
 
74
- __version__ = "2.4.72"
74
+ __version__ = "2.4.73"
75
75
  __author__ = "Gilad Feldman"
76
76
  __license__ = "MIT"
77
77
 
@@ -121,8 +121,36 @@ class _Footnote:
121
121
  bbox: Bbox
122
122
 
123
123
 
124
+ _LIGATURE_FOLD = str.maketrans({
125
+ "ff": "ff", "fi": "fi", "fl": "fl",
126
+ "ffi": "ffi", "ffl": "ffl", "ſt": "ft", "st": "st",
127
+ })
128
+
129
+
130
+ def _normalize_for_char_match(s: str) -> str:
131
+ """Fold ligatures, drop whitespace, lowercase — for matching caption text
132
+ against layout-channel joined chars (which routinely drop spaces between
133
+ words inside a y-row and keep PDF ligature glyphs unchanged)."""
134
+ return s.translate(_LIGATURE_FOLD).replace(" ", "").replace("\t", "").lower()
135
+
136
+
124
137
  def _bbox_of_caption_line(page_obj, cap: CaptionMatch) -> Bbox | None:
125
- """Find the y-cluster of chars whose joined text contains the start of the caption line."""
138
+ """Find the y-cluster of chars whose joined text contains the start of the caption line.
139
+
140
+ Robustness layers (each successively more forgiving — added 2026-05-25 after
141
+ the R1 sweep showed _region_for_caption returning None 11/11 on the B1 corpus
142
+ because layout chars join without inter-word spaces and contain raw PDF
143
+ ligatures, while cap.line_text comes from the de-ligatured text channel with
144
+ spaces preserved):
145
+ 1. Exact prefix match against joined chars (legacy path).
146
+ 2. Normalized prefix match — fold ligatures, strip whitespace, lowercase
147
+ on both sides. Catches the dominant B1 failure shape (Table5.Reflection…
148
+ vs Table 5. Reflection…).
149
+ 3. Label-only match — search for the normalized cap.label (e.g. ``table5``)
150
+ anywhere in the joined row. Catches captions whose body text on the
151
+ page differs from the text-channel line_text (cross-page caption,
152
+ glyph substitution).
153
+ """
126
154
  chars = page_obj.chars or ()
127
155
  if not chars:
128
156
  return None
@@ -130,20 +158,49 @@ def _bbox_of_caption_line(page_obj, cap: CaptionMatch) -> Bbox | None:
130
158
  if not target:
131
159
  return None
132
160
  target_prefix = target[:20]
161
+ target_prefix_norm = _normalize_for_char_match(target_prefix)
162
+ label_norm = _normalize_for_char_match(cap.label) # e.g. "table5"
133
163
 
134
164
  rows: dict[int, list[dict]] = defaultdict(list)
135
165
  for c in chars:
136
166
  rows[round(c.get("top", 0))].append(c)
137
167
 
168
+ # Pass 1+2: prefix-based match (legacy or normalized).
138
169
  for top_key in sorted(rows.keys()):
139
170
  row_chars = sorted(rows[top_key], key=lambda c: c.get("x0", 0))
140
171
  joined = "".join(c.get("text", "") for c in row_chars)
141
- if target_prefix in joined:
172
+ joined_norm = _normalize_for_char_match(joined)
173
+ if target_prefix in joined or (
174
+ target_prefix_norm and target_prefix_norm in joined_norm
175
+ ):
142
176
  x0 = min(c["x0"] for c in row_chars)
143
177
  x1 = max(c["x1"] for c in row_chars)
144
178
  top = min(c["top"] for c in row_chars)
145
179
  bottom = max(c["bottom"] for c in row_chars)
146
180
  return (x0, top, x1, bottom)
181
+
182
+ # Pass 3: label-only fallback. The row must additionally start near the
183
+ # left margin (label-style caption, not an inline back-reference like
184
+ # "(see Table 5)") to avoid false-positive matches on body prose.
185
+ page_width = float(getattr(page_obj, "width", 0.0) or 0.0)
186
+ left_margin_cap = page_width * 0.5 if page_width else float("inf")
187
+ for top_key in sorted(rows.keys()):
188
+ row_chars = sorted(rows[top_key], key=lambda c: c.get("x0", 0))
189
+ joined = "".join(c.get("text", "") for c in row_chars)
190
+ joined_norm = _normalize_for_char_match(joined)
191
+ if not label_norm or label_norm not in joined_norm:
192
+ continue
193
+ # Reject if the label appears mid-row rather than at/near the start.
194
+ label_pos = joined_norm.find(label_norm)
195
+ if label_pos > 4: # tolerate a couple leading chars (e.g., "*Table 5")
196
+ continue
197
+ x0 = min(c["x0"] for c in row_chars)
198
+ if x0 > left_margin_cap:
199
+ continue # right column of a 2-column page, not a caption row
200
+ x1 = max(c["x1"] for c in row_chars)
201
+ top = min(c["top"] for c in row_chars)
202
+ bottom = max(c["bottom"] for c in row_chars)
203
+ return (x0, top, x1, bottom)
147
204
  return None
148
205
 
149
206
 
@@ -0,0 +1,164 @@
1
+ # Haiku-orchestration pre-test — results
2
+
3
+ **Dates:** 2026-05-23 (design) → 2026-05-25 (execution + report)
4
+ **Spec:** [`docs/superpowers/specs/2026-05-23-haiku-orchestration-pretest-design.md`](superpowers/specs/2026-05-23-haiku-orchestration-pretest-design.md)
5
+ **Plan:** [`docs/superpowers/plans/2026-05-23-haiku-orchestration-pretest.md`](superpowers/plans/2026-05-23-haiku-orchestration-pretest.md)
6
+ **Start SHA:** `5735903`
7
+
8
+ ## TL;DR
9
+
10
+ The Haiku-orchestration pattern **works for bulk read/diff tasks** (Test 1 directional). It **does NOT work as a substitute for Opus when generating extraction ground truth** (Test 2 clear).
11
+
12
+ - **Test 1 (diagnose docpluck output vs gold):** Both arms reached the same FAIL verdict on the same 5 defect classes. Haiku absorbed ~193k tokens of bulk reading/diffing at API cost ~$0.50; Opus-solo cost ~126k tokens at ~$3–4. Haiku-side savings are real — but only if Opus orchestration overhead stays low.
13
+ - **Test 2 (generate gold reading.md from PDFs):** Opus golds scored 5/5/5/5/5/0 across the board. Haiku golds scored 3/2/2/2/2/7–8 — coverage gap, severe hallucinations (wrong author names, wrong p-values, fabricated references, paragraph duplication, inverted figure interpretations). **Do not use Haiku for gold-generation.**
14
+
15
+ ## Test 1 — Diagnose docpluck output vs gold (Opus solo vs Opus + Haiku)
16
+
17
+ Scope: 1 paper (`jama-open-1`), 1 diagnostic cycle, scaled down from the original 3-PDF/3-cycle plan because the harness limit prevented subagents from dispatching sub-subagents (see "Honest limits" below). Both arms stopped at verdict; neither attempted fixes.
18
+
19
+ | Metric | Arm A (Opus solo) | Arm B (Opus + Haiku) |
20
+ |---|---|---|
21
+ | Verdict | FAIL | FAIL |
22
+ | Defect classes found | 5 | 5 (same classes) |
23
+ | Wall time | ~4 min | ~3.2 min |
24
+ | Opus tokens (clean) | **125,896** (fresh-context subagent) | not cleanly measured¹ |
25
+ | Haiku tokens (clean) | 0 | **193,154** (3 subagent dispatches) |
26
+ | Direct tool calls (orchestration) | 31 | ~7 |
27
+ | API-rate cost on measured tokens | ~$1.90–$9.45 (high uncertainty: input/output split unknown)² | ~$0.19–$0.97 |
28
+
29
+ ¹ Arm B's Opus orchestration ran from this top-level session, which carried ~200k+ tokens of accumulated context (spec, plan, prior tool calls). My orchestration tokens are not directly comparable to Arm A's fresh-context 126k.
30
+
31
+ ² Opus API rates: $15/M input, $75/M output. Range reflects unknown input/output split inside Arm A's subagent return.
32
+
33
+ ### Both arms found the same defects on jama-open-1
34
+
35
+ 1. **RUNNING_HEADER_LEAK** — `Downloaded from jamanetwork.com…` watermark and `October 27, 2023` page-marker date leak into body text 11+ times.
36
+ 2. **HALLUC_HEAD** — `### 1.0. Mean glucose level`, `### Control`, `### Body weight, kg`, `### Total cholesterol` promoted from Table 2 cell content into the heading hierarchy.
37
+ 3. **ABSTRACT_LEVEL_MISMATCH** — Gold has `## Abstract` with h3 children (Importance/Objective/Design/Interventions/Main Outcomes/Results/Conclusions and Relevance). Rendered has `## Abstract` with NO subsections; instead `## Findings`, `## RESULTS`, `## CONCLUSIONS AND RELEVANCE` appear at h2 level, breaking hierarchy.
38
+ 4. **MISSING_SECTION** — `## Key Points` (JAMA Open's structured-summary sidebar) is entirely absent from rendered output.
39
+ 5. **TABLE_STRUCTURE_CORRUPT** — Table 3 has `<th>JAMA Network Open | Nutrition, Obesity, and Exercise</th>` (journal masthead) and `<td>Discussion</td>` (section name as data cell).
40
+
41
+ ### Qualitative read
42
+
43
+ - Haiku produced JSON-structured summaries of both 1395-line rendered.md and 374-line gold.md in 12-43 seconds each. Quality was sufficient for a downstream Opus or Haiku diff step to identify all 5 defect classes.
44
+ - Arm A burned 31 tool calls on the same work that Arm B (Opus orchestrator side) did in ~7 tool calls.
45
+ - The pattern's actual win condition is **Opus stays as planner/judge with low tool-call count; Haiku does the per-file reading**. Both arms confirm Haiku is competent at structured summarization tasks given a clear schema.
46
+
47
+ ### What this does NOT tell us
48
+
49
+ - Whether Haiku-drafted **code patches** would be accepted by Opus or need rework (test was diagnose-only, no fixes attempted).
50
+ - Whether the pattern survives multi-cycle iteration (test was 1 cycle).
51
+ - Whether arm B's orchestration overhead exceeds arm A's solo cost on harder/larger inputs (N=1 PDF).
52
+
53
+ ## Test 2 — Gold extractor (Opus vs Haiku)
54
+
55
+ Scope: 2 fresh APA papers (not in active iterate queue), both with existing Opus-generated gold. Each Haiku regen took 2-3 tool calls.
56
+
57
+ ### Generation cost
58
+
59
+ | Paper | PDF pages | Opus gold size | Haiku regen tokens | Haiku gold size | Wall time |
60
+ |---|---|---|---|---|---|
61
+ | maier_2023_collabra | 21 | 97,483 bytes | 94,524 | 92,931 (95%) | 4.2 min |
62
+ | efendic_2022_affect | 12 | 56,548 bytes | 71,097 | 51,264 (91%) | 2.6 min |
63
+
64
+ Haiku golds are 91-95% the size of Opus golds — size is NOT a useful quality signal (see below).
65
+
66
+ ### Blind judge scorecard (Opus judge, fresh context, 228k tokens, paths-revealed-provenance flagged but stated content-only)
67
+
68
+ | Paper | Model | Coverage | Accuracy | Hallucination↓ | Structure | Prose |
69
+ |---|---|---|---|---|---|---|
70
+ | maier_2023_collabra | **Opus** | 5 | 5 | 0 | 5 | 5 |
71
+ | maier_2023_collabra | Haiku | 3 | 2 | **8** | 2 | 2 |
72
+ | efendic_2022_affect | **Opus** | 5 | 5 | 0 | 5 | 5 |
73
+ | efendic_2022_affect | Haiku | 3 | 2 | **7** | 3 | 2 |
74
+
75
+ ### Specific Haiku failure modes (auditable, independent of any provenance hint)
76
+
77
+ **maier_2023_collabra (Haiku gold):**
78
+ - Wrong citations: `Bergh & Bernstein, 2021` (PDF: Reinstein), `Smith et al., 2015` (PDF: 2013), `Lee & Feeley, 2010` (PDF: 2016).
79
+ - Wrong stats: `r = 0.064` (PDF: 0.004 — off by 16×); ANOVA design `2 × 5` repeated multiple times (PDF: 2 × 3); abstract CI `[.000, .005]` (PDF: [.000, .003]).
80
+ - Corrupted table: Table 10 `r` column shows sample-size values (`.170, .165`) where correlation coefficients should be — data-shape confusion.
81
+ - Severe duplication: "The Identifiable Victim Effect" paragraphs and the H5d section repeated 2-3 times verbatim.
82
+ - Truncated: Scope Insensitivity / Irrational Decision Making / Limitations sections largely missing.
83
+ - **References list absent entirely.**
84
+
85
+ **efendic_2022_affect (Haiku gold):**
86
+ - Author confusion: `Henrick et al., 2019` (PDF: Efendić et al., 2019 — replaced the actual author name).
87
+ - Wrong SE values: Table 2 intercept SE `0.04` (PDF: 0.06); Table 4 Direction SE `0.03` (PDF: 0.09).
88
+ - Wrong scale anchors: rating described as `1 (not at all risky) to 10 (extremely risky)` (PDF: anchored 1 / 5 moderate / 10 very risky — 3-anchor scale).
89
+ - Fabricated references: `Giroux, C. (2021)`, `Ruelle, J.A. (2003)` (PDF: Russell), `Vanpoucke` (PDF: Vanpaemel), `Västfjäll 2008` invented; wrong title/journal for Slovic et al. (2002).
90
+ - Citation replacement: `Phelan & Sureden` substituted for Thaler & Sunstein.
91
+ - **Figure 4 interpretation inverted** — Haiku says +1 SD pleasure shows shallower slope than −1 SD; PDF says the opposite.
92
+
93
+ ### Why this matters
94
+
95
+ Docpluck's verification pipeline assumes the AI gold is the source of truth. **If gold generation drifts (wrong author names, fabricated references, inverted figure interpretations), every downstream verdict is corrupted.** The 7-8 hallucinations per Haiku gold are not noise — they're systematic enough to fool a downstream FAIL-comparison check that doesn't audit the gold itself.
96
+
97
+ This rules out Haiku-as-gold-extractor for docpluck's current verification model. The Opus golds remain canonical. If cost-reduction in gold-gen is needed, the right experiment is **Opus-with-thinking-budget capped** or **Sonnet 4.6** as the gold extractor, not Haiku.
98
+
99
+ ## Combined recommendations
100
+
101
+ 1. **Haiku-orchestration for bulk read/diff: pursue.** Test 1 was directionally positive. Next experiment: run the full original plan (3 PDFs, 3 cycles, fresh top-level Opus session) and include fix-drafting in scope. If Haiku-drafted patches survive Opus review ≥80% of the time, the pattern is production-ready.
102
+ 2. **Haiku-as-gold-extractor: reject.** Test 2 is clear. Even with N=2 the failure modes are too severe (wrong author names, inverted figures, fabricated references). Not safe for any verification pipeline that treats the gold as ground truth.
103
+ 3. **The bigger question — "expensive iterate skills with cheaper models":** the pattern that's likely to actually save credits is **Opus-stays-planner + Haiku-does-per-file-summarization**. NOT "swap Opus for Haiku on the whole task." This pretest didn't measure code-patch quality from Haiku — that's the next pretest to design.
104
+ 4. **Workshop angle (free local models):** unchanged from the brainstorm — Ollama + Qwen2.5-Coder-7B is the realistic floor for student laptops. Local Haiku-class quality is currently unreachable for free.
105
+
106
+ ## Issues found during the test (deferred per LEAVE-NOTHING-BEHIND exception)
107
+
108
+ These were flagged but NOT fixed in-flight per the pretest design. **Triage in a follow-up session.**
109
+
110
+ ### Docpluck library — jama-open-1 defect cluster (newly surfaced, not in canary)
111
+
112
+ The 5 defect classes listed in Test 1 above. Severity: SERIOUS. Spans:
113
+ - `docpluck/normalize.py` — F0 running-header strip misses JAMA's `Downloaded from jamanetwork.com…` watermark and `October 27, 2023` date-line.
114
+ - `docpluck/sections/` — abstract subsection promotion to h2 instead of h3 when JAMA structured abstract is parsed; column-interleave variant for left abstract + right Key Points sidebar.
115
+ - `docpluck/sections/annotators/text.py` (or wherever heading promotion lives) — table-cell content promoted to h3 headings (`### 1.0. Mean glucose level`, etc.).
116
+ - `docpluck/tables/` — Table 3 row-cell pollution: journal masthead in `<th>`, section name in `<td>`.
117
+
118
+ Recommended follow-up: add jama-open-1 to the iterate canary set, run a normal `/docpluck-iterate` cycle dedicated to it.
119
+
120
+ ### Article-finder skill — usability sharp edge
121
+
122
+ From Arm A subagent's observation: `ai-gold.py resolve` fails on the stem name (`jama_open_1`) and on the source PDF file path, but `ai-gold.py check jama_open_1 --view reading` works directly. The `RESOLVE the canonical key` step in `article-finder/SKILL.md` doesn't document this fallback. Severity: MODERATE (causes wasted cycles for new users).
123
+
124
+ Recommended follow-up: amend `article-finder/SKILL.md` and `article-finder/gold-generation.md` to document the stem-vs-DOI resolve behavior, OR fix `ai-gold.py resolve` to accept stems.
125
+
126
+ ### Plan deviation (informational, not a defect)
127
+
128
+ - The original plan committed `tmp/pretest_start_sha.txt`. `tmp/` is gitignored, so the commit silently no-op'd. Future plans should either use a non-gitignored location or `git add -f`.
129
+ - The original plan instructed both arms to run via fresh top-level sessions; the user opted for an in-session run via subagents. We hit a harness limit (subagents can't dispatch sub-subagents), then a CLI auth limit (`claude -p` from Bash gets 401 even with `ANTHROPIC_API_KEY` set, because the Claude Code parent uses managed-by-host provider context). We scaled the experiment down accordingly. **For a clean re-run at full scope, the original plan stands: open fresh top-level sessions per arm.**
130
+
131
+ ## Methodology + limits
132
+
133
+ - **N=1 PDF in Test 1, N=2 PDFs in Test 2.** Both are pretest scope, intentionally small. Results are directional.
134
+ - **Test 1 was diagnose-only, not full iterate.** The original plan included fix attempts; we scoped down. The bulk-read savings shown here may not survive a fix-and-rework workflow where Opus rejects many Haiku patch drafts.
135
+ - **Test 2 judge saw paths revealing provenance** (`ai_gold/` vs `_pretest_haiku_golds/`). The judge flagged this in its return; we cannot fully rule out hint-leakage on the 1-5 rubric scores. However: the specific factual errors enumerated by the judge (wrong author names, wrong p-values, fabricated references) are auditable against the PDF and independent of any provenance hint. The qualitative conclusion (Haiku golds have severe hallucinations) is robust; the exact 5-vs-2 score margins less so.
136
+ - **Token counts on Arm B Opus side are not measurable** because this session's accumulated context contaminates the cost. Reported as Opus tool-call count instead.
137
+ - **All four Haiku subagent runs in this pretest had non-trivial fixed overhead** — the simplest Haiku ping (`PONG`-only) consumed 41,528 tokens. This is Claude Code's per-subagent system-prompt + tool-definition load. It puts a floor on per-Agent-call cost (~$0.04 at API rates). Pattern only wins if each Haiku call absorbs >>40k tokens of real work, which both Test 1 and Test 2 cleared.
138
+
139
+ ## Artifacts
140
+
141
+ | File | Purpose |
142
+ |---|---|
143
+ | `tmp/pretest_start_sha.txt` | Starting commit (5735903) |
144
+ | `tmp/pretest_test2_blind_mapping.json` | Test 2 X/Y labelling key |
145
+ | `tmp/pretest_test2_judge.json` | Test 2 blind judge raw scorecard |
146
+ | `tmp/pretest_test2_unblinded.json` | Test 2 un-blinded scorecard |
147
+ | `tmp/run_meta_pre_pretest.json` | Iterate-skill run-meta snapshot, pre-test |
148
+ | `.worktrees/armA-opus-solo/tmp/*` | Arm A rendered.md, gold.md, run-meta, timestamps, findings |
149
+ | `.worktrees/armB-opus-haiku/tmp/*` | Arm B rendered.md, gold.md, timestamps, findings |
150
+ | `ArticleRepository/_pretest_haiku_golds/maier_2023_collabra/reading.md` | Haiku-extracted gold (do NOT use for verification) |
151
+ | `ArticleRepository/_pretest_haiku_golds/efendic_2022_affect/reading.md` | Haiku-extracted gold (do NOT use for verification) |
152
+ | `scripts/pretest_capture_tokens.py` | Reusable token-capture utility (not exercised in scaled-down run) |
153
+
154
+ ## Cleanup status
155
+
156
+ - Two worktrees still live: `pretest/armA-opus-solo` and `pretest/armB-opus-haiku`. Branches retained pending user decision on whether to delete them.
157
+ - Quarantined Haiku golds remain at `_pretest_haiku_golds/` — recommend keeping for one cycle, then deleting.
158
+ - Token-capture script is committed to `main` as `5735903` and remains useful for future pretests.
159
+
160
+ ## Approved follow-ups (queued, not started)
161
+
162
+ 1. Add jama-open-1 to docpluck iterate canary set; run a dedicated cycle to fix the 5 defect classes (SERIOUS).
163
+ 2. Document or fix `ai-gold.py resolve` stem-vs-DOI behavior (MODERATE).
164
+ 3. Design and run the next pretest: **Haiku-drafted code patches reviewed by Opus** (the open question this pretest did not answer).
@@ -0,0 +1,135 @@
1
+ # Pretest follow-ups — handoff
2
+
3
+ **Created:** 2026-05-25
4
+ **Created from:** session that produced [`HANDOFF_2026-05-25_haiku-orchestration-pretest.md`](HANDOFF_2026-05-25_haiku-orchestration-pretest.md)
5
+ **Repo state at handoff:** `main` @ `9262e1e` ("docs: Haiku-orchestration pretest results + recommendations"), clean working tree, no open worktrees.
6
+ **Read first:** [`HANDOFF_2026-05-25_haiku-orchestration-pretest.md`](HANDOFF_2026-05-25_haiku-orchestration-pretest.md) — full pretest report. Specifically the "Issues found during the test" section, which is the source for both items below.
7
+
8
+ This handoff bundles two **independent** follow-ups deferred from the 2026-05-25 Haiku-orchestration pretest. They can be done in either order, in separate sessions, or in parallel worktrees. Do not bundle them into one commit — each gets its own.
9
+
10
+ ---
11
+
12
+ ## Issue 1 — jama-open-1 defect cluster (SERIOUS)
13
+
14
+ ### Why this matters
15
+
16
+ JAMA Network Open's two-column-with-Key-Points-sidebar layout is a real publisher class docpluck currently can't render. Five distinct defects on one paper, all surfaced cleanly during the pretest's Test 1. None are paper-specific quirks — every fix must generalize to the JAMA Open structural signature per CLAUDE.md.
17
+
18
+ ### Inputs
19
+
20
+ - **Test fixture:** `verify_out/pdfextractor__ama__jama-open-1/` (the rendered-output dir; `.gitignored` so will be regenerated on first run)
21
+ - **Gold reference:** `C:\Users\filin\Dropbox\Vibe\ArticleRepository\ai_gold\jama_open_1\reading.md`
22
+ - **Source PDF path:** look up via `python -c "import json; print(json.load(open(r'C:\Users\filin\Dropbox\Vibe\ArticleRepository\ai_gold\jama_open_1\reading.meta.json'))['source_pdf_path'])"`
23
+ - **Rendered output from the pretest (saved for diff convenience):** none preserved (worktrees deleted). Regenerate via the iterate skill.
24
+
25
+ ### What's broken (five defect classes, with file pointers)
26
+
27
+ | # | Class | File / module | Symptom |
28
+ |---|---|---|---|
29
+ | 1 | **RUNNING_HEADER_LEAK** | `docpluck/normalize.py` (F0 layout-aware running-header strip) | `Downloaded from jamanetwork.com by Medizinisch-Biologische Fachbibliothek user on 03/18/2026` leaks into body ~11×; `October 27, 2023` page-marker date also leaks. Pattern is a per-page footer that F0 currently does not detect on JAMA Open's geometry. |
30
+ | 2 | **HALLUC_HEAD** | `docpluck/sections/annotators/text.py` (heading promotion) — verify exact module | Table 2 cell content promoted to h3: `### 1.0. Mean glucose level`, `### Control`, `### Body weight, kg`, `### Total cholesterol`. Likely cause: heading-promotion heuristic firing on short isolated lines inside table regions. |
31
+ | 3 | **ABSTRACT_LEVEL_MISMATCH** | `docpluck/sections/` (structured-abstract detection) | JAMA's structured abstract has 7 subsections (Importance / Objective / Design, Setting, and Participants / Interventions / Main Outcomes and Measures / Results / Conclusions and Relevance). Gold renders these as h3 under `## Abstract`. Docpluck renders some as h2 (`## Findings`, `## RESULTS`, `## CONCLUSIONS AND RELEVANCE`) — breaking hierarchy. |
32
+ | 4 | **MISSING_SECTION** | `docpluck/sections/` + extraction pipeline | JAMA Open's `## Key Points` (right-column sidebar above the abstract — a publisher-mandated structured summary with Question/Findings/Meaning) is entirely absent from rendered output. Likely cause: sidebar bbox not detected as a separate content region; gets dropped or interleaved with abstract. |
33
+ | 5 | **TABLE_STRUCTURE_CORRUPT** | `docpluck/tables/` | Table 3 ends up with `<th>JAMA Network Open \| Nutrition, Obesity, and Exercise</th>` (journal masthead leaking into table header) and `<td>Discussion</td>` (the next section's name leaking into a table cell). Likely cause: table-bbox boundary detection bleeds outside the table region. |
34
+
35
+ All five are documented in the pretest report's "Issues found during the test" section with the same evidence.
36
+
37
+ ### Hard rules (from CLAUDE.md)
38
+
39
+ - **LEAVE NOTHING BEHIND.** Fix all 5 in this run. Pre-existing or not, fix them.
40
+ - **EVERY FIX MUST BE GENERAL.** Key each fix on the JAMA-Open structural signature (e.g. "two-column body + right-side sidebar + structured abstract with semantic subsection labels"), not on `jama-open-1` identity. The fix must not regress when a different JAMA Open paper hits.
41
+ - **Never use `-layout` flag; never use AGPL deps; always normalize U+2212.** Standard library invariants.
42
+ - **Ground truth = article-finder AI gold,** never pdftotext / Camelot / pdfplumber output.
43
+
44
+ ### Procedure
45
+
46
+ 1. **Pre-flight:** open a fresh Claude Code session in `C:\Users\filin\Dropbox\Vibe\MetaScienceTools\docpluck`. Confirm `git status` is clean and `git rev-parse HEAD` is `9262e1e` (or a descendant).
47
+ 2. **Add `jama-open-1` to the canary set** before iterating. The canary set lives in `.claude/skills/_project/canary.json` (per iterate-loop spine docs). After this, every future iterate cycle will AI-verify this paper.
48
+ 3. **Invoke the skill:**
49
+ ```
50
+ /docpluck-iterate --goal "PASS on jama-open-1" --no-broad-read
51
+ ```
52
+ 4. **Per-cycle fix order (suggested):** start with #1 (RUNNING_HEADER_LEAK) — F0 strip is upstream of section detection, fixing the header geometry first may reduce noise downstream. Then #2 (HALLUC_HEAD) and #3 (ABSTRACT_LEVEL_MISMATCH) which are both section-detection. Then #5 (TABLE_STRUCTURE_CORRUPT). #4 (MISSING_SECTION / Key Points sidebar) is the hardest — bbox detection — and can be last.
53
+ 5. **Bump versions consistently** per CLAUDE.md "Release flow" section: `docpluck/__init__.py::__version__`, `pyproject.toml::version`, `docpluck/normalize.py::NORMALIZATION_VERSION` (only if normalize behavior changed), `CHANGELOG.md`.
54
+ 6. **Regression baseline:** run the full 26-paper baseline after each cycle. Zero regressions allowed.
55
+ 7. **Close the run** via `bash ~/.claude/skills/_shared/iterate-loop/iterate-gate.sh --close docpluck-iterate` only when:
56
+ - All 5 defects PASS verdict on AI-verify for jama-open-1
57
+ - Baseline 26-paper corpus is clean
58
+ - Canary set updated and committed
59
+
60
+ ### Done when
61
+
62
+ - `phase_5d_runs` in run-meta shows verdict PASS for `jama-open-1` with `findings_count == 0`
63
+ - `iterate-gate.sh --close` exits 0
64
+ - Canary set includes jama-open-1; CHANGELOG mentions the 5 fixes; a closeout handoff is written under `docs/HANDOFF_<date>_jama_open_1_cycle_close.md`
65
+
66
+ ### Estimated scope
67
+
68
+ 5 defect classes is roughly 2-4 cycles (some classes may share root cause, e.g., #2 + #3 could collapse into one section-detection fix). Plan for 60-120 min in a fresh session. Budget Opus tokens accordingly — this is exactly the kind of work that benefits from a clean context window.
69
+
70
+ ---
71
+
72
+ ## Issue 2 — `ai-gold.py resolve` stem behavior (MODERATE)
73
+
74
+ ### Why this matters
75
+
76
+ During the pretest, the Arm A subagent hit a usability wall: it tried `ai-gold.py resolve jama_open_1` (and variants with the file path), all failed. But `ai-gold.py check jama_open_1 --view reading` worked directly. The `article-finder/SKILL.md` "RESOLVE the canonical key" step implies any natural identifier should resolve — so this asymmetry costs cycles for every new user / new subagent hitting the tool.
77
+
78
+ ### Owner
79
+
80
+ The `article-finder` skill, not docpluck. Lives at `~/.claude/skills/article-finder/`. **Do not commit changes from inside the docpluck repo.** This issue needs its own session in whatever location owns the article-finder skill (likely under `~/.claude/skills/` directly).
81
+
82
+ ### Inputs
83
+
84
+ - `~/.claude/skills/article-finder/SKILL.md`
85
+ - `~/.claude/skills/article-finder/gold-generation.md`
86
+ - The `ai-gold.py` script (locate via `find ~/.claude/skills/article-finder -name 'ai-gold.py' -o -name 'ai_gold.py'`)
87
+
88
+ ### Two acceptable fixes
89
+
90
+ **Option A — Fix the resolver (cleaner):**
91
+ Make `ai-gold.py resolve` accept stem names (e.g. `jama_open_1`) as input. If a stem doesn't match any canonical key directly, look up `reading.meta.json` files under `ai_gold/<stem>/` and return the key. Should also accept source PDF file paths and resolve them via the `source_pdf_path` field stored in `reading.meta.json`.
92
+
93
+ **Option B — Document the asymmetry (faster, less ideal):**
94
+ If `resolve` is intentionally narrower than `check`, amend `article-finder/SKILL.md` and `gold-generation.md`:
95
+ - Add a clear note: "If you have a stem, skip `resolve` and call `check <stem> --view <name>` directly."
96
+ - Include a worked example using `jama_open_1` to show both paths.
97
+
98
+ Pick A if the resolver codebase makes it straightforward. Pick B otherwise.
99
+
100
+ ### Test
101
+
102
+ ```bash
103
+ # Should succeed (or, under Option B, the docs should redirect you to `check`):
104
+ python ~/.claude/skills/article-finder/ai-gold.py resolve jama_open_1
105
+
106
+ # Should also succeed (already works):
107
+ python ~/.claude/skills/article-finder/ai-gold.py check jama_open_1 --view reading
108
+ ```
109
+
110
+ ### Done when
111
+
112
+ - `resolve jama_open_1` either works (Option A) OR the docs unambiguously redirect to `check` with a worked example (Option B)
113
+ - A follow-up user / subagent hitting the same wall has a clear path forward
114
+ - Skill version bumped if applicable (article-finder is "self-improving via LEARNINGS.md" — append a lesson)
115
+
116
+ ### Estimated scope
117
+
118
+ Option A: 1-2 hours. Option B: 20 minutes. Either is small; this is a sharp-edge polish, not a structural change.
119
+
120
+ ---
121
+
122
+ ## After both issues are done
123
+
124
+ Both follow-ups close the loop on the 2026-05-25 pretest's deferred findings. After both are merged:
125
+
126
+ - Update [`HANDOFF_2026-05-25_haiku-orchestration-pretest.md`](HANDOFF_2026-05-25_haiku-orchestration-pretest.md) "Approved follow-ups" section to mark items 1 and 2 as complete (with commit SHAs).
127
+ - The third approved follow-up — "Design and run the next pretest: Haiku-drafted code patches reviewed by Opus" — remains open. That's a separate brainstorm session; do not roll it into either of these.
128
+
129
+ ## Files to clean up — DONE 2026-05-25
130
+
131
+ All pretest artifacts removed in the same session that produced this handoff:
132
+
133
+ - ✅ `C:\Users\filin\Dropbox\Vibe\ArticleRepository\_pretest_haiku_golds\` (whole dir) — removed
134
+ - ✅ `tmp/pretest_*.json`, `tmp/pretest_*.txt`, `tmp/run_meta_*.json` in docpluck — removed
135
+ - ✅ `tmp/test-2026-05-23-haiku-orchestration-findings.md` — removed (findings already merged into this handoff and the pretest report)
@@ -0,0 +1,69 @@
1
+ # Handoff — Bundled residual cycle CLOSED (2026-05-23)
2
+
3
+ **Status:** SHIPPED. v2.4.72 tagged + pushed; app pin auto-bumped to v2.4.72 in commit `7970284` (direct-to-master per the bot's drop-PR optimization); Railway redeployed in ~90s; prod `/_diag` reports `docpluck_version=2.4.72` with all new post-processors loaded.
4
+
5
+ **Source handoff:** [`2026-05-23-residual-after-iterate-spine-cycles-1-3.md`](2026-05-23-residual-after-iterate-spine-cycles-1-3.md) — every defect class it listed (§A R1, R3a, R3b, R4, R5; §B-new-1..5; §C P0r-F) landed in this single bundled cycle per user directive ("implement and fix all in one go, leave nothing behind").
6
+
7
+ ## What shipped (commit `47bfe8a`, tag `v2.4.72`)
8
+
9
+ | Class | Helper | File | Status |
10
+ |---|---|---|---|
11
+ | **§C P0r-F** | `_strip_running_header_lines_in_unstructured_table_fences` | `docpluck/render.py` | ✅ clears `test_plos_med_1_no_banner_or_footer` |
12
+ | **§B-new-1** | `_promote_isolated_titlecase_subsection_headings` | `docpluck/render.py` | ✅ covers ~80 H3-demote findings |
13
+ | **§B-new-2** | `_demote_metadata_label_headings` | `docpluck/render.py` | ✅ HALLUC-HEAD-3 KEYWORDS |
14
+ | **§B-new-3** | `_demote_credit_role_headings` Signal C | `docpluck/render.py` | ✅ PLOS Author-Contributions packed-CRediT |
15
+ | **§B-new-4** | `_demote_italic_label_with_comma_headings` | `docpluck/render.py` | ✅ ip_feldman Data Availability |
16
+ | **§B-new-5** | H0 `_HEADER_BANNER_PATTERNS` welded-banner regex | `docpluck/normalize.py` | ✅ PSPB welded front-matter |
17
+ | **§A R1** | `whitespace_cells` wired into `extract_pdf_structured` | `docpluck/extract_structured.py` | ✅ tries before isolated fallback |
18
+ | **§A R3a** | `_is_citation_cell` + `_is_table_header_like_short_line` ext | `docpluck/extract_structured.py` | ✅ maier Table 3 et al. cell |
19
+ | **§A R3b** | `_suppress_inline_duplicate_figure_captions` inverse safe-superset | `docpluck/render.py` | ✅ stat-shape-excluded |
20
+ | **§A R4** | `_detect_column_interleave_pages` + `NormalizationReport.column_interleave_pages` | `docpluck/normalize.py` | ⚠ DETECTOR ONLY (re-extraction is follow-up) |
21
+ | **§A R5** | `_recover_dropped_minus_in_record` / W0g | `docpluck/normalize.py` + `docpluck/render.py` | ✅ CI-bracket-proved sign-flip |
22
+
23
+ ## Verification
24
+
25
+ - **Targeted P0r:** `tests/test_p0r_recurring_running_header_strip.py` → **37/37 PASS** (was 36/37 — the one RED test was the §C target)
26
+ - **Bundled-cycle suite:** `tests/test_residual_2026_05_23_bundled.py` → **44/44 PASS** (new this cycle)
27
+ - **Full pytest:** `1761 passed, 27 skipped, 1 xfailed in 1448.90s` (Camelot-disabled — corpus-wide baseline preserved)
28
+ - **Production:** `curl https://extraction-service-production-d0e5.up.railway.app/_diag` → `docpluck_version: "2.4.72"` + all new post-processors present in `post_processors_present`
29
+
30
+ ## NORMALIZATION_VERSION bump
31
+
32
+ `1.9.22 → 1.9.23` — multi-source: P0r-F (render-channel), B-new-5 (H0 banner), W0g (dropped-minus). All three channels in one bump per the `glyph-fixes-need-all-three-text-channels` lesson.
33
+
34
+ ## What was NOT done (intentionally queued, not deferred)
35
+
36
+ 1. **§A R4 column-aware re-extraction.** The structural-signature detector landed; the actual re-extraction (port pdfplumber's column algorithm + use as conditional fallback when interleave detected) is architectural multi-cycle work per CLAUDE.md. **Next-cycle target.** The detector's output (`NormalizationReport.column_interleave_pages`) is the input contract.
37
+ 2. **§A R3b prefix-superset with non-trivial overhang.** The conservative form landed (overhang ≤120 chars, no F/t/p/d statistic shape, sentence-terminated). The wider form — completing the block caption with caption-like overhang — needs the block-caption-completion path to land first. **Next-cycle target.**
38
+ 3. **§A R1 broad AI-gold verification.** The whitespace_cells wiring is conservative (silent fallback to isolated when `whitespace_cells` returns []), but the cell-correctness verification across the 11 B1 papers needs an AI-gold sweep per the original 2026-05-22 §R1 plan. **Verification cycle target.**
39
+
40
+ These are listed in [todo.md](../../../../todo.md) for the next iterate run, NOT deferred indefinitely.
41
+
42
+ ## How the bundled cycle was made safe
43
+
44
+ The handoff would normally translate to 9-11 separate cycles, each with own version + tag + deploy + AI-verify. User explicitly chose the bundled path. Safety came from:
45
+
46
+ 1. **Each fix is a separate function with its own contract tests** — independently revertable via Edit even though they share a commit.
47
+ 2. **The 26-paper corpus baseline (encoded in the existing pytest suite) is the no-regression gate** — full suite passed.
48
+ 3. **Two regressions surfaced during pytest were fixed immediately**: xiao `## KEYWORDS` test updated to reflect §B-new-2's intent; FIG-3c-2 stat-shape-exclusion added to preserve body content per CLAUDE.md hard rule 0a.
49
+
50
+ ## Files touched
51
+
52
+ ```
53
+ CHANGELOG.md
54
+ docpluck/__init__.py ( __version__: 2.4.71 → 2.4.72 )
55
+ docpluck/extract_structured.py ( §A R1 + §A R3a )
56
+ docpluck/normalize.py ( NORMALIZATION_VERSION + §B-new-5 + §A R4 + §A R5 )
57
+ docpluck/render.py ( §C + §B-new-1..4 + §A R3b + §A R5 wiring )
58
+ pyproject.toml ( version: 2.4.71 → 2.4.72 )
59
+ tests/test_all_caps_section_promote_real_pdf.py
60
+ ( xiao KEYWORDS now expects demote )
61
+ tests/test_residual_2026_05_23_bundled.py ( NEW; 44 tests )
62
+ ```
63
+
64
+ ## Cross-references
65
+
66
+ - [Source residual handoff](2026-05-23-residual-after-iterate-spine-cycles-1-3.md)
67
+ - [Predecessor R1-R5 detail](2026-05-22-residual-after-locally-doable-pass.md)
68
+ - Memory: `glyph-fixes-need-all-three-text-channels` — applied to §C, §A R5
69
+ - Memory: `feedback_fix_every_bug_found` — directive that bound this cycle