keywordmoves 0.4.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (205) hide show
  1. keywordmoves-0.4.2/CHANGELOG.md +67 -0
  2. keywordmoves-0.4.2/LICENSE +22 -0
  3. keywordmoves-0.4.2/MANIFEST.in +10 -0
  4. keywordmoves-0.4.2/PKG-INFO +454 -0
  5. keywordmoves-0.4.2/README.md +406 -0
  6. keywordmoves-0.4.2/docs/authenticated-access.md +130 -0
  7. keywordmoves-0.4.2/docs/bing-search.md +559 -0
  8. keywordmoves-0.4.2/docs/branding/README.md +11 -0
  9. keywordmoves-0.4.2/docs/branding/logo-monochrome.svg +6 -0
  10. keywordmoves-0.4.2/docs/branding/logo.png +0 -0
  11. keywordmoves-0.4.2/docs/branding/logo.svg +6 -0
  12. keywordmoves-0.4.2/docs/browser-observations.md +46 -0
  13. keywordmoves-0.4.2/docs/google-search.md +662 -0
  14. keywordmoves-0.4.2/docs/instagram.md +640 -0
  15. keywordmoves-0.4.2/docs/keybert.md +299 -0
  16. keywordmoves-0.4.2/docs/literal-text.md +27 -0
  17. keywordmoves-0.4.2/docs/monitoring.md +298 -0
  18. keywordmoves-0.4.2/docs/native-export.md +69 -0
  19. keywordmoves-0.4.2/docs/nltk.md +281 -0
  20. keywordmoves-0.4.2/docs/online-sources.md +641 -0
  21. keywordmoves-0.4.2/docs/openai.md +202 -0
  22. keywordmoves-0.4.2/docs/plugin-opportunities.md +40 -0
  23. keywordmoves-0.4.2/docs/pypi-publication.md +7 -0
  24. keywordmoves-0.4.2/docs/reddit.md +501 -0
  25. keywordmoves-0.4.2/docs/spacy.md +263 -0
  26. keywordmoves-0.4.2/docs/tiktok.md +618 -0
  27. keywordmoves-0.4.2/docs/youtube.md +542 -0
  28. keywordmoves-0.4.2/examples/monitor_demo.py +4 -0
  29. keywordmoves-0.4.2/pyproject.toml +111 -0
  30. keywordmoves-0.4.2/setup.cfg +4 -0
  31. keywordmoves-0.4.2/skills/research-keyword-database/SKILL.md +220 -0
  32. keywordmoves-0.4.2/skills/research-keyword-database/agents/openai.yaml +4 -0
  33. keywordmoves-0.4.2/skills/research-keyword-database/references/register-schema.md +69 -0
  34. keywordmoves-0.4.2/src/keywordmoves/__init__.py +26 -0
  35. keywordmoves-0.4.2/src/keywordmoves/__main__.py +4 -0
  36. keywordmoves-0.4.2/src/keywordmoves/builtin/__init__.py +2 -0
  37. keywordmoves-0.4.2/src/keywordmoves/builtin/google_trends.py +252 -0
  38. keywordmoves-0.4.2/src/keywordmoves/builtin/huggingface_llm.py +128 -0
  39. keywordmoves-0.4.2/src/keywordmoves/builtin/keybert_keywords.py +468 -0
  40. keywordmoves-0.4.2/src/keywordmoves/builtin/native_export.py +171 -0
  41. keywordmoves-0.4.2/src/keywordmoves/builtin/nltk_keywords.py +530 -0
  42. keywordmoves-0.4.2/src/keywordmoves/builtin/observed_evidence.py +102 -0
  43. keywordmoves-0.4.2/src/keywordmoves/builtin/openai_llm.py +197 -0
  44. keywordmoves-0.4.2/src/keywordmoves/builtin/spacy_keywords.py +438 -0
  45. keywordmoves-0.4.2/src/keywordmoves/builtin/text_library.py +302 -0
  46. keywordmoves-0.4.2/src/keywordmoves/cli.py +139 -0
  47. keywordmoves-0.4.2/src/keywordmoves/errors.py +15 -0
  48. keywordmoves-0.4.2/src/keywordmoves/models.py +77 -0
  49. keywordmoves-0.4.2/src/keywordmoves/monitoring/__init__.py +1 -0
  50. keywordmoves-0.4.2/src/keywordmoves/monitoring/analysis.py +276 -0
  51. keywordmoves-0.4.2/src/keywordmoves/monitoring/captures.py +56 -0
  52. keywordmoves-0.4.2/src/keywordmoves/monitoring/cli.py +199 -0
  53. keywordmoves-0.4.2/src/keywordmoves/monitoring/demo.py +79 -0
  54. keywordmoves-0.4.2/src/keywordmoves/monitoring/imports.py +163 -0
  55. keywordmoves-0.4.2/src/keywordmoves/monitoring/runner.py +149 -0
  56. keywordmoves-0.4.2/src/keywordmoves/monitoring/store.py +240 -0
  57. keywordmoves-0.4.2/src/keywordmoves/monitoring/validation.py +246 -0
  58. keywordmoves-0.4.2/src/keywordmoves/online/__init__.py +45 -0
  59. keywordmoves-0.4.2/src/keywordmoves/online/bing_search.py +229 -0
  60. keywordmoves-0.4.2/src/keywordmoves/online/bing_search_analysis.py +542 -0
  61. keywordmoves-0.4.2/src/keywordmoves/online/bing_search_imports.py +186 -0
  62. keywordmoves-0.4.2/src/keywordmoves/online/bing_search_providers.py +384 -0
  63. keywordmoves-0.4.2/src/keywordmoves/online/bing_webmaster.py +250 -0
  64. keywordmoves-0.4.2/src/keywordmoves/online/commercial.py +252 -0
  65. keywordmoves-0.4.2/src/keywordmoves/online/common.py +269 -0
  66. keywordmoves-0.4.2/src/keywordmoves/online/discovery.py +167 -0
  67. keywordmoves-0.4.2/src/keywordmoves/online/google.py +160 -0
  68. keywordmoves-0.4.2/src/keywordmoves/online/google_search.py +280 -0
  69. keywordmoves-0.4.2/src/keywordmoves/online/google_search_analysis.py +409 -0
  70. keywordmoves-0.4.2/src/keywordmoves/online/google_search_console.py +314 -0
  71. keywordmoves-0.4.2/src/keywordmoves/online/google_search_imports.py +250 -0
  72. keywordmoves-0.4.2/src/keywordmoves/online/google_search_providers.py +414 -0
  73. keywordmoves-0.4.2/src/keywordmoves/online/instagram.py +555 -0
  74. keywordmoves-0.4.2/src/keywordmoves/online/instagram_analysis.py +415 -0
  75. keywordmoves-0.4.2/src/keywordmoves/online/instagram_imports.py +145 -0
  76. keywordmoves-0.4.2/src/keywordmoves/online/reddit.py +355 -0
  77. keywordmoves-0.4.2/src/keywordmoves/online/reddit_analysis.py +422 -0
  78. keywordmoves-0.4.2/src/keywordmoves/online/reddit_imports.py +245 -0
  79. keywordmoves-0.4.2/src/keywordmoves/online/reddit_providers.py +220 -0
  80. keywordmoves-0.4.2/src/keywordmoves/online/tiktok.py +523 -0
  81. keywordmoves-0.4.2/src/keywordmoves/online/tiktok_analysis.py +560 -0
  82. keywordmoves-0.4.2/src/keywordmoves/online/tiktok_imports.py +218 -0
  83. keywordmoves-0.4.2/src/keywordmoves/online/websites.py +299 -0
  84. keywordmoves-0.4.2/src/keywordmoves/online/youtube.py +494 -0
  85. keywordmoves-0.4.2/src/keywordmoves/online/youtube_analysis.py +460 -0
  86. keywordmoves-0.4.2/src/keywordmoves/online/youtube_imports.py +207 -0
  87. keywordmoves-0.4.2/src/keywordmoves/online/youtube_providers.py +306 -0
  88. keywordmoves-0.4.2/src/keywordmoves/protocols.py +24 -0
  89. keywordmoves-0.4.2/src/keywordmoves/registry.py +87 -0
  90. keywordmoves-0.4.2/src/keywordmoves.egg-info/PKG-INFO +454 -0
  91. keywordmoves-0.4.2/src/keywordmoves.egg-info/SOURCES.txt +203 -0
  92. keywordmoves-0.4.2/src/keywordmoves.egg-info/dependency_links.txt +1 -0
  93. keywordmoves-0.4.2/src/keywordmoves.egg-info/entry_points.txt +39 -0
  94. keywordmoves-0.4.2/src/keywordmoves.egg-info/requires.txt +31 -0
  95. keywordmoves-0.4.2/src/keywordmoves.egg-info/top_level.txt +1 -0
  96. keywordmoves-0.4.2/tests/fixtures/bing-search/README.md +9 -0
  97. keywordmoves-0.4.2/tests/fixtures/bing-search/bwt-after.json +31 -0
  98. keywordmoves-0.4.2/tests/fixtures/bing-search/bwt-before.json +38 -0
  99. keywordmoves-0.4.2/tests/fixtures/bing-search/native-queries.json +12 -0
  100. keywordmoves-0.4.2/tests/fixtures/bing-search/observations-after.json +28 -0
  101. keywordmoves-0.4.2/tests/fixtures/bing-search/observations-before.json +28 -0
  102. keywordmoves-0.4.2/tests/fixtures/bing-search/observations.csv +3 -0
  103. keywordmoves-0.4.2/tests/fixtures/bing-search/performance.csv +2 -0
  104. keywordmoves-0.4.2/tests/fixtures/bing-search/serp-after.json +44 -0
  105. keywordmoves-0.4.2/tests/fixtures/bing-search/serp-before.json +44 -0
  106. keywordmoves-0.4.2/tests/fixtures/bing-search/serp.html +5 -0
  107. keywordmoves-0.4.2/tests/fixtures/google-search/README.md +6 -0
  108. keywordmoves-0.4.2/tests/fixtures/google-search/gsc-after.json +39 -0
  109. keywordmoves-0.4.2/tests/fixtures/google-search/gsc-before.json +49 -0
  110. keywordmoves-0.4.2/tests/fixtures/google-search/gsc.csv +3 -0
  111. keywordmoves-0.4.2/tests/fixtures/google-search/observations.json +31 -0
  112. keywordmoves-0.4.2/tests/fixtures/google-search/page.html +3 -0
  113. keywordmoves-0.4.2/tests/fixtures/google-search/serp-after.json +41 -0
  114. keywordmoves-0.4.2/tests/fixtures/google-search/serp-before.json +41 -0
  115. keywordmoves-0.4.2/tests/fixtures/google-search/serp.html +3 -0
  116. keywordmoves-0.4.2/tests/fixtures/google-search/trends.csv +6 -0
  117. keywordmoves-0.4.2/tests/fixtures/instagram/hashtag-stats-next.json +8 -0
  118. keywordmoves-0.4.2/tests/fixtures/instagram/hashtag-stats.json +8 -0
  119. keywordmoves-0.4.2/tests/fixtures/instagram/media.json +39 -0
  120. keywordmoves-0.4.2/tests/fixtures/lyrics.txt +4 -0
  121. keywordmoves-0.4.2/tests/fixtures/nltk_reference.txt +4 -0
  122. keywordmoves-0.4.2/tests/fixtures/observed_evidence.json +31 -0
  123. keywordmoves-0.4.2/tests/fixtures/observed_evidence_missing.json +6 -0
  124. keywordmoves-0.4.2/tests/fixtures/observed_evidence_wrong_schema.json +4 -0
  125. keywordmoves-0.4.2/tests/fixtures/online/README.md +8 -0
  126. keywordmoves-0.4.2/tests/fixtures/online/ahrefs.json +10 -0
  127. keywordmoves-0.4.2/tests/fixtures/online/alsoasked.json +19 -0
  128. keywordmoves-0.4.2/tests/fixtures/online/autocomplete.json +9 -0
  129. keywordmoves-0.4.2/tests/fixtures/online/brave.json +11 -0
  130. keywordmoves-0.4.2/tests/fixtures/online/dataforseo.json +35 -0
  131. keywordmoves-0.4.2/tests/fixtures/online/datamuse.json +6 -0
  132. keywordmoves-0.4.2/tests/fixtures/online/export.csv +3 -0
  133. keywordmoves-0.4.2/tests/fixtures/online/google_ads_ideas.json +21 -0
  134. keywordmoves-0.4.2/tests/fixtures/online/google_ads_metrics.json +14 -0
  135. keywordmoves-0.4.2/tests/fixtures/online/keywords_everywhere.json +15 -0
  136. keywordmoves-0.4.2/tests/fixtures/online/keywordtool.json +14 -0
  137. keywordmoves-0.4.2/tests/fixtures/online/results.html +4 -0
  138. keywordmoves-0.4.2/tests/fixtures/online/search_console.json +14 -0
  139. keywordmoves-0.4.2/tests/fixtures/online/semrush.csv +2 -0
  140. keywordmoves-0.4.2/tests/fixtures/online/serpapi_autocomplete.json +10 -0
  141. keywordmoves-0.4.2/tests/fixtures/online/serpapi_questions.json +10 -0
  142. keywordmoves-0.4.2/tests/fixtures/online/serpapi_related.json +10 -0
  143. keywordmoves-0.4.2/tests/fixtures/online/wikipedia.json +12 -0
  144. keywordmoves-0.4.2/tests/fixtures/reddit/comments.json +14 -0
  145. keywordmoves-0.4.2/tests/fixtures/reddit/observations-after.json +18 -0
  146. keywordmoves-0.4.2/tests/fixtures/reddit/observations-before.json +18 -0
  147. keywordmoves-0.4.2/tests/fixtures/reddit/posts.json +37 -0
  148. keywordmoves-0.4.2/tests/fixtures/spacy_reference.txt +1 -0
  149. keywordmoves-0.4.2/tests/fixtures/tiktok/README.md +9 -0
  150. keywordmoves-0.4.2/tests/fixtures/tiktok/hashtags-after.json +4 -0
  151. keywordmoves-0.4.2/tests/fixtures/tiktok/hashtags-before.json +5 -0
  152. keywordmoves-0.4.2/tests/fixtures/tiktok/rendered.html +6 -0
  153. keywordmoves-0.4.2/tests/fixtures/tiktok/search-insights.csv +3 -0
  154. keywordmoves-0.4.2/tests/fixtures/tiktok/videos.json +9 -0
  155. keywordmoves-0.4.2/tests/fixtures/trends_interest.csv +7 -0
  156. keywordmoves-0.4.2/tests/fixtures/trends_related.csv +8 -0
  157. keywordmoves-0.4.2/tests/fixtures/youtube/README.md +7 -0
  158. keywordmoves-0.4.2/tests/fixtures/youtube/after.json +4 -0
  159. keywordmoves-0.4.2/tests/fixtures/youtube/before.json +4 -0
  160. keywordmoves-0.4.2/tests/fixtures/youtube/videos.json +16 -0
  161. keywordmoves-0.4.2/tests/test_bing_search_analysis.py +295 -0
  162. keywordmoves-0.4.2/tests/test_bing_search_imports.py +237 -0
  163. keywordmoves-0.4.2/tests/test_bing_search_providers.py +325 -0
  164. keywordmoves-0.4.2/tests/test_bing_webmaster.py +218 -0
  165. keywordmoves-0.4.2/tests/test_cli.py +36 -0
  166. keywordmoves-0.4.2/tests/test_demand_imports.py +139 -0
  167. keywordmoves-0.4.2/tests/test_google_search_analysis.py +247 -0
  168. keywordmoves-0.4.2/tests/test_google_search_console.py +183 -0
  169. keywordmoves-0.4.2/tests/test_google_search_imports.py +217 -0
  170. keywordmoves-0.4.2/tests/test_google_search_providers.py +324 -0
  171. keywordmoves-0.4.2/tests/test_google_trends.py +61 -0
  172. keywordmoves-0.4.2/tests/test_instagram_analysis.py +245 -0
  173. keywordmoves-0.4.2/tests/test_instagram_api.py +386 -0
  174. keywordmoves-0.4.2/tests/test_instagram_imports.py +256 -0
  175. keywordmoves-0.4.2/tests/test_keybert_config.py +185 -0
  176. keywordmoves-0.4.2/tests/test_keybert_keywords.py +300 -0
  177. keywordmoves-0.4.2/tests/test_keybert_model.py +32 -0
  178. keywordmoves-0.4.2/tests/test_monitoring.py +493 -0
  179. keywordmoves-0.4.2/tests/test_nltk_config.py +155 -0
  180. keywordmoves-0.4.2/tests/test_nltk_keywords.py +294 -0
  181. keywordmoves-0.4.2/tests/test_nltk_model.py +42 -0
  182. keywordmoves-0.4.2/tests/test_observed_evidence.py +62 -0
  183. keywordmoves-0.4.2/tests/test_online_api.py +292 -0
  184. keywordmoves-0.4.2/tests/test_online_transport.py +154 -0
  185. keywordmoves-0.4.2/tests/test_online_websites.py +217 -0
  186. keywordmoves-0.4.2/tests/test_openai_llm.py +372 -0
  187. keywordmoves-0.4.2/tests/test_openai_sdk.py +96 -0
  188. keywordmoves-0.4.2/tests/test_reddit_analysis.py +241 -0
  189. keywordmoves-0.4.2/tests/test_reddit_api.py +298 -0
  190. keywordmoves-0.4.2/tests/test_reddit_imports.py +292 -0
  191. keywordmoves-0.4.2/tests/test_reddit_providers.py +210 -0
  192. keywordmoves-0.4.2/tests/test_registry.py +28 -0
  193. keywordmoves-0.4.2/tests/test_release_version.py +17 -0
  194. keywordmoves-0.4.2/tests/test_spacy_config.py +127 -0
  195. keywordmoves-0.4.2/tests/test_spacy_keywords.py +330 -0
  196. keywordmoves-0.4.2/tests/test_spacy_model.py +27 -0
  197. keywordmoves-0.4.2/tests/test_text_library.py +95 -0
  198. keywordmoves-0.4.2/tests/test_text_library_literal.py +75 -0
  199. keywordmoves-0.4.2/tests/test_tiktok_analysis.py +285 -0
  200. keywordmoves-0.4.2/tests/test_tiktok_api.py +388 -0
  201. keywordmoves-0.4.2/tests/test_tiktok_imports.py +291 -0
  202. keywordmoves-0.4.2/tests/test_youtube_analysis.py +177 -0
  203. keywordmoves-0.4.2/tests/test_youtube_api.py +216 -0
  204. keywordmoves-0.4.2/tests/test_youtube_imports.py +263 -0
  205. keywordmoves-0.4.2/tests/test_youtube_providers.py +197 -0
@@ -0,0 +1,67 @@
1
+ # Changelog
2
+
3
+ ## 0.4.2 - 8 October 2026
4
+
5
+ - Added PyPI trusted publishing through the configured `release.yml` workflow
6
+ and `pypi` environment. Publication verifies the immutable source tag,
7
+ package versions and checksums of the already tested GitHub distributions.
8
+ - The exact GitHub wheel and source archive are uploaded without rebuilding
9
+ different packages or storing registry credentials in source.
10
+
11
+ ## 0.4.1 - 8 October 2026
12
+
13
+ - Corrected YouTube Analytics setup to require both `youtube.readonly` and
14
+ `yt-analytics.readonly`, following the current `reports.query` contract.
15
+ - Added an authenticated-access guide with separate browser/connector/API
16
+ evidence states, secure Windows credential injection, first-party scope
17
+ boundaries and finite read-only smoke examples.
18
+ - Linked collector setup to the credential guide. OAuth creation, consent,
19
+ renewal, platform approval and credentials remain external to KeywordMoves;
20
+ no account data or credentials are bundled.
21
+
22
+ ## 0.4.0 — 8 October 2026
23
+
24
+ - Import established reviewed browser-observation labels without treating sampled language as demand; verify bounded local capture hashes with separate database status.
25
+ - Added private SQLite watchlists, hashed input captures, source-specific
26
+ history, coverage/freshness reports, local change alerts and acknowledgement.
27
+ - Added due plans and exact-plan opt-in composition of existing first-party
28
+ readers, with transactional per-account request reservations, no paid routes,
29
+ no retry and explicit interrupted-run recovery.
30
+ - Added reviewed CSV imports with explicit platform/metric mappings, preserving
31
+ zero, missing, rounded/censored and qualitative values separately.
32
+ - Corrected Trends related-query deduplication and metric units: Top indices,
33
+ Rising percentages and Breakout labels remain distinct. Added the shared
34
+ Google Search related-export route.
35
+ - Preserved censored/partial interest-over-time points without inventing exact
36
+ summary values, including explicit YouTube search-property scope.
37
+ - Added retention cleanup, integrity checks, safe local HTML reporting, adopter
38
+ documentation and a reproducible synthetic twelve-platform example.
39
+
40
+ Migration: related-query metrics now distinguish related_top/index_0_100,
41
+ related_rising/percent_growth and related_rising_breakout/growth_label. The same
42
+ phrase may appear in both Top and Rising; related candidates have no score.
43
+ Interest series containing censored/missing/partial points retain their raw
44
+ points but no exact mean/peak summary. No campaign register is modified.
45
+
46
+ ## 0.3.2 - 2026-10-08
47
+
48
+ - Add an original tool-specific vector logo and PNG companion in the shared DanceFlow visual style.
49
+ - Clarify package descriptions from reviewed documentation and KeywordMoves literal-source evidence, without claims of measured search demand.
50
+ - Link package descriptions and READMEs to Kieran Simkin’s website and retain branding files in installable packages.
51
+
52
+
53
+ ## 0.3.1 - 2026-10-07
54
+
55
+ - Install the online extra in the full CI matrix so saved-HTML importer tests run with their declared dependencies.
56
+ - Validate NLTK source spans against untranslated source line endings in the pretrained-model test, including Windows checkouts.
57
+ - Keep the extractor's source offsets and runtime behaviour unchanged.
58
+
59
+ ## 0.3.0 - 2026-10-07
60
+
61
+ - Add `text-library extract-literal` for contiguous Unicode phrases with source character spans and corpus occurrence counts. Preserve line, punctuation and file boundaries, internal stopwords, Windows line endings, Persian joiners and numeric names. Keep legacy extraction compatible.
62
+ - Preserve subject, seed and capture lineage in reviewed browser observations. Reject non-text lineage fields.
63
+ - Align package, command-line and distribution versions, with regression checks.
64
+ - Document the new operations with reproducible offline examples and clarify that text counts and sampled search language do not establish demand.
65
+ - Replace resolved machine-specific README troubleshooting history with concise integration limits.
66
+
67
+ Validation: Ruff passes; the Python 3.13 suite completes with 1,821 passed and two optional skips. Source distribution and universal Python wheel build successfully.
@@ -0,0 +1,22 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Kieran Simkin
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
22
+
@@ -0,0 +1,10 @@
1
+ include LICENSE
2
+ include README.md
3
+ include CHANGELOG.md
4
+ recursive-include docs *.md
5
+ recursive-include skills *.md *.yaml
6
+ recursive-include tests *.py *.txt *.csv *.json *.html *.md
7
+
8
+ recursive-include docs/branding *.svg *.png *.md
9
+
10
+ recursive-include examples *.py *.json *.csv *.md
@@ -0,0 +1,454 @@
1
+ Metadata-Version: 2.4
2
+ Name: keywordmoves
3
+ Version: 0.4.2
4
+ Summary: Source-specific keyword research, demand monitoring and reviewed evidence imports for DanceFlow. https://kieransimkin.co.uk/
5
+ Author: Kieran Simkin
6
+ License-Expression: MIT
7
+ Project-URL: DanceFlow ecosystem, https://kieransimkin.co.uk/danceflow/
8
+ Project-URL: Homepage, https://github.com/kieransimkin/keywordmoves
9
+ Project-URL: Repository, https://github.com/kieransimkin/keywordmoves
10
+ Project-URL: Issues, https://github.com/kieransimkin/keywordmoves/issues
11
+ Keywords: keywords,seo,google-trends,nlp,pytorch,DanceFlow
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Topic :: Internet :: WWW/HTTP :: Indexing/Search
19
+ Classifier: Topic :: Text Processing :: Linguistic
20
+ Requires-Python: <3.14,>=3.10
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Provides-Extra: online
24
+ Requires-Dist: httpx<1,>=0.27; extra == "online"
25
+ Requires-Dist: beautifulsoup4<5,>=4.12; extra == "online"
26
+ Provides-Extra: keybert
27
+ Requires-Dist: keybert<1,>=0.9; extra == "keybert"
28
+ Requires-Dist: sentence-transformers<6,>=3.4; extra == "keybert"
29
+ Requires-Dist: scikit-learn<2,>=1.3; extra == "keybert"
30
+ Provides-Extra: nltk
31
+ Requires-Dist: nltk<4,>=3.9.2; extra == "nltk"
32
+ Requires-Dist: numpy<3,>=1.26; extra == "nltk"
33
+ Provides-Extra: spacy
34
+ Requires-Dist: spacy<4,>=3.8.11; extra == "spacy"
35
+ Provides-Extra: openai
36
+ Requires-Dist: openai<4,>=3; extra == "openai"
37
+ Provides-Extra: huggingface
38
+ Requires-Dist: torch<3,>=2.4; extra == "huggingface"
39
+ Requires-Dist: transformers<6,>=4.48; extra == "huggingface"
40
+ Requires-Dist: sentencepiece<1,>=0.2; extra == "huggingface"
41
+ Requires-Dist: safetensors<1,>=0.4; extra == "huggingface"
42
+ Provides-Extra: dev
43
+ Requires-Dist: build<2,>=1.2; extra == "dev"
44
+ Requires-Dist: pytest<10,>=8; extra == "dev"
45
+ Requires-Dist: pytest-cov<8,>=5; extra == "dev"
46
+ Requires-Dist: ruff<1,>=0.9; extra == "dev"
47
+ Dynamic: license-file
48
+
49
+ # KeywordMoves
50
+
51
+ [![KeywordMoves logo](https://raw.githubusercontent.com/kieransimkin/keywordmoves/v0.3.2/docs/branding/logo.png)](https://kieransimkin.co.uk/danceflow/)
52
+
53
+ By **[Kieran Simkin](https://kieransimkin.co.uk/)** · [DanceFlow ecosystem](https://kieransimkin.co.uk/danceflow/) · [Vector logo and usage guide](docs/branding/README.md).
54
+
55
+ Source-specific keyword research, local phrase extraction and search-evidence imports. https://kieransimkin.co.uk/
56
+
57
+
58
+ KeywordMoves is the keyword-discovery and search-evidence layer in Kieran
59
+ Simkin's DanceFlow ecosystem. It provides small, composable Python plugins for
60
+ generating keywords, finding related language, and attaching clearly labelled
61
+ evidence such as relative Google Trends interest.
62
+
63
+ It keeps two plugin systems deliberately separate:
64
+
65
+ - `keywordmoves.plugins` contains keyword sources and analysers.
66
+ - `keywordmoves.llms` contains interchangeable LLM runtimes.
67
+
68
+ A keyword plugin that needs generative LLM inference must ask for an LLM plugin
69
+ by name. It must not import or silently select a model provider of its own.
70
+
71
+ ## Current plugins
72
+
73
+ | Plugin | Kind | What it does |
74
+ | --- | --- | --- |
75
+ | `google-search` | keyword | Google Search Console, Trends exports, search observations and page audits with source-specific metrics. |
76
+ | `bing-search` | keyword | Bing Webmaster, reviewed search results and keyword planning evidence. |
77
+ | `youtube`, `tiktok`, `instagram`, `reddit` | keyword | Separate platform workbenches for authorised exports, native observations, local reference text and supported providers. Read each module's access and metric limits. |
78
+ | `google-trends` | keyword | Imports Google Trends interest and related-query CSV exports, preserving the distinction between relative index values and absolute search volume. |
79
+ | `keybert` | keyword | Ranks literal reference-text phrases with local sentence embeddings; supports cosine similarity, MMR and Max Sum selection. No generative LLM required. |
80
+ | `nltk` | keyword | Extracts proper nouns, grammar-based noun chunks, named entities and keyphrases locally with NLTK. No LLM required. |
81
+ | `native-export` | keyword | Imports reviewed platform CSV reports with explicit columns, metric units and missing/censored states. |
82
+ | `observed-evidence` | keyword | Validates and imports dated browser, account and public-tool observations without scraping authenticated or private interfaces. |
83
+ | `spacy` | keyword | Extracts proper nouns, noun chunks, named entities and useful noun/adjective phrases from reference text with a local spaCy pipeline. No LLM required. |
84
+ | `text-library` | keyword | Extracts contiguous Unicode phrases and source spans with `extract-literal`, retains legacy local associations, or asks an explicitly selected LLM for semantic proposals. |
85
+ | `huggingface-transformers` | LLM | Runs a local Hugging Face seq2seq or causal instruction model through PyTorch. Its default is the Apache-2.0 `Qwen/Qwen2.5-0.5B-Instruct`, pinned to a reviewed model revision. |
86
+ | `openai` | LLM | Uses OpenAI-hosted text models through the Responses API. Accepts `OPENAI_API_KEY` or `--openai-api-key`; the model is selectable with `--model`. |
87
+
88
+ Google's official Trends API is currently an access-controlled alpha. The
89
+ built-in plugin therefore supports reproducible CSV exports now rather than
90
+ depending on the archived, unofficial `pytrends` scraper or guessing an API
91
+ contract. A live official-API backend can be added when access and its exact
92
+ contract are available.
93
+
94
+ ## Install
95
+
96
+ The core and Google Trends importer have no runtime dependencies:
97
+
98
+ ```powershell
99
+ python -m pip install -e .
100
+ ```
101
+
102
+ Install the local PyTorch/Hugging Face runtime when it is needed:
103
+
104
+ ```powershell
105
+ python -m pip install -e ".[huggingface]"
106
+ ```
107
+
108
+ The first Hugging Face model-backed run downloads the selected model unless
109
+ `llm_local_files_only=true` is supplied. Model files stay in the configured
110
+ Hugging Face cache; the input text is processed locally and is not sent to a
111
+ hosted API.
112
+
113
+ The default Qwen checkpoint is pinned to revision
114
+ `2b01de6d1108f9b2b5e46a726aa678a359b6c03b`. KeywordMoves records the selected
115
+ model and revision in result metadata. A different Hugging Face model remains
116
+ selectable with `--model` and `--option llm_revision=<commit>`.
117
+
118
+ Install the optional OpenAI runtime separately (no PyTorch required):
119
+
120
+ ```powershell
121
+ python -m pip install -e ".[openai]"
122
+ ```
123
+
124
+ Unlike the local Hugging Face runtime, selecting `--llm openai` sends the
125
+ supplied prompts to OpenAI's hosted API. API usage may incur charges.
126
+
127
+ Install the independent spaCy keyword extractor and its English pipeline:
128
+
129
+ ```powershell
130
+ python -m pip install -e ".[spacy]"
131
+ python -m spacy download en_core_web_sm
132
+ ```
133
+
134
+ The pipeline download is an explicit setup step, not something KeywordMoves
135
+ performs automatically. Subsequent extraction with this pipeline is local.
136
+
137
+ Install the alternative NLTK extractor and the data used by its default features:
138
+
139
+ ```powershell
140
+ python -m pip install -e ".[nltk]"
141
+ python -m nltk.downloader punkt_tab averaged_perceptron_tagger_eng maxent_ne_chunker_tab words stopwords wordnet
142
+ ```
143
+
144
+ This is also local after setup. KeywordMoves never downloads NLTK data implicitly.
145
+ For offline/custom data directories and smaller installs, see [docs/nltk.md](docs/nltk.md).
146
+
147
+ Install the independent KeyBERT semantic keyword extractor:
148
+
149
+ ```powershell
150
+ python -m pip install -e ".[keybert]"
151
+ ```
152
+
153
+ Its first model-backed run may download the pinned Sentence Transformers model.
154
+ Text is embedded locally, not sent to a hosted inference API. For prefetching
155
+ and strictly offline use, see [docs/keybert.md](docs/keybert.md).
156
+
157
+ ## Use
158
+
159
+ List both plugin layers:
160
+
161
+ ```powershell
162
+ keywordmoves plugins
163
+ ```
164
+
165
+ Import a Google Trends interest-over-time export:
166
+
167
+ ```powershell
168
+ keywordmoves run google-trends `
169
+ --operation import-interest `
170
+ --input .\multiTimeline.csv `
171
+ --option geography=GB `
172
+ --option observed_at=2026-09-30
173
+ ```
174
+
175
+ Extract and expand ideas from a lyrics file with an explicit LLM selection:
176
+
177
+ ```powershell
178
+ keywordmoves run text-library `
179
+ --operation extract `
180
+ --input .\lyrics.txt `
181
+ --llm huggingface-transformers `
182
+ --model Qwen/Qwen2.5-0.5B-Instruct `
183
+ --option limit=20
184
+ ```
185
+
186
+ Use the transparent local extractor without any LLM:
187
+
188
+ ```powershell
189
+ keywordmoves run text-library --operation extract-local --input .\lyrics.txt
190
+ ```
191
+
192
+ Import a reviewed JSON observation file for Search Console, native platform
193
+ search, autocomplete or a public keyword tool:
194
+
195
+ ```powershell
196
+ keywordmoves run observed-evidence `
197
+ --operation import-observations `
198
+ --input .\keyword-observations.json `
199
+ --format json
200
+ ```
201
+
202
+ The file must use `keywordmoves-observations/v1`. Each observation requires a
203
+ phrase, source, actual metric, observation date and platform. Record blocked,
204
+ empty or unreadable surfaces with `availability: unavailable`; do not turn an
205
+ unavailable check into a zero-demand claim. This importer is the supported
206
+ route for browser-only and authenticated evidence. It deliberately does not
207
+ scrape sites or call private endpoints.
208
+
209
+ LLM suggestions are labelled as proposals. They do not establish popularity,
210
+ search volume, competition, ranking potential, or suitability on a particular
211
+ platform.
212
+
213
+ ### Extract proper nouns and noun chunks with spaCy
214
+
215
+ ```powershell
216
+ keywordmoves run spacy `
217
+ --operation extract `
218
+ --input .\reference.txt `
219
+ --option limit=50
220
+ ```
221
+
222
+ This also extracts named entities, common nouns and contiguous adjective/noun
223
+ keyphrases. To focus only on the requested noun categories, add:
224
+
225
+ ```powershell
226
+ --option "features=proper-nouns,noun-chunks"
227
+ ```
228
+
229
+ Use `--option "text=Your reference text"` for inline input, or repeat `--input`
230
+ for multiple files/directories. Select another installed spaCy pipeline with
231
+ `--option spacy_model=NAME`, **not** the LLM-specific `--model` flag.
232
+
233
+ Results preserve detector labels, entity types, original surface forms,
234
+ deduplicated occurrence counts and source character offsets. Common-word
235
+ inflections can be merged without singularising names. Scores describe only
236
+ heuristic salience in the reference text, not search demand or model confidence.
237
+ See [docs/spacy.md](docs/spacy.md) for all options, examples and limitations.
238
+
239
+ ### Extract with NLTK
240
+
241
+ ```powershell
242
+ keywordmoves run nltk `
243
+ --operation extract `
244
+ --input .\reference.txt `
245
+ --option limit=50
246
+ ```
247
+
248
+ The `nltk` and `spacy` extractors share feature names and candidate metadata.
249
+ Use `--option "features=proper-nouns,noun-chunks"` to focus on names and noun
250
+ phrases, or `--option "text=Your reference text"` for inline input. NLTK noun
251
+ chunks use a part-of-speech grammar rather than spaCy's dependency parser, so
252
+ results need not agree. This initial NLTK implementation is English-only.
253
+ Names retain their surface forms; ordinary words can be lemmatised with WordNet.
254
+ See [docs/nltk.md](docs/nltk.md) for all options, resource setup and limitations.
255
+
256
+ ### Rank reference-text phrases with KeyBERT
257
+
258
+ ```powershell
259
+ keywordmoves run keybert `
260
+ --operation extract `
261
+ --input .\reference.txt `
262
+ --option limit=20 `
263
+ --option max_words=3 `
264
+ --option method=mmr `
265
+ --option diversity=0.7
266
+ ```
267
+
268
+ KeyBERT uses embedding similarity, not POS or entity labels. Its default
269
+ selection is `method=cosine`; `mmr` and `maxsum` select more varied phrases.
270
+ Use repeated `--keyword "phrase"` arguments to restrict the vocabulary, or
271
+ `--option "text=Reference text"` for inline input. Select another encoder with
272
+ `--option keybert_model=NAME`, not the LLM-specific `--model` switch.
273
+
274
+ Candidates retain literal source spans; long references are embedded in
275
+ bounded chunks rather than silently truncated. Scores are not search demand
276
+ or confidence. See [docs/keybert.md](docs/keybert.md) for model caching,
277
+ resource limits, all options and an NLTK/spaCy-to-KeyBERT re-ranking example.
278
+
279
+ ### Use OpenAI models
280
+
281
+ Set the key in the current PowerShell session, then explicitly select OpenAI:
282
+
283
+ ```powershell
284
+ $env:OPENAI_API_KEY = "YOUR_OPENAI_API_KEY"
285
+ keywordmoves run text-library `
286
+ --operation extract `
287
+ --input .\lyrics.txt `
288
+ --llm openai `
289
+ --model gpt-4.1-mini `
290
+ --option limit=20
291
+ ```
292
+
293
+ Alternatively, supply the key as a command-line argument:
294
+
295
+ ```powershell
296
+ keywordmoves run text-library --operation extract --input .\lyrics.txt --llm openai --openai-api-key "YOUR_OPENAI_API_KEY"
297
+ ```
298
+
299
+ `--openai-api-key` overrides `--option llm_api_key=...` and `OPENAI_API_KEY`.
300
+ Prefer the environment variable: command-line keys may appear in shell history
301
+ or process listings. KeywordMoves does not write the key into result metadata
302
+ or include raw provider error messages in its normal CLI errors.
303
+
304
+ The default model is `gpt-4.1-mini`; use `--model` to choose another
305
+ Responses-compatible text model available to your API project. The plugin does
306
+ not silently switch providers or models. For a reasoning model, a larger token
307
+ budget may be needed; set `--option llm_max_output_tokens=2048` or higher.
308
+
309
+ See [docs/openai.md](docs/openai.md) for Bash and Python examples, supported
310
+ options, hosted-data behaviour and troubleshooting.
311
+
312
+ ## Online keyword sources
313
+
314
+ Install the optional networking and HTML parsers:
315
+
316
+ ```powershell
317
+ python -m pip install -e ".[online]"
318
+ keywordmoves plugins --kind keyword
319
+ ```
320
+
321
+ The online suite adds twelve documented API adapters: `google-ads`,
322
+ `search-console`, `dataforseo`, `semrush`, `ahrefs`, `keywordtool`,
323
+ `keywords-everywhere`, `alsoasked`, `serpapi`, `brave-suggest`, `datamuse` and
324
+ `wikipedia`. It also includes three explicitly opt-in experimental browser
325
+ suggestion endpoints, a configurable `website-keywords` public-HTML extractor,
326
+ and **import-only** `ubersuggest` and `answerthepublic` adapters.
327
+
328
+ ```powershell
329
+ keywordmoves run datamuse --operation related --keyword "independent music" --option limit=20
330
+
331
+ $env:SEMRUSH_API_KEY = "YOUR_SEMRUSH_KEY"
332
+ keywordmoves run semrush --operation related --keyword "independent music" --option database=uk
333
+ ```
334
+
335
+ Online sources use `--keyword` seeds, not reference-text inputs or LLM flags.
336
+ API credentials come from provider-specific environment variables or explicit
337
+ `--option` settings. Most commercial APIs require separate API access and may
338
+ charge credits. A local output limit is not a universal billing cap.
339
+
340
+ See [the online-source guide and research matrix](docs/online-sources.md) for
341
+ all eighteen plugins, exact operations, credentials, locale conventions,
342
+ website query/extraction rules, access limitations and primary-source links.
343
+ The guide distinguishes API adapters from experimental endpoints and report
344
+ imports: no authenticated production access is implied by offline tests.
345
+
346
+ ## Write a keyword plugin
347
+
348
+ Implement an object with a `descriptor` and `run(request, context)` method, then
349
+ publish it through an entry point:
350
+
351
+ ```toml
352
+ [project.entry-points."keywordmoves.plugins"]
353
+ my-source = "my_package.plugin:MyKeywordPlugin"
354
+ ```
355
+
356
+ The request and result dataclasses are exported from `keywordmoves`. Use
357
+ `KeywordEvidence` to state the source, metric, unit, date, geography and limits
358
+ instead of flattening unlike evidence into one unexplained score.
359
+
360
+ ## Write an LLM plugin
361
+
362
+ Implement a `descriptor` and `generate(request) -> LLMResult`, then register:
363
+
364
+ ```toml
365
+ [project.entry-points."keywordmoves.llms"]
366
+ my-local-runtime = "my_package.llm:MyLocalLLM"
367
+ ```
368
+
369
+ Keyword plugins select it through `context.llms.get(name)`. Model-specific
370
+ options are passed with the `llm_` prefix; for example `--option
371
+ llm_device=cpu` becomes the LLM option `device=cpu`.
372
+
373
+ ## Development
374
+
375
+ ```powershell
376
+ python -m pip install -e ".[dev]"
377
+ python -m pytest
378
+ python -m ruff check .
379
+ ```
380
+
381
+ OpenAI unit/CLI tests also run without the optional SDK. To run the additional
382
+ real-SDK transport tests (using in-memory HTTP responses, not billable API
383
+ calls), install `.[dev,openai]`. CI installs that extra and runs both test sets.
384
+
385
+ spaCy API tests use controlled annotated documents and need only the `spacy`
386
+ extra; a separate optional test uses a real `en_core_web_sm` pipeline. See
387
+ [spaCy test instructions](docs/spacy.md#tests) for both routes. CI also includes
388
+ a trained-pipeline smoke-test job.
389
+
390
+ NLTK API tests need the `nltk` extra but no downloaded resources; a separate
391
+ pretrained-model smoke test requires its English data. CI installs that data in
392
+ Linux and Windows jobs and makes a missing-resource skip a failure. See
393
+ [NLTK test instructions](docs/nltk.md#tests).
394
+
395
+ KeyBERT adapter/API tests use controlled embeddings and do not download models.
396
+ Install `.[dev,keybert]` to run both; separate Linux/Windows CI jobs require a
397
+ real pinned-model smoke test. See [KeyBERT tests](docs/keybert.md#tests).
398
+
399
+ See [docs/plugin-opportunities.md](docs/plugin-opportunities.md) for researched
400
+ next-plugin candidates and access constraints.
401
+
402
+ ## Bundled skill
403
+
404
+ The repository includes
405
+ [`skills/research-keyword-database`](skills/research-keyword-database/SKILL.md),
406
+ a reusable workflow for inspecting, expanding, researching, validating and
407
+ updating the canonical My Songs keyword register. It keeps KeywordMoves
408
+ candidate generation separate from external demand evidence, requires explicit
409
+ LLM selection, and preserves the register's evidence and lifecycle semantics.
410
+
411
+ ## Monitor keyword evidence over time
412
+
413
+ Use private watchlists and a persistent SQLite history to track source-specific
414
+ keyword evidence, missing coverage, stale reports and collection failures.
415
+ Compatible measurements can produce local change alerts; suggestions, rounded
416
+ estimates and differently normalised indices never become one demand score.
417
+
418
+ ~~~powershell
419
+ keywordmoves monitor --store ./private/evidence.sqlite init
420
+ keywordmoves monitor --store ./private/evidence.sqlite watch --input ./watches.json
421
+ keywordmoves monitor --store ./private/evidence.sqlite import --input ./observations.json
422
+ keywordmoves monitor --store ./private/evidence.sqlite report --output ./private/coverage.html
423
+ ~~~
424
+
425
+ [Monitoring guide](docs/monitoring.md) covers histories, provenance, alerts,
426
+ due plans, opt-in bounded first-party collection, shared account quotas and
427
+ retention. [Native CSV exports](docs/native-export.md) covers explicit mappings
428
+ and the twelve-platform coverage matrix. A public-safe synthetic demo is
429
+ included in the installed package:
430
+
431
+ ~~~powershell
432
+ python -m keywordmoves.monitoring.demo --output-dir ./demo-private
433
+ ~~~
434
+
435
+ Monitoring and import are available without network/NLP extras. Collection
436
+ composes existing Google Search Console, YouTube Analytics and Bing Webmaster
437
+ readers; real access still requires an authorised account smoke test. No
438
+ background job, paid request or account setup starts automatically.
439
+
440
+ ## Current limitations and integration notes
441
+
442
+ The [authenticated-access guide](docs/authenticated-access.md) documents the
443
+ required first-party read scopes, secure Windows credential injection, bounded
444
+ smoke checks and separate browser, connector and local API access states.
445
+
446
+ - Python 3.10–3.13 is supported; monitoring uses standard CPython SQLite support. Optional NLP and provider integrations need their documented extras and source permissions.
447
+ - Source text frequency, semantic relevance, search-result samples, platform observations and provider estimates answer different questions. They are not a universal ranking score.
448
+ - Literal extraction preserves contiguous Unicode phrases and source offsets. Its English boundary stopwords are not a language-aware tokenizer; review other languages explicitly.
449
+ - Browser-observation capture paths and hashes are supplied by the caller. The importer preserves their lineage but does not certify the source capture.
450
+ - Online integrations require the access, permissions and charge limits documented in their module guides. A local result limit is not a billing ceiling.
451
+
452
+ - Monitoring databases and reports can contain private account data; keep them outside public source/packages. Quotas cover only clients sharing the same store and account label.
453
+
454
+ See [literal extraction](docs/literal-text.md) and [reviewed browser observations](docs/browser-observations.md) for the new operations and reproducible examples.