keywordmoves 0.4.2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- keywordmoves-0.4.2/CHANGELOG.md +67 -0
- keywordmoves-0.4.2/LICENSE +22 -0
- keywordmoves-0.4.2/MANIFEST.in +10 -0
- keywordmoves-0.4.2/PKG-INFO +454 -0
- keywordmoves-0.4.2/README.md +406 -0
- keywordmoves-0.4.2/docs/authenticated-access.md +130 -0
- keywordmoves-0.4.2/docs/bing-search.md +559 -0
- keywordmoves-0.4.2/docs/branding/README.md +11 -0
- keywordmoves-0.4.2/docs/branding/logo-monochrome.svg +6 -0
- keywordmoves-0.4.2/docs/branding/logo.png +0 -0
- keywordmoves-0.4.2/docs/branding/logo.svg +6 -0
- keywordmoves-0.4.2/docs/browser-observations.md +46 -0
- keywordmoves-0.4.2/docs/google-search.md +662 -0
- keywordmoves-0.4.2/docs/instagram.md +640 -0
- keywordmoves-0.4.2/docs/keybert.md +299 -0
- keywordmoves-0.4.2/docs/literal-text.md +27 -0
- keywordmoves-0.4.2/docs/monitoring.md +298 -0
- keywordmoves-0.4.2/docs/native-export.md +69 -0
- keywordmoves-0.4.2/docs/nltk.md +281 -0
- keywordmoves-0.4.2/docs/online-sources.md +641 -0
- keywordmoves-0.4.2/docs/openai.md +202 -0
- keywordmoves-0.4.2/docs/plugin-opportunities.md +40 -0
- keywordmoves-0.4.2/docs/pypi-publication.md +7 -0
- keywordmoves-0.4.2/docs/reddit.md +501 -0
- keywordmoves-0.4.2/docs/spacy.md +263 -0
- keywordmoves-0.4.2/docs/tiktok.md +618 -0
- keywordmoves-0.4.2/docs/youtube.md +542 -0
- keywordmoves-0.4.2/examples/monitor_demo.py +4 -0
- keywordmoves-0.4.2/pyproject.toml +111 -0
- keywordmoves-0.4.2/setup.cfg +4 -0
- keywordmoves-0.4.2/skills/research-keyword-database/SKILL.md +220 -0
- keywordmoves-0.4.2/skills/research-keyword-database/agents/openai.yaml +4 -0
- keywordmoves-0.4.2/skills/research-keyword-database/references/register-schema.md +69 -0
- keywordmoves-0.4.2/src/keywordmoves/__init__.py +26 -0
- keywordmoves-0.4.2/src/keywordmoves/__main__.py +4 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/__init__.py +2 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/google_trends.py +252 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/huggingface_llm.py +128 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/keybert_keywords.py +468 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/native_export.py +171 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/nltk_keywords.py +530 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/observed_evidence.py +102 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/openai_llm.py +197 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/spacy_keywords.py +438 -0
- keywordmoves-0.4.2/src/keywordmoves/builtin/text_library.py +302 -0
- keywordmoves-0.4.2/src/keywordmoves/cli.py +139 -0
- keywordmoves-0.4.2/src/keywordmoves/errors.py +15 -0
- keywordmoves-0.4.2/src/keywordmoves/models.py +77 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/__init__.py +1 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/analysis.py +276 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/captures.py +56 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/cli.py +199 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/demo.py +79 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/imports.py +163 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/runner.py +149 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/store.py +240 -0
- keywordmoves-0.4.2/src/keywordmoves/monitoring/validation.py +246 -0
- keywordmoves-0.4.2/src/keywordmoves/online/__init__.py +45 -0
- keywordmoves-0.4.2/src/keywordmoves/online/bing_search.py +229 -0
- keywordmoves-0.4.2/src/keywordmoves/online/bing_search_analysis.py +542 -0
- keywordmoves-0.4.2/src/keywordmoves/online/bing_search_imports.py +186 -0
- keywordmoves-0.4.2/src/keywordmoves/online/bing_search_providers.py +384 -0
- keywordmoves-0.4.2/src/keywordmoves/online/bing_webmaster.py +250 -0
- keywordmoves-0.4.2/src/keywordmoves/online/commercial.py +252 -0
- keywordmoves-0.4.2/src/keywordmoves/online/common.py +269 -0
- keywordmoves-0.4.2/src/keywordmoves/online/discovery.py +167 -0
- keywordmoves-0.4.2/src/keywordmoves/online/google.py +160 -0
- keywordmoves-0.4.2/src/keywordmoves/online/google_search.py +280 -0
- keywordmoves-0.4.2/src/keywordmoves/online/google_search_analysis.py +409 -0
- keywordmoves-0.4.2/src/keywordmoves/online/google_search_console.py +314 -0
- keywordmoves-0.4.2/src/keywordmoves/online/google_search_imports.py +250 -0
- keywordmoves-0.4.2/src/keywordmoves/online/google_search_providers.py +414 -0
- keywordmoves-0.4.2/src/keywordmoves/online/instagram.py +555 -0
- keywordmoves-0.4.2/src/keywordmoves/online/instagram_analysis.py +415 -0
- keywordmoves-0.4.2/src/keywordmoves/online/instagram_imports.py +145 -0
- keywordmoves-0.4.2/src/keywordmoves/online/reddit.py +355 -0
- keywordmoves-0.4.2/src/keywordmoves/online/reddit_analysis.py +422 -0
- keywordmoves-0.4.2/src/keywordmoves/online/reddit_imports.py +245 -0
- keywordmoves-0.4.2/src/keywordmoves/online/reddit_providers.py +220 -0
- keywordmoves-0.4.2/src/keywordmoves/online/tiktok.py +523 -0
- keywordmoves-0.4.2/src/keywordmoves/online/tiktok_analysis.py +560 -0
- keywordmoves-0.4.2/src/keywordmoves/online/tiktok_imports.py +218 -0
- keywordmoves-0.4.2/src/keywordmoves/online/websites.py +299 -0
- keywordmoves-0.4.2/src/keywordmoves/online/youtube.py +494 -0
- keywordmoves-0.4.2/src/keywordmoves/online/youtube_analysis.py +460 -0
- keywordmoves-0.4.2/src/keywordmoves/online/youtube_imports.py +207 -0
- keywordmoves-0.4.2/src/keywordmoves/online/youtube_providers.py +306 -0
- keywordmoves-0.4.2/src/keywordmoves/protocols.py +24 -0
- keywordmoves-0.4.2/src/keywordmoves/registry.py +87 -0
- keywordmoves-0.4.2/src/keywordmoves.egg-info/PKG-INFO +454 -0
- keywordmoves-0.4.2/src/keywordmoves.egg-info/SOURCES.txt +203 -0
- keywordmoves-0.4.2/src/keywordmoves.egg-info/dependency_links.txt +1 -0
- keywordmoves-0.4.2/src/keywordmoves.egg-info/entry_points.txt +39 -0
- keywordmoves-0.4.2/src/keywordmoves.egg-info/requires.txt +31 -0
- keywordmoves-0.4.2/src/keywordmoves.egg-info/top_level.txt +1 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/README.md +9 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/bwt-after.json +31 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/bwt-before.json +38 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/native-queries.json +12 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/observations-after.json +28 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/observations-before.json +28 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/observations.csv +3 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/performance.csv +2 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/serp-after.json +44 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/serp-before.json +44 -0
- keywordmoves-0.4.2/tests/fixtures/bing-search/serp.html +5 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/README.md +6 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/gsc-after.json +39 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/gsc-before.json +49 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/gsc.csv +3 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/observations.json +31 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/page.html +3 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/serp-after.json +41 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/serp-before.json +41 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/serp.html +3 -0
- keywordmoves-0.4.2/tests/fixtures/google-search/trends.csv +6 -0
- keywordmoves-0.4.2/tests/fixtures/instagram/hashtag-stats-next.json +8 -0
- keywordmoves-0.4.2/tests/fixtures/instagram/hashtag-stats.json +8 -0
- keywordmoves-0.4.2/tests/fixtures/instagram/media.json +39 -0
- keywordmoves-0.4.2/tests/fixtures/lyrics.txt +4 -0
- keywordmoves-0.4.2/tests/fixtures/nltk_reference.txt +4 -0
- keywordmoves-0.4.2/tests/fixtures/observed_evidence.json +31 -0
- keywordmoves-0.4.2/tests/fixtures/observed_evidence_missing.json +6 -0
- keywordmoves-0.4.2/tests/fixtures/observed_evidence_wrong_schema.json +4 -0
- keywordmoves-0.4.2/tests/fixtures/online/README.md +8 -0
- keywordmoves-0.4.2/tests/fixtures/online/ahrefs.json +10 -0
- keywordmoves-0.4.2/tests/fixtures/online/alsoasked.json +19 -0
- keywordmoves-0.4.2/tests/fixtures/online/autocomplete.json +9 -0
- keywordmoves-0.4.2/tests/fixtures/online/brave.json +11 -0
- keywordmoves-0.4.2/tests/fixtures/online/dataforseo.json +35 -0
- keywordmoves-0.4.2/tests/fixtures/online/datamuse.json +6 -0
- keywordmoves-0.4.2/tests/fixtures/online/export.csv +3 -0
- keywordmoves-0.4.2/tests/fixtures/online/google_ads_ideas.json +21 -0
- keywordmoves-0.4.2/tests/fixtures/online/google_ads_metrics.json +14 -0
- keywordmoves-0.4.2/tests/fixtures/online/keywords_everywhere.json +15 -0
- keywordmoves-0.4.2/tests/fixtures/online/keywordtool.json +14 -0
- keywordmoves-0.4.2/tests/fixtures/online/results.html +4 -0
- keywordmoves-0.4.2/tests/fixtures/online/search_console.json +14 -0
- keywordmoves-0.4.2/tests/fixtures/online/semrush.csv +2 -0
- keywordmoves-0.4.2/tests/fixtures/online/serpapi_autocomplete.json +10 -0
- keywordmoves-0.4.2/tests/fixtures/online/serpapi_questions.json +10 -0
- keywordmoves-0.4.2/tests/fixtures/online/serpapi_related.json +10 -0
- keywordmoves-0.4.2/tests/fixtures/online/wikipedia.json +12 -0
- keywordmoves-0.4.2/tests/fixtures/reddit/comments.json +14 -0
- keywordmoves-0.4.2/tests/fixtures/reddit/observations-after.json +18 -0
- keywordmoves-0.4.2/tests/fixtures/reddit/observations-before.json +18 -0
- keywordmoves-0.4.2/tests/fixtures/reddit/posts.json +37 -0
- keywordmoves-0.4.2/tests/fixtures/spacy_reference.txt +1 -0
- keywordmoves-0.4.2/tests/fixtures/tiktok/README.md +9 -0
- keywordmoves-0.4.2/tests/fixtures/tiktok/hashtags-after.json +4 -0
- keywordmoves-0.4.2/tests/fixtures/tiktok/hashtags-before.json +5 -0
- keywordmoves-0.4.2/tests/fixtures/tiktok/rendered.html +6 -0
- keywordmoves-0.4.2/tests/fixtures/tiktok/search-insights.csv +3 -0
- keywordmoves-0.4.2/tests/fixtures/tiktok/videos.json +9 -0
- keywordmoves-0.4.2/tests/fixtures/trends_interest.csv +7 -0
- keywordmoves-0.4.2/tests/fixtures/trends_related.csv +8 -0
- keywordmoves-0.4.2/tests/fixtures/youtube/README.md +7 -0
- keywordmoves-0.4.2/tests/fixtures/youtube/after.json +4 -0
- keywordmoves-0.4.2/tests/fixtures/youtube/before.json +4 -0
- keywordmoves-0.4.2/tests/fixtures/youtube/videos.json +16 -0
- keywordmoves-0.4.2/tests/test_bing_search_analysis.py +295 -0
- keywordmoves-0.4.2/tests/test_bing_search_imports.py +237 -0
- keywordmoves-0.4.2/tests/test_bing_search_providers.py +325 -0
- keywordmoves-0.4.2/tests/test_bing_webmaster.py +218 -0
- keywordmoves-0.4.2/tests/test_cli.py +36 -0
- keywordmoves-0.4.2/tests/test_demand_imports.py +139 -0
- keywordmoves-0.4.2/tests/test_google_search_analysis.py +247 -0
- keywordmoves-0.4.2/tests/test_google_search_console.py +183 -0
- keywordmoves-0.4.2/tests/test_google_search_imports.py +217 -0
- keywordmoves-0.4.2/tests/test_google_search_providers.py +324 -0
- keywordmoves-0.4.2/tests/test_google_trends.py +61 -0
- keywordmoves-0.4.2/tests/test_instagram_analysis.py +245 -0
- keywordmoves-0.4.2/tests/test_instagram_api.py +386 -0
- keywordmoves-0.4.2/tests/test_instagram_imports.py +256 -0
- keywordmoves-0.4.2/tests/test_keybert_config.py +185 -0
- keywordmoves-0.4.2/tests/test_keybert_keywords.py +300 -0
- keywordmoves-0.4.2/tests/test_keybert_model.py +32 -0
- keywordmoves-0.4.2/tests/test_monitoring.py +493 -0
- keywordmoves-0.4.2/tests/test_nltk_config.py +155 -0
- keywordmoves-0.4.2/tests/test_nltk_keywords.py +294 -0
- keywordmoves-0.4.2/tests/test_nltk_model.py +42 -0
- keywordmoves-0.4.2/tests/test_observed_evidence.py +62 -0
- keywordmoves-0.4.2/tests/test_online_api.py +292 -0
- keywordmoves-0.4.2/tests/test_online_transport.py +154 -0
- keywordmoves-0.4.2/tests/test_online_websites.py +217 -0
- keywordmoves-0.4.2/tests/test_openai_llm.py +372 -0
- keywordmoves-0.4.2/tests/test_openai_sdk.py +96 -0
- keywordmoves-0.4.2/tests/test_reddit_analysis.py +241 -0
- keywordmoves-0.4.2/tests/test_reddit_api.py +298 -0
- keywordmoves-0.4.2/tests/test_reddit_imports.py +292 -0
- keywordmoves-0.4.2/tests/test_reddit_providers.py +210 -0
- keywordmoves-0.4.2/tests/test_registry.py +28 -0
- keywordmoves-0.4.2/tests/test_release_version.py +17 -0
- keywordmoves-0.4.2/tests/test_spacy_config.py +127 -0
- keywordmoves-0.4.2/tests/test_spacy_keywords.py +330 -0
- keywordmoves-0.4.2/tests/test_spacy_model.py +27 -0
- keywordmoves-0.4.2/tests/test_text_library.py +95 -0
- keywordmoves-0.4.2/tests/test_text_library_literal.py +75 -0
- keywordmoves-0.4.2/tests/test_tiktok_analysis.py +285 -0
- keywordmoves-0.4.2/tests/test_tiktok_api.py +388 -0
- keywordmoves-0.4.2/tests/test_tiktok_imports.py +291 -0
- keywordmoves-0.4.2/tests/test_youtube_analysis.py +177 -0
- keywordmoves-0.4.2/tests/test_youtube_api.py +216 -0
- keywordmoves-0.4.2/tests/test_youtube_imports.py +263 -0
- keywordmoves-0.4.2/tests/test_youtube_providers.py +197 -0
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.4.2 - 8 October 2026
|
|
4
|
+
|
|
5
|
+
- Added PyPI trusted publishing through the configured `release.yml` workflow
|
|
6
|
+
and `pypi` environment. Publication verifies the immutable source tag,
|
|
7
|
+
package versions and checksums of the already tested GitHub distributions.
|
|
8
|
+
- The exact GitHub wheel and source archive are uploaded without rebuilding
|
|
9
|
+
different packages or storing registry credentials in source.
|
|
10
|
+
|
|
11
|
+
## 0.4.1 - 8 October 2026
|
|
12
|
+
|
|
13
|
+
- Corrected YouTube Analytics setup to require both `youtube.readonly` and
|
|
14
|
+
`yt-analytics.readonly`, following the current `reports.query` contract.
|
|
15
|
+
- Added an authenticated-access guide with separate browser/connector/API
|
|
16
|
+
evidence states, secure Windows credential injection, first-party scope
|
|
17
|
+
boundaries and finite read-only smoke examples.
|
|
18
|
+
- Linked collector setup to the credential guide. OAuth creation, consent,
|
|
19
|
+
renewal, platform approval and credentials remain external to KeywordMoves;
|
|
20
|
+
no account data or credentials are bundled.
|
|
21
|
+
|
|
22
|
+
## 0.4.0 — 8 October 2026
|
|
23
|
+
|
|
24
|
+
- Import established reviewed browser-observation labels without treating sampled language as demand; verify bounded local capture hashes with separate database status.
|
|
25
|
+
- Added private SQLite watchlists, hashed input captures, source-specific
|
|
26
|
+
history, coverage/freshness reports, local change alerts and acknowledgement.
|
|
27
|
+
- Added due plans and exact-plan opt-in composition of existing first-party
|
|
28
|
+
readers, with transactional per-account request reservations, no paid routes,
|
|
29
|
+
no retry and explicit interrupted-run recovery.
|
|
30
|
+
- Added reviewed CSV imports with explicit platform/metric mappings, preserving
|
|
31
|
+
zero, missing, rounded/censored and qualitative values separately.
|
|
32
|
+
- Corrected Trends related-query deduplication and metric units: Top indices,
|
|
33
|
+
Rising percentages and Breakout labels remain distinct. Added the shared
|
|
34
|
+
Google Search related-export route.
|
|
35
|
+
- Preserved censored/partial interest-over-time points without inventing exact
|
|
36
|
+
summary values, including explicit YouTube search-property scope.
|
|
37
|
+
- Added retention cleanup, integrity checks, safe local HTML reporting, adopter
|
|
38
|
+
documentation and a reproducible synthetic twelve-platform example.
|
|
39
|
+
|
|
40
|
+
Migration: related-query metrics now distinguish related_top/index_0_100,
|
|
41
|
+
related_rising/percent_growth and related_rising_breakout/growth_label. The same
|
|
42
|
+
phrase may appear in both Top and Rising; related candidates have no score.
|
|
43
|
+
Interest series containing censored/missing/partial points retain their raw
|
|
44
|
+
points but no exact mean/peak summary. No campaign register is modified.
|
|
45
|
+
|
|
46
|
+
## 0.3.2 - 2026-10-08
|
|
47
|
+
|
|
48
|
+
- Add an original tool-specific vector logo and PNG companion in the shared DanceFlow visual style.
|
|
49
|
+
- Clarify package descriptions from reviewed documentation and KeywordMoves literal-source evidence, without claims of measured search demand.
|
|
50
|
+
- Link package descriptions and READMEs to Kieran Simkin’s website and retain branding files in installable packages.
|
|
51
|
+
|
|
52
|
+
|
|
53
|
+
## 0.3.1 - 2026-10-07
|
|
54
|
+
|
|
55
|
+
- Install the online extra in the full CI matrix so saved-HTML importer tests run with their declared dependencies.
|
|
56
|
+
- Validate NLTK source spans against untranslated source line endings in the pretrained-model test, including Windows checkouts.
|
|
57
|
+
- Keep the extractor's source offsets and runtime behaviour unchanged.
|
|
58
|
+
|
|
59
|
+
## 0.3.0 - 2026-10-07
|
|
60
|
+
|
|
61
|
+
- Add `text-library extract-literal` for contiguous Unicode phrases with source character spans and corpus occurrence counts. Preserve line, punctuation and file boundaries, internal stopwords, Windows line endings, Persian joiners and numeric names. Keep legacy extraction compatible.
|
|
62
|
+
- Preserve subject, seed and capture lineage in reviewed browser observations. Reject non-text lineage fields.
|
|
63
|
+
- Align package, command-line and distribution versions, with regression checks.
|
|
64
|
+
- Document the new operations with reproducible offline examples and clarify that text counts and sampled search language do not establish demand.
|
|
65
|
+
- Replace resolved machine-specific README troubleshooting history with concise integration limits.
|
|
66
|
+
|
|
67
|
+
Validation: Ruff passes; the Python 3.13 suite completes with 1,821 passed and two optional skips. Source distribution and universal Python wheel build successfully.
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Kieran Simkin
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
22
|
+
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
include LICENSE
|
|
2
|
+
include README.md
|
|
3
|
+
include CHANGELOG.md
|
|
4
|
+
recursive-include docs *.md
|
|
5
|
+
recursive-include skills *.md *.yaml
|
|
6
|
+
recursive-include tests *.py *.txt *.csv *.json *.html *.md
|
|
7
|
+
|
|
8
|
+
recursive-include docs/branding *.svg *.png *.md
|
|
9
|
+
|
|
10
|
+
recursive-include examples *.py *.json *.csv *.md
|
|
@@ -0,0 +1,454 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: keywordmoves
|
|
3
|
+
Version: 0.4.2
|
|
4
|
+
Summary: Source-specific keyword research, demand monitoring and reviewed evidence imports for DanceFlow. https://kieransimkin.co.uk/
|
|
5
|
+
Author: Kieran Simkin
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: DanceFlow ecosystem, https://kieransimkin.co.uk/danceflow/
|
|
8
|
+
Project-URL: Homepage, https://github.com/kieransimkin/keywordmoves
|
|
9
|
+
Project-URL: Repository, https://github.com/kieransimkin/keywordmoves
|
|
10
|
+
Project-URL: Issues, https://github.com/kieransimkin/keywordmoves/issues
|
|
11
|
+
Keywords: keywords,seo,google-trends,nlp,pytorch,DanceFlow
|
|
12
|
+
Classifier: Development Status :: 3 - Alpha
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Topic :: Internet :: WWW/HTTP :: Indexing/Search
|
|
19
|
+
Classifier: Topic :: Text Processing :: Linguistic
|
|
20
|
+
Requires-Python: <3.14,>=3.10
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
License-File: LICENSE
|
|
23
|
+
Provides-Extra: online
|
|
24
|
+
Requires-Dist: httpx<1,>=0.27; extra == "online"
|
|
25
|
+
Requires-Dist: beautifulsoup4<5,>=4.12; extra == "online"
|
|
26
|
+
Provides-Extra: keybert
|
|
27
|
+
Requires-Dist: keybert<1,>=0.9; extra == "keybert"
|
|
28
|
+
Requires-Dist: sentence-transformers<6,>=3.4; extra == "keybert"
|
|
29
|
+
Requires-Dist: scikit-learn<2,>=1.3; extra == "keybert"
|
|
30
|
+
Provides-Extra: nltk
|
|
31
|
+
Requires-Dist: nltk<4,>=3.9.2; extra == "nltk"
|
|
32
|
+
Requires-Dist: numpy<3,>=1.26; extra == "nltk"
|
|
33
|
+
Provides-Extra: spacy
|
|
34
|
+
Requires-Dist: spacy<4,>=3.8.11; extra == "spacy"
|
|
35
|
+
Provides-Extra: openai
|
|
36
|
+
Requires-Dist: openai<4,>=3; extra == "openai"
|
|
37
|
+
Provides-Extra: huggingface
|
|
38
|
+
Requires-Dist: torch<3,>=2.4; extra == "huggingface"
|
|
39
|
+
Requires-Dist: transformers<6,>=4.48; extra == "huggingface"
|
|
40
|
+
Requires-Dist: sentencepiece<1,>=0.2; extra == "huggingface"
|
|
41
|
+
Requires-Dist: safetensors<1,>=0.4; extra == "huggingface"
|
|
42
|
+
Provides-Extra: dev
|
|
43
|
+
Requires-Dist: build<2,>=1.2; extra == "dev"
|
|
44
|
+
Requires-Dist: pytest<10,>=8; extra == "dev"
|
|
45
|
+
Requires-Dist: pytest-cov<8,>=5; extra == "dev"
|
|
46
|
+
Requires-Dist: ruff<1,>=0.9; extra == "dev"
|
|
47
|
+
Dynamic: license-file
|
|
48
|
+
|
|
49
|
+
# KeywordMoves
|
|
50
|
+
|
|
51
|
+
[](https://kieransimkin.co.uk/danceflow/)
|
|
52
|
+
|
|
53
|
+
By **[Kieran Simkin](https://kieransimkin.co.uk/)** · [DanceFlow ecosystem](https://kieransimkin.co.uk/danceflow/) · [Vector logo and usage guide](docs/branding/README.md).
|
|
54
|
+
|
|
55
|
+
Source-specific keyword research, local phrase extraction and search-evidence imports. https://kieransimkin.co.uk/
|
|
56
|
+
|
|
57
|
+
|
|
58
|
+
KeywordMoves is the keyword-discovery and search-evidence layer in Kieran
|
|
59
|
+
Simkin's DanceFlow ecosystem. It provides small, composable Python plugins for
|
|
60
|
+
generating keywords, finding related language, and attaching clearly labelled
|
|
61
|
+
evidence such as relative Google Trends interest.
|
|
62
|
+
|
|
63
|
+
It keeps two plugin systems deliberately separate:
|
|
64
|
+
|
|
65
|
+
- `keywordmoves.plugins` contains keyword sources and analysers.
|
|
66
|
+
- `keywordmoves.llms` contains interchangeable LLM runtimes.
|
|
67
|
+
|
|
68
|
+
A keyword plugin that needs generative LLM inference must ask for an LLM plugin
|
|
69
|
+
by name. It must not import or silently select a model provider of its own.
|
|
70
|
+
|
|
71
|
+
## Current plugins
|
|
72
|
+
|
|
73
|
+
| Plugin | Kind | What it does |
|
|
74
|
+
| --- | --- | --- |
|
|
75
|
+
| `google-search` | keyword | Google Search Console, Trends exports, search observations and page audits with source-specific metrics. |
|
|
76
|
+
| `bing-search` | keyword | Bing Webmaster, reviewed search results and keyword planning evidence. |
|
|
77
|
+
| `youtube`, `tiktok`, `instagram`, `reddit` | keyword | Separate platform workbenches for authorised exports, native observations, local reference text and supported providers. Read each module's access and metric limits. |
|
|
78
|
+
| `google-trends` | keyword | Imports Google Trends interest and related-query CSV exports, preserving the distinction between relative index values and absolute search volume. |
|
|
79
|
+
| `keybert` | keyword | Ranks literal reference-text phrases with local sentence embeddings; supports cosine similarity, MMR and Max Sum selection. No generative LLM required. |
|
|
80
|
+
| `nltk` | keyword | Extracts proper nouns, grammar-based noun chunks, named entities and keyphrases locally with NLTK. No LLM required. |
|
|
81
|
+
| `native-export` | keyword | Imports reviewed platform CSV reports with explicit columns, metric units and missing/censored states. |
|
|
82
|
+
| `observed-evidence` | keyword | Validates and imports dated browser, account and public-tool observations without scraping authenticated or private interfaces. |
|
|
83
|
+
| `spacy` | keyword | Extracts proper nouns, noun chunks, named entities and useful noun/adjective phrases from reference text with a local spaCy pipeline. No LLM required. |
|
|
84
|
+
| `text-library` | keyword | Extracts contiguous Unicode phrases and source spans with `extract-literal`, retains legacy local associations, or asks an explicitly selected LLM for semantic proposals. |
|
|
85
|
+
| `huggingface-transformers` | LLM | Runs a local Hugging Face seq2seq or causal instruction model through PyTorch. Its default is the Apache-2.0 `Qwen/Qwen2.5-0.5B-Instruct`, pinned to a reviewed model revision. |
|
|
86
|
+
| `openai` | LLM | Uses OpenAI-hosted text models through the Responses API. Accepts `OPENAI_API_KEY` or `--openai-api-key`; the model is selectable with `--model`. |
|
|
87
|
+
|
|
88
|
+
Google's official Trends API is currently an access-controlled alpha. The
|
|
89
|
+
built-in plugin therefore supports reproducible CSV exports now rather than
|
|
90
|
+
depending on the archived, unofficial `pytrends` scraper or guessing an API
|
|
91
|
+
contract. A live official-API backend can be added when access and its exact
|
|
92
|
+
contract are available.
|
|
93
|
+
|
|
94
|
+
## Install
|
|
95
|
+
|
|
96
|
+
The core and Google Trends importer have no runtime dependencies:
|
|
97
|
+
|
|
98
|
+
```powershell
|
|
99
|
+
python -m pip install -e .
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Install the local PyTorch/Hugging Face runtime when it is needed:
|
|
103
|
+
|
|
104
|
+
```powershell
|
|
105
|
+
python -m pip install -e ".[huggingface]"
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
The first Hugging Face model-backed run downloads the selected model unless
|
|
109
|
+
`llm_local_files_only=true` is supplied. Model files stay in the configured
|
|
110
|
+
Hugging Face cache; the input text is processed locally and is not sent to a
|
|
111
|
+
hosted API.
|
|
112
|
+
|
|
113
|
+
The default Qwen checkpoint is pinned to revision
|
|
114
|
+
`2b01de6d1108f9b2b5e46a726aa678a359b6c03b`. KeywordMoves records the selected
|
|
115
|
+
model and revision in result metadata. A different Hugging Face model remains
|
|
116
|
+
selectable with `--model` and `--option llm_revision=<commit>`.
|
|
117
|
+
|
|
118
|
+
Install the optional OpenAI runtime separately (no PyTorch required):
|
|
119
|
+
|
|
120
|
+
```powershell
|
|
121
|
+
python -m pip install -e ".[openai]"
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
Unlike the local Hugging Face runtime, selecting `--llm openai` sends the
|
|
125
|
+
supplied prompts to OpenAI's hosted API. API usage may incur charges.
|
|
126
|
+
|
|
127
|
+
Install the independent spaCy keyword extractor and its English pipeline:
|
|
128
|
+
|
|
129
|
+
```powershell
|
|
130
|
+
python -m pip install -e ".[spacy]"
|
|
131
|
+
python -m spacy download en_core_web_sm
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
The pipeline download is an explicit setup step, not something KeywordMoves
|
|
135
|
+
performs automatically. Subsequent extraction with this pipeline is local.
|
|
136
|
+
|
|
137
|
+
Install the alternative NLTK extractor and the data used by its default features:
|
|
138
|
+
|
|
139
|
+
```powershell
|
|
140
|
+
python -m pip install -e ".[nltk]"
|
|
141
|
+
python -m nltk.downloader punkt_tab averaged_perceptron_tagger_eng maxent_ne_chunker_tab words stopwords wordnet
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
This is also local after setup. KeywordMoves never downloads NLTK data implicitly.
|
|
145
|
+
For offline/custom data directories and smaller installs, see [docs/nltk.md](docs/nltk.md).
|
|
146
|
+
|
|
147
|
+
Install the independent KeyBERT semantic keyword extractor:
|
|
148
|
+
|
|
149
|
+
```powershell
|
|
150
|
+
python -m pip install -e ".[keybert]"
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
Its first model-backed run may download the pinned Sentence Transformers model.
|
|
154
|
+
Text is embedded locally, not sent to a hosted inference API. For prefetching
|
|
155
|
+
and strictly offline use, see [docs/keybert.md](docs/keybert.md).
|
|
156
|
+
|
|
157
|
+
## Use
|
|
158
|
+
|
|
159
|
+
List both plugin layers:
|
|
160
|
+
|
|
161
|
+
```powershell
|
|
162
|
+
keywordmoves plugins
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
Import a Google Trends interest-over-time export:
|
|
166
|
+
|
|
167
|
+
```powershell
|
|
168
|
+
keywordmoves run google-trends `
|
|
169
|
+
--operation import-interest `
|
|
170
|
+
--input .\multiTimeline.csv `
|
|
171
|
+
--option geography=GB `
|
|
172
|
+
--option observed_at=2026-09-30
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
Extract and expand ideas from a lyrics file with an explicit LLM selection:
|
|
176
|
+
|
|
177
|
+
```powershell
|
|
178
|
+
keywordmoves run text-library `
|
|
179
|
+
--operation extract `
|
|
180
|
+
--input .\lyrics.txt `
|
|
181
|
+
--llm huggingface-transformers `
|
|
182
|
+
--model Qwen/Qwen2.5-0.5B-Instruct `
|
|
183
|
+
--option limit=20
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
Use the transparent local extractor without any LLM:
|
|
187
|
+
|
|
188
|
+
```powershell
|
|
189
|
+
keywordmoves run text-library --operation extract-local --input .\lyrics.txt
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
Import a reviewed JSON observation file for Search Console, native platform
|
|
193
|
+
search, autocomplete or a public keyword tool:
|
|
194
|
+
|
|
195
|
+
```powershell
|
|
196
|
+
keywordmoves run observed-evidence `
|
|
197
|
+
--operation import-observations `
|
|
198
|
+
--input .\keyword-observations.json `
|
|
199
|
+
--format json
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
The file must use `keywordmoves-observations/v1`. Each observation requires a
|
|
203
|
+
phrase, source, actual metric, observation date and platform. Record blocked,
|
|
204
|
+
empty or unreadable surfaces with `availability: unavailable`; do not turn an
|
|
205
|
+
unavailable check into a zero-demand claim. This importer is the supported
|
|
206
|
+
route for browser-only and authenticated evidence. It deliberately does not
|
|
207
|
+
scrape sites or call private endpoints.
|
|
208
|
+
|
|
209
|
+
LLM suggestions are labelled as proposals. They do not establish popularity,
|
|
210
|
+
search volume, competition, ranking potential, or suitability on a particular
|
|
211
|
+
platform.
|
|
212
|
+
|
|
213
|
+
### Extract proper nouns and noun chunks with spaCy
|
|
214
|
+
|
|
215
|
+
```powershell
|
|
216
|
+
keywordmoves run spacy `
|
|
217
|
+
--operation extract `
|
|
218
|
+
--input .\reference.txt `
|
|
219
|
+
--option limit=50
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
This also extracts named entities, common nouns and contiguous adjective/noun
|
|
223
|
+
keyphrases. To focus only on the requested noun categories, add:
|
|
224
|
+
|
|
225
|
+
```powershell
|
|
226
|
+
--option "features=proper-nouns,noun-chunks"
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
Use `--option "text=Your reference text"` for inline input, or repeat `--input`
|
|
230
|
+
for multiple files/directories. Select another installed spaCy pipeline with
|
|
231
|
+
`--option spacy_model=NAME`, **not** the LLM-specific `--model` flag.
|
|
232
|
+
|
|
233
|
+
Results preserve detector labels, entity types, original surface forms,
|
|
234
|
+
deduplicated occurrence counts and source character offsets. Common-word
|
|
235
|
+
inflections can be merged without singularising names. Scores describe only
|
|
236
|
+
heuristic salience in the reference text, not search demand or model confidence.
|
|
237
|
+
See [docs/spacy.md](docs/spacy.md) for all options, examples and limitations.
|
|
238
|
+
|
|
239
|
+
### Extract with NLTK
|
|
240
|
+
|
|
241
|
+
```powershell
|
|
242
|
+
keywordmoves run nltk `
|
|
243
|
+
--operation extract `
|
|
244
|
+
--input .\reference.txt `
|
|
245
|
+
--option limit=50
|
|
246
|
+
```
|
|
247
|
+
|
|
248
|
+
The `nltk` and `spacy` extractors share feature names and candidate metadata.
|
|
249
|
+
Use `--option "features=proper-nouns,noun-chunks"` to focus on names and noun
|
|
250
|
+
phrases, or `--option "text=Your reference text"` for inline input. NLTK noun
|
|
251
|
+
chunks use a part-of-speech grammar rather than spaCy's dependency parser, so
|
|
252
|
+
results need not agree. This initial NLTK implementation is English-only.
|
|
253
|
+
Names retain their surface forms; ordinary words can be lemmatised with WordNet.
|
|
254
|
+
See [docs/nltk.md](docs/nltk.md) for all options, resource setup and limitations.
|
|
255
|
+
|
|
256
|
+
### Rank reference-text phrases with KeyBERT
|
|
257
|
+
|
|
258
|
+
```powershell
|
|
259
|
+
keywordmoves run keybert `
|
|
260
|
+
--operation extract `
|
|
261
|
+
--input .\reference.txt `
|
|
262
|
+
--option limit=20 `
|
|
263
|
+
--option max_words=3 `
|
|
264
|
+
--option method=mmr `
|
|
265
|
+
--option diversity=0.7
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
KeyBERT uses embedding similarity, not POS or entity labels. Its default
|
|
269
|
+
selection is `method=cosine`; `mmr` and `maxsum` select more varied phrases.
|
|
270
|
+
Use repeated `--keyword "phrase"` arguments to restrict the vocabulary, or
|
|
271
|
+
`--option "text=Reference text"` for inline input. Select another encoder with
|
|
272
|
+
`--option keybert_model=NAME`, not the LLM-specific `--model` switch.
|
|
273
|
+
|
|
274
|
+
Candidates retain literal source spans; long references are embedded in
|
|
275
|
+
bounded chunks rather than silently truncated. Scores are not search demand
|
|
276
|
+
or confidence. See [docs/keybert.md](docs/keybert.md) for model caching,
|
|
277
|
+
resource limits, all options and an NLTK/spaCy-to-KeyBERT re-ranking example.
|
|
278
|
+
|
|
279
|
+
### Use OpenAI models
|
|
280
|
+
|
|
281
|
+
Set the key in the current PowerShell session, then explicitly select OpenAI:
|
|
282
|
+
|
|
283
|
+
```powershell
|
|
284
|
+
$env:OPENAI_API_KEY = "YOUR_OPENAI_API_KEY"
|
|
285
|
+
keywordmoves run text-library `
|
|
286
|
+
--operation extract `
|
|
287
|
+
--input .\lyrics.txt `
|
|
288
|
+
--llm openai `
|
|
289
|
+
--model gpt-4.1-mini `
|
|
290
|
+
--option limit=20
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
Alternatively, supply the key as a command-line argument:
|
|
294
|
+
|
|
295
|
+
```powershell
|
|
296
|
+
keywordmoves run text-library --operation extract --input .\lyrics.txt --llm openai --openai-api-key "YOUR_OPENAI_API_KEY"
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
`--openai-api-key` overrides `--option llm_api_key=...` and `OPENAI_API_KEY`.
|
|
300
|
+
Prefer the environment variable: command-line keys may appear in shell history
|
|
301
|
+
or process listings. KeywordMoves does not write the key into result metadata
|
|
302
|
+
or include raw provider error messages in its normal CLI errors.
|
|
303
|
+
|
|
304
|
+
The default model is `gpt-4.1-mini`; use `--model` to choose another
|
|
305
|
+
Responses-compatible text model available to your API project. The plugin does
|
|
306
|
+
not silently switch providers or models. For a reasoning model, a larger token
|
|
307
|
+
budget may be needed; set `--option llm_max_output_tokens=2048` or higher.
|
|
308
|
+
|
|
309
|
+
See [docs/openai.md](docs/openai.md) for Bash and Python examples, supported
|
|
310
|
+
options, hosted-data behaviour and troubleshooting.
|
|
311
|
+
|
|
312
|
+
## Online keyword sources
|
|
313
|
+
|
|
314
|
+
Install the optional networking and HTML parsers:
|
|
315
|
+
|
|
316
|
+
```powershell
|
|
317
|
+
python -m pip install -e ".[online]"
|
|
318
|
+
keywordmoves plugins --kind keyword
|
|
319
|
+
```
|
|
320
|
+
|
|
321
|
+
The online suite adds twelve documented API adapters: `google-ads`,
|
|
322
|
+
`search-console`, `dataforseo`, `semrush`, `ahrefs`, `keywordtool`,
|
|
323
|
+
`keywords-everywhere`, `alsoasked`, `serpapi`, `brave-suggest`, `datamuse` and
|
|
324
|
+
`wikipedia`. It also includes three explicitly opt-in experimental browser
|
|
325
|
+
suggestion endpoints, a configurable `website-keywords` public-HTML extractor,
|
|
326
|
+
and **import-only** `ubersuggest` and `answerthepublic` adapters.
|
|
327
|
+
|
|
328
|
+
```powershell
|
|
329
|
+
keywordmoves run datamuse --operation related --keyword "independent music" --option limit=20
|
|
330
|
+
|
|
331
|
+
$env:SEMRUSH_API_KEY = "YOUR_SEMRUSH_KEY"
|
|
332
|
+
keywordmoves run semrush --operation related --keyword "independent music" --option database=uk
|
|
333
|
+
```
|
|
334
|
+
|
|
335
|
+
Online sources use `--keyword` seeds, not reference-text inputs or LLM flags.
|
|
336
|
+
API credentials come from provider-specific environment variables or explicit
|
|
337
|
+
`--option` settings. Most commercial APIs require separate API access and may
|
|
338
|
+
charge credits. A local output limit is not a universal billing cap.
|
|
339
|
+
|
|
340
|
+
See [the online-source guide and research matrix](docs/online-sources.md) for
|
|
341
|
+
all eighteen plugins, exact operations, credentials, locale conventions,
|
|
342
|
+
website query/extraction rules, access limitations and primary-source links.
|
|
343
|
+
The guide distinguishes API adapters from experimental endpoints and report
|
|
344
|
+
imports: no authenticated production access is implied by offline tests.
|
|
345
|
+
|
|
346
|
+
## Write a keyword plugin
|
|
347
|
+
|
|
348
|
+
Implement an object with a `descriptor` and `run(request, context)` method, then
|
|
349
|
+
publish it through an entry point:
|
|
350
|
+
|
|
351
|
+
```toml
|
|
352
|
+
[project.entry-points."keywordmoves.plugins"]
|
|
353
|
+
my-source = "my_package.plugin:MyKeywordPlugin"
|
|
354
|
+
```
|
|
355
|
+
|
|
356
|
+
The request and result dataclasses are exported from `keywordmoves`. Use
|
|
357
|
+
`KeywordEvidence` to state the source, metric, unit, date, geography and limits
|
|
358
|
+
instead of flattening unlike evidence into one unexplained score.
|
|
359
|
+
|
|
360
|
+
## Write an LLM plugin
|
|
361
|
+
|
|
362
|
+
Implement a `descriptor` and `generate(request) -> LLMResult`, then register:
|
|
363
|
+
|
|
364
|
+
```toml
|
|
365
|
+
[project.entry-points."keywordmoves.llms"]
|
|
366
|
+
my-local-runtime = "my_package.llm:MyLocalLLM"
|
|
367
|
+
```
|
|
368
|
+
|
|
369
|
+
Keyword plugins select it through `context.llms.get(name)`. Model-specific
|
|
370
|
+
options are passed with the `llm_` prefix; for example `--option
|
|
371
|
+
llm_device=cpu` becomes the LLM option `device=cpu`.
|
|
372
|
+
|
|
373
|
+
## Development
|
|
374
|
+
|
|
375
|
+
```powershell
|
|
376
|
+
python -m pip install -e ".[dev]"
|
|
377
|
+
python -m pytest
|
|
378
|
+
python -m ruff check .
|
|
379
|
+
```
|
|
380
|
+
|
|
381
|
+
OpenAI unit/CLI tests also run without the optional SDK. To run the additional
|
|
382
|
+
real-SDK transport tests (using in-memory HTTP responses, not billable API
|
|
383
|
+
calls), install `.[dev,openai]`. CI installs that extra and runs both test sets.
|
|
384
|
+
|
|
385
|
+
spaCy API tests use controlled annotated documents and need only the `spacy`
|
|
386
|
+
extra; a separate optional test uses a real `en_core_web_sm` pipeline. See
|
|
387
|
+
[spaCy test instructions](docs/spacy.md#tests) for both routes. CI also includes
|
|
388
|
+
a trained-pipeline smoke-test job.
|
|
389
|
+
|
|
390
|
+
NLTK API tests need the `nltk` extra but no downloaded resources; a separate
|
|
391
|
+
pretrained-model smoke test requires its English data. CI installs that data in
|
|
392
|
+
Linux and Windows jobs and makes a missing-resource skip a failure. See
|
|
393
|
+
[NLTK test instructions](docs/nltk.md#tests).
|
|
394
|
+
|
|
395
|
+
KeyBERT adapter/API tests use controlled embeddings and do not download models.
|
|
396
|
+
Install `.[dev,keybert]` to run both; separate Linux/Windows CI jobs require a
|
|
397
|
+
real pinned-model smoke test. See [KeyBERT tests](docs/keybert.md#tests).
|
|
398
|
+
|
|
399
|
+
See [docs/plugin-opportunities.md](docs/plugin-opportunities.md) for researched
|
|
400
|
+
next-plugin candidates and access constraints.
|
|
401
|
+
|
|
402
|
+
## Bundled skill
|
|
403
|
+
|
|
404
|
+
The repository includes
|
|
405
|
+
[`skills/research-keyword-database`](skills/research-keyword-database/SKILL.md),
|
|
406
|
+
a reusable workflow for inspecting, expanding, researching, validating and
|
|
407
|
+
updating the canonical My Songs keyword register. It keeps KeywordMoves
|
|
408
|
+
candidate generation separate from external demand evidence, requires explicit
|
|
409
|
+
LLM selection, and preserves the register's evidence and lifecycle semantics.
|
|
410
|
+
|
|
411
|
+
## Monitor keyword evidence over time
|
|
412
|
+
|
|
413
|
+
Use private watchlists and a persistent SQLite history to track source-specific
|
|
414
|
+
keyword evidence, missing coverage, stale reports and collection failures.
|
|
415
|
+
Compatible measurements can produce local change alerts; suggestions, rounded
|
|
416
|
+
estimates and differently normalised indices never become one demand score.
|
|
417
|
+
|
|
418
|
+
~~~powershell
|
|
419
|
+
keywordmoves monitor --store ./private/evidence.sqlite init
|
|
420
|
+
keywordmoves monitor --store ./private/evidence.sqlite watch --input ./watches.json
|
|
421
|
+
keywordmoves monitor --store ./private/evidence.sqlite import --input ./observations.json
|
|
422
|
+
keywordmoves monitor --store ./private/evidence.sqlite report --output ./private/coverage.html
|
|
423
|
+
~~~
|
|
424
|
+
|
|
425
|
+
[Monitoring guide](docs/monitoring.md) covers histories, provenance, alerts,
|
|
426
|
+
due plans, opt-in bounded first-party collection, shared account quotas and
|
|
427
|
+
retention. [Native CSV exports](docs/native-export.md) covers explicit mappings
|
|
428
|
+
and the twelve-platform coverage matrix. A public-safe synthetic demo is
|
|
429
|
+
included in the installed package:
|
|
430
|
+
|
|
431
|
+
~~~powershell
|
|
432
|
+
python -m keywordmoves.monitoring.demo --output-dir ./demo-private
|
|
433
|
+
~~~
|
|
434
|
+
|
|
435
|
+
Monitoring and import are available without network/NLP extras. Collection
|
|
436
|
+
composes existing Google Search Console, YouTube Analytics and Bing Webmaster
|
|
437
|
+
readers; real access still requires an authorised account smoke test. No
|
|
438
|
+
background job, paid request or account setup starts automatically.
|
|
439
|
+
|
|
440
|
+
## Current limitations and integration notes
|
|
441
|
+
|
|
442
|
+
The [authenticated-access guide](docs/authenticated-access.md) documents the
|
|
443
|
+
required first-party read scopes, secure Windows credential injection, bounded
|
|
444
|
+
smoke checks and separate browser, connector and local API access states.
|
|
445
|
+
|
|
446
|
+
- Python 3.10–3.13 is supported; monitoring uses standard CPython SQLite support. Optional NLP and provider integrations need their documented extras and source permissions.
|
|
447
|
+
- Source text frequency, semantic relevance, search-result samples, platform observations and provider estimates answer different questions. They are not a universal ranking score.
|
|
448
|
+
- Literal extraction preserves contiguous Unicode phrases and source offsets. Its English boundary stopwords are not a language-aware tokenizer; review other languages explicitly.
|
|
449
|
+
- Browser-observation capture paths and hashes are supplied by the caller. The importer preserves their lineage but does not certify the source capture.
|
|
450
|
+
- Online integrations require the access, permissions and charge limits documented in their module guides. A local result limit is not a billing ceiling.
|
|
451
|
+
|
|
452
|
+
- Monitoring databases and reports can contain private account data; keep them outside public source/packages. Quotas cover only clients sharing the same store and account label.
|
|
453
|
+
|
|
454
|
+
See [literal extraction](docs/literal-text.md) and [reviewed browser observations](docs/browser-observations.md) for the new operations and reproducible examples.
|