answerpath-geo 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Taqi Molavi
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,215 @@
1
+ Metadata-Version: 2.4
2
+ Name: answerpath-geo
3
+ Version: 0.1.0
4
+ Summary: Privacy-first user question mining for GEO and AEO
5
+ Author: Taqi Molavi
6
+ License: MIT
7
+ Project-URL: Homepage, https://molavi.pro
8
+ Project-URL: Repository, https://github.com/tmolavi/answerpath-geo
9
+ Project-URL: Documentation, https://github.com/tmolavi/answerpath-geo#readme
10
+ Project-URL: Issues, https://github.com/tmolavi/answerpath-geo/issues
11
+ Keywords: geo,aeo,seo,answer-engine-optimization,prompt-mining,user-intent,question-mining
12
+ Classifier: License :: OSI Approved :: MIT License
13
+ Classifier: Programming Language :: Python :: 3
14
+ Requires-Python: >=3.10
15
+ Description-Content-Type: text/markdown
16
+ License-File: LICENSE
17
+ Dynamic: license-file
18
+
19
+ # AnswerPath GEO — Discover the Questions People Ask AI Before They Find Your Business
20
+
21
+ **AnswerPath GEO** is a privacy-first, open-source **SEO, GEO and AEO question mining engine**. Give it a business, service, website topic or keyword and it builds a transparent question map showing what people may ask an AI assistant before choosing a provider.
22
+
23
+ It separates **observed questions** extracted from exports and application logs from **generated research prompts**. This distinction matters: generated prompts are hypotheses, not evidence of what people actually asked.
24
+
25
+ [Installation](#installation) · [Quick start](#quick-start) · [Inputs](#supported-inputs) · [Outputs](#outputs) · [GEO/AEO method](#geo-and-aeo-method) · [Privacy](#privacy-and-data-boundaries) · [Integrations](#integrations)
26
+
27
+ **MCP clients:** [Codex, Antigravity, Claude, Cursor and Cloud setup](#mcp-setup)
28
+
29
+ ## Why AnswerPath GEO?
30
+
31
+ Traditional keyword tools show phrases typed into search engines. AI assistants receive longer, conversational questions: “Which agency is reliable for…?”, “What should I compare…?”, and “Is this service worth the price?”. AnswerPath turns the questions you already own into a usable **answer-path map** for content, FAQ, schema and AI visibility research.
32
+
33
+ It is designed for marketers, publishers, agencies and product teams who need to:
34
+
35
+ - discover and normalize user questions from ChatGPT, Claude, Codex, Antigravity, Cursor and chatbot exports;
36
+ - group questions by intent and decision stage;
37
+ - identify recurring questions without silently inventing demand;
38
+ - create an auditable prompt bank for GEO/AEO experiments;
39
+ - keep private conversation data on the operator’s machine.
40
+
41
+ ## Installation
42
+
43
+ Requires Python 3.10+.
44
+
45
+ ```bash
46
+ git clone https://github.com/tmolavi/answerpath-geo.git
47
+ cd answerpath-geo
48
+ python3 -m venv .venv
49
+ .venv/bin/pip install -e .
50
+ ```
51
+
52
+ ## Quick start
53
+
54
+ Create a prompt map for a service or keyword:
55
+
56
+ ```bash
57
+ answerpath "طراحی سایت فروشگاهی"
58
+ ```
59
+
60
+ The command writes `answerpath-output/questions.json` and `questions.csv`.
61
+
62
+ Analyze owned conversation data and add generated discovery prompts:
63
+
64
+ ```bash
65
+ answerpath "مشاوره سئو پزشکی" \\
66
+ --input ~/Downloads/chatgpt-export.zip \\
67
+ --input ./support-chat.json \\
68
+ --out ./research/seo-medical
69
+ ```
70
+
71
+ Keep only questions actually found in your supplied data:
72
+
73
+ ```bash
74
+ answerpath "سرویس حسابداری" --input ./logs --no-generated
75
+ ```
76
+
77
+ ## Supported inputs
78
+
79
+ AnswerPath reads local JSON, JSONL, CSV, directories and ZIP archives. It recognizes common `role=user|human|customer` fields and nested ChatGPT/Claude export structures. It is compatible in principle with exports produced by tools such as [openai_export_parser](https://github.com/temnoon/openai_export_parser), [llm-export-analytics](https://github.com/noah-chelednik/llm-export-analytics), and the local multi-client history model used by [ContextBridgeAI](https://github.com/T-Gojo/ContextBridgeAi).
80
+
81
+ It does not log in to ChatGPT, Claude, Google, or any other account, scrape other users, or obtain API keys.
82
+
83
+ ## Outputs
84
+
85
+ Each normalized question contains:
86
+
87
+ | Field | Meaning |
88
+ |---|---|
89
+ | `text` | Question text as extracted or generated |
90
+ | `evidence` | `observed`, `generated`, or `observed+generated` |
91
+ | `source` | Input file or generation rule |
92
+ | `intent` | `learn`, `compare`, `buy`, `solve`, `trust`, or `discover` |
93
+ | `stage` | Awareness, consideration or decision |
94
+ | `frequency` | Number of similar records merged |
95
+ | `cluster` | Stable output cluster identifier |
96
+
97
+ Frequency is meaningful only for observed records from a defined dataset. Template prompts are clearly labeled and must not be reported as customer demand.
98
+
99
+ ## GEO and AEO method
100
+
101
+ AnswerPath supports a defensible workflow:
102
+
103
+ 1. **Collect** owned exports, support logs or application traces.
104
+ 2. **Extract** user turns and preserve their source path.
105
+ 3. **Normalize** whitespace and common message formats.
106
+ 4. **Classify** intent and decision stage with inspectable rules.
107
+ 5. **Deduplicate** near-identical questions while retaining frequency.
108
+ 6. **Expand** with clearly labeled prompt hypotheses when requested.
109
+ 7. **Publish answers**: create concise answer-first pages, FAQ sections, comparison tables and structured data based on recurring observed questions.
110
+ 8. **Measure** those prompts in a separate GEO benchmark; never mix hypotheses with measured demand.
111
+
112
+ The engine does not claim that a page will rank in Google or be cited by an AI system. Those are outcomes to test with your own content and provider evidence.
113
+
114
+ ## Integrations
115
+
116
+ - **ChatGPT / Claude exports:** pass the downloaded ZIP or extracted JSON to `--input`.
117
+ - **Codex, Antigravity, Cursor and other AI IDEs:** export or copy the local session data first; use a read-only copy as input.
118
+ - **FAQ workflows:** feed `questions.json` into [FAQ Extraction Pipeline](https://github.com/emfrg/faq-extraction-pipeline) for richer issue extraction, embeddings and FAQ synthesis.
119
+ - **GEO prompt discovery:** compare observed questions with generated candidates from projects such as [auto-geo](https://github.com/shadowresearch/auto-geo), keeping the evidence labels separate.
120
+
121
+ ## 🏆 Evidence & Benchmark Contribution
122
+
123
+ AnswerPath GEO generated the **Question Discovery & Intent Stratification Layer** used in the official [GEO, SEO & Digital Marketing Agency Iran 2026 Benchmark](https://github.com/tmolavi/geo-scope/tree/main/benchmarks/geo-seo-digital-agency-iran-2026.1):
124
+
125
+ - **Verified Query Dataset**: [`examples/sample_queries.json`](examples/sample_queries.json)
126
+ - **Standalone Offline Demo**: [`examples/public_demo/`](examples/public_demo/)
127
+ - **Cross-Repository Evidence Map**: [Ecosystem Evidence Flow](https://github.com/tmolavi/geo-scope/blob/main/docs/EVIDENCE_MAP.md)
128
+
129
+ * **Prompts Generated & Stratified**: 30 standardized queries.
130
+ * **Strict Demand Provenance Separation**:
131
+ * **Observed User Demand ($N=15$)**: Extracted from genuine conversational search logs (`source_type: "observed"`).
132
+ * **Exploration Hypotheses ($N=15$)**: Systematic template variations (`source_type: "generated"`).
133
+ * **5 Intent Strata**: `commercial` (general evaluation), `compare` (head-to-head alternatives), `trust` (credibility & contracts), `solve` (technical fixes), and `buy` (procurement & quotes).
134
+ * **Provenance Contract Schema**:
135
+ ```json
136
+ {
137
+ "id": "PRM-IR-001",
138
+ "prompt": "بهترین آژانس دیجیتال مارکتینگ و سئو در ایران کدام است؟",
139
+ "intent": "commercial",
140
+ "source_type": "observed",
141
+ "source_reference": "answerpath",
142
+ "cluster": "general_recommendation"
143
+ }
144
+ ```
145
+ * **Ecosystem Architecture**: See [Benchmark Ecosystem Map](docs/BENCHMARK_ECOSYSTEM.md) for data flow across AnswerPath, GEO-Scope, SAGE, MAVI, and SiteProbe.
146
+
147
+ ## 🏛️ Ecosystem
148
+
149
+ AnswerPath GEO operates as the question discovery component of the **Molavi AI Visibility Stack**:
150
+
151
+ - **Discovery**: [AnswerPath GEO](https://github.com/tmolavi/answerpath-geo)
152
+ - **Measurement**: [GEO-Scope](https://github.com/tmolavi/geo-scope)
153
+ - **Diagnostics**: [SAGE Audit](https://github.com/tmolavi/sage-audit)
154
+ - **Action**: [SiteProbe](https://github.com/tmolavi/siteprobe)
155
+ - **Protocol**: [MCP GEO Server](https://github.com/tmolavi/mcp-geo-server)
156
+
157
+ ## 📖 Runnable Python Example
158
+
159
+ Run the bundled discovery example script:
160
+ ```bash
161
+ python examples/discover_example.py
162
+ ```
163
+ Sample benchmark query payload is available in [`examples/sample_queries.json`](examples/sample_queries.json).
164
+
165
+ ## MCP setup
166
+
167
+ AnswerPath exposes one MCP tool, `discover_questions`. The stdio transport works with local Codex, Antigravity, Claude Desktop, Cursor, Windsurf and other MCP clients:
168
+
169
+ ```json
170
+ {
171
+ "mcpServers": {
172
+ "answerpath": {
173
+ "command": "/absolute/path/to/answerpath-geo/.venv/bin/answerpath",
174
+ "args": ["mcp"]
175
+ }
176
+ }
177
+ }
178
+ ```
179
+
180
+ Ask the client to call `discover_questions` with `topic`, optional owned `inputs` (`text` and `source`), and `include_generated`. Generated prompts are always labeled separately from observed questions.
181
+
182
+ For a private Cloud deployment, run the HTTP transport behind HTTPS and an authentication gateway:
183
+
184
+ ```bash
185
+ answerpath serve-mcp --host 127.0.0.1 --port 8787
186
+ ```
187
+
188
+ The JSON-RPC endpoint is `POST /mcp`. The application deliberately does not implement authentication itself: put it behind your gateway, rate limits and tenant isolation before exposing it publicly. A public URL or a successful protocol handshake does not prove that provider data is available.
189
+
190
+ ## Privacy and data boundaries
191
+
192
+ Processing is local and deterministic. AnswerPath does not transmit input files. Do not place private exports in a public repository or commit generated files containing message content. Remove or hash identifiers before sharing results. Only analyze data for which you have authorization.
193
+
194
+ ## Development & Testing
195
+
196
+ ```bash
197
+ pip install -e .
198
+ pytest tests/ -v
199
+ ```
200
+
201
+ ## 💬 Community & External Collaboration
202
+
203
+ We welcome contributions to query mining, clustering algorithms, and demand stratification:
204
+
205
+ - **Discussions**: [GitHub Discussions](https://github.com/tmolavi/answerpath-geo/discussions)
206
+ - **First Contribution Guide**: [`docs/FIRST_CONTRIBUTION.md`](docs/FIRST_CONTRIBUTION.md)
207
+ - **Research Collaboration**: [`docs/RESEARCH_COLLABORATION.md`](docs/RESEARCH_COLLABORATION.md)
208
+ - **Issues & Bug Reports**: [GitHub Issues](https://github.com/tmolavi/answerpath-geo/issues)
209
+ - **Contribution Standards**: [`CONTRIBUTING.md`](CONTRIBUTING.md) and [`SECURITY.md`](SECURITY.md)
210
+
211
+ ## 👤 Author & License
212
+
213
+ Developed by **Taghi Molavi** — [molavi.pro](https://molavi.pro)
214
+ MIT. See [LICENSE](LICENSE).
215
+
@@ -0,0 +1,197 @@
1
+ # AnswerPath GEO — Discover the Questions People Ask AI Before They Find Your Business
2
+
3
+ **AnswerPath GEO** is a privacy-first, open-source **SEO, GEO and AEO question mining engine**. Give it a business, service, website topic or keyword and it builds a transparent question map showing what people may ask an AI assistant before choosing a provider.
4
+
5
+ It separates **observed questions** extracted from exports and application logs from **generated research prompts**. This distinction matters: generated prompts are hypotheses, not evidence of what people actually asked.
6
+
7
+ [Installation](#installation) · [Quick start](#quick-start) · [Inputs](#supported-inputs) · [Outputs](#outputs) · [GEO/AEO method](#geo-and-aeo-method) · [Privacy](#privacy-and-data-boundaries) · [Integrations](#integrations)
8
+
9
+ **MCP clients:** [Codex, Antigravity, Claude, Cursor and Cloud setup](#mcp-setup)
10
+
11
+ ## Why AnswerPath GEO?
12
+
13
+ Traditional keyword tools show phrases typed into search engines. AI assistants receive longer, conversational questions: “Which agency is reliable for…?”, “What should I compare…?”, and “Is this service worth the price?”. AnswerPath turns the questions you already own into a usable **answer-path map** for content, FAQ, schema and AI visibility research.
14
+
15
+ It is designed for marketers, publishers, agencies and product teams who need to:
16
+
17
+ - discover and normalize user questions from ChatGPT, Claude, Codex, Antigravity, Cursor and chatbot exports;
18
+ - group questions by intent and decision stage;
19
+ - identify recurring questions without silently inventing demand;
20
+ - create an auditable prompt bank for GEO/AEO experiments;
21
+ - keep private conversation data on the operator’s machine.
22
+
23
+ ## Installation
24
+
25
+ Requires Python 3.10+.
26
+
27
+ ```bash
28
+ git clone https://github.com/tmolavi/answerpath-geo.git
29
+ cd answerpath-geo
30
+ python3 -m venv .venv
31
+ .venv/bin/pip install -e .
32
+ ```
33
+
34
+ ## Quick start
35
+
36
+ Create a prompt map for a service or keyword:
37
+
38
+ ```bash
39
+ answerpath "طراحی سایت فروشگاهی"
40
+ ```
41
+
42
+ The command writes `answerpath-output/questions.json` and `questions.csv`.
43
+
44
+ Analyze owned conversation data and add generated discovery prompts:
45
+
46
+ ```bash
47
+ answerpath "مشاوره سئو پزشکی" \\
48
+ --input ~/Downloads/chatgpt-export.zip \\
49
+ --input ./support-chat.json \\
50
+ --out ./research/seo-medical
51
+ ```
52
+
53
+ Keep only questions actually found in your supplied data:
54
+
55
+ ```bash
56
+ answerpath "سرویس حسابداری" --input ./logs --no-generated
57
+ ```
58
+
59
+ ## Supported inputs
60
+
61
+ AnswerPath reads local JSON, JSONL, CSV, directories and ZIP archives. It recognizes common `role=user|human|customer` fields and nested ChatGPT/Claude export structures. It is compatible in principle with exports produced by tools such as [openai_export_parser](https://github.com/temnoon/openai_export_parser), [llm-export-analytics](https://github.com/noah-chelednik/llm-export-analytics), and the local multi-client history model used by [ContextBridgeAI](https://github.com/T-Gojo/ContextBridgeAi).
62
+
63
+ It does not log in to ChatGPT, Claude, Google, or any other account, scrape other users, or obtain API keys.
64
+
65
+ ## Outputs
66
+
67
+ Each normalized question contains:
68
+
69
+ | Field | Meaning |
70
+ |---|---|
71
+ | `text` | Question text as extracted or generated |
72
+ | `evidence` | `observed`, `generated`, or `observed+generated` |
73
+ | `source` | Input file or generation rule |
74
+ | `intent` | `learn`, `compare`, `buy`, `solve`, `trust`, or `discover` |
75
+ | `stage` | Awareness, consideration or decision |
76
+ | `frequency` | Number of similar records merged |
77
+ | `cluster` | Stable output cluster identifier |
78
+
79
+ Frequency is meaningful only for observed records from a defined dataset. Template prompts are clearly labeled and must not be reported as customer demand.
80
+
81
+ ## GEO and AEO method
82
+
83
+ AnswerPath supports a defensible workflow:
84
+
85
+ 1. **Collect** owned exports, support logs or application traces.
86
+ 2. **Extract** user turns and preserve their source path.
87
+ 3. **Normalize** whitespace and common message formats.
88
+ 4. **Classify** intent and decision stage with inspectable rules.
89
+ 5. **Deduplicate** near-identical questions while retaining frequency.
90
+ 6. **Expand** with clearly labeled prompt hypotheses when requested.
91
+ 7. **Publish answers**: create concise answer-first pages, FAQ sections, comparison tables and structured data based on recurring observed questions.
92
+ 8. **Measure** those prompts in a separate GEO benchmark; never mix hypotheses with measured demand.
93
+
94
+ The engine does not claim that a page will rank in Google or be cited by an AI system. Those are outcomes to test with your own content and provider evidence.
95
+
96
+ ## Integrations
97
+
98
+ - **ChatGPT / Claude exports:** pass the downloaded ZIP or extracted JSON to `--input`.
99
+ - **Codex, Antigravity, Cursor and other AI IDEs:** export or copy the local session data first; use a read-only copy as input.
100
+ - **FAQ workflows:** feed `questions.json` into [FAQ Extraction Pipeline](https://github.com/emfrg/faq-extraction-pipeline) for richer issue extraction, embeddings and FAQ synthesis.
101
+ - **GEO prompt discovery:** compare observed questions with generated candidates from projects such as [auto-geo](https://github.com/shadowresearch/auto-geo), keeping the evidence labels separate.
102
+
103
+ ## 🏆 Evidence & Benchmark Contribution
104
+
105
+ AnswerPath GEO generated the **Question Discovery & Intent Stratification Layer** used in the official [GEO, SEO & Digital Marketing Agency Iran 2026 Benchmark](https://github.com/tmolavi/geo-scope/tree/main/benchmarks/geo-seo-digital-agency-iran-2026.1):
106
+
107
+ - **Verified Query Dataset**: [`examples/sample_queries.json`](examples/sample_queries.json)
108
+ - **Standalone Offline Demo**: [`examples/public_demo/`](examples/public_demo/)
109
+ - **Cross-Repository Evidence Map**: [Ecosystem Evidence Flow](https://github.com/tmolavi/geo-scope/blob/main/docs/EVIDENCE_MAP.md)
110
+
111
+ * **Prompts Generated & Stratified**: 30 standardized queries.
112
+ * **Strict Demand Provenance Separation**:
113
+ * **Observed User Demand ($N=15$)**: Extracted from genuine conversational search logs (`source_type: "observed"`).
114
+ * **Exploration Hypotheses ($N=15$)**: Systematic template variations (`source_type: "generated"`).
115
+ * **5 Intent Strata**: `commercial` (general evaluation), `compare` (head-to-head alternatives), `trust` (credibility & contracts), `solve` (technical fixes), and `buy` (procurement & quotes).
116
+ * **Provenance Contract Schema**:
117
+ ```json
118
+ {
119
+ "id": "PRM-IR-001",
120
+ "prompt": "بهترین آژانس دیجیتال مارکتینگ و سئو در ایران کدام است؟",
121
+ "intent": "commercial",
122
+ "source_type": "observed",
123
+ "source_reference": "answerpath",
124
+ "cluster": "general_recommendation"
125
+ }
126
+ ```
127
+ * **Ecosystem Architecture**: See [Benchmark Ecosystem Map](docs/BENCHMARK_ECOSYSTEM.md) for data flow across AnswerPath, GEO-Scope, SAGE, MAVI, and SiteProbe.
128
+
129
+ ## 🏛️ Ecosystem
130
+
131
+ AnswerPath GEO operates as the question discovery component of the **Molavi AI Visibility Stack**:
132
+
133
+ - **Discovery**: [AnswerPath GEO](https://github.com/tmolavi/answerpath-geo)
134
+ - **Measurement**: [GEO-Scope](https://github.com/tmolavi/geo-scope)
135
+ - **Diagnostics**: [SAGE Audit](https://github.com/tmolavi/sage-audit)
136
+ - **Action**: [SiteProbe](https://github.com/tmolavi/siteprobe)
137
+ - **Protocol**: [MCP GEO Server](https://github.com/tmolavi/mcp-geo-server)
138
+
139
+ ## 📖 Runnable Python Example
140
+
141
+ Run the bundled discovery example script:
142
+ ```bash
143
+ python examples/discover_example.py
144
+ ```
145
+ Sample benchmark query payload is available in [`examples/sample_queries.json`](examples/sample_queries.json).
146
+
147
+ ## MCP setup
148
+
149
+ AnswerPath exposes one MCP tool, `discover_questions`. The stdio transport works with local Codex, Antigravity, Claude Desktop, Cursor, Windsurf and other MCP clients:
150
+
151
+ ```json
152
+ {
153
+ "mcpServers": {
154
+ "answerpath": {
155
+ "command": "/absolute/path/to/answerpath-geo/.venv/bin/answerpath",
156
+ "args": ["mcp"]
157
+ }
158
+ }
159
+ }
160
+ ```
161
+
162
+ Ask the client to call `discover_questions` with `topic`, optional owned `inputs` (`text` and `source`), and `include_generated`. Generated prompts are always labeled separately from observed questions.
163
+
164
+ For a private Cloud deployment, run the HTTP transport behind HTTPS and an authentication gateway:
165
+
166
+ ```bash
167
+ answerpath serve-mcp --host 127.0.0.1 --port 8787
168
+ ```
169
+
170
+ The JSON-RPC endpoint is `POST /mcp`. The application deliberately does not implement authentication itself: put it behind your gateway, rate limits and tenant isolation before exposing it publicly. A public URL or a successful protocol handshake does not prove that provider data is available.
171
+
172
+ ## Privacy and data boundaries
173
+
174
+ Processing is local and deterministic. AnswerPath does not transmit input files. Do not place private exports in a public repository or commit generated files containing message content. Remove or hash identifiers before sharing results. Only analyze data for which you have authorization.
175
+
176
+ ## Development & Testing
177
+
178
+ ```bash
179
+ pip install -e .
180
+ pytest tests/ -v
181
+ ```
182
+
183
+ ## 💬 Community & External Collaboration
184
+
185
+ We welcome contributions to query mining, clustering algorithms, and demand stratification:
186
+
187
+ - **Discussions**: [GitHub Discussions](https://github.com/tmolavi/answerpath-geo/discussions)
188
+ - **First Contribution Guide**: [`docs/FIRST_CONTRIBUTION.md`](docs/FIRST_CONTRIBUTION.md)
189
+ - **Research Collaboration**: [`docs/RESEARCH_COLLABORATION.md`](docs/RESEARCH_COLLABORATION.md)
190
+ - **Issues & Bug Reports**: [GitHub Issues](https://github.com/tmolavi/answerpath-geo/issues)
191
+ - **Contribution Standards**: [`CONTRIBUTING.md`](CONTRIBUTING.md) and [`SECURITY.md`](SECURITY.md)
192
+
193
+ ## 👤 Author & License
194
+
195
+ Developed by **Taghi Molavi** — [molavi.pro](https://molavi.pro)
196
+ MIT. See [LICENSE](LICENSE).
197
+
@@ -0,0 +1,2 @@
1
+ from .core import Question, extract_records, mine
2
+ __all__ = ["Question", "extract_records", "mine"]
@@ -0,0 +1,16 @@
1
+ import argparse, sys
2
+ from .core import extract_records, mine, write_outputs
3
+
4
+ def main(argv=None):
5
+ argv=sys.argv[1:] if argv is None else argv
6
+ if argv and argv[0]=="mcp":
7
+ from .mcp import stdio; stdio(); return
8
+ if argv and argv[0] in {"serve-mcp","mcp-http"}:
9
+ from .mcp import http
10
+ p=argparse.ArgumentParser(prog="answerpath serve-mcp"); p.add_argument("--host",default="127.0.0.1"); p.add_argument("--port",type=int,default=8787); a=p.parse_args(argv[1:]); http(a.host,a.port); return
11
+ p=argparse.ArgumentParser(prog="answerpath", description="Mine user questions for GEO/AEO")
12
+ p.add_argument("topic", help="service, business, website or keyword"); p.add_argument("--input", action="append", default=[], help="owned JSON/JSONL/CSV export or directory; repeatable"); p.add_argument("--out", default="answerpath-output"); p.add_argument("--no-generated", action="store_true", help="only keep observed questions")
13
+ a=p.parse_args(argv); rows=[]
14
+ for x in a.input: rows.extend(extract_records(x))
15
+ qs=mine(a.topic, rows, not a.no_generated); out=write_outputs(qs,a.out)
16
+ print(f"Wrote {len(qs)} questions to {out} (observed={sum(q.evidence.startswith('observed') for q in qs)})")
@@ -0,0 +1,120 @@
1
+ """Question mining engine: deterministic, local-first, and explicit about evidence."""
2
+ from __future__ import annotations
3
+ import csv, json, re, zipfile
4
+ from dataclasses import asdict, dataclass
5
+ from pathlib import Path
6
+ from difflib import SequenceMatcher
7
+
8
+ @dataclass
9
+ class Question:
10
+ text: str
11
+ source: str
12
+ evidence: str # observed | generated
13
+ intent: str
14
+ stage: str
15
+ frequency: int = 1
16
+ cluster: int = -1
17
+
18
+ INTENTS = {
19
+ "learn": ["what", "چیست", "چیه", "چطور", "how", "guide"],
20
+ "compare": ["compare", "مقایسه", "بهتر", "vs", "versus", "فرق"],
21
+ "buy": ["price", "قیمت", "خرید", "buy", "cost", "پکیج", "هزینه"],
22
+ "solve": ["مناسب", "حل", "مشکل", "نرم افزار", "service", "خدمت", "حل کردن"],
23
+ "trust": ["review", "نظرات", "اعتماد", "قابل اعتماد", "تجربه", "review"],
24
+ }
25
+
26
+ def _text(value) -> str:
27
+ if isinstance(value, str): return re.sub(r"\s+", " ", value).strip()
28
+ if isinstance(value, list): return _text(" ".join(_text(x.get("text", "") if isinstance(x, dict) else x) for x in value))
29
+ if isinstance(value, dict): return _text(value.get("content", value.get("text", "")))
30
+ return ""
31
+
32
+ def extract_records(path: str | Path) -> list[tuple[str, str]]:
33
+ """Read owned exports/logs; never accesses provider accounts or remote chats."""
34
+ p = Path(path); rows=[]
35
+ files=[]
36
+ if p.suffix.lower()==".zip":
37
+ with zipfile.ZipFile(p) as z:
38
+ for n in z.namelist():
39
+ if n.lower().endswith((".json", ".jsonl", ".csv")): rows.extend(_read_bytes(n, z.read(n), str(p)))
40
+ return rows
41
+ if p.is_dir(): files=[x for x in p.rglob("*") if x.suffix.lower() in {".json", ".jsonl", ".csv"}]
42
+ else: files=[p]
43
+ for f in files: rows.extend(_read_bytes(str(f), f.read_bytes(), str(f)))
44
+ return rows
45
+
46
+ def _read_bytes(name, data, source):
47
+ try: obj=json.loads(data)
48
+ except Exception:
49
+ try: obj=[json.loads(x) for x in data.decode().splitlines() if x.strip()]
50
+ except Exception: obj=None
51
+ if obj is not None: return _walk(obj, source)
52
+ try:
53
+ out=[]
54
+ for row in csv.DictReader(data.decode().splitlines()):
55
+ if str(row.get("role", row.get("speaker", ""))).lower() in {"user","human","customer"}: out.append((_text(row.get("content", row.get("text", ""))), source))
56
+ return out
57
+ except Exception: return []
58
+
59
+ def _walk(obj, source):
60
+ out=[]
61
+ if isinstance(obj, dict):
62
+ role=str(obj.get("role", obj.get("author", obj.get("speaker", "")))).lower()
63
+ if role in {"user","human","customer"}:
64
+ t=_text(obj.get("content", obj.get("text", obj.get("message", ""))))
65
+ if t: out.append((t, source))
66
+ for v in obj.values(): out.extend(_walk(v, source))
67
+ elif isinstance(obj, list):
68
+ for v in obj: out.extend(_walk(v, source))
69
+ return out
70
+
71
+ def classify(text):
72
+ low=text.lower()
73
+ # Commercial and comparison signals should win over broad "how/what"
74
+ # markers (for example, "How much does ... cost?" is a buying question).
75
+ for intent in ("buy", "compare", "trust", "solve", "learn"):
76
+ words = INTENTS[intent]
77
+ if any(w in low for w in words): return intent
78
+ return "discover"
79
+
80
+ def stage(intent): return {"learn":"awareness","compare":"consideration","trust":"consideration","buy":"decision","solve":"decision"}.get(intent,"awareness")
81
+
82
+ def generated_prompts(topic):
83
+ return [
84
+ f"What is the best {topic} for a small business?",
85
+ f"How do I choose a reliable {topic} provider?",
86
+ f"What should I compare before buying {topic}?",
87
+ f"How much does {topic} cost and what is included?",
88
+ f"Which {topic} is suitable for my needs?",
89
+ f"What are the common problems with {topic} and how are they solved?",
90
+ f"Are there trustworthy reviews or examples for {topic}?",
91
+ f"Compare the leading {topic} options for quality and price.",
92
+ ]
93
+
94
+ def mine(topic, inputs=(), include_generated=True, threshold=.88):
95
+ qs=[]
96
+ for text, source in inputs:
97
+ text=_text(text)
98
+ if len(text)<4 or len(text)>1000: continue
99
+ intent=classify(text); qs.append(Question(text, source, "observed", intent, stage(intent)))
100
+ if include_generated:
101
+ for text in generated_prompts(topic):
102
+ intent=classify(text); qs.append(Question(text, "answerpath:template", "generated", intent, stage(intent)))
103
+ groups=[]
104
+ for q in qs:
105
+ match=None
106
+ for existing in groups:
107
+ if SequenceMatcher(None, q.text.lower(), existing.text.lower()).ratio()>=threshold:
108
+ match=existing; break
109
+ if match: match.frequency+=1; match.source += ";"+q.source; match.evidence = "observed+generated" if match.evidence!=q.evidence else match.evidence
110
+ else: groups.append(q)
111
+ for i,q in enumerate(groups): q.cluster=i
112
+ return groups
113
+
114
+ def write_outputs(questions, out):
115
+ out=Path(out); out.mkdir(parents=True, exist_ok=True)
116
+ data=[asdict(q) for q in questions]
117
+ (out/"questions.json").write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")
118
+ with (out/"questions.csv").open("w", newline="", encoding="utf-8") as f:
119
+ w=csv.DictWriter(f, fieldnames=data[0].keys() if data else ["text"]); w.writeheader(); w.writerows(data)
120
+ return out
@@ -0,0 +1,37 @@
1
+ """Minimal dependency-free MCP-compatible JSON-RPC transports."""
2
+ import json, sys
3
+ from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
4
+ from .core import mine
5
+
6
+ TOOLS = [{"name":"discover_questions","description":"Extract and classify observed questions, then optionally add labeled GEO/AEO hypotheses.","inputSchema":{"type":"object","properties":{"topic":{"type":"string"},"inputs":{"type":"array","items":{"type":"object","properties":{"text":{"type":"string"},"source":{"type":"string"}}}},"include_generated":{"type":"boolean","default":True}},"required":["topic"]}}]
7
+
8
+ def call_tool(name, args):
9
+ if name != "discover_questions": raise ValueError(f"unknown tool: {name}")
10
+ rows=[(x.get("text",""), x.get("source","mcp")) for x in args.get("inputs",[]) if isinstance(x,dict)]
11
+ result=[q.__dict__ for q in mine(args["topic"], rows, args.get("include_generated",True))]
12
+ return {"content":[{"type":"text","text":json.dumps(result,ensure_ascii=False)}],"structuredContent":{"questions":result}}
13
+
14
+ def handle(req):
15
+ method=req.get("method"); rid=req.get("id")
16
+ if method=="initialize": result={"protocolVersion":"2024-11-05","capabilities":{"tools":{}},"serverInfo":{"name":"answerpath-geo","version":"0.1.0"}}
17
+ elif method=="tools/list": result={"tools":TOOLS}
18
+ elif method=="tools/call": result=call_tool(req.get("params",{}).get("name"),req.get("params",{}).get("arguments",{}))
19
+ else: return {"jsonrpc":"2.0","id":rid,"error":{"code":-32601,"message":"Method not found"}}
20
+ return {"jsonrpc":"2.0","id":rid,"result":result}
21
+
22
+ def stdio():
23
+ for line in sys.stdin:
24
+ try: print(json.dumps(handle(json.loads(line)),ensure_ascii=False),flush=True)
25
+ except Exception as e: print(json.dumps({"jsonrpc":"2.0","id":None,"error":{"code":-32000,"message":str(e)}},ensure_ascii=False),flush=True)
26
+
27
+ class Handler(BaseHTTPRequestHandler):
28
+ def do_POST(self):
29
+ if self.path.rstrip("/") != "/mcp": self.send_error(404); return
30
+ n=int(self.headers.get("Content-Length","0"));
31
+ try: body=json.loads(self.rfile.read(n)); out=handle(body); raw=json.dumps(out,ensure_ascii=False).encode()
32
+ except Exception as e: raw=json.dumps({"jsonrpc":"2.0","id":None,"error":{"code":-32000,"message":str(e)}},ensure_ascii=False).encode()
33
+ self.send_response(200); self.send_header("Content-Type","application/json"); self.send_header("Content-Length",str(len(raw))); self.end_headers(); self.wfile.write(raw)
34
+ def log_message(self,*args): pass
35
+
36
+ def http(host="127.0.0.1", port=8787):
37
+ ThreadingHTTPServer((host,port),Handler).serve_forever()
@@ -0,0 +1,215 @@
1
+ Metadata-Version: 2.4
2
+ Name: answerpath-geo
3
+ Version: 0.1.0
4
+ Summary: Privacy-first user question mining for GEO and AEO
5
+ Author: Taqi Molavi
6
+ License: MIT
7
+ Project-URL: Homepage, https://molavi.pro
8
+ Project-URL: Repository, https://github.com/tmolavi/answerpath-geo
9
+ Project-URL: Documentation, https://github.com/tmolavi/answerpath-geo#readme
10
+ Project-URL: Issues, https://github.com/tmolavi/answerpath-geo/issues
11
+ Keywords: geo,aeo,seo,answer-engine-optimization,prompt-mining,user-intent,question-mining
12
+ Classifier: License :: OSI Approved :: MIT License
13
+ Classifier: Programming Language :: Python :: 3
14
+ Requires-Python: >=3.10
15
+ Description-Content-Type: text/markdown
16
+ License-File: LICENSE
17
+ Dynamic: license-file
18
+
19
+ # AnswerPath GEO — Discover the Questions People Ask AI Before They Find Your Business
20
+
21
+ **AnswerPath GEO** is a privacy-first, open-source **SEO, GEO and AEO question mining engine**. Give it a business, service, website topic or keyword and it builds a transparent question map showing what people may ask an AI assistant before choosing a provider.
22
+
23
+ It separates **observed questions** extracted from exports and application logs from **generated research prompts**. This distinction matters: generated prompts are hypotheses, not evidence of what people actually asked.
24
+
25
+ [Installation](#installation) · [Quick start](#quick-start) · [Inputs](#supported-inputs) · [Outputs](#outputs) · [GEO/AEO method](#geo-and-aeo-method) · [Privacy](#privacy-and-data-boundaries) · [Integrations](#integrations)
26
+
27
+ **MCP clients:** [Codex, Antigravity, Claude, Cursor and Cloud setup](#mcp-setup)
28
+
29
+ ## Why AnswerPath GEO?
30
+
31
+ Traditional keyword tools show phrases typed into search engines. AI assistants receive longer, conversational questions: “Which agency is reliable for…?”, “What should I compare…?”, and “Is this service worth the price?”. AnswerPath turns the questions you already own into a usable **answer-path map** for content, FAQ, schema and AI visibility research.
32
+
33
+ It is designed for marketers, publishers, agencies and product teams who need to:
34
+
35
+ - discover and normalize user questions from ChatGPT, Claude, Codex, Antigravity, Cursor and chatbot exports;
36
+ - group questions by intent and decision stage;
37
+ - identify recurring questions without silently inventing demand;
38
+ - create an auditable prompt bank for GEO/AEO experiments;
39
+ - keep private conversation data on the operator’s machine.
40
+
41
+ ## Installation
42
+
43
+ Requires Python 3.10+.
44
+
45
+ ```bash
46
+ git clone https://github.com/tmolavi/answerpath-geo.git
47
+ cd answerpath-geo
48
+ python3 -m venv .venv
49
+ .venv/bin/pip install -e .
50
+ ```
51
+
52
+ ## Quick start
53
+
54
+ Create a prompt map for a service or keyword:
55
+
56
+ ```bash
57
+ answerpath "طراحی سایت فروشگاهی"
58
+ ```
59
+
60
+ The command writes `answerpath-output/questions.json` and `questions.csv`.
61
+
62
+ Analyze owned conversation data and add generated discovery prompts:
63
+
64
+ ```bash
65
+ answerpath "مشاوره سئو پزشکی" \\
66
+ --input ~/Downloads/chatgpt-export.zip \\
67
+ --input ./support-chat.json \\
68
+ --out ./research/seo-medical
69
+ ```
70
+
71
+ Keep only questions actually found in your supplied data:
72
+
73
+ ```bash
74
+ answerpath "سرویس حسابداری" --input ./logs --no-generated
75
+ ```
76
+
77
+ ## Supported inputs
78
+
79
+ AnswerPath reads local JSON, JSONL, CSV, directories and ZIP archives. It recognizes common `role=user|human|customer` fields and nested ChatGPT/Claude export structures. It is compatible in principle with exports produced by tools such as [openai_export_parser](https://github.com/temnoon/openai_export_parser), [llm-export-analytics](https://github.com/noah-chelednik/llm-export-analytics), and the local multi-client history model used by [ContextBridgeAI](https://github.com/T-Gojo/ContextBridgeAi).
80
+
81
+ It does not log in to ChatGPT, Claude, Google, or any other account, scrape other users, or obtain API keys.
82
+
83
+ ## Outputs
84
+
85
+ Each normalized question contains:
86
+
87
+ | Field | Meaning |
88
+ |---|---|
89
+ | `text` | Question text as extracted or generated |
90
+ | `evidence` | `observed`, `generated`, or `observed+generated` |
91
+ | `source` | Input file or generation rule |
92
+ | `intent` | `learn`, `compare`, `buy`, `solve`, `trust`, or `discover` |
93
+ | `stage` | Awareness, consideration or decision |
94
+ | `frequency` | Number of similar records merged |
95
+ | `cluster` | Stable output cluster identifier |
96
+
97
+ Frequency is meaningful only for observed records from a defined dataset. Template prompts are clearly labeled and must not be reported as customer demand.
98
+
99
+ ## GEO and AEO method
100
+
101
+ AnswerPath supports a defensible workflow:
102
+
103
+ 1. **Collect** owned exports, support logs or application traces.
104
+ 2. **Extract** user turns and preserve their source path.
105
+ 3. **Normalize** whitespace and common message formats.
106
+ 4. **Classify** intent and decision stage with inspectable rules.
107
+ 5. **Deduplicate** near-identical questions while retaining frequency.
108
+ 6. **Expand** with clearly labeled prompt hypotheses when requested.
109
+ 7. **Publish answers**: create concise answer-first pages, FAQ sections, comparison tables and structured data based on recurring observed questions.
110
+ 8. **Measure** those prompts in a separate GEO benchmark; never mix hypotheses with measured demand.
111
+
112
+ The engine does not claim that a page will rank in Google or be cited by an AI system. Those are outcomes to test with your own content and provider evidence.
113
+
114
+ ## Integrations
115
+
116
+ - **ChatGPT / Claude exports:** pass the downloaded ZIP or extracted JSON to `--input`.
117
+ - **Codex, Antigravity, Cursor and other AI IDEs:** export or copy the local session data first; use a read-only copy as input.
118
+ - **FAQ workflows:** feed `questions.json` into [FAQ Extraction Pipeline](https://github.com/emfrg/faq-extraction-pipeline) for richer issue extraction, embeddings and FAQ synthesis.
119
+ - **GEO prompt discovery:** compare observed questions with generated candidates from projects such as [auto-geo](https://github.com/shadowresearch/auto-geo), keeping the evidence labels separate.
120
+
121
+ ## 🏆 Evidence & Benchmark Contribution
122
+
123
+ AnswerPath GEO generated the **Question Discovery & Intent Stratification Layer** used in the official [GEO, SEO & Digital Marketing Agency Iran 2026 Benchmark](https://github.com/tmolavi/geo-scope/tree/main/benchmarks/geo-seo-digital-agency-iran-2026.1):
124
+
125
+ - **Verified Query Dataset**: [`examples/sample_queries.json`](examples/sample_queries.json)
126
+ - **Standalone Offline Demo**: [`examples/public_demo/`](examples/public_demo/)
127
+ - **Cross-Repository Evidence Map**: [Ecosystem Evidence Flow](https://github.com/tmolavi/geo-scope/blob/main/docs/EVIDENCE_MAP.md)
128
+
129
+ * **Prompts Generated & Stratified**: 30 standardized queries.
130
+ * **Strict Demand Provenance Separation**:
131
+ * **Observed User Demand ($N=15$)**: Extracted from genuine conversational search logs (`source_type: "observed"`).
132
+ * **Exploration Hypotheses ($N=15$)**: Systematic template variations (`source_type: "generated"`).
133
+ * **5 Intent Strata**: `commercial` (general evaluation), `compare` (head-to-head alternatives), `trust` (credibility & contracts), `solve` (technical fixes), and `buy` (procurement & quotes).
134
+ * **Provenance Contract Schema**:
135
+ ```json
136
+ {
137
+ "id": "PRM-IR-001",
138
+ "prompt": "بهترین آژانس دیجیتال مارکتینگ و سئو در ایران کدام است؟",
139
+ "intent": "commercial",
140
+ "source_type": "observed",
141
+ "source_reference": "answerpath",
142
+ "cluster": "general_recommendation"
143
+ }
144
+ ```
145
+ * **Ecosystem Architecture**: See [Benchmark Ecosystem Map](docs/BENCHMARK_ECOSYSTEM.md) for data flow across AnswerPath, GEO-Scope, SAGE, MAVI, and SiteProbe.
146
+
147
+ ## 🏛️ Ecosystem
148
+
149
+ AnswerPath GEO operates as the question discovery component of the **Molavi AI Visibility Stack**:
150
+
151
+ - **Discovery**: [AnswerPath GEO](https://github.com/tmolavi/answerpath-geo)
152
+ - **Measurement**: [GEO-Scope](https://github.com/tmolavi/geo-scope)
153
+ - **Diagnostics**: [SAGE Audit](https://github.com/tmolavi/sage-audit)
154
+ - **Action**: [SiteProbe](https://github.com/tmolavi/siteprobe)
155
+ - **Protocol**: [MCP GEO Server](https://github.com/tmolavi/mcp-geo-server)
156
+
157
+ ## 📖 Runnable Python Example
158
+
159
+ Run the bundled discovery example script:
160
+ ```bash
161
+ python examples/discover_example.py
162
+ ```
163
+ Sample benchmark query payload is available in [`examples/sample_queries.json`](examples/sample_queries.json).
164
+
165
+ ## MCP setup
166
+
167
+ AnswerPath exposes one MCP tool, `discover_questions`. The stdio transport works with local Codex, Antigravity, Claude Desktop, Cursor, Windsurf and other MCP clients:
168
+
169
+ ```json
170
+ {
171
+ "mcpServers": {
172
+ "answerpath": {
173
+ "command": "/absolute/path/to/answerpath-geo/.venv/bin/answerpath",
174
+ "args": ["mcp"]
175
+ }
176
+ }
177
+ }
178
+ ```
179
+
180
+ Ask the client to call `discover_questions` with `topic`, optional owned `inputs` (`text` and `source`), and `include_generated`. Generated prompts are always labeled separately from observed questions.
181
+
182
+ For a private Cloud deployment, run the HTTP transport behind HTTPS and an authentication gateway:
183
+
184
+ ```bash
185
+ answerpath serve-mcp --host 127.0.0.1 --port 8787
186
+ ```
187
+
188
+ The JSON-RPC endpoint is `POST /mcp`. The application deliberately does not implement authentication itself: put it behind your gateway, rate limits and tenant isolation before exposing it publicly. A public URL or a successful protocol handshake does not prove that provider data is available.
189
+
190
+ ## Privacy and data boundaries
191
+
192
+ Processing is local and deterministic. AnswerPath does not transmit input files. Do not place private exports in a public repository or commit generated files containing message content. Remove or hash identifiers before sharing results. Only analyze data for which you have authorization.
193
+
194
+ ## Development & Testing
195
+
196
+ ```bash
197
+ pip install -e .
198
+ pytest tests/ -v
199
+ ```
200
+
201
+ ## 💬 Community & External Collaboration
202
+
203
+ We welcome contributions to query mining, clustering algorithms, and demand stratification:
204
+
205
+ - **Discussions**: [GitHub Discussions](https://github.com/tmolavi/answerpath-geo/discussions)
206
+ - **First Contribution Guide**: [`docs/FIRST_CONTRIBUTION.md`](docs/FIRST_CONTRIBUTION.md)
207
+ - **Research Collaboration**: [`docs/RESEARCH_COLLABORATION.md`](docs/RESEARCH_COLLABORATION.md)
208
+ - **Issues & Bug Reports**: [GitHub Issues](https://github.com/tmolavi/answerpath-geo/issues)
209
+ - **Contribution Standards**: [`CONTRIBUTING.md`](CONTRIBUTING.md) and [`SECURITY.md`](SECURITY.md)
210
+
211
+ ## 👤 Author & License
212
+
213
+ Developed by **Taghi Molavi** — [molavi.pro](https://molavi.pro)
214
+ MIT. See [LICENSE](LICENSE).
215
+
@@ -0,0 +1,13 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ answerpath_geo/__init__.py
5
+ answerpath_geo/cli.py
6
+ answerpath_geo/core.py
7
+ answerpath_geo/mcp.py
8
+ answerpath_geo.egg-info/PKG-INFO
9
+ answerpath_geo.egg-info/SOURCES.txt
10
+ answerpath_geo.egg-info/dependency_links.txt
11
+ answerpath_geo.egg-info/entry_points.txt
12
+ answerpath_geo.egg-info/top_level.txt
13
+ tests/test_core.py
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ answerpath = answerpath_geo.cli:main
@@ -0,0 +1 @@
1
+ answerpath_geo
@@ -0,0 +1,23 @@
1
+ [build-system]
2
+ requires = ["setuptools>=68"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "answerpath-geo"
7
+ version = "0.1.0"
8
+ description = "Privacy-first user question mining for GEO and AEO"
9
+ readme = "README.md"
10
+ requires-python = ">=3.10"
11
+ authors = [{name = "Taqi Molavi"}]
12
+ license = {text = "MIT"}
13
+ keywords = ["geo", "aeo", "seo", "answer-engine-optimization", "prompt-mining", "user-intent", "question-mining"]
14
+ classifiers = ["License :: OSI Approved :: MIT License", "Programming Language :: Python :: 3"]
15
+
16
+ [project.scripts]
17
+ answerpath = "answerpath_geo.cli:main"
18
+
19
+ [project.urls]
20
+ Homepage = "https://molavi.pro"
21
+ Repository = "https://github.com/tmolavi/answerpath-geo"
22
+ Documentation = "https://github.com/tmolavi/answerpath-geo#readme"
23
+ Issues = "https://github.com/tmolavi/answerpath-geo/issues"
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,22 @@
1
+ import json
2
+ from answerpath_geo.core import extract_records, mine, write_outputs
3
+
4
+ def test_observed_and_generated_are_labeled(tmp_path):
5
+ p=tmp_path/'chat.json'; p.write_text(json.dumps({'messages':[{'role':'user','content':'What is the best accounting service?'}]}))
6
+ qs=mine('accounting service', extract_records(p))
7
+ assert any(q.evidence=='observed' for q in qs)
8
+ assert all(q.evidence in {'observed','generated','observed+generated'} for q in qs)
9
+
10
+ def test_no_generated(tmp_path):
11
+ qs=mine('seo', [('How much does SEO cost?', 'x')], include_generated=False)
12
+ assert len(qs)==1 and qs[0].evidence=='observed' and qs[0].intent=='buy'
13
+
14
+ def test_outputs(tmp_path):
15
+ out=write_outputs(mine('web design', [], True), tmp_path/'out')
16
+ assert (out/'questions.json').exists() and (out/'questions.csv').exists()
17
+
18
+ def test_mcp_protocol():
19
+ from answerpath_geo.mcp import handle
20
+ assert handle({'jsonrpc':'2.0','id':1,'method':'initialize','params':{}})['result']['serverInfo']['name']=='answerpath-geo'
21
+ out=handle({'jsonrpc':'2.0','id':2,'method':'tools/call','params':{'name':'discover_questions','arguments':{'topic':'web design','include_generated':False,'inputs':[{'text':'How much does web design cost?','source':'test'}]}}})
22
+ assert out['result']['structuredContent']['questions'][0]['intent']=='buy'