topicscout 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Adecubed
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,151 @@
1
+ Metadata-Version: 2.4
2
+ Name: topicscout
3
+ Version: 0.1.0
4
+ Summary: Find the GitHub repos on a topic, and only the new ones next time. No dependencies, token optional.
5
+ Author: Adecubed
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/adecubed/topicscout
8
+ Keywords: github,scraper,search,discovery,topics
9
+ Requires-Python: >=3.11
10
+ Description-Content-Type: text/markdown
11
+ License-File: LICENSE
12
+ Dynamic: license-file
13
+
14
+ # topicscout
15
+
16
+ [![ci](https://github.com/adecubed/topicscout/actions/workflows/ci.yml/badge.svg)](https://github.com/adecubed/topicscout/actions/workflows/ci.yml)
17
+ [![PyPI](https://img.shields.io/pypi/v/topicscout)](https://pypi.org/project/topicscout/)
18
+
19
+ Find the GitHub repositories on a topic, and next time only the new ones.
20
+
21
+ ```bash
22
+ pip install topicscout # or run it without installing: uvx topicscout run github-scrapers
23
+ topicscout run github-scrapers
24
+ ```
25
+
26
+ ```
27
+ 19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
28
+ ```
29
+
30
+ `new.md` is a table of what appeared since the last run, most stars first: stars,
31
+ last push, language, license, a guess at the interface from the README (MCP, pip,
32
+ npm, REST, Docker...) and the description. `candidates.json` keeps everything seen,
33
+ so a weekly run is a short list, not the same 300 repos again.
34
+
35
+ No dependencies, Python 3.11+. `GITHUB_TOKEN` is optional: without it the searches
36
+ are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.
37
+
38
+ ## Why
39
+
40
+ GitHub search is good at "repos with this topic" and bad at "what's new since I last
41
+ looked, minus the forks, the awesome-lists, the abandoned ones and the repos that
42
+ only tagged themselves". topicscout runs a fixed set of searches, applies your
43
+ filters, remembers what it has seen and what you already have.
44
+
45
+ It was the discovery step of [adebench](https://github.com/adecubed/adebench), a
46
+ benchmark for AI memory systems, which uses it to find new memories to measure.
47
+
48
+ ## Your own topic
49
+
50
+ Copy a built-in profile (they are in [`topicscout/profiles/`](topicscout/profiles/)),
51
+ change the queries and the words a hit must contain, and run it:
52
+
53
+ ```bash
54
+ topicscout run my-topic.json
55
+ ```
56
+
57
+ Run it again next week: `new.md` lists only what appeared in between. Put the repos
58
+ you already know or rejected in `scout-my-topic/known.txt`.
59
+
60
+ ## From a script or an agent
61
+
62
+ `--json` prints the result on stdout, progress stays on stderr:
63
+
64
+ ```bash
65
+ topicscout run github-scrapers --json -q
66
+ ```
67
+
68
+ ```json
69
+ {"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
70
+ ```
71
+
72
+ It is a command-line tool and a Python library (`from topicscout.core import Profile, run`),
73
+ not an MCP server; an agent calls it like any other command.
74
+
75
+ ## Commands
76
+
77
+ ```bash
78
+ topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
79
+ [--known-from GLOB ...] [--json] [-q]
80
+ topicscout profiles # built-in profiles
81
+ topicscout doctor # token and remaining GitHub quota
82
+ ```
83
+
84
+ - `PROFILE` is a built-in name or a path to your own JSON profile.
85
+ - `--out` is the state folder (default `scout-<profile>`): `candidates.json`,
86
+ `new.md`, and your `known.txt`.
87
+ - `--json` prints the summary and the new repos as JSON on stdout; progress and
88
+ errors always go to stderr, so scripts and agents can read stdout as is.
89
+
90
+ ## What counts as known
91
+
92
+ Repos you already have are marked `known` and never listed as new:
93
+
94
+ - `known.txt` in the state folder: one `owner/repo` per line, `#` comments. Put
95
+ rejected repos here too.
96
+ - `--known-from "adapters/*.py"`: the GitHub URLs in the first 20 lines of those
97
+ files. adebench points it at its adapters, so a memory with an adapter drops out.
98
+
99
+ ## Profiles
100
+
101
+ Built in: `ai-memory` (memory systems for AI agents and LLMs) and `github-scrapers`
102
+ (tools that scrape, crawl or mine GitHub). A profile is a JSON file:
103
+
104
+ ```json
105
+ {
106
+ "name": "ai-memory",
107
+ "description": "Memory systems for AI agents and LLMs",
108
+ "queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
109
+ "exclude": "(?i)awesome|curated list|\\bpapers\\b",
110
+ "require": [
111
+ {"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
112
+ {"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
113
+ ],
114
+ "interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
115
+ "min_stars": 10,
116
+ "days": 90
117
+ }
118
+ ```
119
+
120
+ - `queries`: [GitHub repository search](https://docs.github.com/en/search-github/searching-on-github/searching-for-repositories)
121
+ queries. Each one always gets `fork:false archived:false stars:>=min_stars
122
+ pushed:>=today-days`; up to 300 results per query.
123
+ - `exclude`: a regex on name, description and topics; a match drops the repo.
124
+ - `require`: every pattern must match. With `own_words` it must be in the name or
125
+ description, not only in a topic: a database that tagged itself `agent-memory`
126
+ is not a memory. Hyphens and underscores count as spaces.
127
+ - `interfaces`: `[label, regex]` pairs tried on the README of each new repo.
128
+ Leave it out and no README is fetched.
129
+ - `require_language` (default true) drops repos with no language, which are
130
+ usually lists and docs.
131
+
132
+ Patterns are Python regexes; add `(?i)` for case-insensitive.
133
+
134
+ ## Rate limits
135
+
136
+ On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
137
+ stops, saves what it found and says so in `new.md` and on stderr. The next run
138
+ picks up the READMEs it did not get to.
139
+
140
+ ## How it compares
141
+
142
+ - [ghcrawl](https://github.com/pwrdrvr/ghcrawl) goes deep into one repository
143
+ (issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a
144
+ topic.
145
+ - [top-github-scraper](https://github.com/khuyentran1401/top-github-scraper)
146
+ lists the top repos for a keyword, once; topicscout adds filters, profiles and
147
+ memory across runs.
148
+
149
+ ## License
150
+
151
+ MIT
@@ -0,0 +1,138 @@
1
+ # topicscout
2
+
3
+ [![ci](https://github.com/adecubed/topicscout/actions/workflows/ci.yml/badge.svg)](https://github.com/adecubed/topicscout/actions/workflows/ci.yml)
4
+ [![PyPI](https://img.shields.io/pypi/v/topicscout)](https://pypi.org/project/topicscout/)
5
+
6
+ Find the GitHub repositories on a topic, and next time only the new ones.
7
+
8
+ ```bash
9
+ pip install topicscout # or run it without installing: uvx topicscout run github-scrapers
10
+ topicscout run github-scrapers
11
+ ```
12
+
13
+ ```
14
+ 19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
15
+ ```
16
+
17
+ `new.md` is a table of what appeared since the last run, most stars first: stars,
18
+ last push, language, license, a guess at the interface from the README (MCP, pip,
19
+ npm, REST, Docker...) and the description. `candidates.json` keeps everything seen,
20
+ so a weekly run is a short list, not the same 300 repos again.
21
+
22
+ No dependencies, Python 3.11+. `GITHUB_TOKEN` is optional: without it the searches
23
+ are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.
24
+
25
+ ## Why
26
+
27
+ GitHub search is good at "repos with this topic" and bad at "what's new since I last
28
+ looked, minus the forks, the awesome-lists, the abandoned ones and the repos that
29
+ only tagged themselves". topicscout runs a fixed set of searches, applies your
30
+ filters, remembers what it has seen and what you already have.
31
+
32
+ It was the discovery step of [adebench](https://github.com/adecubed/adebench), a
33
+ benchmark for AI memory systems, which uses it to find new memories to measure.
34
+
35
+ ## Your own topic
36
+
37
+ Copy a built-in profile (they are in [`topicscout/profiles/`](topicscout/profiles/)),
38
+ change the queries and the words a hit must contain, and run it:
39
+
40
+ ```bash
41
+ topicscout run my-topic.json
42
+ ```
43
+
44
+ Run it again next week: `new.md` lists only what appeared in between. Put the repos
45
+ you already know or rejected in `scout-my-topic/known.txt`.
46
+
47
+ ## From a script or an agent
48
+
49
+ `--json` prints the result on stdout, progress stays on stderr:
50
+
51
+ ```bash
52
+ topicscout run github-scrapers --json -q
53
+ ```
54
+
55
+ ```json
56
+ {"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
57
+ ```
58
+
59
+ It is a command-line tool and a Python library (`from topicscout.core import Profile, run`),
60
+ not an MCP server; an agent calls it like any other command.
61
+
62
+ ## Commands
63
+
64
+ ```bash
65
+ topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
66
+ [--known-from GLOB ...] [--json] [-q]
67
+ topicscout profiles # built-in profiles
68
+ topicscout doctor # token and remaining GitHub quota
69
+ ```
70
+
71
+ - `PROFILE` is a built-in name or a path to your own JSON profile.
72
+ - `--out` is the state folder (default `scout-<profile>`): `candidates.json`,
73
+ `new.md`, and your `known.txt`.
74
+ - `--json` prints the summary and the new repos as JSON on stdout; progress and
75
+ errors always go to stderr, so scripts and agents can read stdout as is.
76
+
77
+ ## What counts as known
78
+
79
+ Repos you already have are marked `known` and never listed as new:
80
+
81
+ - `known.txt` in the state folder: one `owner/repo` per line, `#` comments. Put
82
+ rejected repos here too.
83
+ - `--known-from "adapters/*.py"`: the GitHub URLs in the first 20 lines of those
84
+ files. adebench points it at its adapters, so a memory with an adapter drops out.
85
+
86
+ ## Profiles
87
+
88
+ Built in: `ai-memory` (memory systems for AI agents and LLMs) and `github-scrapers`
89
+ (tools that scrape, crawl or mine GitHub). A profile is a JSON file:
90
+
91
+ ```json
92
+ {
93
+ "name": "ai-memory",
94
+ "description": "Memory systems for AI agents and LLMs",
95
+ "queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
96
+ "exclude": "(?i)awesome|curated list|\\bpapers\\b",
97
+ "require": [
98
+ {"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
99
+ {"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
100
+ ],
101
+ "interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
102
+ "min_stars": 10,
103
+ "days": 90
104
+ }
105
+ ```
106
+
107
+ - `queries`: [GitHub repository search](https://docs.github.com/en/search-github/searching-on-github/searching-for-repositories)
108
+ queries. Each one always gets `fork:false archived:false stars:>=min_stars
109
+ pushed:>=today-days`; up to 300 results per query.
110
+ - `exclude`: a regex on name, description and topics; a match drops the repo.
111
+ - `require`: every pattern must match. With `own_words` it must be in the name or
112
+ description, not only in a topic: a database that tagged itself `agent-memory`
113
+ is not a memory. Hyphens and underscores count as spaces.
114
+ - `interfaces`: `[label, regex]` pairs tried on the README of each new repo.
115
+ Leave it out and no README is fetched.
116
+ - `require_language` (default true) drops repos with no language, which are
117
+ usually lists and docs.
118
+
119
+ Patterns are Python regexes; add `(?i)` for case-insensitive.
120
+
121
+ ## Rate limits
122
+
123
+ On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
124
+ stops, saves what it found and says so in `new.md` and on stderr. The next run
125
+ picks up the READMEs it did not get to.
126
+
127
+ ## How it compares
128
+
129
+ - [ghcrawl](https://github.com/pwrdrvr/ghcrawl) goes deep into one repository
130
+ (issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a
131
+ topic.
132
+ - [top-github-scraper](https://github.com/khuyentran1401/top-github-scraper)
133
+ lists the top repos for a keyword, once; topicscout adds filters, profiles and
134
+ memory across runs.
135
+
136
+ ## License
137
+
138
+ MIT
@@ -0,0 +1,26 @@
1
+ [project]
2
+ name = "topicscout"
3
+ version = "0.1.0"
4
+ description = "Find the GitHub repos on a topic, and only the new ones next time. No dependencies, token optional."
5
+ readme = "README.md"
6
+ requires-python = ">=3.11"
7
+ license = "MIT"
8
+ authors = [{name = "Adecubed"}]
9
+ dependencies = []
10
+ keywords = ["github", "scraper", "search", "discovery", "topics"]
11
+
12
+ [project.urls]
13
+ Homepage = "https://github.com/adecubed/topicscout"
14
+
15
+ [project.scripts]
16
+ topicscout = "topicscout.cli:main"
17
+
18
+ [build-system]
19
+ requires = ["setuptools>=77"]
20
+ build-backend = "setuptools.build_meta"
21
+
22
+ [tool.setuptools]
23
+ packages = ["topicscout", "topicscout.profiles"]
24
+
25
+ [tool.setuptools.package-data]
26
+ "topicscout.profiles" = ["*.json"]
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,198 @@
1
+ """topicscout with fake HTTP (no network)."""
2
+ from __future__ import annotations
3
+
4
+ import json
5
+ from datetime import date
6
+ from urllib.parse import parse_qs, urlparse
7
+
8
+ import pytest
9
+
10
+ from topicscout import cli, core
11
+ from topicscout.core import Profile
12
+
13
+ TODAY = date(2026, 9, 29)
14
+ MEM = Profile.load("ai-memory")
15
+
16
+
17
+ def repo(name, stars=50, pushed="2026-09-20T10:00:00Z", fork=False, archived=False,
18
+ language="Python", description="Long-term memory for agents", topics=("agent-memory",)):
19
+ return {"full_name": name, "html_url": f"https://github.com/{name}", "stargazers_count": stars,
20
+ "pushed_at": pushed, "fork": fork, "archived": archived, "language": language,
21
+ "description": description, "license": {"spdx_id": "MIT"}, "topics": list(topics)}
22
+
23
+
24
+ class FakeGitHub:
25
+ """Every search query returns the same items; READMEs from a dict."""
26
+
27
+ def __init__(self, items, readmes=None, limited_after=None):
28
+ self.items, self.readmes = items, readmes or {}
29
+ self.limited_after = limited_after
30
+ self.searches, self.readme_calls = [], []
31
+
32
+ def __call__(self, url, accept):
33
+ path = urlparse(url).path
34
+ if path == "/search/repositories":
35
+ if self.limited_after is not None and len(self.searches) >= self.limited_after:
36
+ return 403, {"x-ratelimit-remaining": "0", "x-ratelimit-reset": "9999999999"}, "{}"
37
+ self.searches.append(parse_qs(urlparse(url).query)["q"][0])
38
+ return 200, {}, json.dumps({"items": self.items})
39
+ if path.endswith("/readme"):
40
+ name = path[len("/repos/"):-len("/readme")]
41
+ self.readme_calls.append(name)
42
+ if name in self.readmes:
43
+ return 200, {}, self.readmes[name]
44
+ return 404, {}, ""
45
+ raise AssertionError(url)
46
+
47
+
48
+ def run(tmp_path, gh, profile=MEM, **kw):
49
+ return core.run(profile, tmp_path / "scout", gh, today=TODAY, sleep=lambda s: None, **kw)
50
+
51
+
52
+ def load(tmp_path):
53
+ return json.loads((tmp_path / "scout" / "candidates.json").read_text(encoding="utf-8"))["repos"]
54
+
55
+
56
+ def one_query(**extra):
57
+ return Profile({"name": "t", "queries": ["topic:a"], **extra})
58
+
59
+
60
+ def test_queries_carry_the_filters():
61
+ gh = FakeGitHub([])
62
+ core.search(gh, ["topic:agent-memory"], min_stars=10, days=90, today=TODAY, sleep=lambda s: None)
63
+ q = gh.searches[0]
64
+ assert "topic:agent-memory" in q and "fork:false" in q and "archived:false" in q
65
+ assert "stars:>=10" in q and "pushed:>=2026-07-01" in q
66
+
67
+
68
+ def test_ai_memory_drops_forks_archived_stale_small_lists(tmp_path):
69
+ items = [repo("a/good"), repo("a/fork", fork=True), repo("a/old", archived=True),
70
+ repo("a/stale", pushed="2026-01-01T00:00:00Z"), repo("a/tiny", stars=3),
71
+ repo("a/nolang", language=None), repo("a/awesome-agent-memory"),
72
+ repo("a/papers", description="A curated list of memory papers"),
73
+ repo("a/db", description="Git for data, for agents", topics=("agent-memory", "database")),
74
+ repo("a/claude-mem", description="Persistent context across sessions for every agent")]
75
+ run(tmp_path, FakeGitHub(items))
76
+ assert set(load(tmp_path)) == {"a/good", "a/claude-mem"}
77
+
78
+
79
+ def test_github_scrapers_profile():
80
+ gs = Profile.load("github-scrapers")
81
+ cutoff = date(2025, 9, 29)
82
+ ok = repo("n/github-scraper", description="crawl GitHub web pages", topics=())
83
+ spam = repo("x/OnlySnap", description="Scrape OnlyFans, a GitHub mirror", topics=())
84
+ other = repo("y/web-scraper", description="scrape any website", topics=())
85
+ assert core.keep(ok, gs, min_stars=5, cutoff=cutoff)
86
+ assert not core.keep(spam, gs, min_stars=5, cutoff=cutoff)
87
+ assert not core.keep(other, gs, min_stars=5, cutoff=cutoff)
88
+
89
+
90
+ def test_require_own_words_ignores_topics():
91
+ p = one_query(require=[{"pattern": "(?i)memory", "own_words": True}])
92
+ tagged = repo("a/db", description="Git for data", topics=("agent-memory",))
93
+ said = repo("a/mem", description="memory for agents", topics=())
94
+ cutoff = date(2026, 1, 1)
95
+ assert not core.keep(tagged, p, min_stars=1, cutoff=cutoff)
96
+ assert core.keep(said, p, min_stars=1, cutoff=cutoff)
97
+
98
+
99
+ def test_require_language_can_be_off():
100
+ p = one_query(require_language=False)
101
+ assert core.keep(repo("a/docs", language=None), p, min_stars=1, cutoff=date(2026, 1, 1))
102
+
103
+
104
+ def test_interface_guess_from_readme():
105
+ g = lambda s: core.guess_interface(s, MEM) # noqa: E731
106
+ assert g("Add it to Claude: an MCP server (modelcontextprotocol)") == ["mcp"]
107
+ assert g("pip install foo\n\ncurl http://localhost:8080/v1/memories") == ["pip", "rest"]
108
+ assert g("npm install -g bar; docker compose up") == ["npm", "docker"]
109
+ assert g("A Claude Code plugin: /plugin install x") == ["plugin"]
110
+ assert g("") == []
111
+
112
+
113
+ def test_no_interfaces_no_readmes(tmp_path):
114
+ gh = FakeGitHub([repo("a/one")])
115
+ run(tmp_path, gh, profile=one_query())
116
+ assert gh.readme_calls == [] and load(tmp_path)["a/one"]["interface"] is None
117
+
118
+
119
+ def test_known_from_files_and_known_txt(tmp_path):
120
+ ad = tmp_path / "adapters"
121
+ ad.mkdir()
122
+ (ad / "x.py").write_text('"""Adapter for X (https://github.com/Owner/X-Mem)."""\n', encoding="utf-8")
123
+ (ad / "late.py").write_text("\n" * 30 + "# https://github.com/not/this\n", encoding="utf-8")
124
+ out = tmp_path / "scout"
125
+ out.mkdir()
126
+ (out / "known.txt").write_text("# rejected\nother/thing\n", encoding="utf-8")
127
+ known_from = [str(ad / "*.py")]
128
+ assert core.known_repos(out / "known.txt", known_from) == {"owner/x-mem", "other/thing"}
129
+ run(tmp_path, FakeGitHub([repo("owner/X-Mem"), repo("other/thing"), repo("new/one")]), known_from=known_from)
130
+ repos = load(tmp_path)
131
+ assert repos["owner/X-Mem"]["status"] == "known" and repos["new/one"]["status"] == "new"
132
+ md = (out / "new.md").read_text(encoding="utf-8")
133
+ assert "new/one" in md and "owner/X-Mem" not in md
134
+
135
+
136
+ def test_state_across_runs(tmp_path):
137
+ gh = FakeGitHub([repo("a/one", stars=20)], readmes={"a/one": "pip install one"})
138
+ first = run(tmp_path, gh)
139
+ assert [r["full_name"] for r in first["new"]] == ["a/one"] and gh.readme_calls == ["a/one"]
140
+ gh.items = [repo("a/one", stars=25), repo("b/two", stars=90)]
141
+ second = run(tmp_path, gh)
142
+ assert [r["full_name"] for r in second["new"]] == ["b/two"]
143
+ assert gh.readme_calls == ["a/one", "b/two"] # README only for first sightings
144
+ repos = load(tmp_path)
145
+ assert repos["a/one"]["status"] == "seen" and repos["a/one"]["stars"] == 25
146
+ assert repos["a/one"]["first_seen"] == "2026-09-29" and repos["a/one"]["interface"] == ["pip"]
147
+
148
+
149
+ def test_new_md_most_stars_first(tmp_path):
150
+ run(tmp_path, FakeGitHub([repo("a/small", stars=11), repo("a/big", stars=900)]))
151
+ md = (tmp_path / "scout" / "new.md").read_text(encoding="utf-8")
152
+ assert md.startswith("# New repos for ai-memory") and md.index("a/big") < md.index("a/small")
153
+
154
+
155
+ def test_readme_cap(tmp_path):
156
+ gh = FakeGitHub([repo(f"a/r{i}") for i in range(5)])
157
+ run(tmp_path, gh, readme_max=2)
158
+ assert len(gh.readme_calls) == 2
159
+
160
+
161
+ def test_rate_limit_saves_what_was_found(tmp_path):
162
+ gh = FakeGitHub([repo("a/one")], limited_after=1)
163
+ p = one_query()
164
+ p.queries = ["topic:a", "topic:b", "topic:c"]
165
+ summary = run(tmp_path, gh, profile=p)
166
+ assert summary["stopped"] and "rate limit" in summary["stopped"]
167
+ assert set(load(tmp_path)) == {"a/one"}
168
+
169
+
170
+ def test_profile_from_path_and_unknown(tmp_path):
171
+ f = tmp_path / "mine.json"
172
+ f.write_text(json.dumps({"queries": ["topic:x"]}), encoding="utf-8")
173
+ p = Profile.load(str(f))
174
+ assert p.name == "mine" and p.min_stars == 10 and p.days == 90
175
+ with pytest.raises(FileNotFoundError):
176
+ Profile.load("no-such-profile")
177
+ with pytest.raises(ValueError):
178
+ Profile({"queries": []})
179
+
180
+
181
+ def test_builtin_profiles_all_load():
182
+ names = core.builtin_profiles()
183
+ assert {"ai-memory", "github-scrapers"} <= set(names)
184
+ for n in names:
185
+ Profile.load(n)
186
+
187
+
188
+ def test_cli_run_json(tmp_path, monkeypatch, capsys):
189
+ monkeypatch.setattr(core, "fetch_github", FakeGitHub([repo("a/one")]))
190
+ monkeypatch.setattr(core.time, "sleep", lambda s: None)
191
+ rc = cli.main(["run", "ai-memory", "--out", str(tmp_path / "o"), "--json", "-q"])
192
+ out = json.loads(capsys.readouterr().out)
193
+ assert rc == 0 and out["profile"] == "ai-memory" and [r["full_name"] for r in out["new"]] == ["a/one"]
194
+
195
+
196
+ def test_cli_unknown_profile(capsys):
197
+ assert cli.main(["run", "nope"]) == 2
198
+ assert "no profile" in capsys.readouterr().err
@@ -0,0 +1,2 @@
1
+ """topicscout: find the GitHub repos on a topic, and only the new ones next time."""
2
+ __version__ = "0.1.0"
@@ -0,0 +1,3 @@
1
+ from .cli import main
2
+
3
+ raise SystemExit(main())
@@ -0,0 +1,103 @@
1
+ """topicscout run | profiles | doctor.
2
+
3
+ Results go to stdout (a summary line, or JSON with --json); progress and
4
+ errors go to stderr, so a script or an agent can read stdout as is.
5
+ """
6
+ from __future__ import annotations
7
+
8
+ import argparse
9
+ import json
10
+ import os
11
+ import sys
12
+ from datetime import datetime, timezone
13
+ from pathlib import Path
14
+
15
+ from . import __version__
16
+ from .core import Profile, _stderr, builtin_profiles, rate_limit, run
17
+
18
+
19
+ def _run(args) -> int:
20
+ try:
21
+ profile = Profile.load(args.profile)
22
+ except (FileNotFoundError, ValueError, KeyError) as e:
23
+ _stderr(f"topicscout: {e}")
24
+ return 2
25
+ out = Path(args.out or f"scout-{profile.name}")
26
+ log = (lambda m: None) if args.quiet else _stderr
27
+ if not os.environ.get("GITHUB_TOKEN"):
28
+ log("no GITHUB_TOKEN: searches paced to 10 per minute, READMEs capped at 40")
29
+ s = run(profile, out, min_stars=args.min_stars, days=args.days, readme_max=args.readme_max,
30
+ known_from=args.known_from, log=log)
31
+ if args.json:
32
+ print(json.dumps({**s, "out": str(out)}, indent=2, ensure_ascii=False))
33
+ else:
34
+ print(f"{s['found']} found, {s['kept']} kept, {len(s['new'])} new -> {out / 'new.md'}")
35
+ if s["stopped"]:
36
+ _stderr(f"stopped early: {s['stopped']}")
37
+ return 0
38
+
39
+
40
+ def _profiles(args) -> int:
41
+ ps = builtin_profiles()
42
+ if args.json:
43
+ print(json.dumps(ps, indent=2))
44
+ else:
45
+ for name, desc in ps.items():
46
+ print(f"{name:20} {desc}")
47
+ return 0
48
+
49
+
50
+ def _doctor(args) -> int:
51
+ token = bool(os.environ.get("GITHUB_TOKEN"))
52
+ report = {"version": __version__, "token": token}
53
+ try:
54
+ res = rate_limit()
55
+ except Exception as e: # network, bad token: say it, don't crash
56
+ report["error"] = str(e)
57
+ else:
58
+ for key in ("search", "core"):
59
+ r = res[key]
60
+ reset = datetime.fromtimestamp(r["reset"], timezone.utc).strftime("%H:%M:%S UTC")
61
+ report[key] = {"remaining": r["remaining"], "limit": r["limit"], "reset": reset}
62
+ if args.json:
63
+ print(json.dumps(report, indent=2))
64
+ else:
65
+ print(f"topicscout {__version__}")
66
+ print(f"GITHUB_TOKEN: {'set' if token else 'not set (works, but slower and READMEs capped)'}")
67
+ if "error" in report:
68
+ print(f"GitHub: unreachable ({report['error']})")
69
+ else:
70
+ for key, label in (("search", "search API"), ("core", "READMEs (core API)")):
71
+ r = report[key]
72
+ print(f"{label}: {r['remaining']}/{r['limit']} left, resets {r['reset']}")
73
+ return 1 if "error" in report else 0
74
+
75
+
76
+ def main(argv: list[str] | None = None) -> int:
77
+ ap = argparse.ArgumentParser(prog="topicscout",
78
+ description="Find the GitHub repos on a topic, and only the new ones next time.")
79
+ ap.add_argument("--version", action="version", version=f"topicscout {__version__}")
80
+ sub = ap.add_subparsers(dest="cmd", required=True)
81
+
82
+ r = sub.add_parser("run", help="search GitHub with a profile and list what is new")
83
+ r.add_argument("profile", help="built-in profile name (see `topicscout profiles`) or path to a JSON profile")
84
+ r.add_argument("--out", help="state folder: candidates.json, new.md, known.txt (default scout-<profile>)")
85
+ r.add_argument("--min-stars", type=int, help="override the profile's min_stars")
86
+ r.add_argument("--days", type=int, help="override the profile's days: drop repos with no push in this many days")
87
+ r.add_argument("--readme-max", type=int, help="READMEs read per run (default 40, 300 with GITHUB_TOKEN)")
88
+ r.add_argument("--known-from", action="append", default=[], metavar="GLOB",
89
+ help="files whose first 20 lines name repos you already have (repeatable)")
90
+ r.add_argument("--json", action="store_true", help="print the summary and the new repos as JSON")
91
+ r.add_argument("-q", "--quiet", action="store_true", help="no progress on stderr")
92
+ r.set_defaults(func=_run)
93
+
94
+ p = sub.add_parser("profiles", help="list the built-in profiles")
95
+ p.add_argument("--json", action="store_true")
96
+ p.set_defaults(func=_profiles)
97
+
98
+ d = sub.add_parser("doctor", help="token and GitHub quota")
99
+ d.add_argument("--json", action="store_true")
100
+ d.set_defaults(func=_doctor)
101
+
102
+ args = ap.parse_args(argv)
103
+ return args.func(args)
@@ -0,0 +1,274 @@
1
+ """Search, filter, remember: the whole scout, with no dependencies.
2
+
3
+ A profile says what to look for (GitHub search queries), what a real hit
4
+ must say about itself (require), what to drop (exclude) and how to guess
5
+ the interface from the README. The state is kept across runs, so new.md
6
+ holds only what appeared since the last one.
7
+ """
8
+ from __future__ import annotations
9
+
10
+ import glob
11
+ import json
12
+ import os
13
+ import re
14
+ import sys
15
+ import time
16
+ import urllib.error
17
+ import urllib.request
18
+ from datetime import date, datetime, timedelta, timezone
19
+ from pathlib import Path
20
+ from urllib.parse import urlencode
21
+
22
+ API = "https://api.github.com"
23
+ GITHUB_URL = re.compile(r"github\.com/([\w.-]+)/([\w.-]+)")
24
+
25
+
26
+ class RateLimited(Exception):
27
+ pass
28
+
29
+
30
+ def fetch_github(url: str, accept: str) -> tuple[int, dict, str]:
31
+ """GET with the optional GITHUB_TOKEN; returns (status, lower-case headers, body)."""
32
+ headers = {"Accept": accept, "User-Agent": "topicscout", "X-GitHub-Api-Version": "2022-11-28"}
33
+ if os.environ.get("GITHUB_TOKEN"):
34
+ headers["Authorization"] = f"Bearer {os.environ['GITHUB_TOKEN']}"
35
+ req = urllib.request.Request(url, headers=headers)
36
+ try:
37
+ with urllib.request.urlopen(req, timeout=30) as r:
38
+ return r.status, {k.lower(): v for k, v in r.headers.items()}, r.read().decode("utf-8", "replace")
39
+ except urllib.error.HTTPError as e:
40
+ return e.code, {k.lower(): v for k, v in e.headers.items()}, e.read().decode("utf-8", "replace")
41
+
42
+
43
+ def _stderr(msg: str) -> None:
44
+ print(msg, file=sys.stderr, flush=True)
45
+
46
+
47
+ # ---------------------------------------------------------------- profiles
48
+
49
+ PROFILES = Path(__file__).resolve().parent / "profiles"
50
+
51
+
52
+ def builtin_profiles() -> dict[str, str]:
53
+ """name -> description of the profiles shipped with the package."""
54
+ out = {}
55
+ for p in sorted(PROFILES.glob("*.json")):
56
+ out[p.stem] = json.loads(p.read_text(encoding="utf-8")).get("description", "")
57
+ return out
58
+
59
+
60
+ class Profile:
61
+ """A compiled profile: queries, require/exclude patterns, interface guesses."""
62
+
63
+ def __init__(self, data: dict):
64
+ self.name = data.get("name", "profile")
65
+ self.description = data.get("description", "")
66
+ self.queries = list(data["queries"])
67
+ if not self.queries:
68
+ raise ValueError("a profile needs at least one query")
69
+ self.exclude = re.compile(data["exclude"]) if data.get("exclude") else None
70
+ # own_words: the pattern must be in the repo's name or description, not only in a topic
71
+ self.require = [(re.compile(r["pattern"]), bool(r.get("own_words"))) for r in data.get("require", [])]
72
+ self.interfaces = [(n, re.compile(p)) for n, p in data.get("interfaces", [])]
73
+ self.min_stars = int(data.get("min_stars", 10))
74
+ self.days = int(data.get("days", 90))
75
+ self.require_language = bool(data.get("require_language", True))
76
+
77
+ @classmethod
78
+ def load(cls, name_or_path: str) -> "Profile":
79
+ path = Path(name_or_path)
80
+ if not path.is_file():
81
+ path = PROFILES / f"{name_or_path}.json"
82
+ if not path.is_file():
83
+ raise FileNotFoundError(f"no profile {name_or_path!r}: not a file, and built-in ones are "
84
+ f"{', '.join(builtin_profiles())}")
85
+ data = json.loads(path.read_text(encoding="utf-8"))
86
+ data.setdefault("name", path.stem)
87
+ return cls(data)
88
+
89
+
90
+ # ---------------------------------------------------------------- search
91
+
92
+ def _get(fetch, url, accept, sleep):
93
+ """One request; waits out a short rate limit once, raises RateLimited on a long one."""
94
+ for attempt in (0, 1):
95
+ status, headers, body = fetch(url, accept)
96
+ limited = status == 429 or (status == 403 and (headers.get("x-ratelimit-remaining") == "0"
97
+ or "retry-after" in headers))
98
+ if not limited:
99
+ return status, body
100
+ if "retry-after" in headers:
101
+ wait = int(headers["retry-after"])
102
+ else:
103
+ wait = int(headers.get("x-ratelimit-reset", "0")) - int(time.time())
104
+ if attempt or wait > 70:
105
+ raise RateLimited(f"GitHub rate limit, reset in {max(wait, 0)} s")
106
+ sleep(max(wait, 0) + 1)
107
+ raise AssertionError("unreachable")
108
+
109
+
110
+ def search(fetch, queries, *, min_stars, days, today, sleep, pause=0.0, pages=3,
111
+ log=lambda m: None) -> tuple[dict, str | None]:
112
+ """Run the queries; returns ({full_name: item}, reason it stopped early or None)."""
113
+ since = (today - timedelta(days=days)).isoformat()
114
+ found: dict[str, dict] = {}
115
+ first = True
116
+ for i, q in enumerate(queries, 1):
117
+ full = f"{q} fork:false archived:false stars:>={min_stars} pushed:>={since}"
118
+ for page in range(1, pages + 1):
119
+ if not first and pause:
120
+ sleep(pause)
121
+ first = False
122
+ url = f"{API}/search/repositories?" + urlencode(
123
+ {"q": full, "sort": "stars", "order": "desc", "per_page": 100, "page": page})
124
+ try:
125
+ status, body = _get(fetch, url, "application/vnd.github+json", sleep)
126
+ except RateLimited as e:
127
+ return found, str(e)
128
+ if status != 200:
129
+ return found, f"search failed ({status}) on {q!r}: {body[:200]}"
130
+ items = json.loads(body).get("items", [])
131
+ for it in items:
132
+ found.setdefault(it["full_name"], it)
133
+ log(f"[{i}/{len(queries)}] {q} page {page}: {len(items)} ({len(found)} so far)")
134
+ if len(items) < 100:
135
+ break
136
+ return found, None
137
+
138
+
139
+ def keep(item: dict, profile: Profile, *, min_stars: int, cutoff: date) -> bool:
140
+ if item.get("fork") or item.get("archived"):
141
+ return False
142
+ if profile.require_language and not item.get("language"):
143
+ return False
144
+ if item.get("stargazers_count", 0) < min_stars:
145
+ return False
146
+ if datetime.fromisoformat(item["pushed_at"].replace("Z", "+00:00")).date() < cutoff:
147
+ return False
148
+ own = f"{item['full_name']} {item.get('description') or ''}"
149
+ text = f"{own} {' '.join(item.get('topics') or [])}"
150
+ if profile.exclude and profile.exclude.search(text):
151
+ return False
152
+ own_n, text_n = (s.replace("-", " ").replace("_", " ") for s in (own, text))
153
+ return all(rx.search(own_n if own_words else text_n) for rx, own_words in profile.require)
154
+
155
+
156
+ def guess_interface(readme: str, profile: Profile) -> list[str]:
157
+ return [name for name, rx in profile.interfaces if rx.search(readme)]
158
+
159
+
160
+ def known_repos(known_file: Path, known_from: list[str] = (), head: int = 20) -> set[str]:
161
+ """owner/repo (lower case) from known.txt plus the GitHub URLs in the first
162
+ `head` lines of the files matching the known_from globs."""
163
+ known: set[str] = set()
164
+ for pattern in known_from:
165
+ for f in glob.glob(pattern, recursive=True):
166
+ p = Path(f)
167
+ if not p.is_file():
168
+ continue
169
+ top = "\n".join(p.read_text(encoding="utf-8", errors="replace").splitlines()[:head])
170
+ for owner, name in GITHUB_URL.findall(top):
171
+ name = name.rstrip(".").removesuffix(".git")
172
+ known.add(f"{owner}/{name}".lower())
173
+ if known_file.is_file():
174
+ for line in known_file.read_text(encoding="utf-8").splitlines():
175
+ line = line.split("#", 1)[0].strip()
176
+ if line:
177
+ known.add(line.lower())
178
+ return known
179
+
180
+
181
+ # ---------------------------------------------------------------- run
182
+
183
+ def _record(item: dict) -> dict:
184
+ return {"full_name": item["full_name"], "url": item["html_url"], "stars": item["stargazers_count"],
185
+ "pushed": item["pushed_at"][:10], "license": (item.get("license") or {}).get("spdx_id"),
186
+ "language": item.get("language"), "description": item.get("description") or "",
187
+ "topics": item.get("topics") or []}
188
+
189
+
190
+ def _new_md(new: list[dict], profile: Profile, today: date, stopped: str | None) -> str:
191
+ lines = [f"# New repos for {profile.name}, {today.isoformat()}", ""]
192
+ if stopped:
193
+ lines += [f"> Stopped early: {stopped}. The list is partial.", ""]
194
+ if not new:
195
+ return "\n".join(lines + ["Nothing new since the last run.", ""])
196
+ lines += [f"{len(new)} new.", "", "| repo | stars | last push | language | license | interface | description |",
197
+ "|---|---:|---|---|---|---|---|"]
198
+ for r in new:
199
+ iface = ", ".join(r["interface"]) if r.get("interface") else ("?" if r.get("interface") is None else "-")
200
+ desc = r["description"].replace("|", "\\|").replace("\n", " ")[:160]
201
+ lines.append(f"| [{r['full_name']}]({r['url']}) | {r['stars']} | {r['pushed']} | {r['language'] or '-'} | "
202
+ f"{r['license'] or '-'} | {iface} | {desc} |")
203
+ return "\n".join(lines + [""])
204
+
205
+
206
+ def run(profile: Profile, out: Path, fetch=None, *, today: date | None = None, sleep=None,
207
+ min_stars: int | None = None, days: int | None = None, readme_max: int | None = None,
208
+ pause: float | None = None, known_from: list[str] = (), log=lambda m: None) -> dict:
209
+ fetch = fetch or fetch_github
210
+ sleep = sleep or time.sleep
211
+ today = today or datetime.now(timezone.utc).date()
212
+ min_stars = profile.min_stars if min_stars is None else min_stars
213
+ days = profile.days if days is None else days
214
+ token = bool(os.environ.get("GITHUB_TOKEN"))
215
+ if pause is None:
216
+ pause = 2.0 if token else 7.0 # search API: 30/min with a token, 10/min without
217
+ if readme_max is None:
218
+ readme_max = 300 if token else 40
219
+ out.mkdir(parents=True, exist_ok=True)
220
+ state_file = out / "candidates.json"
221
+ state = json.loads(state_file.read_text(encoding="utf-8")) if state_file.is_file() else {"repos": {}}
222
+ repos: dict[str, dict] = state["repos"]
223
+ known = known_repos(out / "known.txt", known_from)
224
+
225
+ found, stopped = search(fetch, profile.queries, min_stars=min_stars, days=days, today=today,
226
+ sleep=sleep, pause=pause, log=log)
227
+ cutoff = today - timedelta(days=days)
228
+ new_names = []
229
+ for name, item in found.items():
230
+ if not keep(item, profile, min_stars=min_stars, cutoff=cutoff):
231
+ continue
232
+ rec = _record(item)
233
+ old = repos.get(name)
234
+ rec["first_seen"] = old["first_seen"] if old else today.isoformat()
235
+ rec["last_seen"] = today.isoformat()
236
+ rec["interface"] = old.get("interface") if old else None
237
+ if name.lower() in known:
238
+ rec["status"] = "known"
239
+ elif old:
240
+ rec["status"] = "seen"
241
+ else:
242
+ rec["status"] = "new"
243
+ new_names.append(name)
244
+ repos[name] = rec
245
+
246
+ # READMEs: first sightings first, then older ones never read (cap reached last time)
247
+ if profile.interfaces:
248
+ todo = new_names + [n for n, r in repos.items() if r["status"] == "seen" and r.get("interface") is None]
249
+ for name in todo[:readme_max]:
250
+ try:
251
+ status, body = _get(fetch, f"{API}/repos/{name}/readme", "application/vnd.github.raw", sleep)
252
+ except RateLimited as e:
253
+ stopped = stopped or f"{e} (while reading READMEs)"
254
+ break
255
+ repos[name]["interface"] = guess_interface(body, profile) if status == 200 else []
256
+ log(f"READMEs read: {min(len(todo), readme_max)} of {len(todo)}")
257
+
258
+ state = {"profile": profile.name, "updated": today.isoformat(),
259
+ "repos": dict(sorted(repos.items(), key=lambda kv: kv[0].lower()))}
260
+ state_file.write_text(json.dumps(state, indent=2, ensure_ascii=False) + "\n", encoding="utf-8")
261
+ new = sorted((repos[n] for n in new_names), key=lambda r: -r["stars"])
262
+ (out / "new.md").write_text(_new_md(new, profile, today, stopped), encoding="utf-8")
263
+ return {"profile": profile.name, "found": len(found),
264
+ "kept": sum(1 for r in repos.values() if r["last_seen"] == today.isoformat()),
265
+ "new": new, "stopped": stopped}
266
+
267
+
268
+ def rate_limit(fetch=None) -> dict:
269
+ """GitHub's own view of the quota; this call does not count against it."""
270
+ fetch = fetch or fetch_github
271
+ status, _, body = fetch(f"{API}/rate_limit", "application/vnd.github+json")
272
+ if status != 200:
273
+ raise RuntimeError(f"rate_limit failed ({status}): {body[:200]}")
274
+ return json.loads(body)["resources"]
@@ -0,0 +1,59 @@
1
+ {
2
+ "name": "ai-memory",
3
+ "description": "Memory systems for AI agents and LLMs (the adebench list)",
4
+ "queries": [
5
+ "topic:agent-memory",
6
+ "topic:ai-memory",
7
+ "topic:llm-memory",
8
+ "topic:memory-mcp",
9
+ "topic:mcp-memory",
10
+ "topic:long-term-memory",
11
+ "topic:agentic-memory",
12
+ "topic:memory-layer",
13
+ "topic:memory-system",
14
+ "\"agent memory\" in:name,description",
15
+ "\"llm memory\" in:name,description",
16
+ "\"memory layer\" in:name,description",
17
+ "\"mcp memory\" in:name,description",
18
+ "\"long-term memory\" in:description",
19
+ "\"persistent memory\" in:description"
20
+ ],
21
+ "exclude": "(?i)awesome|curated list|reading list|\\bpapers\\b|\\bsurvey\\b",
22
+ "require": [
23
+ {
24
+ "pattern": "(?i)memor|\\bmem\\b|remember",
25
+ "own_words": true
26
+ },
27
+ {
28
+ "pattern": "(?i)\\b(agents?|agentic|llms?|ai|mcp|claude|gpt|rag|assistants?|chatbots?)\\b"
29
+ }
30
+ ],
31
+ "interfaces": [
32
+ [
33
+ "mcp",
34
+ "(?i)\\bmcp\\b|modelcontextprotocol"
35
+ ],
36
+ [
37
+ "pip",
38
+ "(?i)\\bpip3? install\\b|\\buv add\\b|\\bpypi\\b"
39
+ ],
40
+ [
41
+ "npm",
42
+ "(?i)\\bnpm (i|install)\\b|\\bnpx\\b|\\bbun add\\b|\\bpnpm add\\b"
43
+ ],
44
+ [
45
+ "rest",
46
+ "\\bcurl\\b[^\n]*https?://|\\bREST\\b|\\bopenapi\\b|/v1/"
47
+ ],
48
+ [
49
+ "docker",
50
+ "(?i)\\bdocker\\b"
51
+ ],
52
+ [
53
+ "plugin",
54
+ "(?i)/plugin install|\\.?claude-plugin|claude code plugin"
55
+ ]
56
+ ],
57
+ "min_stars": 10,
58
+ "days": 90
59
+ }
@@ -0,0 +1,55 @@
1
+ {
2
+ "name": "github-scrapers",
3
+ "description": "Tools that scrape, crawl or mine GitHub itself",
4
+ "queries": [
5
+ "topic:github-scraper",
6
+ "topic:github-crawler",
7
+ "topic:github-scraping",
8
+ "topic:github-mining",
9
+ "\"github scraper\" in:name,description",
10
+ "\"github crawler\" in:name,description",
11
+ "\"scrape github\" in:name,description",
12
+ "\"github scraping\" in:name,description",
13
+ "\"scrapes github\" in:description",
14
+ "\"crawl github\" in:name,description",
15
+ "\"github repository scraper\" in:description",
16
+ "\"github repo scraper\" in:description"
17
+ ],
18
+ "exclude": "(?i)awesome|curated list|onlyfans",
19
+ "require": [
20
+ {
21
+ "pattern": "(?i)scrap|crawl|mining|miner|harvest"
22
+ },
23
+ {
24
+ "pattern": "(?i)github"
25
+ }
26
+ ],
27
+ "interfaces": [
28
+ [
29
+ "cli",
30
+ "(?i)\\busage:|\\$ \\w+ --|\\bcommand line\\b|\\bcli\\b"
31
+ ],
32
+ [
33
+ "pip",
34
+ "(?i)\\bpip3? install\\b|\\buv add\\b|\\bpypi\\b"
35
+ ],
36
+ [
37
+ "npm",
38
+ "(?i)\\bnpm (i|install)\\b|\\bnpx\\b"
39
+ ],
40
+ [
41
+ "action",
42
+ "(?i)uses: [\\w.-]+/[\\w.-]+@|github action"
43
+ ],
44
+ [
45
+ "docker",
46
+ "(?i)\\bdocker\\b"
47
+ ],
48
+ [
49
+ "token",
50
+ "(?i)GITHUB_TOKEN|personal access token"
51
+ ]
52
+ ],
53
+ "min_stars": 5,
54
+ "days": 365
55
+ }
@@ -0,0 +1,151 @@
1
+ Metadata-Version: 2.4
2
+ Name: topicscout
3
+ Version: 0.1.0
4
+ Summary: Find the GitHub repos on a topic, and only the new ones next time. No dependencies, token optional.
5
+ Author: Adecubed
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/adecubed/topicscout
8
+ Keywords: github,scraper,search,discovery,topics
9
+ Requires-Python: >=3.11
10
+ Description-Content-Type: text/markdown
11
+ License-File: LICENSE
12
+ Dynamic: license-file
13
+
14
+ # topicscout
15
+
16
+ [![ci](https://github.com/adecubed/topicscout/actions/workflows/ci.yml/badge.svg)](https://github.com/adecubed/topicscout/actions/workflows/ci.yml)
17
+ [![PyPI](https://img.shields.io/pypi/v/topicscout)](https://pypi.org/project/topicscout/)
18
+
19
+ Find the GitHub repositories on a topic, and next time only the new ones.
20
+
21
+ ```bash
22
+ pip install topicscout # or run it without installing: uvx topicscout run github-scrapers
23
+ topicscout run github-scrapers
24
+ ```
25
+
26
+ ```
27
+ 19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
28
+ ```
29
+
30
+ `new.md` is a table of what appeared since the last run, most stars first: stars,
31
+ last push, language, license, a guess at the interface from the README (MCP, pip,
32
+ npm, REST, Docker...) and the description. `candidates.json` keeps everything seen,
33
+ so a weekly run is a short list, not the same 300 repos again.
34
+
35
+ No dependencies, Python 3.11+. `GITHUB_TOKEN` is optional: without it the searches
36
+ are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.
37
+
38
+ ## Why
39
+
40
+ GitHub search is good at "repos with this topic" and bad at "what's new since I last
41
+ looked, minus the forks, the awesome-lists, the abandoned ones and the repos that
42
+ only tagged themselves". topicscout runs a fixed set of searches, applies your
43
+ filters, remembers what it has seen and what you already have.
44
+
45
+ It was the discovery step of [adebench](https://github.com/adecubed/adebench), a
46
+ benchmark for AI memory systems, which uses it to find new memories to measure.
47
+
48
+ ## Your own topic
49
+
50
+ Copy a built-in profile (they are in [`topicscout/profiles/`](topicscout/profiles/)),
51
+ change the queries and the words a hit must contain, and run it:
52
+
53
+ ```bash
54
+ topicscout run my-topic.json
55
+ ```
56
+
57
+ Run it again next week: `new.md` lists only what appeared in between. Put the repos
58
+ you already know or rejected in `scout-my-topic/known.txt`.
59
+
60
+ ## From a script or an agent
61
+
62
+ `--json` prints the result on stdout, progress stays on stderr:
63
+
64
+ ```bash
65
+ topicscout run github-scrapers --json -q
66
+ ```
67
+
68
+ ```json
69
+ {"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
70
+ ```
71
+
72
+ It is a command-line tool and a Python library (`from topicscout.core import Profile, run`),
73
+ not an MCP server; an agent calls it like any other command.
74
+
75
+ ## Commands
76
+
77
+ ```bash
78
+ topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
79
+ [--known-from GLOB ...] [--json] [-q]
80
+ topicscout profiles # built-in profiles
81
+ topicscout doctor # token and remaining GitHub quota
82
+ ```
83
+
84
+ - `PROFILE` is a built-in name or a path to your own JSON profile.
85
+ - `--out` is the state folder (default `scout-<profile>`): `candidates.json`,
86
+ `new.md`, and your `known.txt`.
87
+ - `--json` prints the summary and the new repos as JSON on stdout; progress and
88
+ errors always go to stderr, so scripts and agents can read stdout as is.
89
+
90
+ ## What counts as known
91
+
92
+ Repos you already have are marked `known` and never listed as new:
93
+
94
+ - `known.txt` in the state folder: one `owner/repo` per line, `#` comments. Put
95
+ rejected repos here too.
96
+ - `--known-from "adapters/*.py"`: the GitHub URLs in the first 20 lines of those
97
+ files. adebench points it at its adapters, so a memory with an adapter drops out.
98
+
99
+ ## Profiles
100
+
101
+ Built in: `ai-memory` (memory systems for AI agents and LLMs) and `github-scrapers`
102
+ (tools that scrape, crawl or mine GitHub). A profile is a JSON file:
103
+
104
+ ```json
105
+ {
106
+ "name": "ai-memory",
107
+ "description": "Memory systems for AI agents and LLMs",
108
+ "queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
109
+ "exclude": "(?i)awesome|curated list|\\bpapers\\b",
110
+ "require": [
111
+ {"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
112
+ {"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
113
+ ],
114
+ "interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
115
+ "min_stars": 10,
116
+ "days": 90
117
+ }
118
+ ```
119
+
120
+ - `queries`: [GitHub repository search](https://docs.github.com/en/search-github/searching-on-github/searching-for-repositories)
121
+ queries. Each one always gets `fork:false archived:false stars:>=min_stars
122
+ pushed:>=today-days`; up to 300 results per query.
123
+ - `exclude`: a regex on name, description and topics; a match drops the repo.
124
+ - `require`: every pattern must match. With `own_words` it must be in the name or
125
+ description, not only in a topic: a database that tagged itself `agent-memory`
126
+ is not a memory. Hyphens and underscores count as spaces.
127
+ - `interfaces`: `[label, regex]` pairs tried on the README of each new repo.
128
+ Leave it out and no README is fetched.
129
+ - `require_language` (default true) drops repos with no language, which are
130
+ usually lists and docs.
131
+
132
+ Patterns are Python regexes; add `(?i)` for case-insensitive.
133
+
134
+ ## Rate limits
135
+
136
+ On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
137
+ stops, saves what it found and says so in `new.md` and on stderr. The next run
138
+ picks up the READMEs it did not get to.
139
+
140
+ ## How it compares
141
+
142
+ - [ghcrawl](https://github.com/pwrdrvr/ghcrawl) goes deep into one repository
143
+ (issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a
144
+ topic.
145
+ - [top-github-scraper](https://github.com/khuyentran1401/top-github-scraper)
146
+ lists the top repos for a keyword, once; topicscout adds filters, profiles and
147
+ memory across runs.
148
+
149
+ ## License
150
+
151
+ MIT
@@ -0,0 +1,15 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ tests/test_topicscout.py
5
+ topicscout/__init__.py
6
+ topicscout/__main__.py
7
+ topicscout/cli.py
8
+ topicscout/core.py
9
+ topicscout.egg-info/PKG-INFO
10
+ topicscout.egg-info/SOURCES.txt
11
+ topicscout.egg-info/dependency_links.txt
12
+ topicscout.egg-info/entry_points.txt
13
+ topicscout.egg-info/top_level.txt
14
+ topicscout/profiles/ai-memory.json
15
+ topicscout/profiles/github-scrapers.json
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ topicscout = topicscout.cli:main
@@ -0,0 +1 @@
1
+ topicscout